Music and text to speech
Generate original music tracks from a description with Music Generation, and turn any text into natural-sounding spoken audio with Text to Speech.
Two audio apps cover most of what a team needs for launches, videos and presentations. Music Generation writes an original track from a description: background music for a product video, a jingle, a loop. Text to Speech reads your text aloud in the voice and language you choose: a voice-over, an announcement, an audio version of an update. Both run from Apps and give you a file to download.
Music generation
- Open Music Generation
Go to Apps and choose Music Generation. The New track screen opens.
- Describe the music
In PROMPT, describe the mood, the instruments, the tempo and where it will be used:
Upbeat, optimistic product-launch background track: warm electric piano, tight modern drums, plucky synth bass, subtle handclaps, 112 BPM, confident and clean, no vocals. Good under a 30-second product video. - Choose the model
Model Best for Eleven Music v2 (default) Studio-grade music from natural language prompts. Lyria 3 Pro Full-length songs with verses, choruses, bridges, vocals and timed lyrics. Lyria 3 Clip Short 30-second music clips and loops. - Set the output and generate
Under OUTPUT SETTINGS, the Output format is MP3. Set the Duration from 3 to 300 seconds, and switch on Instrumental if you want no vocals. Then select Generate track.
When the track is Done, its page has an audio player, the LYRICS (or <instrumental> for a track without words), your PROMPT, both with Copy, and DETAILS: model, duration, format, whether it is instrumental, the cost in credits and the date. Select Download to save the MP3.
Tips for music prompts
- Say where it will play. “Under a 30-second product video” or “hold music” shapes the arrangement more than any genre label.
- Give a tempo in BPM when the timing matters, for example to match video cuts.
- Name three or four instruments, not ten.
- Say “no vocals” in the prompt as well as using Instrumental, so the description and the setting agree.
- Pick the model for the length. Use Lyria 3 Clip for short clips and loops, and Lyria 3 Pro when you want a full song with vocals and lyrics.
Text to speech
- Open Text to Speech
Go to Apps and choose Text to Speech. The New speech screen opens.
- Paste or type the text
Put the script in TEXT. A character count under the box shows how long it is.
- Choose the model, voice and language
- MODEL: Eleven v3 is the default, for emotionally rich delivery. Gemini text-to-speech models such as Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS are also available.
- VOICE: pick a voice from the list, for example Darian - Warm Grounded Storyteller.
- LANGUAGE: leave it on Auto-detect, or set it when the text mixes languages or is short.
- Generate
Select Generate speech.
The finished speech has a player, the TEXT with Copy, and DETAILS: model, voice, language, cost in credits and date. Select Download to save the audio.
Tips for speech
- Read the script aloud once yourself before generating. If you stumble, the voice will too.
- Spell out numbers, abbreviations and names the way they should sound.
- Match voice to purpose. A warm storyteller voice suits an internal update; a crisp voice suits a product demo.
- Split long scripts into sections, so you can regenerate one part without redoing the rest.