Music and text to speech

Generate original music tracks from a description with Music Generation, and turn any text into natural-sounding spoken audio with Text to Speech.

3 min read · Updated 28 September 2026

Two audio apps cover most of what a team needs for launches, videos and presentations. Music Generation writes an original track from a description: background music for a product video, a jingle, a loop. Text to Speech reads your text aloud in the voice and language you choose: a voice-over, an announcement, an audio version of an update. Both run from Apps and give you a file to download.

Music generation

  1. Open Music Generation

    Go to Apps and choose Music Generation. The New track screen opens.

  2. Describe the music

    In PROMPT, describe the mood, the instruments, the tempo and where it will be used:

    Upbeat, optimistic product-launch background track: warm electric piano,
    tight modern drums, plucky synth bass, subtle handclaps, 112 BPM,
    confident and clean, no vocals. Good under a 30-second product video.
  3. Choose the model
    ModelBest for
    Eleven Music v2 (default)Studio-grade music from natural language prompts.
    Lyria 3 ProFull-length songs with verses, choruses, bridges, vocals and timed lyrics.
    Lyria 3 ClipShort 30-second music clips and loops.
  4. Set the output and generate

    Under OUTPUT SETTINGS, the Output format is MP3. Set the Duration from 3 to 300 seconds, and switch on Instrumental if you want no vocals. Then select Generate track.

When the track is Done, its page has an audio player, the LYRICS (or <instrumental> for a track without words), your PROMPT, both with Copy, and DETAILS: model, duration, format, whether it is instrumental, the cost in credits and the date. Select Download to save the MP3.

Tips for music prompts

  • Say where it will play. “Under a 30-second product video” or “hold music” shapes the arrangement more than any genre label.
  • Give a tempo in BPM when the timing matters, for example to match video cuts.
  • Name three or four instruments, not ten.
  • Say “no vocals” in the prompt as well as using Instrumental, so the description and the setting agree.
  • Pick the model for the length. Use Lyria 3 Clip for short clips and loops, and Lyria 3 Pro when you want a full song with vocals and lyrics.

Text to speech

  1. Open Text to Speech

    Go to Apps and choose Text to Speech. The New speech screen opens.

  2. Paste or type the text

    Put the script in TEXT. A character count under the box shows how long it is.

  3. Choose the model, voice and language
    • MODEL: Eleven v3 is the default, for emotionally rich delivery. Gemini text-to-speech models such as Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS are also available.
    • VOICE: pick a voice from the list, for example Darian - Warm Grounded Storyteller.
    • LANGUAGE: leave it on Auto-detect, or set it when the text mixes languages or is short.
  4. Generate

    Select Generate speech.

The finished speech has a player, the TEXT with Copy, and DETAILS: model, voice, language, cost in credits and date. Select Download to save the audio.

Tips for speech

  • Read the script aloud once yourself before generating. If you stumble, the voice will too.
  • Spell out numbers, abbreviations and names the way they should sound.
  • Match voice to purpose. A warm storyteller voice suits an internal update; a crisp voice suits a product demo.
  • Split long scripts into sections, so you can regenerate one part without redoing the rest.