“Audio generation” can mean a voice reading a script, a music bed under a demo, or the sound of a door closing. Those are different jobs. A great speech model does not automatically make good music.
The shortlist
| Model | Category | Good starting use | Available in Dexto |
|---|---|---|---|
| Cartesia Sonic 3.6 | Speech | Natural-sounding interactive speech | No |
| Inworld Realtime TTS-2 | Speech | A streaming voice application | No |
| Eleven v3 | Speech | Expressive narration and voiceover | Yes |
| Fish Audio S2.1 Pro | Speech | Directed speech with usage-based pricing | No |
| Eleven Flash v2.5 | Speech | A latency-focused alternative to v3 | No |
| Kokoro 82M | Speech | A compact open-weight baseline | No |
| Stable Audio 3 Small | Music | Trying several instrumental directions | Yes |
| Lyria 2 | Music | A short instrumental cue | Yes |
| Eleven Music | Music | A separately generated soundtrack | No |
| ElevenLabs Sound Effects v2 | Effects | Foley, ambience, and interface sounds | No |
The Dexto column shows available media models. Browse the model catalog for current options.
What the speech benchmarks measure
The Artificial Analysis Provider Voice Arena compares speech using providers' native voices. Snapshot for the models above:
| Model | Elo | 95% interval |
|---|---|---|
| Sonic 3.6 | 1282 | ±17 |
| Realtime TTS-2 | 1252 | ±18 |
| Eleven v3 | 1175 | ±11 |
| Fish Audio S2.1 Pro | 1141 | ±13 |
| Flash v2.5 | 1080 | ±11 |
| Kokoro 82M v1.0 | 1060 | ±11 |
Scores measure listener preference using each provider’s native voices. Results apply to the exact model versions listed. Source: Artificial Analysis, September 8, 2026.
1. Cartesia Sonic 3.6
Sonic 3.6 is a strong candidate when you are evaluating natural speech for an interactive application. Cartesia documents emotional direction, voice cloning, multilingual output, and pronunciation dictionaries. Those controls matter when a pleasant generic voice still gets your product name wrong.
Test a short response containing a name, a date, and a correction. Listen for whether the phrasing makes the correction clear. Then measure time to first audible output through your actual application, rather than treating a provider's model-latency claim as end-to-end latency.
Pricing is subscription-based: the page lists a $5/month Pro plan with 100,000 model credits and a $49/month Startup plan with 1.25 million. Credits are shared across supported products, so work from the TTS allowance and your expected usage instead of assuming every credit is interchangeable with a character.
2. Inworld Realtime TTS-2
Inworld's Realtime TTS-2 is worth comparing when speech must start quickly and respond to delivery instructions. Its streaming focus makes it relevant to voice applications, where a slightly better reading may be less useful if the user waits too long to hear it.
Current on-demand pricing is $25 per million characters for TTS-2. Subscription tiers reduce the usage rate; the Creator plan lists $20 per million alongside its monthly commitment. The separate Flash model has different pricing and should be evaluated separately.
For a 10,000-character script, the on-demand TTS-2 rate gives $0.25 of synthesis usage. That does not include the language model, transcription, or other parts of a voice agent.
3. Eleven v3
Eleven v3 is the practical starting point for voiceover inside Dexto. It has a built-in preset, and the model supports expressive delivery through audio tags.
Start with a clean script. Add direction only where it changes the meaning: a quieter sentence, a pause, a more excited announcement. Generate a plain version too. More instructions do not necessarily create a better performance.
The fal endpoint lists $0.10 per 1,000 characters, so 10,000 characters cost $1 at that provider rate.
Review names, numbers, pacing, and the last sentence of each paragraph. For a demo, measure the reading against the video timeline before generating the entire narration. A polished sixty-second read is not useful if the edit only has forty seconds for it.
4. Fish Audio S2.1 Pro
Fish Audio S2.1 Pro is another speech option to evaluate when you want to direct a performance and pay by usage. Fish's developer page lists the model and streaming API options.
Its pricing has an easy-to-miss detail: the paid rate is $15 per million UTF-8 bytes, not universally per million characters. Plain ASCII text uses one byte per character; many other characters use more. A 10,000-byte input is $0.15, but a 10,000-character multilingual script can be larger. Check the pricing and rate limits.
Try the same passage with restrained and expressive direction. Listen for unwanted additions as well as missed words.
5. Eleven Flash v2.5
Flash v2.5 gives you a useful comparison against Eleven v3 when quick response matters more than a highly directed performance. ElevenLabs documents the models separately, with different latency and capability tradeoffs.
The API pricing page lists $0.05 per 1,000 characters for Flash and Turbo, versus $0.10 for v2/v3. A 10,000-character input is therefore $0.50 in synthesis usage at the listed Flash rate.
ElevenLabs notes that Flash does not normalize numbers by default. Test dates, currencies, and phone numbers explicitly. Use short utterances in your evaluation: confirmations, questions, and a response that must be interrupted. An audiobook paragraph alone will not tell you how a voice behaves in a conversation.
6. Kokoro 82M
Kokoro is a compact open-weight speech model. It belongs in a shortlist when you want a smaller baseline or need to explore running speech generation yourself.
Hosting changes the economics. The American English endpoint on fal lists $0.02 per 1,000 characters, or $0.20 for 10,000 characters. Running the weights yourself has a different cost: hardware time, deployment, and maintenance. “Open weights” does not mean zero operating cost.
Compare pronunciation and long-paragraph stability before deciding that its size or price makes it sufficient. Use the exact voice and language endpoint you intend to ship.
7. Stable Audio 3 Small
For music inside Dexto, Stable Audio 3 Small is an inexpensive way to explore a few directions. The music endpoint describes stereo compositions up to two minutes and lists $0.0217 per generated audio output.
That unit matters: it is per output, not a per-minute rate. Ten generations are $0.217 at the quoted provider price.
Give it a production brief: mood, instruments, approximate tempo, duration, and whether vocals are wanted. For narration, ask for a restrained instrumental cue with space in the middle frequencies. Then listen under the actual voice track. Music that sounds good alone can make every spoken word harder to hear.
8. Lyria 2
Lyria 2 is the other music preset available in Dexto. It is a straightforward candidate for a short cue when you want to compare a second musical interpretation of the same brief.
The fal endpoint lists $0.10 per thirty seconds. Budget for the actual output units, then include retries. Generating five alternatives is a different expense from getting one usable cue on the first attempt.
Try a specific prompt such as “a calm instrumental bed for a product walkthrough, soft percussion, warm keys, no dramatic build.” Review the opening, the ending, and the sections where narration is busiest. Those are usually more important to an edit than whether the first few seconds are impressive.
9. Eleven Music
Eleven Music generates soundtracks from a brief. The fal music endpoint offers a separate music-generation route.
Provider choice changes the quote. fal lists $0.60 per output minute, rounded up to the next whole minute. A thirty-second output is billed as one minute there. ElevenLabs' direct API page lists a different music rate, $0.15 per minute, with its own plan and metering terms. Compare the route you will actually use.
Evaluate the arrangement over its full length. Does it leave room for a voice? Does the ending work without an awkward cut? If you need vocals, include lyric intelligibility in the listening test.
10. ElevenLabs Sound Effects v2
Sound effects are a separate generation task from narration and music.
ElevenLabs documents
an endpoint for generating effects and ambience from descriptions, with
duration and prompt controls. The
API reference
identifies the default model as eleven_text_to_sound_v2.
Use it for a specific event: a short mechanical click, footsteps on gravel, or a room-tone loop. Describe the material, distance, environment, and timing. “A satisfying button sound” gives the model less useful information than “a short, dry mechanical click with no reverb.”
The API pricing page lists $0.12 per minute for sound effects. It also describes generation-based metering, so check the actual billable unit and selected duration before extrapolating tiny per-second costs for very short effects.
For evaluation, judge how cleanly the effect starts and stops and how well it aligns with the event in your video.
Compare the billable units before comparing prices
| Pricing unit | What to record |
|---|---|
| Characters | Exact script length, including repeated generations |
| UTF-8 bytes | Encoded input size, especially for multilingual scripts |
| Model credits | Plan, TTS conversion, included allowance, and overages |
| Output seconds or minutes | Generated duration and any rounding rule |
| Per generation | Number of outputs, including rejected candidates |
Speech APIs commonly bill for characters or bytes; music APIs may bill for output duration or each generation. Keep those units explicit when estimating a project.
For a project, use total audio spend ÷ accepted minutes. If a pronunciation issue makes you regenerate the whole script three times, all three attempts belong in the cost. Save the approved pronunciation and delivery instructions so the next run does not repeat that work.
A small listening test you can repeat
For speech, prepare five passages: a normal paragraph, a sentence with names, one with dates and amounts, an emotional line, and a longer paragraph. Use native listeners for the languages you will ship. Hide model names during listening where practical.
For music, give every candidate the same brief and listen under the same edit at similar loudness. Check arrangement, unwanted vocals, repetition, and whether the ending is usable. For effects, check the attack, tail, background noise, and synchronization.
Record acceptance, generation time, and total cost separately. A preference score and a fast demo are useful clues, but neither tells you how many retries your project will need.
Start with the workflow in Dexto
If you are making a narrated product video in Dexto, start with Eleven v3 for the voice, then compare Stable Audio 3 Small and Lyria 2 for the music. Approve the script before generating speech. Keep the voice and music as separate assets so you can revise one without regenerating both.
Ask the agent to save the brief, selected model, and output with the project. Put recurring pronunciation and music directions in a skill. The video model comparison covers the visual side, including native audio and what it adds to the price.