Top text-to-speech APIs in 2026
This guide compares the 12 best TTS APIs in 2026, covering their voice quality, latency, pricing, and ideal use cases to help you choose the right solution for your project.



A text-to-speech API converts written text into spoken audio over an HTTP or WebSocket call. For 2026 the strongest general-purpose options are ElevenLabs and Google Cloud TTS for voice quality, Cartesia and Rime for sub-200ms streaming, and Amazon Polly and Microsoft Azure for enterprise SSML control. If you're building a voice agent, the TTS choice matters less than the speech-to-text layer feeding it.
That last sentence is the one most comparison posts skip, and it's the one that decides whether your agent works. We'll come back to it.
Text-to-speech API comparison, 2026
| Vendor | Model | Pricing shape | Latency | Best for | Price last verified |
|---|---|---|---|---|---|
| Rime | Mist v2, Arcana, Coda | Per million characters, with a free tier | Streaming-first; on-prem option | Conversational voice agents | Verify at rime.ai/pricing |
| ElevenLabs | Not named in vendor-neutral sources; confirm current family | Subscription tiers with a character allowance, then overage | Streaming available | Voice cloning, dubbing, expressive narration | Verify at elevenlabs.io/pricing |
| OpenAI | GPT-4o mini TTS — 11 built-in voices | Usage-based, inside existing OpenAI billing | Moderate | Teams already on the OpenAI platform | Voice count verified Aug 19, 2026; rate not verified |
| Google Cloud TTS | Standard, WaveNet, Neural2, Chirp families — confirm current set | Per character, tiered by voice family, monthly free allowance | Moderate | Multilingual apps, GCP-native stacks | Verify at cloud.google.com/text-to-speech/pricing |
| Microsoft Azure | Neural, Standard, Custom Neural Voice | Per character by voice class; custom voice quoted | Low to moderate | Widest language coverage, viseme and lip-sync work | Verify at azure.microsoft.com pricing |
| Amazon Polly | Neural and Standard engines | Per character; AWS free-tier terms changed structure in 2025 | Moderate | AWS-native stacks, batch synthesis | Verify at aws.amazon.com/polly/pricing |
| Murf.ai | Studio platform plus API | Subscription tiers | Not a real-time product | Marketing and e-learning voiceover | Verify at murf.ai/pricing |
| Play.ht | Long-form voice models | Subscription tiers with character allowances | Moderate | Audiobooks, podcasts, article narration | Verify at play.ht/pricing |
| Telnyx | Natural and NaturalHD, plus hosted third-party engines | Per character, varies by engine; bundled agent rate available | Streaming on owned telephony network | Phone-based agents, one bill across engines | Verify at telnyx.com/pricing |
| Cartesia | Sonic — confirm current version (Sonic 3.6 shipped in beta) | Per character with volume discounts | Among the lowest published time-to-first-audio | Latency-critical agents, edge deployment | Verify at cartesia.ai/pricing |
| MiniMax | Speech 2.8, HD and Turbo variants | Per character, token-based plan available | Low on the Turbo variant | Multilingual content with emotional range | Verify at minimax.io pricing |
| Soniox | TTS v2 | Per generated hour | Streaming, character-level timestamps | Mid-sentence language switching, cloning | Verify at soniox.com/pricing |
The pattern in that table is worth naming. Half of these companies price per character, a quarter price per subscription tier, and the rest price per generated hour or per minute of call. Those three units are not convertible without knowing your average utterance length, which is why "which is cheapest" has no answer until you've measured your own traffic.
What is a text-to-speech API?
A text-to-speech API is a service that converts written text into spoken audio using neural voice models. You send text to an HTTP or WebSocket endpoint and get synthesized speech back, either as streaming audio chunks or as a complete file.
Under the hood, most systems run through four stages:
- Text normalization expands abbreviations, numbers, and symbols into speakable words, so "$3.50" becomes "three dollars and fifty cents".
- Grapheme-to-phoneme conversion turns characters into phonetic representations.
- Prosody prediction decides stress, intonation, and rhythm.
- A neural vocoder generates the final waveform.
Most APIs also support SSML, an XML-like markup for controlling pronunciation, pauses, emphasis, and speaking rate.
The distinction that actually changes your architecture is streaming versus batch. Streaming TTS returns audio as it's generated, which is the only workable option for anything interactive. Batch synthesis builds the whole file first, which is fine for podcasts and audiobooks and disastrous for a phone call.
Key TTS terms
- TTFB (time-to-first-byte)—the gap between sending text and receiving the first audio frame. For real-time use cases this is the only latency number that matters.
- Neural vocoder—the model that generates the waveform, replacing older concatenative and parametric methods.
- Voice cloning—building a synthetic replica of a specific voice from a short sample.
- Prosody—rhythm, stress, and intonation. It's what separates "reading" from "talking".
- Word boundaries—timestamps that tell you which word the synthesizer was on. If your agent can be interrupted, you need these to know how much of its answer the caller actually heard.
Whichever TTS you pick, your agent is only as good as what it heard. Get a free API key and test Universal-3.6 Pro Realtime on your own audio.
How do you choose a text-to-speech API?
The right provider depends entirely on what you're building. Seven criteria, roughly in the order they eliminate candidates:
Latency. For voice agents you need sub-300ms time-to-first-byte. For pre-recorded content it barely matters. This single factor removes most of the list from consideration for one use case or the other, so decide it first.
Voice quality. The gap between providers is still audible, and buyers know it: in our 2026 Voice Agent Report—455 responses, fielded Q4 2025 to Q1 2026—66% of respondents named natural-sounding voice as a priority when choosing a voice agent stack, and 37.5% reported a robotic or unnatural voice as a frustration they have actually experienced. So test with your actual content, not with the demo script the vendor tuned against, and test it on the audio path your users have. Synthetic speech that sounds warm over studio monitors can sound uncanny through a phone codec.
Language support. Check which languages have neural voices, not which languages are listed. The quality gap between a provider's primary and secondary languages is often larger than the gap between providers.
Pricing model. Per character, per minute, per hour, or per subscription. A provider that's cheapest at 100K characters a month can be the most expensive at 10M.
Customization. Voice cloning, custom lexicons, and emotional control are not universal, and the providers that offer them price them differently.
Compliance. If you're in healthcare, finance, or government, filter for SOC 2, data residency, and a signable BAA early. Retrofitting compliance is miserable.
Ecosystem fit. If your stack already lives on AWS or inside OpenAI, the native option removes a vendor relationship. Don't underestimate what that's worth.
Use-case recommendations
- Voice agents: prioritize streaming and time-to-first-byte. Rime, Cartesia, and Telnyx are built for it.
- Audiobooks and podcasts: consistency over speed. ElevenLabs, Play.ht, and MiniMax.
- Customer service at scale: multilingual coverage plus uptime. Azure and Google Cloud.
- Global apps: Azure has the broadest language list of anyone here.
- Prototyping: whichever one your stack already authenticates against.
Which text-to-speech APIs are free?
There is no genuinely free commercial TTS API at production volume. There are three things people mean when they ask, and they're different.
| Option | What "free" means | Catch | Last verified |
|---|---|---|---|
| Web Speech API (browser) | Genuinely free, no key, no server | Voices come from the user's OS, so output is inconsistent across devices and you can't record it server-side | W3C spec, stable |
| Self-hosted open source (Coqui TTS and successors) | No per-character cost | You pay in GPU time, ops, and latency tuning instead | Verify project status before adopting |
| Commercial free tiers (Polly, Google Cloud, Rime and others) | A monthly or first-year character allowance | Allowances and eligibility windows change often, and no free tier carries an SLA | Verify on each vendor's pricing page |
The practical answer: prototype on a free tier or in the browser, and price the paid tier before you write the integration, not after. If what you actually need is free speech-to-text rather than speech synthesis, that's a different list—see our roundup of free speech-to-text APIs and open-source engines.
The 12 best text-to-speech APIs in 2026
1. Rime
Rime's angle is that most TTS models are trained to sound like a narrator, and a narrator is the wrong reference for a conversation. Their models are built around how real people actually speak, which shows up as cadence and stress variation rather than as raw fidelity. Three model tiers ship today: Mist v2 for general-purpose work, Arcana for expressive output, and the newer Coda.
Best for conversational agents, regulated industries that need SOC 2 and a BAA, and on-premise deployments where the network hop is the latency budget.
2. ElevenLabs
If you've heard a synthetic voice that genuinely fooled you, this is usually who made it. Emotional range, natural pauses, and micro-expressions are where the model spends its capacity. Cloning from a short sample is the headline capability, and it's the reason most media and dubbing teams end up here.
One architectural note for agent builders: ElevenLabs' conversational product has historically enforced a hard concurrency cap on simultaneous agents. If you're planning hundreds of concurrent sessions, confirm the current limit before you design around it.
3. OpenAI
This entry has been wrong on a lot of listicles, including an earlier version of this one, which said six voices. It's 11, and the model has a name.
Simplicity is the actual feature. There's no model zoo and no parameter surface to learn: send text, pick a voice, get audio. For teams already authenticated against OpenAI, that's often worth more than a few points of expressiveness.
4. Google Cloud TTS
Where Google wins is precision. Custom pronunciation dictionaries, sub-sentence emphasis, and fine-grained speed control give you a level of authority over the output that most competitors don't expose at all. If you're synthesizing product names, drug names, or anything with a pronunciation your users will notice you getting wrong, this matters more than voice quality does.
5. Microsoft Azure TTS
The language breadth is not a checkbox—it's the reason Azure is the default for genuinely global deployments. Viseme generation is the other differentiator, and it's a niche that no one else really serves: if you're driving an avatar's mouth, you need frame-level phoneme timing, and Azure gives it to you.
6. Amazon Polly
Polly plugs into Lambda, S3, and CloudFront without extra auth plumbing, offers custom lexicons for pronunciation control, and has an asynchronous synthesis mode that's genuinely useful for large batch jobs. Queue it, walk away, collect the output.
7. Murf.ai
This is the entry that doesn't belong in a developer comparison and belongs on the list anyway, because it's what a lot of teams actually need. Marketing, instructional design, and internal comms don't want an SDK. They want drag-and-drop timing, visual pitch adjustment, and a review workflow. Murf has all three, and an API for when the workflow eventually needs automating.
8. Play.ht
Long-form is a different problem from conversational synthesis. Over 40 minutes, small rhythmic monotony that's invisible in a 10-second demo becomes unbearable. Play.ht optimizes for the long tail of that, and bundles enough of the publishing pipeline that you can go from text to a hosted episode without leaving the platform.
9. Telnyx
The interesting part is the routing model. Telnyx offers its own Natural and NaturalHD voices alongside hosted third-party engines and bring-your-own-key options, so you pick the engine per use case and keep one bill. Removing the network hop between synthesis and playback is a real latency win in telephony, and it's a win you can't buy from a pure TTS vendor.
Long-form audio is outside its lane. The design center is real-time agents on phone calls.
10. Cartesia
If you obsess over milliseconds, this is the vendor that shares your obsession. WebSocket streaming means audio starts arriving almost immediately, and the edge story means you can put synthesis on-device or on local infrastructure instead of round-tripping to a region.
11. MiniMax
The emotion control is the differentiator worth testing. Most providers give you a handful of style presets. MiniMax lets you dial specific emotions, which matters if you're localizing narrative content across markets. The HD-versus-Turbo split is a useful knob: trade a small amount of quality for meaningfully lower latency, per use case rather than per account.
12. Soniox
Two things make this entry worth watching. Per-generated-hour pricing is unusual in a per-character market and makes cost modeling much easier for conversational workloads, where utterance lengths vary wildly. And character-level timestamps are the same primitive that makes interruption handling tractable—you can't correctly resume a sentence you don't know the caller only heard half of.
Test streaming transcription speed and accuracy on your own audio before you pick a TTS vendor. No code, no card.
How do you call a text-to-speech API from JavaScript or Python?
The cheapest possible TTS integration doesn't involve an API key at all. The browser's Web Speech API is a W3C standard and every modern browser implements the synthesis half:
const utterance = new SpeechSynthesisUtterance(
"Your call is being transcribed in real time."
);
// Voices are supplied by the user's operating system,
// so the list differs per device. Pick defensively.
const speakNow = () => { const voices = speechSynthesis.getVoices();
utterance.voice = voices.find((v) => v.lang === "en-US") ?? voices[0];
utterance.rate = 1.0;
utterance.pitch = 1.0;
speechSynthesis.speak(utterance); }; if (speechSynthesis.getVoices().length) speakNow(); else speechSynthesis.onvoiceschanged = speakNow;That's free and instant, and it's the right answer for accessibility features, reading modes, and prototypes. It's the wrong answer the moment you need a consistent voice across devices or need to store the audio, because the synthesis happens on the user's machine and you never see the waveform.
For the server-side half, use each vendor's own Python quickstart rather than a copy on a third-party blog. Per-character request shapes drift, and a stale snippet is worse than no snippet.
What's worth writing down is the other half of the loop. Here's the input side of a voice agent against the Python SDK, version 1.0.0 (released August 14, 2026), which unified the async, streaming, and sync clients behind one shape:
Python (AssemblyAI Python SDK 1.0.0):
import os
from assemblyai.streaming.v3 import (
RealTimeEvents,
RealTimeParameters,
RealTimeTranscriber,
RealTimeTranscriberOptions,
TurnEvent,
)
# api_key is a kwarg on the transcriber; the options object carries
# api_host and terminate_timeout.
client = RealTimeTranscriber(
RealTimeTranscriberOptions(terminate_timeout=30.0),
api_key=os.environ["ASSEMBLYAI_API_KEY"],
)
def on_turn(client: RealTimeTranscriber, event: TurnEvent):
if event.end_of_turn:
# Hand the finalized turn to your LLM, then to your TTS vendor.
print(event.transcript)
client.on(RealTimeEvents.Turn, on_turn)
# Pin the model explicitly rather than relying on defaults.
client.connect(
RealTimeParameters(
sample_rate=16000,
speech_model="universal-3-6-pro", # singular on streaming
mode="balanced", # min_latency | balanced | max_accuracy
voice_focus="near-field", # near-field | far-field
)
)
# The SDK doesn't capture audio — supply 16-bit PCM chunks yourself.
try:
for chunk in microphone_stream():
client.stream(chunk)
finally:
client.disconnect(terminate=True) # billing is on session duration
If you'd rather skip the SDK, the raw endpoint is wss://streaming.assemblyai.com/v3/ws with your key in an Authorization header, no Bearer prefix.
Pass the agent's own question back in through agent_context and short replies resolve far more reliably—across a benchmark of 20,000 voice-agent audio files, agent context cut word error rate by 10.2%.
TTS APIs for voice agents
Text-to-speech is the last mile of the voice agent pipeline. Speech-to-text captures what the caller said, an LLM decides what to say back, TTS speaks it. Every millisecond in the TTS step is a millisecond of silence, and silence in a conversation feels roughly twice as long as it is.
So what does a TTS provider need for agent work?
- Sub-300ms time-to-first-byte, ideally under 200ms.
- Streaming output—you can't wait for a full response to synthesize.
- Conversational prosody, not narration prosody.
- Word-level timing, so you know how much of a sentence the caller heard before interrupting.
From the list above, Rime, Cartesia, and ElevenLabs are the usual finalists, with Telnyx worth a look if your traffic is telephony.
But here's where it gets interesting. Almost every team that A/B tests TTS vendors on an agent discovers the same thing: the voice was never the bottleneck.
Ask builders what breaks their agents and the answer comes back as transcription accuracy, not synthesis quality. Ask end users what makes them give up and the answer is having to repeat themselves. Voice quality matters—two in three buyers say so—but it is a preference, and mishearing an account number is a failure. Our writeup on AI voice agents covers why that asymmetry persists.
That's an argument about where to spend your evaluation time, and it's measurable. On AssemblyAI's English voice-agent benchmark of 12,460 scripted voice-agent scenarios, Universal-3.6 Pro Realtime records a 5.19% word error rate against Deepgram Flux EN at 13.50%, ElevenLabs Scribe v2 at 7.78%, and Deepgram Nova-3 at 8.64%. On entity error rate—names, addresses, phone numbers, account IDs, the things a caller has to repeat—it's 14.4% against Deepgram Flux EN at 30.1%. On the independent Pipecat open STT benchmark of real voice-agent conversations, it posts a 0.96% pooled semantic word error rate.
Those are the turns where agents die. A caller spells an email address, the model splits it across two turns, and now your agent is asking a clarifying question about something the caller already said perfectly clearly.
AssemblyAI's streaming speech-to-text handles that input side, with mid-sentence code-switching across 32 languages, diarization with revision for up to 10 speakers, and voice_focus to suppress background speech in a car or a drive-thru.
If you'd rather not manage three vendors at all, the Voice Agent API puts speech-to-text, LLM reasoning, and speech synthesis behind a single WebSocket at a flat $4.50 per hour, with roughly one second of end-to-end latency and no SDK to install. It's a complete agent runtime, not a standalone TTS product: if what you need is narration for an audiobook, use one of the twelve above. If what you need is an agent that answers the phone, one connection replaces three integrations, three latency budgets, and three failure modes. Pricing across the rest of the platform is on the pricing page, billed per second with no minimums.
For the wider decision, we've written up how to choose your STT foundation and what changes when you build agents that speak more than one language, and the voice agents solutions page covers the deployment patterns.
What we'd actually do
Pick two vendors, not five. Synthesize the fifty ugliest sentences in your product—the ones with account numbers, product SKUs, foreign names, and prices—and listen to them on a phone speaker, not headphones. That test eliminates more candidates in twenty minutes than a week of reading comparison tables, including this one.
Then do the thing almost nobody does: measure your time-to-first-byte from your own region, at your own concurrency, at 9am on a Monday. Published latency figures are measured under conditions that don't resemble production, and the vendor with the best benchmark number is regularly not the vendor with the best p95 on your traffic.
And here's the part worth carrying away. The TTS market is converging fast—twelve vendors, all improving, all within a perceptual hair of each other on clean text. What's not converging is what happens upstream, where the model has to decide whether the caller said "fifteen" or "fifty". That gap is still enormous, it's still measurable, and it's still where agent conversations break. Choose your voice carefully. Choose your ears first.
Whatever TTS you pick, the speech-to-text layer sets your agent's ceiling. Start free with Universal-3.6 Pro Realtime and hear the difference on entity-heavy audio.
Frequently asked questions
What is a TTS API?
A TTS (text-to-speech) API converts written text into spoken audio through an HTTP or WebSocket request. You send text and a voice identifier, and the service streams back audio in a format like MP3, WAV, or PCM. Most providers bill per character of input text; latency ranges from under 100ms for streaming-first models to a few seconds for batch synthesis.
Which is the best free text to speech API?
For genuinely free production use, the browser-native Web Speech API and self-hosted open-source engines carry no per-character cost, though quality and voice choice are limited. Among commercial providers, Amazon Polly, Google Cloud TTS, and Rime all publish free tiers, but allowances and eligibility windows change often—check each provider's current pricing page before you build against one.
Is Google text to speech API free?
Google Cloud Text-to-Speech is not free, but it publishes a monthly free character allowance, with Standard voices allotted more characters than the premium WaveNet, Neural2, and Chirp voice families. Beyond the allowance you pay per character, and premium voices cost several times the Standard rate. Confirm current tiers at cloud.google.com/text-to-speech/pricing before estimating spend.
How much does the ElevenLabs text-to-speech API cost?
ElevenLabs prices by subscription tier with a monthly character allowance, then charges for overage; enterprise pricing is quoted separately. Because ElevenLabs revises tier names and character allowances frequently, we don't restate the numbers here—check elevenlabs.io/pricing for the current figures.
Is there a free API for speech-to-text?
Yes. AssemblyAI's speech-to-text API is free to start with no credit card, and after that Universal-3.5 Pro costs $0.21 per hour of audio for pre-recorded transcription and $0.45 per hour base for Universal-3.6 Pro Realtime streaming, billed per second with no minimums. Open-source Whisper is also free if you're willing to run and scale the inference yourself.
Which text-to-speech API is best for voice agents?
For voice agents, pick the TTS engine with the lowest time-to-first-audio rather than the richest voice library—Cartesia and Rime are both built around streaming synthesis. But TTS is rarely the bottleneck: if the speech-to-text layer mishears an account number, no voice quality recovers the turn. AssemblyAI's Voice Agent API sidesteps the integration entirely by bundling STT, LLM, and TTS behind one WebSocket at a flat $4.50 per hour.
Lorem ipsum dolor sit amet, consectetur adipiscing elit, sed do eiusmod tempor incididunt ut labore et dolore magna aliqua. Ut enim ad minim veniam, quis nostrud exercitation ullamco laboris nisi ut aliquip ex ea commodo consequat. Duis aute irure dolor in reprehenderit in voluptate velit esse cillum dolore eu fugiat nulla pariatur.




