If you're adding a voice to an app — a support agent that talks back, a transcription feature, a voice assistant — you'll eventually pick a speech API, and two names that keep coming up are Deepgram and Rime. They're not quite the same kind of product: Deepgram does both speech-to-text and text-to-speech under one bill, while Rime is a text-to-speech specialist. If you only need one direction — turning text into a voice — this comparison is for you.
Deepgram — one API, one bill, for both directions
Deepgram covers speech-to-text and text-to-speech together, plus a bundled Voice Agent API that combines STT, TTS and LLM orchestration into one integration — useful if you're building a full conversational agent, not just a narration feature.
For who: developers building voice AI applications who want transcription and speech generation from a single vendor.
Price: $200 free credit. STT ~$0.0048-$0.0065/min. TTS ~$0.015-$0.030/1K characters. Voice Agent API ~$0.075/min. Growth plans from ~$4,000/year. Enterprise custom.
Forces: speech-to-text and text-to-speech under one API and one invoice, a bundled Voice Agent API for full conversational pipelines, built-in audio intelligence (summarization, sentiment, diarization).
Limites: not the clear latency or expressiveness leader against specialists like Cartesia or Rime, and the Voice Agent API's $0.075/min adds up fast at high call volume.
Verdict: the simpler choice if you need both STT and TTS and don't want to manage two vendors — you trade some voice quality for that convenience.
Rime — a TTS specialist built by linguists, for real conversation
Rime does one thing — text-to-speech — and goes deep on it: 600+ voices across 50+ languages and regional dialects, deterministic pronunciation controls for proper nouns and technical terms, and sub-100ms latency with on-premise/VPC deployment for regulated industries.
For who: teams in fintech, healthcare and hospitality that need natural-sounding, multilingual conversational voice specifically, not transcription.
Price: $0.05 per 1,000 characters. 3,000 free minutes for new accounts. Custom volume pricing for enterprise.
Forces: the widest voice and dialect selection of the three, deterministic pronunciation control that generic TTS engines don't offer, sub-100ms latency with regulated-industry deployment options.
Limites: narrower scope than a bundled platform like Deepgram — it's TTS-focused, not a full pipeline — and like competitors, its latency claims are self-reported and worth independently testing.
Verdict: the better pick when voice quality, dialect coverage, or pronunciation control is the priority and you're handling transcription separately (or not at all).
Where Cartesia fits into this
Both Deepgram and Rime name Cartesia as a direct alternative to each other, and it's worth knowing about as a third option: Cartesia is also a real-time voice specialist, built on a State Space Model architecture specifically for low-latency speech, with cloud, on-premise and on-device deployment and a very cheap $5/month entry tier. If neither Deepgram's bundled approach nor Rime's dialect breadth is the deciding factor, Cartesia is the third name worth testing before you commit.
Side-by-side
| Deepgram | Rime | |
|---|---|---|
| Scope | STT + TTS + Voice Agent API | TTS only |
| Voices/languages | Multiple, general-purpose | 600+ voices, 50+ languages/dialects |
| Latency | Not the latency leader | Sub-100ms claimed |
| Price | $200 free credit, TTS ~$0.015-0.030/1K chars | $0.05/1K chars, 3,000 free minutes |
| Best for | One vendor for both STT and TTS | Dialect coverage + pronunciation control |
Pick Deepgram if you need transcription and speech generation together and would rather manage one API than two. Pick Rime if text-to-speech quality, language/dialect coverage, or precise pronunciation control is the actual requirement — and pair it with a separate STT tool, or Deepgram itself, for the other direction. Either way, test on your own scripts before committing: latency and naturalness claims across this whole category are vendor-reported, and the only benchmark that matters is how your specific audio sounds.