Any product that listens to a caller and talks back — a voice AI agent, a transcription feature, a real-time dubbing tool — needs speech recognition and speech generation fast enough that the conversation doesn't feel laggy. Deepgram and Cartesia both sell that layer, and they list each other (along with Rime) as their closest competitors. The real difference isn't raw speed — both chase sub-second latency — it's how much of the stack each one wants to own.
Deepgram: one API, one bill, the full pipeline
Deepgram puts speech-to-text and text-to-speech under the same API and the same invoice, and adds a bundled Voice Agent API that wires STT, an LLM and TTS together into a working conversational pipeline. It also ships audio intelligence features — summarization, sentiment, speaker diarization — built in, not bolted on.
Pricing: $200 in free credit to start. STT runs about $0.0048–$0.0065/minute, TTS about $0.015–$0.030 per 1,000 characters, and the bundled Voice Agent API about $0.075/minute. Paid plans start around $4K/year, with custom enterprise pricing above that.
Pick Deepgram if: you want speech-to-text, text-to-speech and agent orchestration from a single vendor and a single bill, and you value built-in audio intelligence (sentiment, diarization) over squeezing out the last bit of latency.
Watch out for: Deepgram itself acknowledges it isn't the clear latency or expressiveness leader against specialists like Cartesia and Rime — the tradeoff is breadth over being the sharpest tool in any one dimension, and the $0.075/min Voice Agent rate adds up fast at real call volume.
Cartesia: purpose-built for the fastest, most natural voice layer
Cartesia is narrower on purpose: real-time voice AI models built on a State Space Model architecture designed specifically for low-latency, natural-sounding speech, meant to sit underneath someone else's voice agent pipeline rather than replace it. It also offers cloud, on-premise and on-device deployment, which matters for regulated industries that can't send audio to a third-party cloud.
Pricing: free tier with 20K credits/month. Pro $5/month for 100K credits, Startup $49/month for 1.25M credits, Scale $299/month for 8M credits, custom Enterprise above that. Voice agent calls run about $0.06/minute, telephony about $0.014/minute.
Pick Cartesia if: you're already assembling your own pipeline (or building on a platform like Vapi or Retell AI) and want the most natural-sounding, lowest-latency audio layer specifically, especially if on-premise or on-device deployment is a requirement.
Watch out for: its own leaderboard/ranking claims are self-reported marketing worth verifying independently, and its feature set is narrower than Deepgram's bundled platform — you're buying one very good layer, not a full pipeline.
Side-by-side
| Deepgram | Cartesia | |
|---|---|---|
| Scope | STT + TTS + full Voice Agent API bundle | Focused real-time voice model layer |
| Entry price | $200 free credit, then usage-based | Free tier; Pro from $5/mo |
| Voice agent cost | ~$0.075/min (bundled) | ~$0.06/min agent, ~$0.014/min telephony |
| Deployment | Cloud API | Cloud, on-premise or on-device |
| Extra features | Audio intelligence (summarization, sentiment, diarization) | Very cheap entry tier to test before committing |
| Best for | Teams that want one vendor for the whole voice stack | Teams assembling their own pipeline who want the sharpest audio layer |
The honest verdict: if you'd rather manage one API and one bill for speech recognition, speech generation and agent orchestration together, Deepgram's bundle is the simpler starting point. If you're building or already have your own pipeline and specifically need the most natural, lowest-latency audio layer — especially with on-premise or on-device deployment for a regulated industry — Cartesia is the more specialized fit. Both also compete against Rime, a third TTS-focused specialist with 600+ voices across 50+ languages and sub-100ms latency claims of its own, so it's worth a look if multilingual coverage or deterministic pronunciation control (getting proper nouns and technical terms right) is your specific pain point. Whichever you pick, test with your own audio and call patterns before committing — every published latency number in this category is self-reported.