Comparatifs

Deepgram vs Cartesia: Which Voice AI API Should You Build On?

Deepgram bundles speech-to-text, text-to-speech and a full voice agent pipeline under one bill. Cartesia bets everything on being the fastest, most natural-sounding layer underneath someone else's pipeline.

Any product that listens to a caller and talks back — a voice AI agent, a transcription feature, a real-time dubbing tool — needs speech recognition and speech generation fast enough that the conversation doesn't feel laggy. Deepgram and Cartesia both sell that layer, and they list each other (along with Rime) as their closest competitors. The real difference isn't raw speed — both chase sub-second latency — it's how much of the stack each one wants to own.

Deepgram: one API, one bill, the full pipeline

Deepgram puts speech-to-text and text-to-speech under the same API and the same invoice, and adds a bundled Voice Agent API that wires STT, an LLM and TTS together into a working conversational pipeline. It also ships audio intelligence features — summarization, sentiment, speaker diarization — built in, not bolted on.

Pricing: $200 in free credit to start. STT runs about $0.0048–$0.0065/minute, TTS about $0.015–$0.030 per 1,000 characters, and the bundled Voice Agent API about $0.075/minute. Paid plans start around $4K/year, with custom enterprise pricing above that.

Pick Deepgram if: you want speech-to-text, text-to-speech and agent orchestration from a single vendor and a single bill, and you value built-in audio intelligence (sentiment, diarization) over squeezing out the last bit of latency.

Watch out for: Deepgram itself acknowledges it isn't the clear latency or expressiveness leader against specialists like Cartesia and Rime — the tradeoff is breadth over being the sharpest tool in any one dimension, and the $0.075/min Voice Agent rate adds up fast at real call volume.

Cartesia: purpose-built for the fastest, most natural voice layer

Cartesia is narrower on purpose: real-time voice AI models built on a State Space Model architecture designed specifically for low-latency, natural-sounding speech, meant to sit underneath someone else's voice agent pipeline rather than replace it. It also offers cloud, on-premise and on-device deployment, which matters for regulated industries that can't send audio to a third-party cloud.

Pricing: free tier with 20K credits/month. Pro $5/month for 100K credits, Startup $49/month for 1.25M credits, Scale $299/month for 8M credits, custom Enterprise above that. Voice agent calls run about $0.06/minute, telephony about $0.014/minute.

Pick Cartesia if: you're already assembling your own pipeline (or building on a platform like Vapi or Retell AI) and want the most natural-sounding, lowest-latency audio layer specifically, especially if on-premise or on-device deployment is a requirement.

Watch out for: its own leaderboard/ranking claims are self-reported marketing worth verifying independently, and its feature set is narrower than Deepgram's bundled platform — you're buying one very good layer, not a full pipeline.

Side-by-side

DeepgramCartesia
ScopeSTT + TTS + full Voice Agent API bundleFocused real-time voice model layer
Entry price$200 free credit, then usage-basedFree tier; Pro from $5/mo
Voice agent cost~$0.075/min (bundled)~$0.06/min agent, ~$0.014/min telephony
DeploymentCloud APICloud, on-premise or on-device
Extra featuresAudio intelligence (summarization, sentiment, diarization)Very cheap entry tier to test before committing
Best forTeams that want one vendor for the whole voice stackTeams assembling their own pipeline who want the sharpest audio layer

The honest verdict: if you'd rather manage one API and one bill for speech recognition, speech generation and agent orchestration together, Deepgram's bundle is the simpler starting point. If you're building or already have your own pipeline and specifically need the most natural, lowest-latency audio layer — especially with on-premise or on-device deployment for a regulated industry — Cartesia is the more specialized fit. Both also compete against Rime, a third TTS-focused specialist with 600+ voices across 50+ languages and sub-100ms latency claims of its own, so it's worth a look if multilingual coverage or deterministic pronunciation control (getting proper nouns and technical terms right) is your specific pain point. Whichever you pick, test with your own audio and call patterns before committing — every published latency number in this category is self-reported.