Deepgram
Speech recognition and voice-generation API that lets developers add accurate, fast transcription and natural-sounding speech to any app.
🔗 Visit DeepgramDescription
Turning spoken words into text — and text back into natural-sounding speech — sounds simple until you actually try building it: real conversations have background noise, accents, interruptions, and need to be transcribed in a fraction of a second for a live phone call to feel natural. Deepgram is a ready-made engine for both directions of that problem, built specifically so developers don't have to train their own speech models from scratch.
Deepgram provides speech-to-text and text-to-speech APIs built for real-time and batch use, powering voice AI applications, contact centers and conversational AI products. Its Voice Agent API bundles transcription, text-to-speech and LLM orchestration into a single integration point, and the platform includes automatic language detection across roughly ten languages, audio intelligence features like summarization and sentiment analysis, and speaker diarization with redaction for compliance-sensitive use cases. Pricing is pay-as-you-go with $200 in free credit to start: speech-to-text runs roughly $0.0048-$0.0065 per minute, text-to-speech about $0.015-$0.030 per 1,000 characters, and the bundled Voice Agent API around $0.075 per minute, with a Growth plan starting near $4,000/year and custom Enterprise pricing above that.
💬 Our review
The short version: Deepgram is one of the more established, developer-first names in speech AI, and having both transcription and speech generation under one API — plus a bundled Voice Agent endpoint — reduces the number of vendors a team building a voice product has to stitch together.
It competes in a genuinely crowded 2026 field: AssemblyAI focuses more narrowly on transcription accuracy and audio intelligence, while newer specialists like Cartesia and Rime push harder on ultra-low-latency, highly expressive speech generation specifically for real-time voice agents. Deepgram's advantage is breadth — one vendor, one bill, for both STT and TTS — rather than being the single best option on either axis alone. The per-minute and per-character pricing is competitive but not obviously cheaper than specialists, so the honest reason to pick Deepgram over a best-of-breed combination is integration simplicity, not a clear cost or quality edge on either individual capability.
💰 Pricing
📊 Global score
🤖 AI-enriched data
$200 free credit. STT ~$0.0048-$0.0065/min. TTS ~$0.015-$0.030/1K chars. Voice Agent API ~$0.075/min. Growth from ~$4K/year. Enterprise custom.
Pros
Both speech-to-text and text-to-speech under one API and one bill
Bundled Voice Agent API combining STT, TTS and LLM orchestration
Audio intelligence features (summarization, sentiment, diarization) built in
Cons
Not the clear latency or expressiveness leader against specialists like Cartesia/Rime
Pricing is competitive but not a standout discount versus focused competitors
Real-world cost with Voice Agent API ($0.075/min) adds up for high call-volume products
