Vapi
Developer platform that wires together speech recognition, an AI model and speech generation into a working phone-answering voice agent.
🔗 Visit VapiDescription
Building a voice agent that can actually answer a phone call means connecting several separate pieces — something to turn speech into text, something to decide what to say back, something to turn that back into speech, and a phone system to carry the call — and making sure the whole chain responds fast enough that the caller doesn't notice any lag. Vapi packages that entire chain into one developer platform, so building a voice agent looks more like configuring a service than assembling five different vendors.
Vapi lets developers configure and deploy voice AI agents that combine speech-to-text, an LLM, text-to-speech and PSTN (regular phone network) telephony integration, with infrastructure aimed at sub-500ms response latency. It includes real-time call monitoring and analytics, AI guardrails to reduce hallucinated responses during a live call, and compliance certifications (SOC 2, HIPAA, PCI) for regulated use cases. Pricing on the Build tier is usage-based — around $0.05/minute for hosting plus $10 per concurrent line per month — with underlying model provider costs (for STT, the LLM, and TTS) either passed through or waived if you bring your own API keys; a Scale tier with annual contracts and negotiated per-minute rates is available for higher-volume enterprise customers.
💬 Our review
The short version: Vapi's value is orchestration — it doesn't try to be the best speech-to-text or text-to-speech provider itself, it's the layer that wires best-in-class components together into a working phone agent, which saves real integration work compared to building that pipeline from scratch.
It competes directly with Retell AI and Bland AI in the "voice agent platform" category, all three solving a similar orchestration problem with broadly comparable latency targets and compliance certifications. The honest differentiator between them tends to be smaller things — Vapi's ability to bring your own model API keys to avoid provider markup, specific integration ecosystem, and pricing structure (per-concurrent-line rather than purely per-minute) — rather than a fundamental capability gap. For a team that wants control over which STT/LLM/TTS providers sit underneath the agent (to optimize cost or quality per component), Vapi's bring-your-own-key option is a real advantage over more locked-down competitors; for a team that just wants the fastest path to a working phone agent without those decisions, any of the three leading platforms will get the job done.
💰 Pricing
📊 Global score
🤖 AI-enriched data
Build tier: ~$0.05/min hosting + $10/concurrent line/month, model costs passed through or waived with own API key. Scale tier: annual contract, custom per-minute rates.
Pros
Orchestrates STT, LLM, TTS and PSTN telephony into one working pipeline
Sub-500ms latency target for natural-feeling phone conversations
Bring-your-own API key option to control underlying model costs/quality
Cons
Competes closely with Retell AI and Bland AI with limited fundamental differentiation
Per-concurrent-line pricing ($10/line/month) plus usage adds real cost at scale
Underlying model provider costs are a separate variable to budget for
