Gradium
An API for making computers talk and listen in a natural-sounding voice — text-to-speech, transcription, voice cloning, and live translation — built by a team that split off from a well-known French AI lab.
🔗 Visit GradiumDescription
Building a voice feature into an app — a voice assistant, a dubbed video, a customer support bot that actually sounds human — usually means stitching together separate transcription, translation, and speech-synthesis services from different vendors. Gradium offers all of it as one platform: text-to-speech, speech-to-text, voice cloning from as little as a few seconds of audio, and real-time voice-to-voice translation, currently across English, French, German, Spanish and Portuguese.
It's built by a Paris-based team spun out of the Kyutai research lab, founded by Neil Zeghidour, a former Google DeepMind voice researcher, and it emphasizes speed — sub-300ms latency for streaming responses, positioned for live conversational use rather than just batch audio generation. Pricing is credit-based: a free tier gives 45,000 credits a month (roughly an hour of TTS generation, but restricted from commercial use), with paid tiers from $13/month up to $1,615/month for heavier volume, plus custom enterprise pricing. The company raised $100M in seed funding as of July 2026, with Nvidia among the investors — a large round for a young company, reflecting how competitive the AI voice space currently is.
💬 Our review
The short version: Gradium is a credible, well-funded challenger to ElevenLabs in the same crowded voice-AI space, with real technical pedigree (Kyutai, ex-DeepMind founder) but a much shorter public track record and narrower language coverage.
Against ElevenLabs or the cloud giants' speech services (Google, AWS Polly, Azure), Gradium's differentiators are its live translation feature and low-latency streaming focus, which matter specifically for real-time conversational products rather than pre-recorded audio generation. The honest limitation is scope: five supported languages today is meaningfully fewer than ElevenLabs' broader coverage, and the credit-based pricing means costs can climb quickly in production — worth modeling your expected usage against the credit-per-character/second math before committing budget. The $100M seed round with Nvidia's backing suggests real staying power, but it's still a 2026-launch product without years of production hardening behind it. For a team specifically building low-latency, real-time voice translation or conversational agents, it's worth evaluating directly against ElevenLabs; for simple one-off narration or dubbing where language breadth matters more than latency, an established competitor may serve you better today.
💰 Pricing
📊 Global score
🤖 AI-enriched data
Gratuit : 45k crédits/mois, usage non commercial. XS : 13$/mois (225k crédits). S : 43$/mois (900k crédits). M : 340$/mois (9M crédits). L : 1 615$/mois (45M crédits). Enterprise : sur devis.
Pros
Latence très basse (<300ms), pensé pour la conversation en temps réel
Traduction voix-à-voix en direct — fonctionnalité rare chez les concurrents
Fondé par un ex-chercheur voix de Google DeepMind, adossé à Kyutai, 100M$ levés (Nvidia investisseur)
Cons
Seulement 5 langues supportées contre une couverture plus large chez ElevenLabs
Consommation de crédits qui peut devenir coûteuse en production intensive
Produit très récent (2026), historique limité