Zro
Lets an AI coding assistant (like Claude Code or Cursor) run on open-weight language models instead of a big-name API, with a guarantee that none of your code ever gets stored or used to train anything — like renting a private, no-questions-asked engine r
🔗 Visit ZroDescription
When you plug an AI coding assistant into a model provider, you're usually sending your codebase through someone else's servers — and it's not always clear what happens to that data afterward. Zro's whole pitch is a private lane: your prompts and code get processed and then genuinely discarded, not logged or used to improve anyone else's model.
Zro is a private inference endpoint built specifically for coding agents (Claude Code, Cursor, Cline, and similar tools), serving open-weight models like MiniMax M3, GLM-5.2, and Kimi K2.7 Code across multiple regions on AMD, NVIDIA, and TPU hardware. It uses an OpenAI/Anthropic-compatible API, so it drops into existing agent tooling without custom integration work, and applies its own "HyperQuant" compression with custom attention kernels aimed at long-context, multi-turn coding sessions. Pricing is subscription-based: Pro at $20/month for roughly 300M tokens, Max at $60/month for roughly 1.5B tokens, with Enterprise available on custom terms and smaller usage packs ($5-$100, 90-day expiration) for lighter needs. There is no free tier.
💬 Our review
The short version: if you're already running an AI coding agent and want the raw model swapped out from a closed frontier API to a cheaper, privacy-first open-weight backend without re-plumbing your tools, Zro is built exactly for that swap — it just isn't free to try.
The zero-retention, zero-training promise is the actual differentiator here, not raw model quality — MiniMax M3, GLM-5.2, and Kimi K2.7 Code are solid open-weight coding models but they're not going to out-reason Claude or GPT-class frontier models on hard problems. What Zro is really competing with is other inference resellers (Together AI, Fireworks, Groq) rather than the big labs directly, and its edge there is the explicit agent-coding framing plus the compatible API that means no rewiring of Claude Code/Cursor/Cline configs. At $20/month entry with no free tier, it's a harder sell for someone who just wants to kick the tires — you're committing before you've seen the token-for-token quality tradeoff versus what you're already paying a frontier lab. Teams with real data-residency or client-confidentiality constraints will find the privacy guarantee worth the premium over a cheaper but less transparent reseller.
📊 Global score
🤖 AI-enriched data
Pro 20$/mois (~300M tokens) ; Max 60$/mois (~1,5B tokens) ; Enterprise sur devis ; packs à l'usage 5$-100$ (expiration 90 jours) ; aucun palier gratuit
Pros
Zéro rétention des requêtes, zéro entraînement sur les données client
API compatible OpenAI/Anthropic — s'intègre sans réécriture des outils existants
Optimisé spécifiquement pour les sessions de code longues et multi-tours
Infrastructure multi-région sur AMD, NVIDIA et TPU
Cons
Aucun palier gratuit pour tester avant de payer
Modèles open-weight servis, pas les modèles frontière (Claude, GPT) — capacités de raisonnement en retrait
Marché concurrentiel des revendeurs d'inférence (Together AI, Fireworks, Groq)
Date de fondation non communiquée
