Marker

Marker

A grading and testing platform for voice AI agents — it listens to how your phone bot actually performs and scores it against rules you define, instead of you finding out it went off-script from an angry customer call.

🔗 Visit Marker
📁 Monitoring & Observability🗣️ English📅 July 29, 2026

Description

Voice agents are unusually hard to trust in production: a text chatbot's mistakes are visible in a transcript, but a voice agent can mumble, misjudge a pause, or drift off-script in ways that only show up once real callers hit it. Marker exists to catch that before it becomes a support escalation, by systematically simulating and scoring calls the way a QA team would, just automated and continuous.

Marker ingests live call transcripts along with audio and timing signals, keeping visibility into tool calls and trace context so you can see not just what the agent said but what it did. It supports version-pinned batch testing to catch regressions when you change a prompt or swap models, uses an LLM judge for evaluation (returning boolean, numeric, or category-based scores), and tracks alignment between machine and human labels to keep automated grading honest. It evaluates agents built on Claude, GPT-4o mini, Gemini, or Grok, and ships with API, CLI, and MCP interfaces, plus deployment options ranging from hosted to customer VPC, on-premises, or fully air-gapped — with SAML/OIDC auth, signed container images, and an SBOM for security-conscious buyers.

💬 Our review

The short version: if you're running voice agents at any real call volume, Marker's combination of multi-signal analysis (audio + transcript + execution trace) and enterprise-grade deployment options (VPC, on-prem, air-gapped) makes it a serious tool for catching regressions before customers do — the main unknown is cost, since pricing isn't published anywhere.

What stands out is the breadth of evaluation signal: most agent-testing tools only look at the transcript, while Marker also tracks tool-call execution and audio/timing, which matters specifically for voice where tone and pacing are part of the failure mode. The air-gapped and on-premises deployment options, plus SAML/OIDC and SBOM support, signal this is built for security-sensitive enterprise buyers rather than solo developers. The catch is the same one you hit with most enterprise infra tools: no public pricing means you can't compare cost against competitors without a sales call, and the deployment flexibility (customer VPC, air-gapped) suggests this is priced and packaged for larger contracts, not a quick trial for a small team.

💰 Pricing

Sur devisTarification non publiée, contact commercial requis
Enterprise Sur devis

📊 Global score

45Average
🌐Availability15/100Faible

1 language · 0 platform

📄Profile75/100Bien

Profile completeness

🤖 AI-enriched data

💰 Pricing model
💳 Sur devis

Aucune tarification publique — nécessite un contact commercial.

👥 Target audienceÉquipes construisant des agents vocaux IA à l'échelle, notamment dans des environnements réglementés ou sensibles à la sécurité
🗣️ Languagesen
🌍 Target countriesInternational
👍

Pros

Analyse multi-signal (audio, conversation, exécution des outils)

Tests par batch version-pinned pour détecter les régressions

Options de déploiement étendues : hébergé, VPC client, on-premises, air-gapped

SAML/OIDC, images signées et SBOM pour les besoins de conformité

👎

Cons

Aucune tarification publique

Packaging orienté grands comptes, peu adapté à un essai rapide en petite équipe

Différenciation face aux outils d'éval agent plus larges (LangSmith) à valider au cas par cas

❓ Frequently asked questions

What is Marker in one sentence?
How much does it cost?
Which voice/LLM models can it evaluate?
How does it grade an agent's performance?
Who is this built for?
Is it worth the money compared to alternatives?
Which tool should you pick for your case?