Over and above technical metrics to evaluate whether an agent can hold a call. For a scaled voice AI orchestrator running evaluations for 20+ agents across 8+ languages and 5+ industries.
Listening to real calls and applying judgment is how a team learns what breaks a conversation. That works at small volume. Once the platform is handling serious volume across many deployed agents, there's no bandwidth to sample at a meaningful scale, so issue detection turns sporadic and reactive.
Needs: a steady, representative review that scales with volume, not a heroic listen.
Quality was being read through latency, barge-in and similar measures. Those miss what actually matters: how the person on the other end perceives the call, or how the client would judge it if they sampled one.
Needs: the subjective sense of a good call turned into metrics that can each be scored.
The team tried interns, then a small internal reviewer group. Beyond the operational overhead, the job is not as simple as hiring people. It takes real investment to break the task into reviewable subtasks, build the tooling to capture judgment efficiently, and train and monitor reviewers to keep quality high.
Needs: rubric design, workflow tooling and reviewer calibration as one system.
Design the metrics, diagnose where the agent fails, build targeted data to fix it, monitor reliability as things change
We turned a single subjective sense of quality into metrics that can each be scored on their own. Each metric gets the evaluation it actually needs: a machine, a human, or both.
Does the whole call sound like a real person the customer would trust?
Dead air and slow responses that make the agent feel laggy or cut people off.
Does the extracted outcome, like intent to pay, match what really happened?
Overall word accuracy: did the agent hear what was actually said?
Critical fields (order ID, phone, amount) captured exactly right.
Digits, dates and alphanumeric codes, where one wrong token breaks the call.
Names, places and brands the model most often mishears.
Mixed-language speech transcribed cleanly, not garbled into nonsense.
Does the agent stay in the language the caller is actually speaking?
Did the call actually do what the user rang in to get done?
Does the agent hold on to details the caller gave earlier?
Repeating itself or looping without ever moving the call forward.
Internal tool or function names spoken out loud to the caller.
Does the agent stop talking the moment the caller interrupts?
Wrong accents and mispronounced names, places and terms across languages.
Right words, wrong emotion: flat, robotic, or mismatched to the moment.
The agent records the wrong value for a field the outcome hinges on
Agent misses the goal, drifts language, loops, or leaks its own plumbing to the caller.
Tone and pronunciation live outside the transcript entirely. Only a listener catches them, and they decide whether the agent sounds human.
Disposition, often a call's most important output, comes from an LLM judge that had never been measured against a human. We compute its accuracy and plug humans in where it's most likely to be wrong.
The engagement becomes a business-as-usual pipeline that reviews ~10% of calls across agent types, languages and industries. As fresh problems surface we design the way to track each one at scale, and build a human or an automated evaluation depending on what the metric actually needs.
A representative sample scored continuously, so regressions show up as models, prompts and rules change.
Agent reliability read through inter-rater agreement and comparison against a ground-truth set.
The rubric design, workflow tooling and reviewer training deploy across new use-cases with high confidence.
Send us the calls your agent actually handles. We come back with a metric breakdown, the failures that matter, and a monitoring pipeline that scales.