The number almost every voice AI deployment reports first is the one that tells you least. Automation rate (the share of calls the agent handled without a human) goes up reliably, looks like progress, and can rise at the same time your customers are getting a worse experience.
It rises because it is easy to move. Narrow what counts as a handled call, or let the agent keep talking when it should transfer, and the percentage improves. Neither change helped anyone who called you.
This is a short guide to what to measure instead, aimed at the operations leads who have to defend these numbers to people who do not care how the model works.
Containment is not resolution
A contained call is one that did not reach a human. A resolved call is one where the customer got what they rang for. Those are different sets, and the gap between them is where the real problem lives.
The measurement that closes the gap is repeat contact. Take the calls the agent closed, and count how many of those callers came back within a few days about the same thing. A call the agent handled and the customer had to make again is not a success, and it is invisible in every containment dashboard.
Report resolution as containment minus repeat contact. It is a less flattering number and it is the one that tracks reality.
Read every metric per intent
Aggregates hide the failure that matters. An agent can hold a strong overall resolution rate while handling one specific request badly, and if that request is the expensive one (a complaint, a cancellation, someone who is already annoyed), the average is actively misleading.
Break every number down by call reason. Volume, resolution, escalation, and latency, per intent, read separately. A resolution rate that slips from 80% to 74% tells you almost nothing. A resolution rate that is healthy everywhere except "reschedule an existing appointment" tells you what to fix this week.
This is also the only view where you can see the agent getting worse at something it used to be good at, which is the failure mode that aggregate reporting is structurally unable to surface.
Escalation quality, not escalation rate
Most teams track how often the agent hands off. Fewer track what the human receives when it does.
A handoff is good when the person picking up has the transcript, the classified intent, and whatever the agent already collected, and can continue the conversation rather than restart it. A handoff is bad when the customer explains themselves twice. Both count identically in an escalation-rate metric, and they are completely different experiences.
Measure the second one directly: sample transferred calls and check whether the customer had to repeat information. If they did, the automation moved work rather than removing it, and the escalation rate you were optimising was measuring the wrong thing.
A related number worth watching is how quickly a caller who asks for a person gets one. That should be the shortest path in the system.
Latency, and what callers actually notice
Response speed is not a technical vanity metric in voice. It is the difference between a conversation and a bad phone line.
ALLO answers in around 1.2 seconds. What matters operationally is not the average but the tail — the calls where something upstream was slow. Averages smooth over exactly the calls that went badly. Track the distribution and watch the slow end, because when responses stretch, callers start talking over the agent, which produces overlapping audio, which degrades recognition, which slows things further. The failure compounds inside one call and never registers as an error.
The quality signals you already have
Every call ALLO handles is transcribed, topic-classified, sentiment-analysed, quality-scored, recorded and archived. That produces a set of measurements most contact centres never had for their human agents, and it is routinely left unused.
Sentiment is most useful as a trend per intent rather than a headline figure. A single negative call means nothing; the same intent trending negative over a fortnight means the script is wrong. Quality scores work the same way — the value is in the drift, not the absolute number.
None of it works without someone actually reviewing calls. An agent nobody listens to develops the failure modes the dashboards were not built to catch, and the archive only pays for itself if it is somebody's job to open it.
Safety and policy, counted honestly
Incorrect actions, answers outside policy, and anything the agent said that it should not have — count these individually rather than as a rate, because at low volumes a rate makes a serious incident look negligible.
Set a threshold where one incident triggers a review rather than waiting for a pattern. The categories that never reach the agent at all should be fixed in advance and never renegotiated by whoever is measured on containment that quarter.
What to put in front of leadership
Four things, in this order. Resolution net of repeat contact, per intent. Escalation quality, sampled rather than inferred. The latency tail. And safety incidents as a count.
Automation rate can appear on the page. It should not be the number anyone is managed against, because it is the easiest one to move without improving anything.
If you cannot explain a metric to an operations lead in one sentence, it is measuring the model rather than the service.
For the engineering side of these measurements, production hardening for AI integrations covers permissions, per-intent instrumentation and handover design. If you are still planning the rollout, our IVR-to-voice-AI migration plan sets the gates each phase has to pass.







