Customer-facing agents that don't suck
Almost everyone has met a bad one. The reasons they are bad are well documented, consistent, and fixable — but only in that order.
The problem was never the model. It was everything around it.
Deploying an agent in front of customers is the most exposed thing a business can do with AI. Get it right and it answers at three in the morning with the patience of someone who has read every policy document. Get it wrong and it becomes the brand — publicly, at scale, on a screenshot.
We only build these on top of something real, and we say no when there isn't one.
What the record actually shows
One more number worth holding: hallucination accounts for only 0.34% of AI-handled tickets, yet 71% of CX leaders rank it a top-three governance risk. They are right to. The cost of an error here is not the ticket — it is the screenshot.
Why they suck — the five failure modes
- It doesn't know anything. Wired to a stale FAQ instead of the systems that hold the answer, so it cannot see the customer's order, policy or case. It is not a support agent; it is a search box with manners.
- It won't let go. No escalation path, or one deliberately hidden to suppress contact volume. 89% of consumers say there should always be a route to a human — and the businesses that hide it are the reason the other 79% flinch.
- Nobody defined what it must never do. No scope, no refusal behaviour, no approved-claims list. So it invents a policy, quotes a price, or promises a refund — confidently.
- It was never tested properly. Signed off on a demo of ten happy-path questions. Real customers arrive angry, ambiguous, mid-transaction and in a second language.
- It is measured on the wrong thing. Deflection rate. Which rewards a bot that fails to help so successfully that the customer gives up — and it will optimise straight into that.
The standard we build to
Answers from the client's own truth
Wired to the real systems and the real documents, with permissions intact — the customer's actual order, policy, case or invoice. If the answer is not in a source it can cite, it does not have one, and it says so.
This is why the layer underneath matters. CompanyOS is what makes it possible; without it you get a well-spoken guesser.
A written list of what it must never do
Approved claims, prohibited topics, no pricing improvisation, no policy interpretation, no commitments it cannot honour. Written before a line of it is built, signed off by whoever owns the risk.
A narrow agent that is right is worth more than a broad one that is plausible.
Hands over well, and early
An always-visible route to a human, plus automatic escalation on frustration, repetition, complaint language, vulnerability signals or anything financial. And it hands over with the context — the customer never repeats themselves.
Handover quality is the single biggest driver of whether the experience is remembered as good.
Tested against real conversations
A standing evaluation set built from the client's actual history — and this is where Ground Truth pays for itself, because a mapped review corpus is a ready-made bank of the awkward questions customers really ask.
Every change is scored against that set before it ships. Accuracy, refusal behaviour, escalation, tone.
Resolution, not deflection
Did the customer's problem get solved, would they use it again, did it cost more or less than the human path, and how many escalations arrived worse than they started. Deflection rate is banned as a primary metric.
Reported monthly against the baseline agreed before launch.
Internal first, then a narrow slice
Run it behind the service team as a copilot first — same questions, same knowledge, no customer exposure. When it beats the team's own accuracy on the eval set, put it in front of one narrow, low-risk topic. Widen only on the numbers.
Nobody's brand should be the pilot.
When we say no
- No source of truth. If the systems cannot answer the question, the agent cannot either. Fix the layer first — that is a CompanyOS conversation, not an agent one.
- Deflection is the actual goal. If the brief is to reduce contact volume rather than resolve contacts, the project will succeed on its metric and damage the business. We would rather not build it.
- Regulated advice. Anything that constitutes advice, an underwriting decision, a medical or legal opinion, or a binding commitment stays with a qualified human. The agent routes; it does not rule.
- No appetite for the eval work. A client unwilling to fund the testing is buying the 40% outcome, and we should say so before they buy it from someone else.
Who it is for
- Who buys it: a customer-service or commercial director with real contact volume, real repetition in it, and a board asking why service costs what it does.
- Prerequisites: systems that hold the answer, an owner for the scope document, and a budget line for evaluation.
- Sells with Ground Truth: the review corpus supplies both the questions the agent must handle and the evaluation set it is scored on.
- How it is judged: resolution rate, escalation quality, cost per resolved contact, and customer-reported satisfaction against the human baseline.