Equals Five.IdeasGround TruthThe InstallAgentsCompanyOSGrowth ModelDelivery MultiplierAI VisibilityGold Digger ProProfit X-Ray

Customer-facing agents that don't suck

Almost everyone has met a bad one. The reasons they are bad are well documented, consistent, and fixable — but only in that order.

The play

The problem was never the model. It was everything around it.

Deploying an agent in front of customers is the most exposed thing a business can do with AI. Get it right and it answers at three in the morning with the patience of someone who has read every policy document. Get it wrong and it becomes the brand — publicly, at scale, on a screenshot.

We only build these on top of something real, and we say no when there isn't one.

What the record actually shows

Customers have been trained to dread it
79%
Consumers who strongly prefer speaking to a human rather than an AI agent; 84% believe humans give more accurate support, and 81% believe businesses deploy AI to save money rather than improve service. Your agent starts the conversation in a hole somebody else dug.
And often with reason
46%
Consumers who say AI-powered service rarely or never reaches a successful outcome. Around 20% of people who used an AI service agent got zero benefit from the interaction — roughly four times the failure rate of AI use generally.
It is getting better where it is done properly
57%
Consumers reporting a positive recent AI service experience — up from 38% in 2024. The distribution is bimodal, not bad: well-built agents are pulling the average up while badly-built ones keep dragging it down. Which side you land on is a delivery decision.
The projects themselves fail
40%+
Agentic AI projects Gartner expects to be cancelled by end of 2027 — escalating cost, unclear business value, inadequate risk controls — from a poll of 3,400+ organisations. It also flags "agent washing": roughly 130 vendors offer genuine agentic capability out of thousands claiming it.

One more number worth holding: hallucination accounts for only 0.34% of AI-handled tickets, yet 71% of CX leaders rank it a top-three governance risk. They are right to. The cost of an error here is not the ticket — it is the screenshot.

Why they suck — the five failure modes

Every one of those is an operating decision, not a technical limitation. Which is why this is a delivery problem and why an embedded specialist team is the right shape to solve it.

The standard we build to

1 · Grounded

Answers from the client's own truth

Wired to the real systems and the real documents, with permissions intact — the customer's actual order, policy, case or invoice. If the answer is not in a source it can cite, it does not have one, and it says so.

This is why the layer underneath matters. CompanyOS is what makes it possible; without it you get a well-spoken guesser.

2 · Scoped

A written list of what it must never do

Approved claims, prohibited topics, no pricing improvisation, no policy interpretation, no commitments it cannot honour. Written before a line of it is built, signed off by whoever owns the risk.

A narrow agent that is right is worth more than a broad one that is plausible.

3 · Escalating

Hands over well, and early

An always-visible route to a human, plus automatic escalation on frustration, repetition, complaint language, vulnerability signals or anything financial. And it hands over with the context — the customer never repeats themselves.

Handover quality is the single biggest driver of whether the experience is remembered as good.

4 · Evaluated

Tested against real conversations

A standing evaluation set built from the client's actual history — and this is where Ground Truth pays for itself, because a mapped review corpus is a ready-made bank of the awkward questions customers really ask.

Every change is scored against that set before it ships. Accuracy, refusal behaviour, escalation, tone.

5 · Measured honestly

Resolution, not deflection

Did the customer's problem get solved, would they use it again, did it cost more or less than the human path, and how many escalations arrived worse than they started. Deflection rate is banned as a primary metric.

Reported monthly against the baseline agreed before launch.

6 · Staged

Internal first, then a narrow slice

Run it behind the service team as a copilot first — same questions, same knowledge, no customer exposure. When it beats the team's own accuracy on the eval set, put it in front of one narrow, low-risk topic. Widen only on the numbers.

Nobody's brand should be the pilot.

When we say no

Who it is for