Ground Truth
What a buyer finds when they check you, and what your own customers have already told you. One offer, two stages: the outside-in audit that closes the proof gap, then the corpus mine that becomes your comms programme.
You are assessed on evidence you have never read, from sources you do not own.
A buying committee checks a supplier when nobody from that supplier is in the room — reviews, directories, ratings, employee pages, third-party lists. Most mid-market firms have never audited that surface, never benchmarked it against the people they lose to, and never noticed it is now the layer AI models quote from.
And the ones who do have a review corpus are usually sitting on tens of thousands of unprompted customer statements they have never read past the star average.
Stage one finds out whether the proof exists. Stage two reads what it says.
Why this matters more than it did two years ago
Reviews were always useful. What changed is that they stopped being a page a buyer clicks and became the evidence an AI model quotes — to a buyer who never tells you they asked.
Stage one — the proof gap
Outside-in. Works with no review corpus at all, which is why it is the entry point. Four questions, in order.
What is actually there
Every trust signal that exists, where it lives, at what volume, recency, rating and response rate. Own review platform, Google, Trustpilot, sector directories, app stores, the employee surface, third-party listicles, forum threads, and the ratings markup on their own site.
Inventory before opinion. Most of the argument disappears once there is a list.
Against three named competitors
Same inventory, same method, for the firms they actually lose to. Absolute numbers mean nothing — 200 reviews is excellent or dismal depending entirely on who else is in frame.
The gap is the finding, and it is the slide that gets circulated.
Which gaps cost deals
Ranked, not listed. No reviews since March. Fourteen unanswered negatives. Absent from the two directories the engines actually cite in this category. No named-client proof. No ratings markup. An employee page describing leadership churn to anyone running supplier due diligence.
Each with an owner, an effort estimate and the reason it matters.
Turn the fix into visibility
A review-generation flow producing a steady stream rather than a one-off push. Response standards and templates. Directory presence engineered rather than hoped for. AggregateRating and Review schema. Question-shaped pages written from real reviewer language.
Cheapest available route into AI answers, because engines weight independent evidence over anything a vendor says about itself.
Why the fix is worth more than it sounds
Stage two — the corpus mine
Inside-out. Needs a real corpus, which stage one either finds or builds. Not a survey, not a workshop, not a sample — the whole book, classified, then read.
- Take the whole corpus. Every review, not a representative slice, because the interesting cohorts are small and a sample kills them.
- Classify every review. Theme, sentiment, whether it names a concrete fixable problem, and a one-line reason. Done by model at a cost per review that rounds to zero — which is what makes "all of them" affordable in the first place.
- Analyse deterministically. Distributions, confidence intervals, theme × sentiment crosstabs, rating-versus-text concordance, movement over time. Python, not vibes, so the findings survive a finance director.
- Split the output two ways. The demand half goes to marketing and comms. The defect half goes to operations. Different audiences, different meetings.
- Leave it running. The same pipeline runs live afterwards, flagging at-risk reviews within hours rather than at the next quarterly readout.
The questions customers actually ask
Every complaint is a question somebody typed into a search box first. Clustered by volume and written in the customer's own words, the themes are the query set — no keyword guesswork, no persona workshop.
That becomes the content and comms calendar, the answer-shaped pages engines extract from, the FAQ spine, the schema, the pre-scored testimonial bank, ad and landing copy, and the objection handling sales uses.
Where service actually breaks
Recurring defects ranked by frequency and sentiment cost, with the verbatim evidence attached. Not anecdotes from the last angry customer the MD happened to hear about.
That becomes a prioritised fix list for the teams owning the journey, a win-back queue of nameable at-risk customers, retention scripting armed with proven strengths, and a live alert on any theme whose share starts rising.
It is not our opinion
The findings are what tens of thousands of customers already said, unprompted, at a sample size no research budget would buy. Nothing for a sceptical board to argue with except the data.
And the operations half is what buys credibility. A marketing supplier handing the COO a ranked defect list is not the supplier the CFO cuts first.
The worked example — a live insurance corpus
Run on a Gulf motor and medical insurer with roughly 37,000 reviews and about 830,000 customers. The pilot classified 1,000 reviews spanning 9 March to 31 May 2026 — at that book's rate of ~362 new reviews a month, that is effectively the most recent twelve weeks rather than a sample of the past. Labelled on a local model at zero marginal cost; analysis in Python. Everything below is from that run.
Two mechanics worth naming, because they generalise. Detractors write an average of 40.3 words; promoters write 9.9 — negativity is longer and more specific, which is why it is minable and why the fix list has detail in it. And the rating distribution is a J-curve: a wall of five stars with a hard one-star pocket larger than two- and three-star combined. Lovers and a committed detractor minority, not a soft middle — which changes what the retention programme should aim at.
What the same run produced for the comms programme
- Answer-engine content. Each theme cluster converted into the actual question a customer types — "why is my renewal more expensive", "how long does a claim take" — published answer-first, forty to sixty words at the top, detail below, marked up as
FAQPageso it extracts cleanly. - Ratings proof, machine-readable.
AggregateRatingandReviewschema carrying a real 4.48 from verified customers. The 3.2× structural lever, at the cost of a developer's afternoon. - A video series from the complaints. The most concentrated grievance — pricing — became the card deck for a WIRED-Autocomplete-style explainer: the grievance is the question, the answer is the content, and the same script feeds social cut-downs and the help centre.
- ~692 pre-scored five-star verbatims ready for testimonial and landing-page use without anyone reading a spreadsheet to find them.
- A measurement loop. A monthly probe of those same questions across the four engines, tracking whether the brand is named and cited — where this play hands over to AI Visibility.
What it does not do
Said plainly, because the credibility of the rest depends on it.
| Limit | Why | The fix |
|---|---|---|
| It cannot prove churn | Review exports carry no retention outcome. The satisfied-but-complaining cohort is latent risk, not measured churn. | Join the review's order ID to the renewal system. Cheap, and then it is measured. |
| It cannot tell you response rates | Most standard exports have no merchant-reply or reply-timestamp field. | Re-export with those fields included. |
| It cannot cut by product or branch | Not in the review data unless the order ID is decoded. | The client's own product-code mapping. |
| Twelve weeks reads direction only | Too short for seasonality or year-on-year. | Run the full book, then keep it live. |
| It is observational | Reviewers self-select. The stage-one response-rate and recency figures are associations, not causal proof. | Findings framed as testable hypotheses, and the claims presented as reasons to act rather than guaranteed lift. |
| Platform coverage varies by sector | G2 and Capterra are dense in software, thin elsewhere; Trustpilot dominates consumer-facing finance, travel and telecoms. | Part of stage one is establishing which platforms matter in this category. |
Two things we will not do: gate reviews or solicit selectively — it breaches every platform's terms and is the fastest way to lose a listing. And we cannot fix a firm people dislike. If the reviews are bad because the service is bad, this surfaces it quickly and the work that follows is operational.
Who buys it, and what it sits next to
- Who: the MD, commercial director or marketing director of a B2B firm losing deals it never knew it was in. Stage one is diagnostic-shaped and cheap enough to start a relationship; stage two needs volume, so it suits insurance, financial services, healthcare, hospitality, e-commerce — anything with a customer service function.
- The sequence: stage one always first. It either finds a corpus worth mining or builds the flow that creates one, in which case stage two follows in year two.
- Where it fits: inside the Strategy & Planning Model as the evidence base, or standalone as a first paid piece of work.
- Sells with AI Visibility: Ground Truth builds and mines the sources the engines cite; AI Visibility measures whether they cite them. One supplies, the other scores.
- How it is judged: stage one on the gap against three named competitors closing — volume, recency, rating, response rate, directory presence. Stage two on content published from the findings and cited by the engines, and a named operations fix shipped off the back of it.