Equals Five.IdeasGround TruthThe InstallAgentsCompanyOSGrowth ModelDelivery MultiplierAI VisibilityGold Digger ProProfit X-Ray

Ground Truth

What a buyer finds when they check you, and what your own customers have already told you. One offer, two stages: the outside-in audit that closes the proof gap, then the corpus mine that becomes your comms programme.

The play

You are assessed on evidence you have never read, from sources you do not own.

A buying committee checks a supplier when nobody from that supplier is in the room — reviews, directories, ratings, employee pages, third-party lists. Most mid-market firms have never audited that surface, never benchmarked it against the people they lose to, and never noticed it is now the layer AI models quote from.

And the ones who do have a review corpus are usually sitting on tens of thousands of unprompted customer statements they have never read past the star average.

Stage one finds out whether the proof exists. Stage two reads what it says.

Why this matters more than it did two years ago

Reviews were always useful. What changed is that they stopped being a page a buyer clicks and became the evidence an AI model quotes — to a buyer who never tells you they asked.

The race ends before you know it started
68%
B2B buyers who already have a front-runner in mind at the very start of the process — and that front-runner wins roughly 80% of the time. Meanwhile G2's 2026 survey of 1,000+ B2B buyers found 71% now use AI chatbots for research, up from ~60% seven months earlier.
Where the trust went
23,700
Citations G2 accumulated in Perplexity alone — plus 12,000 in ChatGPT, 10,900 in Google AI Mode, 8,800 in AI Overviews, 5,200 in Gemini. Two years ago it had none, and its own organic traffic fell over the same window. The clicks left; the influence moved upstream into the answer.
Peer proof beat analyst proof
74% vs 13%
Buyers using peer reviews versus analyst reports — analyst use down roughly 63% since 2022. The signals that decide B2B deals now live on third-party platforms, not in a paid quadrant.
Citation has decoupled from ranking
38%
Share of Google AI Overview citations coming from pages in the top ten. G2 ranks 48th organically for "zendesk alternative" and is still cited in the AI answer for it. Structured markup makes a page 3.2× more likely to be cited, and 80.9% of listicle citations go to third-party lists rather than brand-authored ones.
Read together: the shortlist forms anonymously, from peer evidence, on platforms the client does not own — and those same platforms are what the AI answer is assembled from. A firm that has never audited that surface is competing blind in both channels at once.

Stage one — the proof gap

Outside-in. Works with no review corpus at all, which is why it is the entry point. Four questions, in order.

1 · Track

What is actually there

Every trust signal that exists, where it lives, at what volume, recency, rating and response rate. Own review platform, Google, Trustpilot, sector directories, app stores, the employee surface, third-party listicles, forum threads, and the ratings markup on their own site.

Inventory before opinion. Most of the argument disappears once there is a list.

2 · Benchmark

Against three named competitors

Same inventory, same method, for the firms they actually lose to. Absolute numbers mean nothing — 200 reviews is excellent or dismal depending entirely on who else is in frame.

The gap is the finding, and it is the slide that gets circulated.

3 · Extract

Which gaps cost deals

Ranked, not listed. No reviews since March. Fourteen unanswered negatives. Absent from the two directories the engines actually cite in this category. No named-client proof. No ratings markup. An employee page describing leadership churn to anyone running supplier due diligence.

Each with an owner, an effort estimate and the reason it matters.

4 · Generate

Turn the fix into visibility

A review-generation flow producing a steady stream rather than a one-off push. Response standards and templates. Directory presence engineered rather than hoped for. AggregateRating and Review schema. Question-shaped pages written from real reviewer language.

Cheapest available route into AI answers, because engines weight independent evidence over anything a vendor says about itself.

Why the fix is worth more than it sounds

Silence is read as an answer
88% vs 47%
Consumers who would choose a business replying to all its reviews, against one replying to none — a 41-point gap. Multiple unanswered negatives correlate with conversion drops of 15–25%: read as indifference or a systemic problem. Responding to 80%+ correlates with a 25% lift in loyalty; 81% expect a reply inside a week.
A big old review bank counts for little
32%
Consumers who read only reviews written in the previous two weeks. Recency beats volume — which is why a firm with 400 reviews and none since March is, in practice, a firm with no reviews.

Stage two — the corpus mine

Inside-out. Needs a real corpus, which stage one either finds or builds. Not a survey, not a workshop, not a sample — the whole book, classified, then read.

  1. Take the whole corpus. Every review, not a representative slice, because the interesting cohorts are small and a sample kills them.
  2. Classify every review. Theme, sentiment, whether it names a concrete fixable problem, and a one-line reason. Done by model at a cost per review that rounds to zero — which is what makes "all of them" affordable in the first place.
  3. Analyse deterministically. Distributions, confidence intervals, theme × sentiment crosstabs, rating-versus-text concordance, movement over time. Python, not vibes, so the findings survive a finance director.
  4. Split the output two ways. The demand half goes to marketing and comms. The defect half goes to operations. Different audiences, different meetings.
  5. Leave it running. The same pipeline runs live afterwards, flagging at-risk reviews within hours rather than at the next quarterly readout.
Output A · Marketing

The questions customers actually ask

Every complaint is a question somebody typed into a search box first. Clustered by volume and written in the customer's own words, the themes are the query set — no keyword guesswork, no persona workshop.

That becomes the content and comms calendar, the answer-shaped pages engines extract from, the FAQ spine, the schema, the pre-scored testimonial bank, ad and landing copy, and the objection handling sales uses.

Output B · Operations

Where service actually breaks

Recurring defects ranked by frequency and sentiment cost, with the verbatim evidence attached. Not anecdotes from the last angry customer the MD happened to hear about.

That becomes a prioritised fix list for the teams owning the journey, a win-back queue of nameable at-risk customers, retention scripting armed with proven strengths, and a live alert on any theme whose share starts rising.

Why it wins the room

It is not our opinion

The findings are what tens of thousands of customers already said, unprompted, at a sample size no research budget would buy. Nothing for a sceptical board to argue with except the data.

And the operations half is what buys credibility. A marketing supplier handing the COO a ranked defect list is not the supplier the CFO cuts first.

The worked example — a live insurance corpus

Run on a Gulf motor and medical insurer with roughly 37,000 reviews and about 830,000 customers. The pilot classified 1,000 reviews spanning 9 March to 31 May 2026 — at that book's rate of ~362 new reviews a month, that is effectively the most recent twelve weeks rather than a sample of the past. Labelled on a local model at zero marginal cost; analysis in Python. Everything below is from that run.

The headline everyone watches
4.48 / 5
Mean rating. 78.2% positive by text (±2.6%), 16.5% negative (±2.3%), n=1,000 at 95% confidence. A number reported to the board every quarter that tells it almost nothing.
What the headline hides
55
Reviews scoring 4 or 5 stars while writing a clearly negative comment — 40 at four stars, 15 at five. Satisfied-but-complaining. A CSAT average cannot see this cohort, and no sample of two hundred would have found it.
The counter-intuitive finding
32 / 1,000
Share of the conversation about claims — the thing an insurer assumes is its problem, and mostly fine. The fixable money sits in renewal (45 actionable complaints, the longest queue) and pricing, where 70% of every mention is an actionable complaint.
The evidenced positioning
45%
Share of all reviews about representatives and turnaround speed, overwhelmingly positive. The "human service and speed" claim stops being aspirational and becomes a thousand customers saying it first — usable in copy, in renewal calls, and as schema-marked proof.

Two mechanics worth naming, because they generalise. Detractors write an average of 40.3 words; promoters write 9.9 — negativity is longer and more specific, which is why it is minable and why the fix list has detail in it. And the rating distribution is a J-curve: a wall of five stars with a hard one-star pocket larger than two- and three-star combined. Lovers and a committed detractor minority, not a soft middle — which changes what the retention programme should aim at.

What the same run produced for the comms programme

What it does not do

Said plainly, because the credibility of the rest depends on it.

LimitWhyThe fix
It cannot prove churnReview exports carry no retention outcome. The satisfied-but-complaining cohort is latent risk, not measured churn.Join the review's order ID to the renewal system. Cheap, and then it is measured.
It cannot tell you response ratesMost standard exports have no merchant-reply or reply-timestamp field.Re-export with those fields included.
It cannot cut by product or branchNot in the review data unless the order ID is decoded.The client's own product-code mapping.
Twelve weeks reads direction onlyToo short for seasonality or year-on-year.Run the full book, then keep it live.
It is observationalReviewers self-select. The stage-one response-rate and recency figures are associations, not causal proof.Findings framed as testable hypotheses, and the claims presented as reasons to act rather than guaranteed lift.
Platform coverage varies by sectorG2 and Capterra are dense in software, thin elsewhere; Trustpilot dominates consumer-facing finance, travel and telecoms.Part of stage one is establishing which platforms matter in this category.

Two things we will not do: gate reviews or solicit selectively — it breaches every platform's terms and is the fastest way to lose a listing. And we cannot fix a firm people dislike. If the reviews are bad because the service is bad, this surfaces it quickly and the work that follows is operational.

Nothing above needs a platform, a client login, or software to maintain. It is a piece of work our specialists do inside the client's business, and the client's own team runs the live loop afterwards.

Who buys it, and what it sits next to