Three well-known brands came to me in the last six weeks with the same problem wearing different clothes.

One found ChatGPT quoting a pricing tier they killed in 2024. Another found ChatGPT crediting them with a compliance certification they never held (which is a fun conversation to have with legal). The third found Gemini comparing their current product against a competitor's two-generation-old model (and losing).

Here's the kicker: their visibility numbers looked great. They were showing up, getting cited, winning the dashboards. The answers were just wrong.

For two years AEO has been about showing up. Do you appear? How often? Ahead of whom? Right questions, still. But there's a harder one underneath: when AI talks about you, is what it says true?

Most brands answer that with a shrug. Here's why that shrug is expensive, and the first tool I've seen that replaces it with a number.

The numbers

  • The 5W Hallucination Index asked AI models 25 core business questions per brand. 20% of answers contained material errors. One in five prospects heard something false about you this week. 

  • Profound analyzed 50,000 prompts across seven industries and found that nearly half of AI response content, 47%, was unsolicited. Comparisons, opinions, and "pro tips" nobody asked for, pulled from whatever the engine found across the web. 

  • NewsGuard tracked how often chatbots decline to answer when they don't have good information. That rate fell from 31% to 0% in one year. The engines answer everything now. When they don't know, they make something up and say it with a straight face.

And the stat that ends the "not our problem" debate: 21% of consumers blame the brand when AI gets a fact wrong about it, even when the error came from nowhere near the brand's own properties. You own the answer whether you wrote it or not.

Why engines get you wrong

Four patterns cover nearly every case I've audited to date:

  1. The model's memory of you is old. It learned your business at a snapshot in time. You've launched new products and changed pricing since. It hasn't noticed, and it answers like it's still last year.

  2. It trusts a page that's wrong. When the engine searches the live web, it lands on a confidently outdated third-party page and repeats it as gospel.

  3. Mistaken identity. You almost share a name with another company. Their facts bleed into your answer. Their bad reviews can too.

  4. Zombie listings. A directory profile someone on your team made in 2019 is still feeding engines your old address and old product line. Nobody owns it. Nobody remembers it exists.

One finding makes all four worse. Profound analyzed 50,000 AI answers and found 47% of the content was unsolicited. Ask a simple question, get commentary, comparisons, and opinions nobody requested. Half of what gets said about you is stuff the engine volunteered on its own, and that volunteered half is where the errors pile up.

Where the errors actually live

Those four patterns showed up in every audit I've run. Profound just released new data on this exact thing. Two weeks ago they published an analysis of 158,000 AI claims about real brands, each one graded true or false against verified records. Three findings stand out.

The first one hurts. 54% of brands had at least one false claim traced back to their own website. Not a competitor. Not a stale directory. Their own pages, contradicting their own records. Before you email a single journalist, check your CMS. There's a decent chance your first correction is in there. Having spent so much time at Webflow in the CMS market, I can tell you firsthand, CMS debt is real. 

Third-party coverage is the other big offender, and it doesn't work the way you'd hope. The errors follow a long tail. The ten worst domains account for just 13% of earned media inaccuracies, and the top hundred only get you to 40%. You can't fix this by leaning on a few big publishers. But here's the useful part: for any given brand, the errors typically trace back to about four specific sites. Short list. Different for every brand. Which is why nobody finds theirs by guessing.

And the thing engines get wrong most? Your price. Pricing and billing made up 12% of the claims tested but 24% of the errors, double its weight. Among brands with enough pricing claims to measure, it was the worst topic for two out of three. Makes sense. Nothing on your site changes more often, and nothing across the web goes stale faster. The most important number in your business is the one AI is most likely to misquote.

The free version: audit it yourself this week

  1. Write the 15 questions a real prospect asks before buying. Pricing, integrations, security, "you vs. competitor." Pull them from sales calls, not your imagination.

  2. Run each through ChatGPT, Perplexity, Gemini, Copilot. Log every factual claim in a spreadsheet.

  3. Grade each claim true or false against current reality. No partial credit. "Mostly right" is wrong.

  4. For every false claim, open the citations and find the page feeding it.

Block two hours. Do this and you'll know more than 95% of your competitors, because almost nobody has bothered.

You'll also hit the wall: it's one snapshot, on your phrasings, graded by hand. Buyers phrase things thousands of ways, answers shift week to week, and you can't staff a team to re-run this forever. 

What FactCheck does

Profound launched FactCheck in July. It's the first system that measures AI accuracy about a brand at scale, and the mechanics are the story. I’ve been using it in more and more clients’ Profound instances. Here’s what it does: 

It reads every answer sentence by sentence. A claim detection model pulls apart AI responses and isolates the checkable facts: prices, features, policies, specs. Opinions get set aside. Facts get tested.

It checks every claim against your source of truth. You maintain a Knowledge Base of verified facts about your business. Every claim gets graded against it, right or wrong.

It shows you the page to blame. Every wrong claim comes with the truth that disproves it and the exact URL feeding it. Finding errors without finding sources is just anxiety. This is a fix list with addresses.

It matches the fix to the source. Wrong info on a page you own gets refreshed. Wrong info on a competitor's comparison page gets answered with counter-content. Wrong info in a third-party publication gets an outreach request to the author. Different source, different play, and the platform's agents run each one.

Early proof: WHOOP turned it on and found AI misrepresenting them 11% of the time within the first week. A recurring pattern was engines comparing rivals' newest wearables against older WHOOP generations, apples to oranges at the exact moment a buyer decides. They traced it, fixed the sources, and tracked the recovery. At WHOOP's query volume, 11% is millions of wrong answers reaching real customers.

The caveat I'd want flagged for me

The system is only as good as your knowledge base. A stale product page in your source of truth makes the checker bless an old claim or reject a correct new one. Treat the knowledge base like your pricing page: one owner, a schedule for updating it, an approval step before changes go in. Garbage in, confident garbage out applies to your side too.

Practical note: FactCheck is included on Profound ent plans at no extra cost and runs against prompts you already track. The extra work is tagging your brand-critical prompts and standing up the knowledge base…it takes an afternoon. 

Whose job is this?

I get asked this question quite a bit…who owns the fix? 

The tempting answer is one team. The honest answer is that the errors don't come from one place, so the fixes can't either. Engines stitch answers together from your site, trade publications, review platforms, Reddit threads, LinkedIn posts, and a directory listing an intern made in 2019. Every source has a different fix. Every fix belongs to a different marketer.

Content and web own your pages. A wrong price on your own site is the fastest fix in the building, and it happens more than anyone wants to admit.

PR owns everyone else's pages. When the false claim is coming from a trade publication or an analyst writeup, someone has to email the author. That's media relations. The old reflex is a press release, which fixes exactly nothing an engine cites. Profound's Aim shows which publications are actually feeding the errors, so PR can go correct three articles instead of announcing one.

Social owns the places engines listen in. If AI keeps butchering your pricing, put the real numbers where engines actually read: LinkedIn, Reddit, the threads where your category gets argued about. That's not brand air cover anymore. That's evidence.

One thing can't be split up: the insight. Somebody has to own the knowledge base, run the check on a schedule, route each error to the right team, and hand leadership one number: what share of AI claims about us this month were true. Name that person in writing. Everyone's job is nobody's job…and this problem compounds quietly while nobody's watching.

Where this goes

AI search built in layers. Visibility: are you in the answer? Sentiment: how does the answer feel about you? Accuracy is the third layer, and it holds up the other two. High visibility for an answer that misquotes your pricing is worse than not showing up at all. You paid to get cited, and the citation is working against you.

My bet: within 18 months, AI accuracy rate sits in brand health dashboards next to NPS and share of voice. The brands measuring now get a 12-month head start on corrected sources while their competitors are still finding out from angry customers.

The engines will keep talking about you either way. The only open question is whether you know what they're saying, and whether it's true.

Keep Reading