An $8 T-shirt and a $1,650 T-shirt are not competing for the same customer.
In research Profound published on September 9, representative T-shirts from the brands recommended to low-income shoppers ranged from $8 to $24. Middle-income shoppers got brands in the $10 to $58 range. High-income shoppers got $140 to $1,650. Those are prices from the recommended brands rather than products named in each answer, but the spread tells you what happened: the question stayed the same, the person changed, and the AI sent them into a different market.
Now picture reporting your brand's performance across all of those answers as one visibility number. The math can be correct and the picture of your business can still be wrong, because the number never asked who was buying.
More and more of the people who come to StackedGTM asking us to review their AEO program arrive with the same thing: a dashboard that's up and to the right, and a question about why the pipeline isn't following. Visibility is a leading indicator. It's supposed to move before revenue does. But it only leads somewhere useful if it's measuring the buyers who'll eventually show up in the pipeline, and most of the dashboards we're asked to look at can't say who it's measuring. A dashboard without that context doesn't just miss detail. It lies to you, and it does it with clean numbers.
That's the question I'd bring into the next AEO review: are we measuring whether AI recommends us to the customers we want, or whether it recommends us to anyone at all?

What Profound found
This is the most useful data I've seen published on the question. Profound's research team, in a study written by Allison Huang, ran the same base prompts across ChatGPT, Claude, and Gemini over fourteen days in August and collected 71,147 responses. The prompts came from three consumer categories: clothing, furniture, and credit cards. Each one was prefixed with a single detail about the person asking, so "best affordable furniture companies" became "I am in my 30s. best affordable furniture companies." The details tested were gender, four age brackets, three income levels, and five kinds of occupation.
Five things came out of it that matter for this argument.
The shortlist changes. Ask the same question twice for the same person and the answers share about 40% of their recommended brands. Change the person and that drops to about 25%. That comparison is the careful part of the study, because these systems already vary from run to run; the effect of changing the person sits on top of that noise rather than being explained by it. The brands themselves moved in ways you'd recognize: Patagonia and Ralph Lauren led clothing answers for men, Reformation and Free People for women, and low-income clothing answers surfaced resale sites like Poshmark and Depop while high-income answers filled up with luxury labels.
The shortlist gets longer for richer buyers. High-income shoppers averaged 7.5 brand mentions and 4.8 cited sources per answer. Middle-income got 6.8 and 4.4. Low-income got 5.8 and 4.1. Some buyers are handed more options than others for the same question, which means raw mention counts run higher for some segments before you've done anything.
The sources change, and so does where they come from. Cited websites overlapped 34% for repeated same-person answers and 20% across different people. In Gemini's clothing answers, high-income responses pulled 24 percentage points more of their citations from brand-owned sites, 17 points fewer from press and third-party coverage, and 8 points fewer from institutional sources than low-income responses did. If you sell upmarket, your own site is doing more of the work. If you sell downmarket, coverage from other people is.
The AI rewrites the question before it searches. Profound could see the actual searches ChatGPT and Claude ran (Gemini doesn't expose them). Both engines routinely turned income into budget words, searching for "luxury" options for high-income shoppers and "affordable" ones for low-income shoppers, across every category tested. The buyer never typed those words. The AI decided they belonged in the search.
The AI fills in details nobody supplied. Across ChatGPT and Claude, clothing searches for people in their 50s were about 39 percentage points more likely to contain the word "women" than searches for other age groups. The prompt supplied an age. The search query added a gender. The researchers saw the same thing with guesses about work, retirement, accessibility, and stage of life.
Two limits, stated once so I don't have to keep repeating them: the study tested details written into the prompt, not account memory or whatever these products do behind the scenes, and it covered three consumer categories, not business software. Profound's own conclusion is the right one: measure visibility for each audience you care about, not just one combined score, and make your intended customer explicit in what you publish. They've built their persona tracking around exactly that. What I'd add is about who inside a company owns that work, because the tool can show you the answers by audience, but it can't tell you which audience you're trying to win.
You can't be recommended to a buyer you never defined
Every measurement program makes choices: which questions get tracked, how much context those questions carry, which markets and situations they represent, how each answer gets weighted. Those choices define the buyer your reporting describes. A set of generic category questions is a sample of possible buying situations, and calling the result "our AI visibility" makes the sample sound more representative than it is.
Take a hypothetical software company that sells mostly to regulated enterprises. Its reporting shows mentions climbing across category prompts and competitor share falling, and the trend looks good. Now suppose the gains sit in answers written for small teams that want a low price and a fast setup, and when the question includes complex permissions, a security review, or a painful migration off the tool they already use, the company isn't there. Its measured visibility went up. Whether its business got stronger is a question a combined score cannot answer, because nobody told it which buyer to care about.
We already know how this works in every other channel. More leads can hide fewer qualified deals. More traffic can hide weaker intent. Cheaper customers can hide customers who leave in ninety days. When I owned a revenue number at Webflow, nobody cared how many people we reached; they cared whether the people we reached bought and stayed. AEO hasn't earned an exemption from that, and Profound's data is the first clean evidence of how much room there is to fool yourself.
The question the AI actually asks is not the one your buyer typed
The search-query finding is the one I'd spend the most time on, because it changes what "showing up" even means.
A buyer gives one piece of context. The AI expands it into a more specific reading of who they are, rewrites the search, and then goes looking. Sometimes the rewrite is useful. Sometimes it inserts a requirement the buyer never expressed, or a budget, or a gender. Either way, your content isn't competing on the question the buyer asked. It's competing on the question the AI built from it.
For business software, nobody has published the equivalent yet, and the open questions are large. Does "founder" get rewritten as small budget? Does "enterprise" get rewritten as pick the big name everyone already uses? Does "procurement" push the search toward cost controls when the real concern is whether the rollout will go badly? If the rewrite misreads the buying situation, one more comparison page doesn't fix it. You'd be optimizing for a buyer the AI doesn't think exists.
The only way to know is to run it, which is why the next section is about testing rather than publishing.
Product marketing needs to own the buyer definition
An AEO program needs a clear account of why a specific customer should pick you, and that forces decisions a lot of companies have been able to leave vague: who gets the most value from this product, what usually happened inside their business before they started looking, which constraints make you the strong choice and which make you the wrong one, and what evidence backs any of that up.
Those are positioning questions. They live in customer interviews, win/loss reviews, and sales calls. Technical, content, and analytics teams still do most of the execution in AEO, but the definition of who the work is for has to come from the people who talk to customers, and right now in most companies it comes from nowhere.
"An intuitive platform for growing businesses" leaves almost the entire buying situation unspecified. "A reconciliation platform for multi-entity operators that need to keep their existing accounting system" gives a reader a problem, a customer, and a constraint, and it gives a search query something to match. I'd want to test whether that kind of specificity improves the recommendations that matter. The study doesn't establish it as a ranking factor and I'm not claiming it is. But it's a far better direction for experiments than another round of powerful, seamless, and easy to use.
What I'd ask my team to do next
Start with three buying situations drawn from real customers, not job titles. For a software business that might be a first-time buyer with no implementation team, a department replacing a tool it already has, and an enterprise expanding across regions.
Keep one base question. Add one piece of context at a time. Run it enough times that a single answer can't become a strategy, and log the model, the date, the prompt, and the context so a change six weeks from now is interpretable. Where the AI exposes its searches, capture those too; they're the closest thing you'll get to seeing how it read your buyer.
Then read past whether the brand showed up:
Was it recommended for the right reasons?
Did the answer describe the product accurately?
Did the search add requirements the buyer never stated?
Which competitors appeared under which conditions?
What evidence did the answer lean on, and was it yours or someone else's?
The output should be specific enough to assign. "We need more visibility" isn't work. "We disappear when the buyer mentions switching off their current tool, the AI rewrites that as 'easy setup,' and none of the cited sources explain our migration process" is work. Someone owns it, the evidence gets fixed, the test gets rerun.
The B2B study I want to see
Hold the software category constant. Vary company size, buyer seniority, budget, existing tools, and rollout constraints. Capture the rewritten searches alongside the answers. Those dimensions were outside this study's scope, and they're the ones a B2B marketer lives in. Profound has the setup and the discipline to run it well, and I'd read it the day it came out.
I'd want to know which of them move the shortlist the most, whether the brands gaining overall visibility are gaining with their intended customers or with everyone else, and whether the source mix flips the way it did in the consumer data. That would tell a CMO where to invest, show product marketing where the market misunderstands the product, and give agencies a more honest way to report progress.
A better definition of winning
A combined visibility score is still worth having. It's a starting point and a trend line. The mistake is letting it answer a question it was never built for: are we more likely to win the business we want than we were last quarter?
So before the next uptick gets celebrated, ask to see the answers that match your best customers' real buying situations. Look at who gets recommended, read the reasons, and check what the AI assumed. You can grow your share of answers without growing your share of buyers.