Your AI Visibility Score Is Only as Honest as Your Prompt Index
Why the most important input in AEO is sitting in a settings panel you opened once.
Open ChatGPT. Ask it to name the best brands in your category. Screenshot what it says. Now do that ninety-nine more times.
You will get close to a hundred different lists. Order shuffles, length changes, the name that led one answer goes missing from the next. It looks random but really mostly isn’t and the part that holds still is the entire game.
I have spent the last 2 years auditing AEO strategies for companies paying real money to watch a number on a dashboard, and I open every engagement the same way: I read the prompts the tool is tracking and hold them against the questions I am watching buyers actually ask. Nine times out of ten the tool is measuring a market that stopped existing months ago, and nobody in the building knows, because the report keeps coming back clean. This is not a vendor bug you can file a ticket on. It is what these systems are, by design.
Which raises the question almost every team skips on the way to buying a dashboard: if the answer is different every time you ask, what exactly is the number on your AI visibility report measuring?
It is measuring a sample. Most teams read it like a census, and the gap between those two things is where AEO budgets go to die. The one input that decides whether the sample is any good is your prompt index, the curated universe of questions your tools ask the engines on your behalf, organized by topic, persona, and intent. Get it right and your visibility data describes the market you actually compete in. Get it wrong and you will spend two years optimizing against a mirage with a clean decimal point stamped on it.
When I audit an index, the questions I test it against don’t come from a dashboard. I run agents across the places buyers actually talk, the Reddit threads, the G2 reviews, the community forums, and pull the raw language people use while they are still deciding. That is buyer intent captured upstream, before anyone opens a chatbot. It almost never lines up with what the tool is tracking. Sometimes the set was wrong on day one. More often it was fine at setup and then nobody touched it again, frozen in whatever the market looked like that week while the buyers moved on without it. Either way, the dashboard on top kept reporting confident numbers the whole time.
The index isn’t a setup step you knock out in week one and never open again. It’s the denominator under every metric you will ever report, and a CMO who signs off on an AEO strategy is signing off on a prompt index whether they know it or not.
Here is what took me too long to see, even while I was running a global growth org full time: a good prompt index is barely a measurement asset. It is the clearest read you will ever get on what your market is trying to buy. When my team built our AEO program, the prompts that mattered were never the ones with our name in them. They were the unbranded ones, the “what is the best headless CMS for a marketing team that can’t lean on engineers” questions, because that is the exact moment a buyer is deciding and hasn’t picked yet. Get that set of questions right and you stop merely measuring AI visibility. You start holding a live transcript of your buyers’ intent, in their own words, refreshed continuously. That transcript should be the input every other GTM decision keys off. Most teams have it sitting in a settings panel they opened once and never reopened.
TL;DR
Every number on your AI visibility dashboard is a sample, not a census. A set of questions called your prompt index generates it, and that set is almost never audited after setup.
AI answers are wildly inconsistent run to run, but a brand’s appearance rate across many runs is stable and measurable. The catch is that it is only as honest as the questions you chose to ask.
The errors split in two. Depth, meaning how many times you run a prompt, is recoverable with more runs. Breadth, meaning which intents, personas, and engines you cover, is not. A gap in coverage is a blind spot no volume of data fixes, and your dashboard looks just as confident with the hole as without it.
Wording barely matters. The models read intent, not phrasing. Coverage of real buyer intent is what moves the number.
A good prompt index is not a reporting asset. It is the sharpest read you get on what your market wants to buy, which makes it a CMO decision and a GTM input, not an SEO setting. The revenue leaders who treat it that way are the ones who will win. This is my personal most important belief in this whole piece.
A five-minute gut check, and how to build an index that survives scrutiny, are at the bottom.
The randomness is the easy part
Last winter Rand Fishkin ran the experiment everyone in this space should have run before spending a dollar. Working with Patrick O’Donnell at Gumshoe, he had 600 volunteers run twelve prompts across ChatGPT, Claude, and Google’s AI nearly three thousand times, and logged every brand list that came back. The result was brutal for anyone selling “rankings in AI.” Less than a one in a hundred chance of getting the same list of brands twice. Less than one in a thousand of getting the same order. These tools produce a fresh answer on every run, so any product promising you a stable rank position inside an AI answer is selling something that does not exist.
If that were the whole story, the category would be snake oil and you'd be right to bail (don’t get me wrong, there is an unfair amount of BS in this category). It isn't, and the reason is where most teams stop reading.
When Rand’s team stopped staring at individual answers and started counting how often each brand appeared across all the runs, the noise resolved into signal. Ask Google’s AI to recommend digital marketing agencies with e-commerce expertise ninety-five times, and one firm, Smartsites, shows up in eighty-five of the answers. An 89% appearance rate. The rank inside any single answer bounced around almost at random, but whether a brand appeared at all held remarkably steady. One answer is a slot-machine pull. The pattern across hundreds of them is the thing you can actually measure.
So a visibility score isn’t a reading off a meter, it’s a statistic. Ask a population of questions enough times, count how often you turn up, divide. That makes it a poll, and every leader already knows the first law of polling even if they have never run one. The result is only as good as the sample. You can poll a thousand people and land further from the truth than someone who polled a hundred, if the thousand were the wrong people. Precision tells you nothing about whether the number is true. Your prompt index is that sample, and almost nobody audits it.
The math that makes the number real
Almost no dashboard will tell you this, but the randomness is a solved problem statistically. You just have to respect the unit of measurement.
Run a single prompt once and you have close to nothing. Run it enough times and the share of answers you show up in becomes a number you can actually estimate, and at low run counts the error bars are embarrassingly wide. This is just binomial math, the same math behind every political poll. A brand tracked on a single prompt, sampled the handful of times most tools manage, carries a margin of error wide enough that a "40% visibility" reading could honestly be sitting in the twenties or the fifties, and the tidy number on the screen would never tell you. Roll dozens of prompts into a topic and track them across weeks, not days, and that error tightens to low single digits. Nothing changed except the unit you measured and how often you looked. One prompt is jitter you can't act on. Roll enough of them into a topic and you have a number that survives a board meeting.
This is exactly why the serious platforms organize the world as topics with prompts rolled up underneath, instead of a flat list of questions. A topic is a broad buyer need, something like “cloud security solutions.” A prompt is the specific question underneath it, something like “compare the top three cloud security tools for a small business.” The topic is what you can attach a confidence interval to. The prompts are the samples feeding it. Leave prompts floating in a flat list with no topic layer and you have nothing you can measure with any rigor, which is where most tools leave them.
The depth side has hard rules of thumb now, and they are cheap to apply. A 2026 arXiv analysis of source stability found you need at least seven or eight runs per prompt before per-brand detection rates settle, and that daily source churn runs around 65%, so a single day or even a single week is too short a window to separate signal from noise. To be clear…that is seven or eight runs accumulated over weeks, not a prompt you hammer twenty times every morning. That would be expensive, and it misses the point. Depth only has to clear the noise floor. Once it does, every extra dollar buys you more by widening the net than by deepening the same hole.
None of this is fringe anymore. In May 2026 AMEC, the standards body that governs how communications gets measured, put out its GEO Principles, and the fourth one draws the line I've been drawing on whiteboards for two years: what one tool returns to one prompt, at one moment, in one market, is directional at best. Real measurement means a documented set of prompts, run repeatedly, with the variation shown rather than buried. That last part is where the whole category fails. Almost no AEO dashboard shows you the variation. They hand you a percentage and let you assume it's a fact.
There is a tradeoff buried in that math worth internalizing. Ten prompts run thirty-nine times each and fifty prompts run eight times each land at roughly the same confidence, around five points of error, for about the same total number of queries. The work is similar. But spreading the budget across more distinct prompts buys you something repetition never can, an estimate of the whole topic instead of one phrasing. That is the rule to carry out of the math. Once each prompt clears the minimum runs to stabilize, more prompts beats more repeats every time. You clear depth once. Breadth is where the game is actually played.
Run your dashboard against those numbers and most setups fail at least two of them. That is not a small thing. It means a meaningful share of the week-over-week “wins” teams report are smaller than their own margin of error, which is the analytical equivalent of re-polling forty people every morning and calling the jitter a trend.
Why none of that math saves you
Everything above fixes depth, and none of it touches breadth. Breadth is where the errors you can’t take back live.
Go back to Rand. After the brand-list experiment he asked his volunteers to write their own prompts, in their own words, for one shared intent: the best headphones for a family member who travels. A hundred and forty-two people produced a hundred and forty-two wildly different prompts. When his team scored how similar those prompts were to each other, the number came back at 0.081 on a scale where 1.0 is identical. Fishkin likened the average pair to Kung Pao chicken and peanut butter, two foods with peanuts in them and otherwise nothing in common. People do not search the way they Googled. They do not compress intent into three keywords. They ramble. They pile on context nobody asked for, then get oddly specific about the one thing that actually matters to them.
And the brands still converged. Across nearly a thousand responses to those scattershot prompts, Bose, Sony, Sennheiser, and Apple each turned up in 55 to 77% of answers. The phrasing was all over the place. The intent underneath was identical, and the models read the intent under the words and returned a stable set of brands.
Conductor’s research team ran the harder version of this experiment and landed in the same place. They fired 14,000 API calls across ten industries, seven intent types, and four engines, fifty runs per prompt, and found AI recommendations are not random at all. They are predictably inconsistent, and the predictor is not the industry, it is the intent behind the question. Purchase prompts, where a buyer is ready to spend, were the least stable in the study. Out of every ten brands that surfaced across any two runs, only four showed up in both. Comparison prompts, where the brands are already named in the question, were the most stable by a wide margin, the same brand leading 91% of the time. Same market, same buyer, wildly different reliability depending only on the shape of the question. That is the breadth problem stated as a law. Over-index on the easy, stable intents and skip the volatile ones, and your dashboard looks calmest exactly where your exposure is highest.
That should change how you build an index. The specific words barely matter, because the engine responds to intent, not syntax. Obsessing over whether you phrased a prompt the way a “real buyer” would is wasted effort. What matters is whether your index spans the full range of intents your buyers carry. Miss a phrasing and you lose nothing. Miss an intent, a persona, a journey stage, or an engine, and you have a hole no volume of additional runs will ever fill, because the observations that would fill it were never collected. Depth you can always go back and buy. Breadth, once you have decided what to ask, is locked, and the dashboard looks exactly as confident with the hole as without it.
That single distinction reframes the whole job. Building a prompt index is a coverage problem, not a phrasing exercise, and coverage means deciding which people, which journey stages, and which engines you are willing to be measured on. Those are strategy calls, and they belong to whoever owns the strategy.
The three indexes that lie to you
Coverage fails in three common ways, and each one feels responsible while you are doing it.
The first is the borrowed panel. Many tools do not build your index from your business at all. They infer prompts from third-party data, often scraped from browser extensions, and sell the size of that pool as a strength. The problem is who is in the pool. ChatGPT alone fields something like 2.5 billion prompts a day, and a panel assembled from the narrow, self-selected slice of people who install tracking extensions is a fraction of a percent of that, and skewed on top of it, because the extension-installer is no stand-in for your CFO buyer or your clinical director or your procurement lead. You are sampling the people nearest the exit and calling it the electorate. There is a quieter risk too. Data sourced this way can vanish overnight when an app store rule or a privacy law shifts, pulling the floor out from under your measurement without your consent. Panels are not worthless. You just have to know when one is the foundation of your number rather than a supplement to it.
The second is the wish list, the all-human index where a smart team writes down the questions they believe customers ought to be asking. It feels rigorous because it took effort, and two biases undercut it anyway. The first is reach. No group of people can enumerate the real variety of how a market asks, so the long tail goes uncovered, and the long tail is often where the engines cite most generously. The second is vanity. Teams write the prompts they want to win, the ones that flatter the product, and quietly drop the ones where they come off badly. What comes back is an instrument calibrated to make you feel good.
The third is the generic model dump, where you ask a general chatbot to list prompts for a topic. Fast, free, and hollow, because the model does not know who you are. It does not know your content, your architecture, your personas, or where your demand sits, so it hands you the average of the internet, a description of a generic competitor in your category rather than of you. The president of Optimizely said it well when his team went looking for an AI search data partner: most visibility tools build their prompts in a vacuum, with no grounding in what people are actually searching for. Same disease as the borrowed panel, newer clothes.
All three share one root cause. The index has been cut off from the only two anchors that keep it honest, your own business and real demand. Reconnect those two and most of the failure modes sort themselves out.
What a good index looks like, concretely
You can actually inspect a good index, point at each part and say why it’s there. That is the whole line between infrastructure and a vibe.
It is built from your own business as the source of truth, not from a generic outside list. The most rigorous version I have seen starts from your actual website content and site architecture. This is the core of how Conductor approaches it. Your site is stored in a vector database, and the system reads your specific pages and the semantic relationships between them to generate prompts grounded in what you have real authority on, then validates those against real organic search demand drawn from more than a decade of search data, so the questions are not just on-topic but genuinely in demand, in the geographies you care about. Your content is where the relevance comes from. Real demand is what keeps the set representative instead of merely plausible. A brainstorm gives you neither. A chatbot hands you a confident average of everyone who isn’t your buyer.
It is generated, not hand-typed, because the scale demands it. The method that matters here is synthetic prompt generation, using AI to produce a large, realistic set of conversational queries that mirror how real people ask, built around intent and context rather than keyword volume. Done badly, that is just a chatbot listing a hundred questions, which drops you right back at the generic dump. Done well, it is grounded in your site, weighted by real demand, and segmented before a single prompt is written. That gap, between a list of plausible questions and an index grounded in your business, is the thing to push on when you evaluate any tool.
It covers personas and intent on purpose. A hands-on practitioner and a C-suite buyer ask different questions in different language about the identical topic, and a good index holds both rather than averaging them into a person who does not exist. Better still when personas are built into generation from the start and can be mapped to specific products, and when the index deliberately walks the buying journey instead of camping in the easy middle. The lifecycle is the frame I find most useful: education and “what is” questions, recommendation and “best of” discovery, head-to-head comparison, pricing and value, branded navigation, high-intent purchase, and post-purchase support. Each stage is a different visibility fight, frequently against a different set of competitors. Conductor lets you choose which stages to generate against, and which personas and engines to track, before you create a single prompt, then surfaces the result as a persona-and-intent grid where the empty cells are the uncaptured demand. The grid view that matters is the one that shows you the holes instead of hiding them
Filled cells are demand you are capturing. Empty cells are demand you are not. Most teams have never seen their own index laid out like this.
It spans engines, not just one. Independent analyses of large prompt sets have found that different engines cite startlingly different sources for the same questions, with overlap between some platforms down in the single digits. Being the answer in ChatGPT and being invisible in Perplexity or Google’s AI are separate outcomes, and a single-engine index will happily hide one behind the other. Track one engine and you are measuring a slice of your visibility and guessing the rest.
It is mostly unbranded. Roughly three quarters of prompts should carry no brand name. Branded prompts confirm you own your own name, which you should, and which teaches you almost nothing. Unbranded prompts tell you whether you are reaching people who have never heard of you, which is the entire reason AI search matters as a discovery channel. An index stuffed with branded queries produces high scores and confirms what you already knew.
And it is deep enough to clear the noise floor, on the depth rules above, so the movements you report are real and not the dice resetting.
Why this lands on the CMO’s desk
Everything I just walked through sounds like configuration work, somebody in a settings panel checking boxes after they bought the tool. The prompt index is where an organization encodes the definition of its own market, which makes it a leadership decision wearing a technical costume.
I couldn’t be more passionate about this belief.
Trace what the index silently sets. It fixes your competitive set, because the only rivals who can appear beside you are the ones your chosen questions summon, which means a sloppy index can have you benchmarking against companies you don’t compete with and feeling good about beating them. It sets the baseline for the sentiment and share-of-voice figures headed into board decks and budget fights. It decides which content gaps look urgent, since a gap can only surface if a prompt existed to reveal it. A topic you forgot to index is a market you are blind to, and blindness does not feel like blindness from the inside. It feels like a clean dashboard and a good night’s sleep.
This is why “the SEO person handles the prompts” is an expensive sentence. It is the equivalent of saying the analyst handles the definition of revenue. The mechanics can live with a specialist, gladly. The strategic question underneath, are we measuring the market we actually fight in, belongs to the CMO, and deserves the scrutiny you would give any change to how you count pipeline. It also imposes a discipline most teams skip. If your reported quarter-over-quarter movement is smaller than your margin of error, you have learned nothing, and calling that wobble a trend is how good teams talk themselves into bad bets. When a tool flags a competitor newly claiming a high-value topic, or sentiment sliding on a specific source, those are not tickets for the SEO queue. They are early signals about your position in a market, worth exactly as much as the index that produced them and not a cent more. Conductor frames its AI Search Performance report as a system of record for this reason. The dashboard was never the point. The point is a defensible, audience-aligned account of where you stand that leadership can trust, because the sample under it was built to hold weight.
The view you can actually take to a board: benchmarked share of the conversation, where the gaps are, and what the competitive exposure looks like.
The stakes here are not abstract. AI platforms drove over a billion referral visits in a single month last year, up around 357% year over year, and those visitors tend to arrive already pre-qualified by the model, which is why teams keep finding they convert at multiples of ordinary organic traffic. The channel is still small as a share of total traffic and growing faster than anything else on the web. The brands building an honest measurement foundation now are the ones who will be able to act on it while it still counts as an advantage rather than table stakes.
The index pays off well past marketing
Once you hold an index that truly maps how your market asks questions, you have something more valuable than a visibility tool. You have a structured, continuously refreshed picture of buyer intent, and that picture is useful to functions that have never heard the term AEO.
I saw this firsthand at Webflow. The work that earned us a 64% answer share in the CMS category we didn’t even own 2% market share in was not a clever schema trick. It started with the index telling us, plainly, which comparison and objection questions buyers were bringing to the models, and a lot of those were questions our product marketing had never framed an answer to. The index was doing demand research before it was doing measurement. That is the part most teams leave on the table.
Product can read the comparison and pricing prompts that keep recurring as a live feed of the objections and feature expectations in the market, in customers’ own words rather than filtered through a roadmap survey. Sales and enablement can mine the same prompts for the exact language buyers use, which beats most win-loss programs for discovery and objection-handling. Communications can watch sentiment move at the level of individual prompts and trace a shift back to its source fast enough to act while it still matters, turning reputation work from quarterly archaeology into something close to real time. Competitive intelligence falls out almost for free, since the index constantly shows which rivals own which ground in front of which audience.
None of that holds if the index is a borrowed panel or a wish list, because then it is a detailed map of the wrong country and everyone navigating by it is lost in the same confident way. Build it right and it stops being a marketing utility and becomes shared infrastructure, one honest definition of what your market is asking that every customer-facing team can build on.
The agentic turn is already here
This matters more every quarter, not less, and there is a reason. Search is going agentic. The same engines that answer questions are starting to take actions, and the platforms are racing to wire AI insight straight through to execution. When Optimizely built out its AEO platform this year, it chose Conductor to power the prompt and visibility layer underneath its autonomous agents, the ones that find content gaps, benchmark share of voice, and route work without a human in the loop. Strip away the announcement language and the logic is plain. An agent that acts on your behalf is only as good as its read of the market, and its read of the market is the prompt index. The foundation does not get less important when machines start acting on it. It gets more important, because now the errors propagate at machine speed.
A five-minute audit you can run this week
You do not need a project to find out whether your foundation is sound. Take your current AEO setup into your next review and put six questions to it. Each one has a failing answer, and the failing answers are common.
How was our index built? If the answer is a third-party panel, a brainstorm from last spring, or a chatbot session, you do not have a measurement program. You have a number that moves.
How many prompts per topic, and how many topics? Under fifteen prompts a topic and you are below the floor where the curve even starts to flatten. A flat list with no topic layer means you have no unit you can attach a confidence interval to.
What share of prompts is unbranded? Below roughly three quarters and you are mostly measuring whether you own your own name.
Which personas, intent stages, and engines are covered, and which are not? If nobody can name the gaps, the gaps are running your strategy. Ask to see the empty cells, not the full ones.
How many times is each prompt run, and over what window? Fewer than seven runs, or a window shorter than a couple of weeks, and your numbers are inside the noise.
Does our reported movement exceed our margin of error? If you can’t answer this, you can’t tell a win from a wobble, and you are probably reporting both as wins.
Fail two or more, and most teams do, and the fix is not a better dashboard. It is a better sample.
Where to start
I’ve built this index the hard way, and I’ve reviewed enough of other people’s to know the pattern never changes. The dashboard looks fine. The foundation underneath it is a guess.
So rebuild it from the two anchors that carry real weight: your own content and authority, and real search demand. Cover the personas, journey stages, and engines you actually compete in on purpose, not by luck. It’s real work, and there’s no shortcut around it. The tools that do this well, Conductor among them, use your site as the source of truth and ground thousands of prompts in real demand instead of a generic list that flatters your product. Whatever you run, run it with your guard up and hold it to the six questions above.
Get the index right and everything on top of it gets honest. Get it wrong and you keep walking into rooms with confident slides about a market that was never there. I’ve been in those rooms. I’d rather you weren’t.




