Here is a typical week inside one of our AEO programs. The company is a composite. The sequence is exactly how it runs.

Monday, the decay agent flags a drop. Coverage on "best data pipeline tools for Snowflake" fell from four cited pages to one. The three lost citations went to a competitor's integrations page, published nine days earlier.

The displacement agent reads the source. Competitor, not publisher. Routes it to positioning, not PR.

The diagnosis: their page lists 40 named integrations with setup docs for each. Ours says "integrates with the modern data stack." The model picked the one it could verify.

Tuesday, the client's PMM ships a real integrations page. Forty-two connectors, each with a doc link.

Friday, the agent reruns the prompts. Three of four citations back. The fourth went to a review site, so that one goes to customer marketing.

Five days, one fix, no content calendar involved.

That week teaches the thing most AEO programs miss. The model picked the page it could check. Writing quality had nothing to do with it. Forty named integrations with docs beats "the modern data stack" every time, because a machine can confirm the first one and can only shrug at the second.

Once you see citations that way, the job changes. Almost every team can publish. Very few can tell you, on a given Tuesday, which claim lost, to whom, and why. Most AEO dashboards I get shown are screensavers. Pretty, updated weekly, connected to nothing.

The five agents below answer the Tuesday question. None of them write a word.

Observe → Diagnose → Act → Verify

Each section covers what the agent does, then how we build it. The build notes assume you can run prompts against ChatGPT, Google AI Mode, and Perplexity on a schedule and log the cited URLs. We use Profound for that layer. A scraper and a cron job also work.

1. Citation decay agent

Coverage erodes on its own. A source gets updated. A competitor publishes something more specific. The model quietly stops picking your page, and nothing in your analytics tells you until pipeline does.

This agent watches coverage across the prompts and pages that matter and flags drops worth a human's time. When coverage falls, it pulls which prompts we disappeared from, which URLs took the citation, and what changed in those sources.

The third one is where the answer lives. A citation is a claim plus a source the model trusts. When you lose one, either your claim went stale or someone else's got more checkable.

Refresh because the market moved. Not because a 90-day reminder fired.

How we build it

Start with a fixed prompt set. Fifty to two hundred prompts, each tagged with the page you expect to be cited and the funnel stage it maps to. This list is the strategy. Everything else is plumbing.

Run every prompt three times per model per day. Answers vary run to run, so a single run tells you nothing. Coverage for a prompt is the share of runs where any of your URLs appear.

Compare trailing 7-day coverage to the prior 14 days, per prompt. We flag a drop of more than 30 points that holds for three consecutive days. We started at 15 points and one day. Within two weeks the Slack channel was muted by everyone who mattered. Set the threshold so a flag means something, or nobody will read the second one.

When a flag fires, the agent lists the URLs that appeared in place of yours, fetches the current version of each, diffs it against the copy it stored last time, and opens a ticket with the prompt, the coverage curve, the replacing URLs, and the diff attached.

A person reads the diff and decides the fix. That's the one step we haven't automated, and I'm not sure it should be.

After the fix ships, the agent reruns the prompt for 14 days. It closes the ticket if coverage recovers and reopens it if not.

2. Fanout gap agent

Take a prompt like this:

"What's the best AEO platform for enterprise SaaS?"

The model doesn't answer that as one question. It breaks it into the questions a buyer would ask underneath it. Security. Integrations. Pricing. Customer proof. International support. Comparisons.

This agent maps those questions and grades us on each one: strong, weak, or absent.

The pattern that shows up almost every time: the company is strong on the questions it likes answering and absent on the ones it doesn't. Pricing pages that say "contact us." Security pages with a badge and no detail. Comparison pages that don't exist because nobody wanted to name a competitor.

The model has no such reluctance. It answers the pricing question anyway, using someone else's page. Usually a competitor's comparison post that quotes your pricing wrong.

How we build it

Generate sub-questions for each head prompt from two sources. Ask the model directly: "What would you need to know to answer this well?" Then pull the sub-queries Google AI Mode shows and the follow-ups ChatGPT and Perplexity suggest. Dedupe. You'll land on eight to fifteen per head prompt.

Run each sub-question through the same daily loop. Record three things: were we cited, who was, and did the answer name us at all.

Grade each one. Strong means cited in most runs. Weak means mentioned but the citation goes elsewhere. Absent means neither.

For every weak or absent sub-question, the agent searches our own site and docs for the question and scores the best match. This surfaces the "product does the thing but marketing buried it" cases without anyone reading the whole site.

The output is a grid. Head prompts down the side, sub-questions across, cells colored by grade, each cell linking to our best existing page or left blank. A PMM can read it in two minutes and know what to write next.

Then the harder check. Where we have a page and still lose the citation, the agent pulls the source that won and classifies it: our site, competitor, or third party. If third parties are winning on a question we answer ourselves, that's an authority gap, and it goes to a different owner.

3. Citation displacement agent

The first agent catches the drop. This one decides what happens next.

It looks at the source that replaced us and routes the response to the team that can fix it:

Who took the citation

Where it goes

Competitor

Evidence and positioning

Publisher

PR

Review site

Customer marketing

Outdated third party

Outreach

Our own content

Fix it

Product limitation

Product

The routing is the whole value. A review site outranking you and a competitor outranking you are different problems owned by different people, and a content team can't fix either one by writing more.

If the system only tells you what happened, you built a dashboard. If it figures out why and triggers the next action, you built something useful.

How we build it

The classifier is simpler than it sounds. Maintain a domain list with types: your domains, each competitor, known publishers, review sites, partners. Most replacing URLs match the list. For the rest, the agent fetches the page and asks the model to classify it into one of the six types with a one-line reason. A person reviews the uncertain ones, which is a few a week.

Each type maps to a column on one board. The agent files the ticket into the right column with the prompt, the losing URL, the winning URL, and the diff.

The product limitation route needs a guard. It fires when the winner is a competitor and the diff shows a capability we lack. The agent suggests it. A human confirms it before it reaches product, because a false "we lack this feature" ticket burns trust fast.

Run this for a quarter and count tickets by column. That distribution tells you something about the company. Mostly positioning means marketing has a specificity problem. Mostly PR and customer marketing means an authority problem. Mostly product means no amount of AEO will help and someone should say so. We put that count in the monthly review. It's the slide that makes the room go quiet, because it's the first time anyone has shown a marketing team that half their citation losses are a product or PR problem wearing a content costume.

4. Brand truth agent

Checks what ChatGPT, Google, and Perplexity say about the company against what's true.

Pricing. Features. Integrations. Security. ICP. Markets. Positioning. Limitations.

Then it traces every wrong answer to its source.

Our site is wrong? Fix it. A publisher is outdated? Contact them. No credible source supports the correct answer? Create one.

The third case is the one teams underestimate. When the model says something wrong and no page anywhere states the right thing plainly, the model is doing its job. It answered from the best evidence available, and the best evidence is stale. You can't correct an answer by disagreeing with it. You correct it by publishing something more checkable than what it's using now.

Share of voice means nothing if the model is telling prospects the wrong thing about you. Being cited for a price you no longer charge is worse than not being cited.

How we build it

First, write down the truth. One document, owned by PMM, with current facts across each category: pricing tiers and what's in them, the full feature list, every integration, security certifications with dates, target customer, markets served, and the things the product deliberately doesn't do. If nobody owns this, the agent is comparing the model against nothing.

Second, the question set. About forty questions a prospect would ask about you specifically. "How much does X cost?" "Does X integrate with Salesforce?" "Is X SOC 2 certified?" "Who is X for?" Run them daily across the three models.

Profound's new FactCheck does the comparison step well and is where we'd start. If you're building your own, hand the model's answer and the relevant section of the truth document to a second model call and ask: does the answer contradict the truth, and on which claim? Log every contradiction with its cited source.

Third, trace. Check each cited URL against the domain list from the displacement agent and route the same way. Our own page has the wrong price? Ticket to web. A review site has our old tier names? Ticket to customer marketing with the specific line to correct. No source cited, or the source doesn't actually say it? That's the create-one case, and it goes to content with the exact claim the page needs to make.

Fourth, verify. Rerun the question for 14 days after the fix. Some corrections land in days. Some take weeks, because the model is still leaning on a cached copy of the old page. Track time to correction per source type. That number tells you which fixes are worth chasing.

5. Commercial shortlist agent

My favorite, because it sits closest to revenue.

It runs the prompts that produce a shortlist:

  • "Best X"

  • "Best X for enterprise"

  • "X vs Y"

  • "Alternatives to X"

Then it pulls the answer apart. Who got recommended? What evidence supported each one? Why were we included or excluded?

Then it classifies the loss. Most teams skip this. It's the part that matters.

They have something we don't. Product problem.

We have it, but nobody can tell. Positioning problem.

We say it, but nobody credible backs it. Authority problem.

The evidence exists, but the model isn't finding it. AEO problem.

Four losses. Four owners. Four fixes. Treating them all as "we need more content" is how teams spend a quarter publishing and end the quarter in the same spot on the shortlist.

The part nobody likes hearing: when we run this, the fourth category is the smallest. Most shortlist losses are positioning or authority. The evidence isn't hidden from the model. It doesn't exist, or it only exists on the company's own site, which the model treats the way a buyer treats a sales deck.

How we build it

Cross the four templates with your categories, segments, and top five competitors. "Best X" becomes "best X for enterprise," "best X for mid-market," "best X for [industry]." Do the same for "X vs Y" and "alternatives to X" per competitor. You'll end up with sixty to a hundred prompts. Run them daily.

For every run, extract the recommended list in order and, for each vendor, the reason given and the URL cited for it. Ask the model to return a table: vendor, position, reason, source. Store every run.

Diagnosis happens only when we're excluded or ranked below a competitor. The agent asks four questions in order and stops at the first yes.

Does the winner's stated reason describe a capability we lack? Check it against the truth document. If yes, product loss.

Do we have it, but no page on our site states it plainly? Search our site for the capability. If nothing matches, positioning loss.

Do we state it, but the winner's cited source is a third party and ours would be our own site? Authority loss.

Otherwise we state it, third parties back it, and the model still isn't finding it. That's the AEO loss, and it's the only one where the fix is technical: structure, crawlability, the claim buried in a PDF, whatever it turns out to be.

Each loss type files to its owner on the displacement board. The monthly output is two numbers per prompt: our shortlist position, and the loss type when we're not first. Plot those over time and you have the closest thing to an AEO revenue metric I've found.

Where it still breaks

The decay agent over-flags on prompts with unstable answers. Some "best X" queries reshuffle citations every run no matter what anyone publishes. It took weeks of watching to learn which prompts are stable enough to act on. Three runs a day helps. It doesn't fix it.

The displacement agent gets the routing wrong when a source is two things at once. A publisher with a review section. A partner that's also a competitor. A human still checks the route before anything triggers.

The truth document goes stale. Pricing changes, nobody updates it, and for two weeks the brand truth agent files tickets against a model that's right and a PMM who's wrong. The fix is a rule: no pricing or packaging change ships without a truth document update in the same ticket. Getting that rule to stick is a management problem, not a technical one.

The system

Less than people assume. A fixed set of prompts that map to revenue. A scheduled run against each model. A diff on cited sources between runs. A domain list with types. A truth document. One board with six columns. A rerun after every fix.

The hard part is the prompt set. Choosing the hundred that matter is a strategy decision, and it's the one most teams never make. They monitor 3,000 prompts, act on none of them, and call it visibility.

What changed? Why? What should we do? Did it work?

Observe. Diagnose. Act. Verify. Run it again.

Production is cheap now. Anyone can ship 40 pages a month. Diagnosis is the scarce skill, and it's the one nobody's hiring for yet.

If you're building something like this, reply and tell me where your loop breaks. I'll compile the patterns and publish them.

Go build some cool shit this week. If you want a second set of eyes on it, or you just want to humble brag about an automation that worked, send it my way. I accept both.

Keep Reading