Should You Fire Your AEO Agency?
The dark side of AEO right now...
I’ve had some version of this call more than twenty times in the last month. Let me compress all twenty into one.
It’s a Tuesday. A founder shares their screen. There’s a slide with a number that went from 34 to 51, arrow pointing up, and underneath it a screenshot of an answer where their company got named first. They’ve been paying between $8k and $40k a month for a year to make that slide exist.
They look at it for a second. Then they say the thing all twenty of them say, and it isn’t “we got scammed.” It’s a bit more careful than that.
“I think we’re doing the work. I just can’t tell if it’s the right work.”
That sentence is why this exists.
Why I’m the one saying it
I didn’t want to be a voice in this market.
Still forever have imposter syndrome. I’m a high school dropout. No degree doing quiet work in the background, no school anybody recognized, no network that arrived free with it. What I had was the willingness to show up and put a number next to my name that nobody in the room could argue with. Credentials were never available to me. Proof was the only currency I had, which is where the rule came from and I held it too hard: those who can’t do, teach. It’s also why a number on a slide that nobody can reproduce makes the hair stand up on my neck.
Then this thing broke, and for most of a year the loudest conversation in the field was whether to call it AEO, GEO or AIO. An industry naming itself before it had learned anything. Meanwhile my team at Webflow was inside it, building and measuring and putting revenue against signups arriving from assistants, with answers to questions nobody in public had thought to ask because we’d been forced to find them to hit a number.
So I swallowed my pride, broke my own rule, and started publishing the work with the results attached whether they were good or not.
What I didn’t expect is that when you become someone people bring things to, they bring you everything. Roadmaps. Scopes of work. Monthly reports. Renewal decks. Dashboards with the password typed into the chat. I read them the way an operator reads them, which is by asking what each line item is supposed to do to the business.
That’s the part that’s hard to sit with. This is not a hit piece. Sometimes the agency is the problem, sometimes the client is, and I’ll take both apart. But I’ve read enough of these documents to say plainly that the average one is not good, and the people signing them are sharp operators with no way to know that yet, because there’s no reference point in a field this young. Some agencies and advisors are exceptional at this. Props to them, and none of what follows is aimed their way. When I’m at capacity I send people to them, which is the only endorsement in this business that costs the endorser anything. Every question below is one they would enjoy being asked.
Where my bias is
Loud and up front, not buried in a footer.
I own a business that does this work. StackedGTM.AI is a media and advisory business, and the companies reading this are the companies I sell to. Which means there’s an obvious version of this article: a list of reasons your agency is failing you, a sympathetic nod, and a phone number at the bottom.
This is not that, and here’s how you hold me to it. There’s no call to action at the end. Every diagnostic below is something you run yourself in an afternoon, with no tools and nobody new on payroll, including me. Several of them will make my own sales conversations harder, because a buyer who has run this test walks in asking much better questions than one who hasn’t.
That’s the trade and I’m making it on purpose. If the standard I’m holding agencies to doesn’t apply to me first, none of it counts.
First, run the calendar
The urgency broke about eighteen months ago. Not the idea, the urgency. Buyers started opening a chat window instead of a search tab, pipeline arrived from places nobody had a dashboard for, GA4 got strange, and every board deck grew a slide about it.
Agencies spent the next two quarters building an offering. Sit with what that actually means. You have a services business with a P&L built on retainers, a delivery team trained on keyword clusters and content briefs, and a sales team fielding a question they cannot answer with “we don’t do that.” So they don’t. You rename the practice, write a point of view, build a dashboard, put AEO on the site, and go sell it.
Contracts got signed about a year ago. The first ninety days worked. Genuinely worked, and I’ll come back to that.
We’re now a year into the doing. The first real cohort of AEO retainers is up for renewal, the founders and CMOs who signed them are being asked what they got, and the honest answer for a lot of them is a great first quarter followed by nine months of a flat line dressed up as volatility.
Most of these engagements didn’t fail. That’s what everybody gets wrong. They succeeded once, and the agency mistook a one-time correction for a repeatable method.
That distinction is the difference between firing somebody and fixing something.
Why the increment died
If your curve went vertical and then went sideways, you’re the pattern, not the outlier. Here’s what sits underneath it.
The first quarter was a correction. A year and a half ago, essentially every company was in a state of basic neglect. Bot access misconfigured. No comparison pages. Entity data a mess: wrong founding year, three different company descriptions, a product renamed a year earlier that half the web still called the old thing. Content that only existed after JavaScript ran. Nobody had done the hygiene because until then there was no reason to.
You hired someone, they fixed it, and you got a step function. Everybody got a step function. It looked like the playbook working. It was the baseline being unoccupied.
Then the room filled up. Over the following few quarters, every vendor in your category hired their own version of your agency and ran the same corrections. Presence stopped being differentiation the moment it became universal. Your absolute position looks similar. Your relative position is gone. A report that tracks only your own trend line will never show you that, and nobody grades on a curve except the market.
The fight moved from retrieval to synthesis, and the playbook only covers retrieval. Nearly all of that first-quarter work was about getting pulled in. Crawlable, structured, present. That layer is table stakes across your whole category now. What’s left is the harder half: what the model says about you once it has everybody’s documents in hand. That gets decided by third-party consensus, by recent real signal, by what your customers say in places you don’t own. Different problem, different levers. Most agencies never built the second half, because the first half is where the deliverables lived.
Vendor-authored content is worth less than it was. This is a pattern read rather than a documented mechanism, so weigh it accordingly, but I’d bet on it. When fifteen vendors all have equally well-optimized owned surfaces making equally confident claims, something has to break the tie, and it won’t be the marketing copy. It’ll be whatever reads as independent. Which makes the core deliverable of most retainers a depreciating asset. Same monthly fee, less bought every quarter.
Volatility gives everyone cover. Because output shifts run to run, a flat trend is indistinguishable from noise unless somebody is sampling properly and wants to know. Six months of nothing gets reported as holding steady in a turbulent environment. That framing isn’t always dishonest. It is always convenient, and the person delivering it may not have noticed which.
And the renewal needs a story. So the pitch becomes expansion. More prompts, more markets, more engines. Horizontal growth of a playbook that stopped compounding, sold to you as ambition. Nobody in that meeting is lying. They’re just never going to be the ones to tell you the engine stalled.
Hold all of that in your head and the list below stops being red flags. It becomes a test of whether your partner has a second act.
One thing before we start. Some of what looks like a bad agency is a bad client. That section is near the bottom and I’d read it before you fire anybody. I’ve watched teams torch a workable partnership because they wanted a quarterly result out of a job that takes three quarters, then go hire somebody worse who promised it faster.
1. They’re a content shop with a new name on the door
Read the scope of work before you read a single report. If the deliverables are twelve articles a month, a few landing pages and a schema audit, you didn’t buy AEO. You bought content marketing with a different acronym on the invoice.
The most expensive misunderstanding in this market, and I’ve watched it burn entire annual budgets: more owned content does not move the answer. Your website is a thin slice of what decides what gets said about you, and it’s the slice with the most reason to be discounted, because the system knows you’re the vendor.
Pull a real answer to a question your buyer would ask and look at where it’s sourcing from. Reddit threads. A review profile. A listicle from a media company you’ve never heard of. A YouTube transcript. A competitor’s comparison page where you’re filed under alternatives. A forum post from 2023. Somebody’s newsletter. Count how many of those you own. Usually one. Often zero.
Now open the last report they sent and count what’s being measured. Pages you published. Mentions on your domain. Your comparison page. Your glossary. They are measuring the one surface that isn’t deciding the outcome.
That mismatch is the business model, not an oversight. A shop built on publishing can only invoice for publishing. The work that moves the answer sits off your domain, in places you don’t control, and it’s slow and awkward and involves talking to people who don’t work for you. It doesn’t fit inside a monthly content calendar, so it never gets proposed.
A real plan has line items that make a content team visibly uncomfortable. Getting into the review corpus with enough recent, specific entries that a model has something usable to pull. Getting practitioners to write about you somewhere that isn’t your blog. Going after the comparison page your competitor owns instead of being dignified about it. Cleaning up the entity mess, meaning the six ways your company gets described across the web, the wrong founding year, the funding round nobody updated, the product you renamed that half the internet still calls the old name.
None of that is content. All of it moves the answer more than your next twelve posts, and that gap widens every quarter.
Ask: “What share of the sources cited across our category’s answers sit on a domain we own, and what’s the plan for the rest?”
If they can’t answer the first half, they have never looked at where the answer comes from. They’ve been optimizing in the dark and reporting on the light.
2. They talk in visibility and never in revenue
This one gets under my skin, probably because I carried a revenue number for a decade. It should get somebody fired, and nobody pushes on it, because the number is right there and it’s going up and it feels like progress.
Let me be careful here, because the lazy version of this critique is wrong. Attribution here is genuinely hard. Some assistants pass a referrer and some don’t, and none of them capture the bigger pattern, which is somebody reading an answer, closing the window, and typing your brand into Google two days later, where it lands in your reporting as branded organic. Anybody selling you deterministic attribution is lying with more confidence than the people selling you visibility.
But hard attribution is not a reason to stop at visibility. Visibility is a leading indicator, and leading indicators are only worth paying for if somebody can name the thing they lead to.
Here’s the chain. Every link is measurable, none are precise, and that’s fine.
Presence on prompts that have money behind them. Weighted, not total. Share on a question nobody with a budget asks is worth nothing, and an agency that doesn’t weight by buyer intent is grading itself on the easy questions.
Branded demand. Branded search and direct traffic are the shadow metric for this whole channel. If presence climbs three straight months and branded demand is dead flat, the thesis is wrong or the prompts are. That’s a finding, not an anomaly to bury.
Self-reported attribution. The crudest instrument available and the most honest one. One free-text question on the demo form, plus one asking whether they used an assistant while researching. Messy, underreported, and still the only place your buyer tells you directly.
What sales is hearing. Are buyers arriving pre-informed, fluent in the objection handling, already comparing you against two competitors you never sent them to? That moves before anything in the CRM does.
Win rate and cycle length on deals where an assistant was involved. Rough and directional beats precise and irrelevant.
Then the thing that binds all five together, and the reason most of these engagements can never prove anything either way.
You can’t attribute a deal. You can compare two windows.
Take a clean period before the work started and a matched period after, and read the whole chain across both. Then hold the obvious confounders still: seasonality, paid spend, a funding announcement, a launch, a competitor going quiet. If you changed four things and one number moved, you learned nothing and you should say so.
The stronger version runs a control. Compare your movement against the category’s, or against a set of prompts you deliberately left alone. Category-wide lift is weather. Lift against a flat control is work.
There’s a diagnostic hiding inside all of it. Pre and post requires a before. If nobody captured a baseline at kickoff, whether any of this worked is permanently unanswerable, by anyone, forever. Baseline capture is a day-one deliverable and I have almost never seen one in a scope of work. Go look for yours. If it isn’t there, that tells you what the first ninety days were really for.
Ask: “Show me the baseline you captured at kickoff and the same measures today, with whatever you controlled for. Then, if this works perfectly for twelve months, which number on our P&L moves, roughly how much, and when do I see the first evidence?”
You’re not asking for a guarantee. You’re asking whether they’ve thought past the dashboard. Somebody serious hands you a chain with a confidence level attached and names the links they can’t see. Somebody who hasn’t will tell you attribution is impossible and change the subject back to visibility. Attribution being hard is true. It’s also the most convenient sentence in this industry, and you should notice who reaches for it first.
3. They’ve never shown you a competitor
Go back through a year of reporting and count how many times you saw another vendor’s name attached to a number.
This is a share game. Presence stopped being the metric the moment your whole category became present, and you don’t get a date for when that happened. Nobody announces it. You find out by looking, which is the point. What matters now is the frequency table: across your twenty highest-intent prompts, which vendors get named, how often, and which direction each one is moving.
Without that, your own trend line can’t be read. Flat could mean you stalled. Flat could also mean you held position while two competitors surged past you, which is a much worse situation reported as the identical picture. Up could mean the work is working, or it could mean the category rose together and you paid a retainer for weather.
There’s a second reason this beats any audit. The competitor who keeps getting named on the prompts you lose is running an experiment on your behalf, in public, for free. Go look at what sits behind their answers. It’s almost never their website. It’s a review profile with ninety recent entries, or a founder who has spent two years answering questions in the same three communities, or one comparison page on a publication everyone in your space reads.
Ask: “Show me the vendor frequency table for our top twenty prompts, us against the four competitors we actually lose to, over the last six months. Then tell me what the winner has that we don’t.”
4. Their proof is a screenshot and a score you can’t reproduce
Run the same prompt five times in five fresh chats. You’ll get up to five different answers. Different sources, different order, sometimes a different set of vendors. Reword it slightly and it moves. Logged in versus logged out, it moves. Different day, model version or country, it moves.
So a report leading with one good screenshot isn’t evidence. It’s a lottery ticket somebody found and framed. And this is how these engagements survive a flat year: any agency can produce a favorable screenshot for almost any client on almost any prompt if they’re willing to run it enough times and only show you the winner. I’m not accusing anybody specifically. I’m telling you the tooling makes it trivial and you would never know.
The composite score has the same disease in a nicer suit. Everyone has an index now. Visibility Score, Answer Share, AI Presence. A number out of a hundred that went from 34 to 51. The slide from the top of this piece.
Three questions. What’s the formula. What’s the exact prompt set, and who chose those prompts. Can my team reproduce the number without your help.
If the prompt set is theirs and they won’t hand it over, the metric can’t be falsified. It can only go up, because the people choosing the inputs are the people being graded on the output. I’m not against composite metrics, I’ve built them. But a metric you can’t audit is a story, and you’re paying retainer for stories.
Ask: “What’s your sampling methodology: runs per prompt, models, versions, logged in or out? Send me the raw logs including the losses, and the prompt set behind the index.”
Somebody serious is quietly pleased you asked. Somebody who isn’t will explain why sampling doesn’t matter.
5. Nothing they do is a test
Everything ships everywhere at once. Twelve pages, a schema pass, a technical fix and an outreach push all land in the same six weeks, the number moves, and the report tells you the number moved.
Which of those four did it? Nobody knows. Not you, and here’s the uncomfortable part, not them either.
An agency that never isolates a variable has no idea which of its own tactics works. It has a bundle. Maybe two things in that bundle do all the lifting and the other eight are funded by them. A year in they still can’t tell you which is which, so they keep selling you all ten. That isn’t dishonesty. It’s what happens when nobody ever designed a test.
Real work has a shape you can recognize. A hypothesis written down before anybody started. A subset it gets applied to. Something held back as a control, even a crude one, even just a comparable set of prompts left untouched. A timebox. And a readout capable of returning the word no.
That last part is the whole thing. If no possible outcome would cause them to stop doing an activity, it was never a test. It was a deliverable with a story attached.
Ask: “Name one change you made last quarter, the hypothesis behind it, what you deliberately left alone as a comparison, and what it did. Then name one thing you’ve killed.”
6. They talk about “ranking”
There’s no ranking you can climb. There’s a retrieval step, where documents get pulled and scored against each other, and a synthesis step, where the model decides what to say about the ones it pulled. Two problems, two sets of levers, and your category already commoditized the first.
You can be retrieved and destroyed in the synthesis. Your page gets pulled, and the answer explains that you’re the expensive option for enterprises with a procurement department, sourced from your own pricing page. You were cited into a loss, and every citation-counting report on earth files that as a win.
The word is a tell, but the plan is the proof. Go through the scope of work and sort every line item into the two buckets. Access, structure, crawlability and publishing are all retrieval. Then look at what’s left. If the synthesis column is empty, that’s your gap, and it’s the expensive one, because retrieval is where your competitors already caught you.
The levers on that side are different and they’re slower. What independent sources say when they describe you. How recent that material is. Whether the specific objection that kills your deals has a credible public answer anywhere outside your own domain. Whether the descriptions of your product across the web agree with each other. None of that is a publishing task, which is why it rarely appears next to an owner and a due date.
Ask: “Sort your last two quarters of work into retrieval and synthesis for me. Then show me one case where we got retrieved and the answer still didn’t recommend us, and tell me what you changed.”
7. llms.txt was a deliverable
I want to be fair to the file before I use it as a murder weapon. It costs ten minutes, it’s harmless, it might matter later. Ship it.
But if it showed up as a line item with an owner and a due date and a place in the monthly report, you learned something about how that plan got built.
As of writing, no major assistant has confirmed llms.txt as a meaningful retrieval signal. It’s a proposal, not a standard. What it reliably is: easy to produce, easy to verify, easy to invoice. It photographs beautifully. You can put it on a slide with a green check next to it.
That’s a category, and once you can see it you’ll find it all over your scope of work. Call it the checkable artifact. Things chosen because they can be marked complete inside a billing cycle rather than because they move anything. An llms.txt file. FAQ schema bolted onto pages nobody reads. An “AI readiness audit” that turns out to be a slide deck. A summary paragraph added to two hundred old posts. Every one of those ships, screenshots and closes out. None of them are why you’re absent from the answer.
The reason this category dominates isn’t malice, it’s structural. Retainers need deliverables that complete on a monthly rhythm. The work that moves the answer doesn’t complete inside a month and doesn’t produce an artifact, so it loses to the file. Every time, in every agency, unless somebody deliberately stops it from losing.
Now the part that should make you put your coffee down.
An llms.txt file is a note you leave for a guest you’ve locked out of the building.
I’ve been on more than one call where a company was paying a five-figure monthly retainer to appear in a specific assistant while blocking that assistant’s crawler at the edge, and had been for months, and nobody on either side had checked. The file was shipped. It was in the report. The crawler meant to read it was getting a 403.
So when they tell you the file is done, go the other direction and ask about access at every layer, because there are more layers than people think and they’re owned by different teams. robots.txt is the one everybody checks. Then there’s your CDN’s bot rules, your WAF, and the bot-management product somebody in security bought two years ago without telling marketing. Any one of those can quietly refuse the exact crawlers you’re paying to court.
Then keep going down the fundamentals, in order of what they cost you.
Access. Which crawlers are allowed, at every layer, verified rather than assumed.
Rendering. If your content only exists after JavaScript executes, assume part of the retrieval layer never sees it. This one silently eats entire product sites.
Gating. Your best material, the implementation detail and the pricing logic and the real comparison, is usually behind a form, a login or a cookie wall. A model cannot cite what it cannot read, and your competitor with open docs is eating you alive on exactly the questions that close deals.
Plumbing. Status codes, redirect chains, whether the sitemap reflects reality, whether canonicals point somewhere sensible.
That’s the hierarchy. Access, rendering, gating, plumbing, and the optional files at the very bottom. Anybody working the last item before verifying the first is decorating.
Ask: “Walk me through every layer between the public internet and our content where a crawler could be blocked, and what you found at each one. Then tell me which of our highest-intent pages are gated.”
If they shipped you an llms.txt file and can’t answer that, you’re paying for the note, not the access.
8. Nobody is maintaining anything
Every plan I read is creation. Almost none have a maintenance line, and that gap gets more expensive every month.
Recency carries weight, and everything you own is quietly rotting. Your review profile is impressive and most of it is from 2023, which reads as a company people used to like. The roundup you fought to get into gets refreshed annually and you’re not in the new version. The comparison page that made you look good describes a pricing model you retired. Your docs cover a product two releases back. A third-party page cites you accurately for a feature you’ve since renamed, so now there are two names for the same thing and confidence in neither.
None of that requires anybody to do anything wrong. It happens through pure inertia and it eats a real share of whatever the first year bought you. Some meaningful part of your plateau is gains leaking out the back while somebody publishes new pages at the front.
The work is unglamorous and it’s mostly a list. Who owns review velocity as an ongoing number rather than a one-time push. Who re-pitches the roundups every cycle. Who keeps an inventory of external pages that describe you and gets the wrong ones corrected. Who checks that what the internet says you cost is what you actually cost.
Ask: “What’s on our refresh calendar for the next two quarters, and who owns getting third-party pages updated when the product changes?”
9. They count citations and never read the sentence
Being mentioned is not being recommended.
You can be cited as the expensive one. The one with the learning curve. The one that’s overkill for a team your buyer’s size. The one people migrate off. Every one of those is a citation and every one goes in the report as a win.
Almost nobody reads the sentence. No sentiment layer, no split between recommended and merely listed and actively warned about, no read on what role you play in the story. Just a count going up.
This matters more the deeper into the plateau you are. Once presence is universal in your category, how you get described is the only thing left that moves a deal.
Ask: “Of last month’s mentions, how many recommended us, how many just listed us, how many were negative? Show me the negative ones.”
10. Their prompt set has no buyer in it
If it’s “best [category] software” and forty variations, they’re tracking head terms because head terms are what they know how to track.
Nobody types “best CRM” in the week they’re about to spend eighty grand. They type the thing they’d be slightly embarrassed to ask a salesperson. Is this worth it for a team of twenty. Can it handle our compliance situation. Why do people leave this. What breaks at scale. Which of these two is easier to rip out in a year. What does implementation actually cost, not the sticker price.
Those prompts have a person inside them. They’re where the deal is won and lost, they’re harder to track, so they get skipped.
Every prompt set I’ve seen also stops at turn one, which is the more expensive half. Nobody asks one question and buys. They ask, they get three vendors back, then they push. Which of these is cheaper to run. Which will I regret in a year. What’s the catch with the first one you named.
You can be named in the opening answer and eliminated two turns later. The later turns are where the objection handling lives, where the model reaches past your surface copy for anything it can find about limitations, and where your competitor’s unhappy customers get their say. If your agency has only ever shown you opening answers, they’re reporting on the handshake and calling it the meeting.
One more thing belongs here. Nobody should be spreading effort evenly across five assistants. Your buyers concentrate somewhere. If the plan weights every engine the same, nobody asked where your pipeline comes from.
Ask: “Show me a full three-turn conversation for our two most valuable prompts, on the engine our buyers actually use.”
11. They offered to “handle Reddit”
Stop.
Community moves more of this than almost anything you can buy, which is exactly what makes it dangerous. There’s a bright line between being worth mentioning and manufacturing mentions, and the second one is a brand liability with a long fuse. Communities are very good at catching it. Moderators keep records. Screenshots outlive everybody.
You’re not buying visibility. You’re buying a headline about your company astroturfing, at a discount, on delay.
The legitimate version is slower and barely looks like marketing. Your employees participating as themselves over months. Your customers having something specific and good to say. Showing up where your category already talks and being useful there.
There’s no question to ask on this one. If an agency pitches you the fast version, you already have your answer, and it’s about what they think of your judgment.
12. Their answer to the plateau is more of the same, wider
This is the one you’ll meet in the renewal conversation.
You raise the flat line. They come back with expansion. Double the tracked prompts. Add three more engines. Extend into two more markets. Build a hub for an adjacent category. Every item is more of the exact thing that produced your step function in quarter one and nothing since.
Horizontal expansion of a plateaued playbook isn’t a strategy. It’s a renewal.
Here’s the shape of a real answer, and you can hear it in the first thirty seconds. A second act names something they’re going to stop doing. It moves budget from one column to another rather than adding a column. It includes at least one thing that can fail, and one thing that requires your team rather than theirs. If every item on the new plan is additive, nothing was learned in twelve months, because a year of real work always kills something.
Ask: “Our results were front-loaded. What’s structurally different about your next six months versus your first six?”
If the answer is a bigger version of the same list, you have your answer.
13. They’ve never retired a tactic
Ask which of their own tactics they’ve stopped using in the last year, and why. Then ask for the account where they did the work and the answers didn’t move.
Anybody with real reps has both and gives them up without much prompting, because it’s the most interesting thing they know. This field is a year and a half old. Every serious operator has watched something that worked in early 2025 quietly stop working and has an opinion about why. Nobody has a clean record, nobody has ten years of experience, and if somebody implies otherwise that’s its own answer.
An agency with no losses and nothing retired has no memory or no honesty. Both are expensive.
The test to run before your next call
Twenty minutes, no tools. Do it yourself, don’t delegate it, and don’t tell your agency you’re doing it.
Write ten prompts a real buyer would type in the week before they sign. Not category terms. The messy, specific ones they’d be slightly embarrassed to ask a salesperson.
Run each five times. Fresh chat every time. Logged out where you can. Note the model and the date.
For each answer, log three things. Were you mentioned. Were you recommended, merely listed, or warned about. And every source cited.
Tally the sources by domain and work out what share of the cited corpus you own.
Every time I’ve walked a team through this, the owned share comes back small. Usually much smaller than they expected. And the third-party sources are almost never the ones anybody has a plan for.
Then take the tally into your next call and ask three questions.
“What’s our plan for the sources we don’t own?”
“Which number on our P&L is this supposed to move, and when?”
“Our results were front-loaded. What’s structurally different about your next six months?”
The first tells you whether they understand the channel. The second, whether they understand the business. The third, whether they have anything left. You’ll know inside a minute on all three.
Before you fire them, check whether it’s you
I’ve watched this go the wrong direction enough times to put it in writing.
You gave a three-quarter job ninety days. Entity cleanup, review velocity, earned coverage, community presence. None of it lands in a quarter. If you signed expecting a Q1 result, the person who oversold you was in the room. So were you.
You didn’t hand over the keys. No engineering time. No access to product marketing. No permission to touch the review sites or ask customers for anything. No PR support. You hired a partner and gave them a blog. You’re going to get blog results.
You never told them which number mattered. If you never named the outcome, you can’t be furious that they optimized the one they could put on a slide. Visibility became the metric because nobody in the room asked for a better one.
You asked them to fix a reputation problem. This is the uncomfortable center of all of it. These systems are largely summarizing what real people already say about you. If your reviews say onboarding is a nightmare and the answers say onboarding is a nightmare, nothing is broken. No agency alive can talk a model out of a consensus your customers built. That’s a product conversation and a support conversation, and it sits above the agency’s pay grade.
What good actually looks like
Not credentials. There aren’t any yet. It comes down to one thing: will they show you their work?
The good ones hand over raw logs including the runs where you lost. They name the number they’re trying to move and tell you honestly which links in the chain they can’t see. They’ll tell you which of their own tactics they’ve stopped believing in and roughly when it stopped working. They have a written point of view on what they don’t know. They give you homework, because a real plan requires things only your team can do. They’ll tell you when a problem isn’t theirs to solve. And they’ll say “I don’t know, let’s test it” out loud, in front of you, without flinching.
The bad ones sell certainty. In a field this young, certainty is the loudest possible signal that somebody isn’t paying attention.
That list describes how I work, and I wrote it, so apply it to me too. That’s the reason to put a standard in public instead of keeping it in your head.
The easy half of this is finished. Every competitor you have is present, structured, crawlable and confident, and none of that is worth what it was worth last year. What’s left is the slow work of being the vendor that independent people actually recommend, in the places you don’t own, on the questions that cost money.
Some agencies were built for that. Most were built for the ninety days that are already over.
The founder on that Tuesday call told me they couldn’t tell whether they were doing the right work. That was the only thing they were wrong about.
You can tell. It takes twenty minutes.


