A few weeks ago Instinct started showing up in my workflow. Then Muse. And somewhere between ordering DoorDash, shopping for another car (being a collector is why I am still working), hearing about a friend use it to find a private chef for date night, and running half my task management through it, something clicked. I wasn't searching anymore. I was delegating.
When you search, you do the work. You evaluate results, click through, compare, come back, refine. When you delegate, the agent does that work and you decide whether to act on what it brings back. One conversation, from first question to recommendation. That's the window your published information has to survive now.

What these things actually do
Worth being specific here as "AI is changing search" has been said enough times that it stopped meaning anything.
Meta launched Muse on September 8. It runs on its own secure virtual machine with its own browser and works across the apps you connect to it, learning from conversations as it goes. Give it a goal and it builds a plan, then advances the work on its own: opening a browser, filling out forms, negotiating. For longer jobs it keeps running after you close the app and comes back when something changes or it needs your sign-off. Meta's own examples include selling a car for more and getting a bill lowered.
Instinct came out of Spear Street Technology, registered in April 2026 by Noah Shinn, who was previously a research scientist at Sierra. There's no app. You text it or you call it, and it uses a phone and a computer the way a person would. What set it apart in a crowded market was how far it pushes. Early users describe an assistant that books the restaurant, negotiates the bill, fills out the paperwork, and follows up without being asked twice.
Two things about that matter if you sell tech.
They don't stop. The evaluation isn't a session that dies when the tab closes. It runs, waits, and picks back up when something changes.
They act. Muse checks out through Link built by Stripe and is the first agent covered by Link's purchase protections. Instinct's founder says people buying through it average over $1,300 a month. That's his number, not an audited one, but money is already moving through these things.
Both are consumer products today. Instinct is explicitly a personal-life agent rather than an AI employee, and one review notes it's a bad fit for teams under SOC2 or GDPR. Nobody is running enterprise procurement through it this quarter.
I'd still pay attention. Buying habits form in consumer and walk into work. Somebody who spent six weeks letting an agent handle their bills, their travel, and their car hunt is going to want the same thing when a vendor evaluation lands on their desk. The behavior shows up before the enterprise-grade tooling does.
The handoff is disappearing
For most of the history of B2B software, marketing's job ended at the handoff. Generate interest, capture the hand raise, give it to sales. The buyer showed up with questions and a human answered them. That was the design. The information gap created the meeting.
Agents that research, evaluate, and negotiate collapse that whole sequence. The buyer reaches a conclusion before anyone on your team knows they exist.
AEO as most programs run it is built for the first half: get cited, show up in the answer, own the category question. That work matters and I'd keep doing it. But visibility is where the evaluation starts. The moment a buyer tells an agent to go deeper, you stop being described and start being evaluated. Getting cited can get you considered. Earning the recommendation means answering what comes next.
Where I'd expect this to break first
Start with tier fit. It behaves differently from every other evaluation question. It's binary. Either the feature works on the plan the buyer can afford or it doesn't. There's no way for an agent to split the difference, so it either gets a clean confirmation or it works with uncertainty.
Now think about where tier information actually lives. The pricing page is current because somebody owns it. The comparison page was built under a previous pricing model and nobody flagged it when things moved. The help docs still use a feature name from before it jumped a tier. The integration directory lists a connector and says nothing about which plan unlocks it. Four surfaces, four slightly different answers, nobody responsible for reconciling them. An agent that finds all four has to do something with that. It flags the discrepancy, hedges the recommendation, or keeps looking. Cleaner information could tip the call toward a competitor even when your product is the better fit.
The second place I'd look is implementation evidence. Not docs on how the product works. Evidence of what it takes to go live. Why one customer was running in three weeks and another took six months. What the fast ones had in common, which is usually something concrete like an existing data structure or a dedicated admin. What the slow ones needed from the customer side that nobody surfaced during the sales cycle. That knowledge sits with customer success and almost never gets published, because it's nobody's job to publish it. An agent goes looking and finds a help article and a case study with the hard parts scrubbed out. Neither answers what a serious buyer is asking, which isn't "can this work" but "what is this going to cost me in time and people."
The third one I haven't seen anybody connect yet. Agents are bad at walls. Early Instinct testers hit CAPTCHAs, two-factor prompts, and a high-demand purchase that crashed on them repeatedly. If your best evidence sits behind a form or a login, it doesn't exist during an agent-run evaluation. Every proof point you put behind lead capture is a proof point the agent can't reach at the exact moment it's deciding. That's a real trade-off now and most demand gen teams haven't done the math on it.
All three start inside the company, well before anyone sits down to write.
Your best answers never make it out of sales calls
Marketing has the overview. Product has the spec. Sales has learned the real explanation from giving it a hundred times. Customer success knows what happens after the contract is signed.
The rep who can answer the tier question in thirty seconds. The SE who knows exactly why that one implementation took six months. The CSM who could name the three conditions that make deployment fast. None of that lives anywhere a buyer can reach before the first call.
Fine when the buyer had to come to you for it. Not fine when they're forming an opinion before they'll take the meeting. The agent works with what's public, and if the thing that would move them only exists inside your building, it doesn't move them.
Three diagnoses, three owners
If your product does the thing but your documentation never explains it, that's an evidence problem and the fix is publishing. If the workflow only works on a plan they can't afford, that's a fit problem and more content won't touch it, though better tier docs will at least stop you burning cycles on deals that were never going to close. If they can afford it but your trial kills the feature they need to evaluate, the trial is the problem and writing about the feature is not a substitute for letting them use it.
None of those is a blog post. Inconsistent tier docs are product and marketing. Missing implementation evidence is customer success and marketing. A trial that can't validate the core use case is product.
This is where most AEO programs hit the ceiling. Running agents through real buying scenarios turns up things marketing can't fix alone, and if those findings can't reach the people who own the fix, you're producing content that closes nothing. Before you start, get agreement from whoever runs product and customer success that what you find lands on a roadmap instead of in a report nobody reads.
The negotiation question nobody has priced in
Here's the part I keep chewing on. Both of these agents negotiate. Meta puts lowering a bill on the marketing page. Instinct users describe it taking the vendor call. That's live today in consumer.
Run it forward. If a buyer's agent negotiates for them, your published pricing turns into an opening position instead of a fact. If agents are doing that on behalf of a lot of buyers in your category, the terms you've been quietly giving away become a pattern somebody outside your company can read. Discounting norms that survived on information asymmetry get harder to hold.
I don't know how that plays out and I'd be careful with anyone who says they do. But if your pricing model depends on buyers not comparing notes, you have less runway than you think.
Context makes this harder than one content strategy solves
"Find a CRM" means something completely different for a five-person team starting clean than for a company with custom integrations, years of history, and a contract expiring in six weeks. A connected agent may already have that context from the buyer's own systems. Or the buyer just tells it. Either way the evaluation gets specific fast. The small team cares about setup time and admin overhead. The six-week team needs migration evidence and needs it verified. Both could land on you, and what makes you the right answer is different for each.
You find out what matters by following real deals, not by guessing. Go through sales calls, implementation conversations, support tickets, win-loss interviews. Look for the moments where somebody said "okay but will it work if..." That's your question list. Pay attention to the answers that got things unstuck, especially the ones your best reps worked out on their own, because those are the least likely to exist anywhere else. Fifteen to twenty questions that keep coming back, not a catalog of everything anyone ever asked.
Then go check what somebody outside your company can actually find for each one. Is there a current answer? Does it explain the conditions? Does it hold up against what third parties are saying? Can it be verified without taking your word for it?
A real implementation story helps. So does a technical walkthrough, a practitioner review, or a comparison that doesn't soften the trade-offs. Your docs establish what a feature does. A customer story establishes what it took to make it work. Both need to be reachable without a form.
Brand gets you into the research. It doesn't close the evaluation.
Someone can like your company and walk anyway because the integration question never got a clean answer. Preference earns you a seat in the consideration set. It doesn't earn you the recommendation. Keep doing the work that builds preference. Just don't ask it to do a job it was never built for.
Measure the second half
Citation and visibility tracking tells you whether you're getting into the research and which sources are carrying weight. Keep it. It's the floor, not the program.
Then build evaluation tasks around buyers you actually understand, meaning ones you've closed or lost, not personas. Give the agent their systems, budget, timeline, and constraints, and run it far enough that a real requirement enters the conversation. Test whether it answers the tier question cleanly. Test whether it finds implementation evidence past a scrubbed case study. Test whether it hits the conflict between your pricing page and your comparison content, and watch what it does when it does. Test what happens when it runs into something gated.
Write down what you can see: sources it touched, claims it made, questions it resolved, questions it left open, and the exact point where it stalled. Run the same task again under the same conditions before you change anything, because one run tells you a problem exists and nothing about whether it's consistent.
The thing to watch for is whether the agent couldn't find the information or found it and had a reason to go somewhere else. Those look identical in a report and lead to completely different conversations internally. The first is a publishing gap. The second might be a real fit problem you'd rather know about now.
Here's what I'd do at your next AEO review
Pick one buying situation where you know you're the right fit. Run an agent through it, past the shortlist, until it stalls. Then go find the person who could have answered that question on a call without thinking about it. That's the gap. Not your citation count, not your share of voice. The distance between what your company knows and what a buyer can find.
Figure out why that answer hasn't made it out of the building.
StackedGTM update: Services has been ripping. We're running AEO for some of the best growth companies out there right now, and the pattern in this piece is the one we keep running into. Getting the answers out of the building is most of the work. That's Forward Deployed Marketing. https://www.stackedgtm.ai/services