How to Vet an AEO or GEO Agency Against Google’s Own Documentation

WordyPlus article cover: How to vet an AEO or GEO agency, with a timeline marking three dated changes to Google's guidance.
Three changes to Google’s own documentation gave buyers a public standard to hold AI search vendors to. Most vetting checklists still in circulation predate all three.

To vet an AEO or GEO agency, check their advice against Google’s own generative-AI guidance and ask them to show your citation baseline before you sign anything. Google’s Do you need an SEO? documentation, updated 5 June 2026, names AEO and GEO directly and tells buyers to verify vendor claims against official guidance.

  • Google named AEO and GEO in its hiring documentation in June 2026. You now have a neutral public standard instead of one agency’s opinion about another agency.
  • “Show me your client case studies” is a filter on how long an agency has existed, not on whether it can do the work.
  • The test that cannot be faked in a deck is asking for your own citation baseline before the contract.
  • An agency selling llms.txt files or AI-specific markup for Google surfaces is arguing against Google’s published position.
  • An agency quoting “it’s still SEO” as a reason to have no separate method for ChatGPT and Perplexity has made the opposite mistake.

Why can’t you tell a real AI search agency from a relabelled one?

Because there is nothing to check. No certification exists for answer engine optimisation or generative engine optimisation, no shared methodology exists, and a PR firm publishing its own screening list noted in June 2026 that a meaningful share of the agencies pitching it were doing conventional SEO eighteen months earlier (Proper Propaganda, June 2026). That is not disqualifying. Plenty of good practitioners came from SEO, and the craft overlaps heavily. It means the label tells you nothing, so every pitch sounds the same.

So you fall back on the filter everyone recommends: ask for client case studies with before-and-after citation data. Read six guides on choosing an AEO agency, and that is criterion one in all six.

It is a reasonable request and a bad first filter. It passes a large SEO shop that renamed a service line in January and had thirty existing accounts to point at. It fails a two-person specialist who started in March and is better at the work. You end up sorting on agency age. There is a better test, and since June 2026, there is a neutral standard behind it.

What did Google publish about AEO and GEO in 2026?

Google published two things. On 15 May 2026, it released its first official guide to optimising for generative AI features in Search, which lists tactics site owners can ignore. On 5 June 2026, it updated the hiring guide businesses use to vet an SEO, naming AEO and GEO by acronym for the first time and telling buyers to check vendor advice against that guidance.

Timeline of three dated changes to Google's guidance: 7 May, FAQ rich results retired; 15 May, generative AI optimisation guide published; 5 June, hiring guide names AEO and GEO.
The caveat line at the bottom is the part most vendors skip.

The May guide sits in Search Central documentation as Optimizing your website for generative AI features on Google Search, announced the same day on the Search Central blog. Its position is that generative AI features run on the same ranking and quality systems as the rest of Search, so SEO best practices still apply. Search Engine Journal read the mythbusting section as the consequential part because it tells site owners to skip tactics a growing AEO and GEO services industry has been selling.

Three items on that skip list are worth memorising before any vendor call. Google says you do not need llms.txt files, AI text files, special markup or Markdown to appear in Google Search, including its generative AI features, because Search does not use them. It says AI-specific rewriting and content chunking are unnecessary. And it warns that producing separate content for every variation of how people might search, including fan-out queries, primarily to influence rankings or generative AI responses, runs into its scaled content abuse spam policy. Query fan-out programmes are a live product in this market. Google has now put a spam-policy label near them.

The June update is the one that gives you leverage. Google’s Do you need an SEO? page now asks whether a practitioner’s AI-experience advice lines up with official guidance and whether they cite official Google documentation as supporting evidence. A separate page published alongside it states plainly that third-party tools have no access to Google’s internal ranking data and cannot guarantee performance.

One limit, and it cuts hard in the other direction. All of this covers Google’s own surfaces, AI Overviews and AI Mode. ChatGPT, Perplexity, and Claude retrieve and cite by their own logic, and Google’s documentation does not speak for them. An agency that quotes “Google says it’s still SEO” as a reason to have no distinct method for those engines has used a Google-scoped document to justify ignoring the engines Google does not run. That is the second failure mode, and it is at least as common as the first.

What is the one question that separates a real agency from a rebadge?

Ask them to show you your own citation baseline before you sign anything. Which prompts they would track, which engines, where you appear today, and where a named competitor appears instead. This tests measurement design, prompt selection, and honesty about blind spots at the same time, and none of it can be answered from a slide deck.

The reason this works is that the prompt set is the hard part, and it is the part nobody can fake in a meeting. Anyone can run your brand name through ChatGPT and screenshot the answer. Choosing the twenty prompts your actual buyer types on the way to a purchase decision requires understanding your category, your competitors, and where in the funnel the AI answer intervenes. Watch how they build that list. If they ask you good questions while building it, that is most of your answer.

A good response sounds concrete and slightly uncomfortable. Named prompts. Named engines. A date, because these answers change week to week. Some prompts where you already appear, some where a competitor owns the answer, and an admission of what they cannot see. A weaker response arrives as a proposal, or as a traffic dashboard, or as a rankings screenshot with the word “AI” in the title slide. I would treat a rankings screenshot as close to conclusive. It means the measurement layer does not exist yet.

Two honest caveats on my own test. It favours agencies with tooling budgets, since running a baseline across four engines for free costs them real time. And a baseline is cheap to produce and hard to interpret, so a vendor can hand you a competent-looking one and still have no plan behind it. It is a better first filter than case studies. It is not the whole interview.

What are the twelve questions to ask an AEO or GEO agency?

Twelve questions across three areas: how they work, how they measure, and what you are asked to sign. Each is written so a vague answer is itself informative. Take the whole list into the call.

Twelve agency vetting questions in three groups: Method, Measurement and Commercials – four questions each.
Ordered by how quickly a weak answer shows up. Method questions fail fastest.

Method

  1. Can you point me to the page in Google’s documentation that supports this recommendation? Google’s own hiring guide now asks buyers to check exactly this. A practitioner who works from published guidance can find the page in under a minute. Worrying answer: proprietary AI ranking factors that only they understand.
  2. Which of your tactics apply only to Google, and what is your separate method for ChatGPT and Perplexity? Google’s guidance is scoped to its own features. Worrying answer: treating all engines as one thing, in either direction. “It’s all the same” and “we have a secret LLM playbook” are both wrong, just in different ways.
  3. Do you sell llms.txt, AI-specific markup, or query fan-out content programs? Google lists the first two as unnecessary for its features and puts fan-out content near its scaled content abuse policy. Worrying answer: yes, with no evidence tied to a specific non-Google engine. A good answer names the engine and the reasoning.
  4. Do you still recommend FAQPage schema for rich results? Google retired FAQ rich results from Search on 7 May 2026, with Search Console reporting removed in June and API support in August (The HOTH, May 2026). The markup is still a valid schema and may still help machines parse a page. The SERP feature is gone. Worrying answer: pitching FAQ schema as a visibility win. It dates their knowledge to before May.
Measurement
  1. Will you show me my citation baseline before I sign? Covered above. Worrying answer: a proposal arrives instead.
  2. Who writes the prompt list? How many prompts are there, and how often is it re-run? The prompt set is the measurement instrument. If it drifts, every number after it drifts too. Worrying answer: “Our tool handles that.” Ask to see the list.
  3. What is the difference between a mention and a citation in your reporting, and which am I paying for? A mention means your name appeared somewhere in an answer. A citation means you were the source. Worrying answer: using the two words interchangeably, or reporting a count without saying which it counts.
  4. How does a citation become a record in my CRM? This is where most AI visibility programmes stop, and it is where the money is. Worrying answer: they report visibility and hand the attribution problem back to you.
Commercials
  1. What can you not control? Google states that third-party tools cannot see its internal ranking data and cannot guarantee performance. Any honest practitioner has a list. Worrying answer: “nothing”, or a guaranteed lift in citations.
  2. Who actually does the work? Name the people who will touch your content and your schema, not the people on the call. Worrying answer: reluctance, or an org chart instead of names.
  3. If I leave in month three, what do I keep? The prompt set, the tracking configuration, the content, the schema. Worrying answer: the agency owns the prompt set. That converts your measurement layer into a hostage.
  4. What would you have to see to tell me to stop paying you? The hardest question on the list and the most revealing. A practitioner who has run real programmes has watched one fail and can describe the signal. Worrying answer: no answer.

Which AI visibility metrics mean something and which are decoration?

Two metrics belong in a monthly report: citation share across a fixed prompt set and prompts where a named competitor wins and you do not. Everything else is context or theatre. The test is whether the number moving would change what you do next month.

MetricWhat it tells youWhat it hidesIn the report?
Citation share, fixed prompt setHow often you are the source on the questions your buyer asksNothing, if the prompt set is fixed and visibleYes
Competitor gapWhich prompts a rival owns and you do notWhether those prompts matter commerciallyYes
Mention countYour name appeared somewhereWhether it helped, hurt, or sat in a list of nine alternativesContext only
SentimentHow the model characterises youMoves on model updates you had nothing to do withContext only
AI referral sessionsPeople who clicked throughMost AI answers end without a click, so this undercounts badlyContext only
Keyword rankingsTraditional organic positionSays nothing directly about whether you get citedSeparate report
Vendor “visibility score”A composite the vendor definedThe formula, usually. Ask for the inputsNo
Assessment as of August 2026. Metric names vary between vendors; ask what each one counts.

Referral sessions get used dishonestly in both directions. Low AI referral traffic is not evidence that AI visibility does not matter, because the answer often resolves the question without a click. Rising referral traffic is not a growth number either, because the denominator is invisible. Weak signal. Keep the decision on citation share.

What should you refuse to sign?

Four contract terms are worth walking away over: a guaranteed number of citations, any claim of being Google-approved, a twelve-month lock-in agreed before a baseline exists, and agency ownership of the prompt set or the tracking configuration. The first two contradict Google’s published position. The second two are structural, and they are the ones that actually cost you.

Guarantees are the clearest signal. Google’s page on third-party tools says those tools cannot see its ranking data and cannot guarantee performance. A vendor guaranteeing citation counts either has not read the documentation or is counting on you not reading it. Google also warns against firms implying they are Google-approved, so treat badge language in a proposal as a flag, not a credential.

The lock-in is different. A twelve-month term is fair for this work, because entity and authority signals do not move in six weeks. Agreeing one before anybody has measured where you start is not. Sign the baseline, then sign the term.

Sample monthly report layout with four rows and blank values, marked as a layout example rather than a client result.
Values are blank because WordyPlus has no published client results yet. Ask any agency to show you theirs with the values filled in.

Where this framework is unfair, including to us

Every vetting checklist on this topic, including the questions above about what a vendor has watched fail, quietly rewards agencies that have been running a while. Experience is real, so that is not wrong. It does mean a buyer following the standard advice filters out competent new agencies and keeps incompetent old ones, and nobody selling you a checklist has a reason to mention it.

WordyPlus fails the standard first filter. We have no published client case studies with before-and-after citation data. Apply criterion one from any of the guides ranking for this query and we come off your shortlist. By that rule you would be right to remove us.

So here is what to demand from a newer agency instead, and hold us to it as hard as anyone else. Make them run the baseline before the contract, at their cost, so you are buying evidence of method rather than evidence of tenure. Start with a fixed-scope diagnostic instead of a retainer, so the downside is capped at one invoice. Read their published method and check it against Google’s documentation, because a public method is falsifiable in a way a case study is not. And get it in writing that you keep the prompt set, the tracking and the content if you leave in month two.

If a new agency will not accept all four, the newness was not the problem.

Get the baseline before you sign with anyone

The AI Visibility Audit runs the test this post tells you to demand. Twenty buyer prompts, five named competitors, your citation share across ChatGPT, Gemini, Perplexity and Google AI Overviews, plus a prioritised fix list. Fixed scope, roughly ten working days, yours to keep whoever you hire afterwards.

Still comparing approaches rather than vendors? The digital marketing line covers where AI search optimisation sits against paid ads and content and which one your stage actually calls for.

Frequently asked questions

Is GEO just SEO rebranded?

For Google’s surfaces, largely yes. Google’s May 2026 guide states its generative AI features run on the same core ranking and quality systems as Search, so the same work drives both. For ChatGPT, Perplexity and Claude, it is less settled because those engines cite their own logic, and Google’s documentation does not cover them. The foundations are shared. The measurement is not.

Do I need a separate AEO agency if I already have an SEO agency?

Usually not a separate agency. Usually a conversation with the one you have. Ask your current agency the twelve questions above. If they handle the method and measurement ones, a second vendor mostly adds coordination cost. If they cannot answer the measurement questions at all, that is a capability gap, and the honest fix is training them or replacing them rather than layering a specialist on top.

How long before I can tell whether it is working?

You can tell whether the measurement works in the first month, because either the baseline exists or it does not. Whether the work moves or citation share takes longer, anyone giving you a confident number for that is guessing. At 30 days you can hold an agency to a fixed prompt set, a documented starting position and a named first intervention. Judge the process early and the outcome later.

What should a first engagement cost?

Less than the retainer and scoped so that the deliverable stands alone if you never spend another rupee or dollar on them. A first engagement should be a fixed-price diagnostic, not a discounted month of an ongoing programme, because those two things create very different incentives about what the findings say. Our own pricing sits on the pricing page.

Share your love
Sakshi
Sakshi

Co-Founder & Marketing Head at WordyPlus. I write about marketing, SEO, business, digital growth, branding, and the ideas shaping how businesses connect with people in the digital world. With 4+ years of experience in SEO and digital marketing, I combine practical experience with a writer’s perspective to turn complex ideas into clear, useful, and engaging content.

Articles: 3

Newsletter Updates

Enter your email address below and subscribe to our newsletter

Leave a Reply

Your email address will not be published. Required fields are marked *