What GEO actually is, and why the platforms can't agree on a definition
GEO, short for Generative Engine Optimization, is the practice of making your content easy for AI engines to find, quote, and cite when they answer a question.
Those engines include ChatGPT, Microsoft Copilot, Claude, Perplexity, Google’s Gemini, and the AI answers now sitting on top of Google Search (AI Overviews and AI Mode).
This new field of search optimization feels overwhelming to SEOs, because it is. And the (not so) fun part is that the companies building these engines do not even agree on whether GEO is a real discipline.
Google’s official position is that optimizing for its AI features is ordinary SEO, and it explicitly tells you to ignore several tactics the industry sells as GEO essentials.
Microsoft’s official guidance for Bing and Copilot does almost the opposite, and walks you through chunking, parsing, and snippability.
So I wanted to dig deeper and start to write an independent guide on my own to understand whether GEO is even a real thing, and what we can actually do to optimize websites for AI.
I also tested the theory on a live site. Same brand, same prompts, same week, seven engines: 29.2% visibility on Google’s AI Mode, and a flat zero on Claude. Those field notes are at the end of this article.
This first part is about the actual definition and the differences between the platforms and their engines.
What do we know for sure about GEO?
That content crawlability is the first non-negotiable requirement.
Underneath the disagreements, all platforms describe the same mechanism: an AI answer is stitched together from a set of web pages the engine pulled in real time, on top of whatever the model already learned during training.
The pulling-in step runs on search infrastructure (a crawler, an index, a ranking pass), and that is why a decade of SEO habits still matter. If a page cannot be crawled, rendered, and judged useful, it cannot be cited, no matter which acronym you optimize for.
What has changed is the shape of the result. In classic search you compete for a position, and the user sees ten links. In AI search there is no fixed list: the engine reads across several sources, writes one answer, and decides who to mention.
You are now competing for a mention (your brand or product gets named) and a citation (your website gets linked), and the outcome is probabilistic. Two identical prompts can produce two different sets of sources.
Is GEO the same as AEO, GAIO, AIO, LLMO?
Yes. GEO, GAIO, AEO, AIO, LLMO are different acronyms that describe one practice. The differences are cosmetic.
Vendors and writers picked different words for the same shift in search, and none of them won.
Generative Engine Optimization (GEO) and Generative AI Optimization (GAIO) frame it around the generative model. Answer Engine Optimization (AEO) frames it around the answer box. AI Optimization (AIO) and Large Language Model Optimization (LLMO) frame it around the AI field in general, with different nouns.
When someone insists these are distinct disciplines, ask them for the distinct techniques. There usually aren’t any.
The platforms themselves add to the mess. Google refuses the new vocabulary and folds everything back into SEO. Microsoft uses GEO, AIO, and SEO almost interchangeably in the same post.
So just pick one term, define it once for your reader or stakeholders, and move on. I use GEO just because it’s short and easy to pronounce.
What Google says, and what Bing says
The two biggest search companies published official guidance within months of each other, and they point in different directions.
Google’s AI optimization guide (last updated July 10, 2026) is very direct: there is no special AI SEO. The same fundamentals that earn you rankings earn citations in AI features.
Google states that to be eligible for its generative AI features a page must be indexed and eligible to appear in Google Search with a snippet, and the site must also be included in Search generative AI features in Search Console.
Google names the two techniques behind its AI features:
- Retrieval-augmented generation, which Google calls grounding, means the model looks up fresh pages from the search index while it writes the answer.
- Query fan-out takes one prompt and quietly runs many related searches behind it, then merges what comes back.
Both are the reason SEO still works: the index feeding the AI is the search index.
Google’s guide also lists what it says does not help. Its mythbusting section names five things you can ignore for Google Search:
- llms.txt files and other special markup.
- Chunking content.
- Rewriting content just for AI systems.
- Seeking inauthentic mentions.
- Overfocusing on structured data.
Microsoft’s guidance for AI search answers tells a different story. Bing powers Copilot and Microsoft’s other AI surfaces, and Microsoft describes the retrieval step as parsing: the engine breaks a page into passages and assembles an answer from passages across several sources.
Everything Microsoft recommends follows from that:
- Align the title, H1, and meta description so the machine reads one consistent claim.
- Write descriptive H2s and H3s that work like chapter titles.
- Phrase key points as questions and answers.
- Use lists and tables.
- Add schema.
- Make each passage self-contained, so a chunk lifted out of context still makes sense on its own.
Microsoft goes further into formatting than Google ever does. It warns against:
- Walls of text.
- Hiding key information behind tabs or accordions that need a click to reveal.
- Burying facts inside images or PDFs where the parser cannot reach them.
- Overusing em dashes and decorative symbols (like emojis) that can confuse a machine trying to segment a sentence.
So who is right? Both, for their own engines.
Google can afford to say “just do SEO” probably because it owns the largest search index on earth and its AI features drink straight from it.
Microsoft’s guidance focuses on parsing because Copilot needs to extract and reuse passages out of your page to build its answer.
So the takeaway is: treat Google’s advice as the foundation and Bing’s as the ceiling.
Do the SEO Google demands, then add the structure Bing rewards. Nothing in Bing’s list hurts your Google performance anyway.
How AI answers actually get built
Every AI answer draws on two sources of knowledge: training data and real-time retrieval.
The first source is training data, the snapshot of text the model learned from before it shipped. When ChatGPT names the CEO of a well-known company without searching, that comes from training. It is broad and baked in, and stale the moment it is finalized.
You influence training data slowly, by being written about widely and consistently enough that the pattern gets absorbed. You cannot edit it on demand.
The second source is real-time retrieval. When a question needs fresh or specific information, the engine searches the live web, pulls back a set of pages, reads them, and writes an answer grounded in what it found. This is the part you can influence in a relatively short time, because it runs on crawling and indexing you already know how to influence.
Google calls this grounding: the answer gets anchored in pages pulled from the search index at answer time.
Microsoft’s parsing is one step inside that same loop, where the engine breaks each retrieved page into passages it can lift and reassemble.
Grounding is the whole process, parsing is how a page gets read once it arrives.
Retrieval rarely searches for your exact prompt. It uses query fan-out: one question becomes many.
“Plan a five-day trip to Japan in November” fans into dozens of smaller searches running in parallel, on neighborhoods, weather, rail passes, and so on, then the engine merges the results into one answer. Google confirms fan-out in its own documentation, so this is not a theory.
The practical consequence is that you rarely rank for the headline prompt (trip to Japan) but you get pulled in because you answered one of the invisible sub-questions well (best ryokans to visit in autumn).
There is a trap here: building a separate page for every possible query variation, fan-out queries included, in order to manipulate rankings or AI responses violates Google’s scaled content abuse spam policy. Google also argues it does not work, because its systems can judge a page’s relevance without an exact keyword match.
So, try to answer the sub-questions inside genuinely useful pages. Do not farm them.
Of course you still need consistency across sources, freshness, and existing search authority to raise the odds a page gets cited. But two identical prompts can surface two different source sets. That is how the engine is built.
What AI understands about your company
Page optimization moves retrieval. It decides whether a crawler can reach a URL, whether the passage is clean, whether the answer to a fan-out query is sitting there ready to be lifted. That work is real and it still matters.
But retrieval is only half of the job. The other half is what the engine already believes about your company before it searches anything.
That belief is assembled from the entire web: your site, yes, but also what other people write about you, where you appear, and whether the story stays consistent across all of it.
Google says this out loud. Its guide notes that generative AI features can surface what is being said about products and services across blogs, videos and forum discussions. The picture the engine has of your brand is built from everywhere you appear, and most of those places are not yours.
So the question worth asking is: what would an AI say about your company if it never visited your website?
That answer is your entity. And it is the thing you are actually optimizing.
How to work on the entity
Google says that seeking inauthentic mentions is a dead end.
Manufacturing brand name-drops across the web is a tactic that its ranking systems reward with nothing and its spam systems may punish. So this is not a link-building program with a new hat.
What holds up is consistency and evidence:
- Say the same thing everywhere. Same company name, same description, same founding facts, same claims, on every profile, directory, and platform where you appear. Contradictions give the model nothing to converge on.
- Claim the surfaces Google actually reads. Google’s own guide points to Google Business Profile and Merchant Center for local and product data. That is the company telling you where it looks.
- Make the entity machine-legible on your own site. Organization schema, a real About page, named authors with credentials. Cheap to do, and it removes any ambiguity about who you are.
- Earn the coverage. Being genuinely written about, by people with no incentive to flatter you, is the slow lever. It is also the only one that survives a spam update.
Three crawlers, three jobs
Every major platform now runs multiple crawlers. Treating them as one “AI bot” is the most common way sites accidentally delete themselves from AI answers.
- There is a training crawler, which collects content to improve future models.
- There is a search or indexing crawler, which builds the index that powers real-time answers and citations.
- There is a user-triggered fetcher, which grabs a specific page in the moment a person asks the assistant about it.
Each one is controlled separately, usually through its own robots.txt rule, and the settings are independent. Blocking the training crawler does not block the search crawler, and vice versa.
This matters because the three jobs have different stakes. Blocking the training crawler is a data-and-IP decision with a slow, diffuse effect. Blocking the search or user-facing crawler is a visibility decision with an immediate effect: it removes you from the answers people see today.
I wrote about what happens when you block GoogleOther, which is a good example of how easy it is to cut off the wrong crawler without meaning to.
User-triggered fetches are the exception to this model. Depending on the platform, robots.txt may not control them in the same way as automatic crawlers, although firewall and access rules can still affect whether requests succeed.
The table below maps the current names and functions, drawn from each platform’s own documentation.
| Platform | Training | Search / index | User-triggered fetch |
|---|---|---|---|
| OpenAI (ChatGPT) | GPTBot | OAI-SearchBot | ChatGPT-User |
| Anthropic (Claude) | ClaudeBot | Claude-SearchBot | Claude-User |
| Perplexity | (not used for training) | PerplexityBot | Perplexity-User |
| Google (Gemini) | Google-Extended token | Googlebot / search index | (served from index) |
| Microsoft (Copilot) | Bingbot family / index | Bingbot / search index | (served from index) |
Perplexity says outright that it does not crawl to train foundation models, so it has no training bot in the usual sense.
And Google does not expose a separate search-indexing crawler for its AI features, because those features reuse Googlebot’s existing search index.
The only extra control is Google-Extended, which governs training and grounding while Googlebot handles the indexing.
How each engine ranks and cites
The engines share mechanics but not preferences. Below is what each platform’s official documentation says about how it finds and cites content.
OpenAI: ChatGPT
ChatGPT sees your site through four separate agents, and only some of them affect visibility.
- OAI-SearchBot builds the index behind ChatGPT search; if you opt it out, OpenAI says your site will not be shown in ChatGPT search answers, though it may still appear as a plain navigational link.
- ChatGPT-User fetches a page live when a user or a custom GPT asks about it, and OpenAI notes this agent is not used to decide whether your content appears in search.
- GPTBot is the training crawler.
- OAI-AdsBot validates ad-related pages.
The settings are independent, robots.txt changes take roughly a day to register, and OpenAI publishes IP ranges as JSON for verification. The one you cannot afford to block by accident is OAI-SearchBot.
Microsoft: Copilot
Copilot runs on Bing’s index, so its citation behavior is Bing’s parsing behavior. Microsoft’s guidance rewards content that segments cleanly: consistent title and headings, question-and-answer phrasing, lists, tables, and self-contained passages that survive being lifted out of context.
If a page needs a click to reveal its content, or hides its facts in an image, Copilot’s parser tends to miss it. Optimizing for Bing and optimizing for Copilot are the same task.
Anthropic: Claude
Anthropic runs three bots, documented on its crawler page.
- ClaudeBot collects training data.
- Claude-SearchBot indexes content so Claude can surface and cite it in search-style answers.
- Claude-User fetches a page when someone asks Claude about it directly.
Each is controllable per user agent through robots.txt, and Anthropic publishes its crawler IP ranges for verification. To let Claude cite you, keep Claude-SearchBot and Claude-User unblocked.
Perplexity
Perplexity is the answer engine built around citations: every answer ships with numbered, clickable source links, which makes it particularly valuable for referral traffic when it cites you.
Its official crawler docs list two agents.
- PerplexityBot builds the index that powers those cited answers, and Perplexity states plainly that it is not used to train foundation models.
- Perplexity-User fetches a page on demand when a live user’s question calls for it.
Perplexity publishes IP ranges as JSON and recommends allowing its agents by both user agent and IP at the firewall level.
Google: Gemini, AI Overviews, and AI Mode
All three are Google, all three drink from the same search index… but all three cite differently.
Gemini, the standalone app, grounds its answers in Google’s search index. The only publisher control specific to it is the Google-Extended token, which decides whether your content can be used to train future Gemini models and for grounding in Gemini and Vertex AI.
Google-Extended is a control token, so it does no crawling of its own: Google’s existing crawlers do the fetching, and the robots.txt token only tells them what they may use the content for. It is not a ranking signal and does not affect Google Search inclusion.
AI Overviews are part of Google Search and run on the live search index, using retrieval-augmented generation and query fan-out.
Also, blocking Google-Extended does not remove you from AI Overviews. AI Overviews run on the live search index, and Google-Extended only touches training and grounding data.
To disappear from AI Overviews you would have to block Googlebot itself and lose Google Search entirely, which almost nobody wants.
AI Mode is Google’s fuller conversational search surface with heavier query fan-out. Even though it is a Google product built on the same index, it does not cite like AI Overviews.
Ahrefs’ December 2025 analysis put the citation overlap between AI Overviews and AI Mode at 13.7%, despite the two producing answers that were about 86% semantically similar. Same company, same index, similar words, different sources (that figure is Ahrefs’ own measurement; Google has not published one).
The takeaway across Google’s three surfaces is that there is no single “rank in Google’s AI” lever.
Solid SEO makes you eligible everywhere, but which surface actually cites you varies, and you cannot assume winning one wins the others.
What every official doc agrees on
If I ignore the disagreements, there is still a short list of shared requirements. If you do nothing else, do these five:
- Be crawlable by the right agent. A page the search or user crawler cannot reach cannot be cited. Check your robots.txt and firewall rules.
- Be indexed and eligible. Google requires a page to be indexed, snippet-eligible, and included in Search before it can appear in AI features at all.
- Write non-commodity content. This is Google’s strongest claim in the entire guide: unique, compelling, useful content will influence your presence in generative AI search more than anything else it recommends. Google’s own example contrasts commodity content like “7 Tips for First-Time Homebuyers” with a first-hand piece like “Why We Waived the Inspection and Saved Money.” First-hand experience is the moat.
- Answer the question fully and accurately. These engines favor complete sources over partial ones.
- Keep the SEO foundation intact. Retrieval reuses search infrastructure, and authority still carries weight.
None of this is new. It is the same discipline pointed at a new kind of result.
What about JavaScript?
This is one of the most interesting points I found.
Google can process content inside JavaScript as long as the script is not blocked, and Google explicitly says perfectly semantic HTML is not a requirement. So the “AI can’t read JavaScript” claim is inaccurate for Google.
But it holds up elsewhere. The crawlers behind ChatGPT, Claude, and Perplexity are documented as fetching and reading page markup, with no equivalent public commitment to rendering JavaScript the way Googlebot does.
Server-rendered HTML is the safe default across all seven engines, and if you serve a JavaScript-heavy site you may be visible to Google and invisible to everyone else.
The “GEO tactics” platforms can’t agree on
Three tactics are in open dispute between the official sources.
Here is where each platform actually stands, so you can decide for your own site on the evidence.
- llms.txt. Google says outright that it does not use llms.txt. No other major platform has committed to reading it in official documentation either. It costs little to publish, but do not expect Google’s systems to consult it.
- Content chunking. Google says you do not need to pre-chunk your content for its models. Microsoft’s whole approach assumes chunking and rewards content that is already broken into clean, self-contained passages. Writing in well-segmented sections satisfies Bing without hurting Google, so the safe move is to structure clearly and let Google ignore the parts it does not use.
- Structured data. Google says schema is not required to appear in its AI features, while remaining useful for rich results in classic search. Microsoft actively recommends schema for AI answers. Keep your structured data. It helps on Bing’s side and costs you nothing on Google’s.
Measuring GEO
Google’s guide now points to the Generative AI performance report in Search Console, which shows how your content is performing in generative AI features across Google Search and Discover. I do not have it enabled on my own sites yet, so I cannot report on how useful it is in practice.
Google also cautions readers to be wary of third-party tools that promise ranking success or claim to use internal Google metrics, and states flatly that no third-party tool has access to its internal ranking or AI systems. Use them if they help your workflow, Google says, but check their advice against the official guidance.
The visibility dashboards, share-of-voice scores and citation trackers you find on the market can’t see inside Google (just like SEO tools can’t). They are all sampling outputs from the outside and inferring. That does not make them useless, but you should treat every number they hand you as an estimate.
Preparing for agentic experiences
AI agents are autonomous systems that act on someone’s behalf: booking a table, comparing product specifications, completing a task.
Browser agents reach your site and gather what they need by analyzing visual renderings like screenshots, inspecting the DOM structure, and interpreting the accessibility tree.
Here is the ironic part. Semantic HTML is the thing Google says you do not strictly need for search, and it is exactly what feeds the accessibility tree an agent reads. The markup we have all been deprioritizing is the markup agents depend on.
Accessibility work and agent-readiness turn out to be the same job.
Google points to agent-friendly website best practices for anyone who wants to prepare, and notes that protocols like the Universal Commerce Protocol are emerging to let Search agents do more.
This is early, and Google frames it as something to explore if you have spare time. I am flagging it because the direction of travel is clear: the next thing reading your site will not be a crawler building an index but a machine trying to complete a task.
What I’m seeing on my own site
Everything above comes from official documentation. This part comes from one live project, and it is purely qualitative. Treat it as a field observation.
For a quantitative study, go to Ahrefs’ study of 730,000 query pairs.
The site I am monitoring is zenacademy.it, a meditation courses website that also publishes free guides and articles, in Italian. It has ranked well in traditional search for years, which is exactly why it makes a useful test bed: the SEO foundation was already solid, so anything I build on top could be judged against a stable baseline.
I started by monitoring 11 prompts the target audience would plausibly type. They are in Italian, and they are variations on one question: where to find meditation courses.
On this heatmap I tracked mentions, meaning an LLM explicitly naming the brand or the product inside its answer.

The spread across engines is telling, considering it is the same entity, the same prompts, the same week:
| Engine | Visibility |
|---|---|
| AI Mode | 29.2% |
| AI Overview | 25.1% |
| Copilot | 24.1% |
| Gemini | 15.9% |
| ChatGPT | 4.4% |
| Perplexity | 4.4% |
| Claude | 0% |
Google’s AI Mode and AI Overviews lead in visibility. Copilot follows close behind.
This is not surprising considering the good SEO baseline (Google and Bing).
On the other hand, the engines with their own retrieval stacks (ChatGPT and Perplexity) barely register any visibility. And Claude never mentions the brand at all.
So my SEO foundation carried straight into the engines built on search infrastructure, and did close to nothing for the rest.
The second thing worth looking at is Gemini at 15.9%, roughly half the AI Mode figure, on the same entity. Three Google surfaces show three different numbers despite coming from the same platform.
Ahrefs measured 13.7% citation overlap between AI Overviews and AI Mode across 540,000 query pairs. This is what that divergence looks like from the inside of a single site.
Being mentioned and being cited are two different events
The heatmap above averages everything. If we zoom into a single one of those 11 prompts we can see an interesting spread across LLMs.

On this prompt, two engines named the brand: Copilot and AI Mode. The other five wrote their answer without it.
When it comes to citations, three engines linked the site: Copilot at 1 of 7 sources, AI Mode at 1 of 21, and Perplexity at 1 of 10.
That is also why Perplexity can sit at 4.4% in the heatmap while linking the site here: the heatmap counts mentions (names in the answer), and a link is not a mention but a citation.
Meanwhile AI Overview had ten citation slots to fill and gave the site none of them, despite scoring 25.1% visibility across the full 11 prompts.
Again, to recap: a mention builds brand recall with no click, a citation offers a click with no recall. That’s why it’s important to track both metrics.
Any dashboard reporting one blended “visibility” number is flattening all of this into a single figure that describes neither outcome.
And remember what Google says about mentions: chasing inauthentic ones is a dead end, because the ranking systems reward quality while other systems block spam.
So the mention count is worth watching as a signal, and worth nothing as a target. The moment you start manufacturing mentions to move that number, you are optimizing for the metric instead of the thing the metric was measuring.
How do I explain GEO to my colleagues and stakeholders who are not SEO experts?
The simplest way I could find is: “For years we competed for a spot on a list of ten links. Now the answer arrives pre-written, and we compete to be one of the sources it gets written from.”
The tough conversation is the one about KPIs, because we can’t rely on impressions and clicks anymore.
Three different metrics must be tracked to assess the impact of GEO tactics:
- Mentions (the brand or product name is mentioned in the answer). Brand recall, no click.
- Citations (the website is linked). Now a click is possible.
- Referral traffic from AI engines. This is when someone actually clicked on a citation, and what your analytics sees. I recommend creating a custom channel grouping for AI traffic.
All three move independently, and you have just seen it: one engine linked the site without ever naming it, and another named the brand across the prompt set while handing it zero links on the prompt I zoomed into.
If you report them as three separate lines, you will spend less time explaining why traffic did not follow visibility.
”Can you promise us a ranking?”
No, and say so early. There is no position 1 in an AI answer. Two identical prompts can return two different sets of sources.
What you can promise is eligibility and better odds.
”The tool says we’re at 24% visibility. Is that real?”
Treat it as an estimate. Google states plainly that no third-party tool can see inside its ranking or AI systems.
Every dashboard on the market samples from the outside and infers.
”So what do we do differently on Monday?”
Less than they fear:
- Be crawlable by the right bots, because blocking the wrong one deletes you from the answers.
- Write the things only you can write. Expert content and first-hand experiences are irreplaceable.
- Structure pages so a machine can cite a clean passage out of them.
What comes in part 2
Now that I have a baseline for mentions and citations, I want to find out what GEO tactics actually move the needle.
I will work through each of the “GEO best practices” on the live site, and then run deliberate experiments: change one variable at a time, watch the seven engines, and record what moves.
The goal is a list of what actually works and what actually does nothing.
The official docs disagree with each other. Part 2 is where I will stop reading documentation and start breaking things on purpose :)
Stay tuned!
Sources
- Google Search Central, AI optimization guide (updated July 10, 2026)
- Google Search Central, Google-Extended and common crawlers
- Microsoft Advertising, Optimizing your content for inclusion in AI search answers (October 2025)
- OpenAI, Overview of OpenAI crawlers
- Anthropic, Does Anthropic crawl data from the web
- Perplexity, Perplexity crawlers
- Ahrefs, AI Overviews vs AI Mode: citation overlap study (December 2025)