SearchEngines.Net logo — an independent reference on search enginesSearchEngines.NetWho runs which index

The shift to AI search

AEO and GEO explained

Two names for a young commercial field, and the gap between what its vendors sell and what the platforms document.

What the terms claim to mean

Answer Engine Optimization (AEO) and Generative Engine Optimization (GEO) are names for the same proposition: that appearing as a cited source inside an AI-generated answer is an outcome distinct from ranking in a list of search results, and that it can be pursued as its own discipline.

The same proposition also circulates as LLMO, AI SEO, LLM SEO, GAIO and "AI visibility". A field with a settled body of knowledge usually settles on a word for itself; five competing names in three years is a reasonable indicator of how young and how unsettled this one is.

It helps to separate two claims that are routinely bundled together, because one is straightforwardly true and the other is contested.

  • The descriptive claim: AI answers cite a different distribution of sources than the top ten links for the same query. This is plainly true and measurable. Some publications appear in generated answers far more often than their search positions would predict, and some pages that rank well are never cited at all.
  • The prescriptive claim: that specific, repeatable actions on a page reliably change whether it is chosen for citation. This is what the field sells, and it is the part that platform documentation, measurement difficulty and the absence of a published ranking model all bear on.

This page describes the field and what is knowable about it. It is a description, not an instruction manual; how search engines and answer engines work is the subject here, and the practice of trying to influence them is a different one.

Where the names came from

AEO is the older term and it predates generative AI. It circulated in marketing writing when "the answer" meant a featured snippet, a People Also Ask entry or a passage read aloud by a voice assistant. Those were extraction features: the engine lifted a specific passage from a specific page and displayed it more or less verbatim. The mapping from a page to an answer was direct, observable, and attributable — a site could see the snippet, see which page produced it, and see the referral in its logs. That is a materially easier object to reason about than a paragraph a language model composed from several sources.

GEO is the newer term and is generally traced to a 2023 academic preprint of that name, which proposed a benchmark for measuring how visible a source is inside a generated answer and reported that certain content changes moved that visibility. The paper is widely cited as the origin of the phrase. Its reported effect sizes are quoted heavily in vendor material and have not been verified against the primary document for this page; anyone relying on them should read the paper's own methodology rather than a summary of it, because what was measured, and on what system, is the whole question.

From roughly 2024 onward the terms were absorbed into commercial tooling — "AI visibility" dashboards, prompt-panel tracking, share-of-citation metrics — which is the point at which a research idea became a product category.

What the platforms actually say

Google's documentation on AI features in Search, last updated 10 December 2025, addresses the question directly and negatively:

"There are no additional requirements to appear in AI Overviews or AI Mode, nor other special optimizations necessary."

"You don't need to create new machine readable files, AI text files, or markup to appear in these features."

Google's stated position is that the material eligible to appear in AI Overviews and AI Mode is the material already indexed for Search, that the ordinary preview controls — nosnippet, data-nosnippet, max-snippet and noindex — are what limit what is shown, and that the existing guidance for Search applies unchanged. That is consistent with the architecture: AI Overviews are grounded on Google's ordinary index, crawled by the same Googlebot, rather than on a separate AI crawl. Google separately documents a robots.txt token, Google-Extended, which has no user-agent of its own and governs whether content is used to train and ground Gemini models in Google's other systems; Google states it does not affect inclusion in Search. See Google's AI features documentation for the current wording.

Two cautions about reading that as settled. Google is not a neutral narrator of its own system: it has an obvious interest in discouraging attempts to manipulate its output, and it has published versions of "make good content" for two decades while operating a ranking system with hundreds of undisclosed signals. Equally, a platform saying "nothing new is required" is a statement about requirements, not about outcomes — it does not claim that all indexed pages are equally likely to be cited, and no engine publishes how that selection is made.

The parts that are documented, and the parts that are not

There is a clean line between two questions that vendor material usually blurs. Whether a page can appear at all in a given product is documented, mechanical and checkable. Which of the eligible pages is chosen for a particular generated sentence is documented by nobody.

On the first question, the published record is unusually specific:

  • OpenAI separates its crawlers by purpose. OAI-SearchBot is the search-index crawler, and OpenAI states that sites opted out of it will not be shown in ChatGPT search answers. GPTBot is the model-training crawler and governs training rather than search visibility. ChatGPT-User fetches a page when a user asks the assistant to look at it, and OpenAI's documentation says robots.txt rules may not apply to it. OAI-AdsBot arrived with the advertising product. The two are confused constantly in both directions.
  • Perplexity documents PerplexityBot as its indexing crawler, and Perplexity-User as a user-initiated fetcher which, by Perplexity's own documentation, generally ignores robots.txt. Cloudflare published research in August 2025 alleging additional undeclared crawlers, which Perplexity disputed.
  • Microsoft Copilot has no crawler at all. Its answers are grounded on the Bing search service, so the controls that apply are Bingbot's. Microsoft also documents that Copilot writes its own search query from the user's prompt before sending it — so the text a page is matched against is not the text anyone typed.
  • Yahoo Scout, launched 27 January 2026, grounds on Microsoft Bing's grounding API with Anthropic's Claude as the model. Its citable universe is therefore Bing's index.

The consequence is that the structural questions have answers — which crawler governs which product, and whose index a given answer engine can draw from at all — while the selection question does not. No engine publishes a model of why one eligible source was cited and another was not.

llms.txt, specifically

llms.txt is a proposed convention: a plain-text or markdown file at a site's root offering a curated, simplified guide to the site intended for consumption by language models rather than by browsers. It has circulated since 2024 and is widely published, partly because it costs almost nothing to add.

What the public record shows as of 19 August 2026:

  • Google's AI features documentation states that new machine readable files, AI text files or markup are not needed to appear in AI Overviews or AI Mode. That is about as direct a comment on the format as a platform has made.
  • No published statement was located for this page, from Google, Microsoft, OpenAI, Perplexity or Anthropic, that llms.txt is read as an inclusion or ranking input to a search or answer product. Absence of a published statement is not proof that no system reads the file — it is simply the whole of the public record.
  • The convention has no standards-body status and no registered semantics comparable to robots.txt, which is a decades-old crawler-control convention that the major crawlers document themselves as honouring.

The honest summary is that llms.txt is a proposal with real adoption on the publishing side and no confirmed adoption on the consuming side. Cheapness explains its spread better than evidence does.

Why this field is unusually hard to measure

The methodological problems are not incidental; they are the reason the disagreement persists.

  • Answers are non-deterministic. The same question asked twice can produce different wording and different citations. There is no equivalent of "position 4" to record.
  • Different users get different indexes. Independent measurement of ChatGPT's retrieval stack published on 17 August 2026 found the free tier served largely from OpenAI's own index and the paid tier largely from purchased scraped Google results. A test run on one tier does not describe the other.
  • The model underneath changes. Gemini 3 became the default model behind AI Overviews globally on 27 January 2026. Any before-and-after measurement spanning that date is partly measuring a model swap.
  • There are no logs. An answer that satisfies the reader produces no visit, so the outcome is largely invisible in a site's own records. This is why measurement in this field is done with synthetic query panels rather than server data, and panel design determines the result.
  • Nobody publishes the selection model. Search ranking is undisclosed but at least produces a stable, ordered, inspectable artefact. Citation selection produces neither order nor stability.

None of this makes the field unserious. It does mean that a reported effect size is a property of a study design first and of the world second, and that the study design is the part worth reading.

The tension, stated plainly

Vendors of AEO and GEO services have a commercial interest in the discipline existing. Platforms have a commercial interest in it not existing, or at minimum in not documenting a lever that could be pulled. Neither party is a disinterested witness, and treating either one's account as settled is a mistake.

What is not in dispute is narrower and more interesting than either side's framing. Generated answers demonstrably cite a different set of sources than ranked results do. The citable universe of any given product is fixed by whose index it grounds on, which is a structural fact rather than an editorial one — Copilot and Yahoo Scout can only cite what Bing indexed, and ChatGPT's free and paid tiers do not draw from the same pool.

And for a meaningful share of the sources that appear most often, the route in runs through a contract rather than a technique. OpenAI's ChatGPT search launched with content licensing agreements naming the Associated Press, Axel Springer, Condé Nast, the Financial Times, Le Monde, News Corp, Reuters and The Atlantic among others. Yahoo joined Microsoft's Publisher Content Marketplace pilot when Scout launched in January 2026. Perplexity, having been sued by News Corp, Reddit and three Japanese newspaper publishers, runs a revenue-sharing programme with cited publishers. That commercial layer is the least-discussed part of the subject and arguably the most consequential one, because it is the part where the answer to "who gets cited" is decided in a negotiation rather than by a model.

Frequently asked questions

What is Answer Engine Optimization (AEO)?

AEO is a marketing term for the idea that appearing inside an answer — historically a featured snippet or voice readout, now an AI-generated summary — is a distinct outcome from ranking in a list. It predates generative AI. In its original sense the answer was an extracted passage from one identifiable page, which is a far more tractable object than a paragraph composed from several sources.

What is Generative Engine Optimization (GEO)?

GEO is the newer term for the same proposition applied to generative systems: influencing whether a source is cited inside an AI-generated answer. It is generally traced to a 2023 academic preprint of that name, which proposed a benchmark for measuring source visibility in generated answers. From 2024 onward it was absorbed into commercial tooling sold as AI visibility tracking.

Are AEO and GEO the same thing?

In practice they are used interchangeably, along with LLMO, LLM SEO, GAIO and AI visibility. The nominal distinction is that AEO covers answer surfaces generally, including pre-AI ones such as featured snippets, while GEO is specific to generative systems. Five competing names for one alleged discipline is itself a fair indication of how unsettled the field is.

What does Google say about optimizing for AI Overviews?

Google's AI features documentation, last updated 10 December 2025, states that there are no additional requirements to appear in AI Overviews or AI Mode and no special optimizations necessary, and that new machine readable files, AI text files or markup are not needed. It is consistent with the architecture, since AI Overviews are grounded on Google's ordinary Search index rather than a separate AI crawl.

Does llms.txt do anything?

No published statement was located, as of 19 August 2026, from Google, Microsoft, OpenAI, Perplexity or Anthropic confirming that llms.txt is read as an inclusion or ranking input. Google's documentation says AI text files are not needed. Unlike robots.txt, the convention has no standards status and no documented honouring by major crawlers. Absence of a statement is not proof of non-use, but it is the entire public record.

Does blocking GPTBot remove a site from ChatGPT's answers?

No, and this is the most common error in the area. OpenAI documents GPTBot as its model-training crawler and OAI-SearchBot as its search-index crawler, stating that sites opted out of OAI-SearchBot will not be shown in ChatGPT search answers. The two controls do different jobs, and publishers frequently get them backwards in both directions.

Can AI visibility be measured?

Only indirectly and imperfectly. Generated answers are non-deterministic, free and paying users of the same product can be served from different indexes, the underlying model changes without notice, and an answer that satisfies the reader leaves no referral in server logs. Measurement is therefore done with synthetic query panels, and the panel design largely determines what is reported.

Do AI answer engines rank pages the way search engines do?

Not in a form anyone outside can inspect. A search engine produces a stable ordered list that can be recorded and compared over time. An answer engine produces a variable paragraph in which a source is either cited or not, with no position and no published model of why one eligible source was chosen over another. The citable pool is also fixed by whose index the product grounds on.

Sources

Top