SearchEngines.Net logo — an independent reference on search enginesSearchEngines.NetWho runs which index

The shift to AI search

AEO and GEO explained

Two names for a young commercial field, and the gap between what its vendors sell and what the platforms document.

What the terms claim to mean

Answer Engine Optimization (AEO) and Generative Engine Optimization (GEO) are names for the same proposition: that appearing as a cited source inside an AI-generated answer is an outcome distinct from ranking in a list of search results, and that it can be pursued as its own discipline.

The same proposition also circulates as LLMO, AI SEO, LLM SEO, GAIO and "AI visibility". A field with a settled body of knowledge usually settles on a word for itself; five competing names in three years is a reasonable indicator of how young and how unsettled this one is.

It helps to separate two claims that are routinely bundled together, because one is straightforwardly true and the other is contested.

  • The descriptive claim: AI answers cite a different distribution of sources than the top ten links for the same query. This is plainly true and measurable. Some publications appear in generated answers far more often than their search positions would predict, and some pages that rank well are never cited at all.
  • The prescriptive claim: that specific, repeatable actions on a page reliably change whether it is chosen for citation. This is what the field sells, and it is the part that platform documentation, measurement difficulty and the absence of a published ranking model all bear on.

The descriptive claim has now been measured. Researchers at the New Jersey Institute of Technology and Rutgers University, in work presented at SIGIR '26, compared the sources cited in a Google AI Overview against the traditional search results for the same query across a benchmark of 14,212 queries. The overlap was 0.18 Jaccard similarity — a set-similarity measure running from 0 for two sets with nothing in common to 1 for two identical sets — and 0.23 rank-biased overlap, a variant that weights agreement near the top of each list more heavily. The two source sets differ in composition as well as in membership: Reddit appears 27.88 percent less often in AI Overviews than in the traditional results for the same queries, and YouTube 17.09 percent more often. Ranking on the first page and being cited in the Overview above it are, on that evidence, substantially different outcomes drawn from the same index.

This page describes the field and what is knowable about it. It is a description, not an instruction manual; how search engines and answer engines work is the subject here, and the practice of trying to influence them is a different one.

Where the names came from

AEO is the older term and it predates generative AI. It circulated in marketing writing when "the answer" meant a featured snippet, a People Also Ask entry or a passage read aloud by a voice assistant. Those were extraction features: the engine lifted a specific passage from a specific page and displayed it more or less verbatim. The mapping from a page to an answer was direct, observable, and attributable — a site could see the snippet, see which page produced it, and see the referral in its logs. That is a materially easier object to reason about than a paragraph a language model composed from several sources.

GEO is the newer term and is generally traced to a 2023 academic preprint of that name, which proposed a benchmark for measuring how visible a source is inside a generated answer and reported that certain content changes moved that visibility. The paper is widely cited as the origin of the phrase. Its reported effect sizes are quoted heavily in vendor material and have not been verified against the primary document for this page; anyone relying on them should read the paper's own methodology rather than a summary of it, because what was measured, and on what system, is the whole question.

From roughly 2024 onward the terms were absorbed into commercial tooling — "AI visibility" dashboards, prompt-panel tracking, share-of-citation metrics — which is the point at which a research idea became a product category.

What the platforms actually say

Google's documentation on AI features in Search, last updated 10 December 2025, addresses the question directly and negatively:

"There are no additional requirements to appear in AI Overviews or AI Mode, nor other special optimizations necessary."

"You don't need to create new machine readable files, AI text files, or markup to appear in these features."

Google's stated position is that the material eligible to appear in AI Overviews and AI Mode is the material already indexed for Search, that the ordinary preview controls — nosnippet, data-nosnippet, max-snippet and noindex — are what limit what is shown, and that the existing guidance for Search applies unchanged. That is consistent with the architecture: AI Overviews are grounded on Google's ordinary index, crawled by the same Googlebot, rather than on a separate AI crawl. Google separately documents a robots.txt token, Google-Extended, which has no user-agent of its own and governs whether content is used to train and ground Gemini models in Google's other systems; Google states it does not affect inclusion in Search. See Google's AI features documentation for the current wording.

Two cautions about reading that as settled. Google is not a neutral narrator of its own system: it has an obvious interest in discouraging attempts to manipulate its output, and it has published versions of "make good content" for two decades while operating a ranking system with hundreds of undisclosed signals. Equally, a platform saying "nothing new is required" is a statement about requirements, not about outcomes — it does not claim that all indexed pages are equally likely to be cited, and no engine publishes how that selection is made.

The parts that are documented, and the parts that are not

There is a clean line between two questions that vendor material usually blurs. Whether a page can appear at all in a given product is documented, mechanical and checkable. Which of the eligible pages is chosen for a particular generated sentence is documented by nobody.

On the first question, the published record is unusually specific:

  • OpenAI separates its crawlers by purpose. OAI-SearchBot is the search-index crawler, and OpenAI states that sites opted out of it will not be shown in ChatGPT search answers. GPTBot is the model-training crawler and governs training rather than search visibility. ChatGPT-User fetches a page when a user asks the assistant to look at it, and OpenAI's documentation says robots.txt rules may not apply to it. OAI-AdsBot arrived with the advertising product. The two are confused constantly in both directions.
  • Perplexity documents PerplexityBot as its indexing crawler, and Perplexity-User as a user-initiated fetcher which, by Perplexity's own documentation, generally ignores robots.txt. Cloudflare published research in August 2025 alleging additional undeclared crawlers, which Perplexity disputed.
  • Microsoft Copilot has no crawler at all. Its answers are grounded on the Bing search service, so the controls that apply are Bingbot's. Microsoft also documents that Copilot writes its own search query from the user's prompt before sending it — so the text a page is matched against is not the text anyone typed.
  • Yahoo Scout, launched 27 January 2026, grounds on Microsoft Bing's grounding API with Anthropic's Claude as the model. Its citable universe is therefore Bing's index.

The consequence is that the structural questions have answers — which crawler governs which product, and whose index a given answer engine can draw from at all — while the selection question does not. No engine publishes a model of why one eligible source was cited and another was not.

llms.txt, specifically

llms.txt is a proposed convention: a plain-text or markdown file at a site's root offering a curated, simplified guide to the site intended for consumption by language models rather than by browsers. It has circulated since 2024 and is widely published, partly because it costs almost nothing to add.

What the public record shows as of 19 August 2026:

  • Google's AI features documentation states that new machine readable files, AI text files or markup are not needed to appear in AI Overviews or AI Mode. That is about as direct a comment on the format as a platform has made.
  • No published statement was located for this page, from Google, Microsoft, OpenAI, Perplexity or Anthropic, that llms.txt is read as an inclusion or ranking input to a search or answer product. Absence of a published statement is not proof that no system reads the file — it is simply the whole of the public record.
  • The convention has no standards-body status and no registered semantics comparable to robots.txt, which is a decades-old crawler-control convention that the major crawlers document themselves as honouring.

The honest summary is that llms.txt is a proposal with real adoption on the publishing side and no confirmed adoption on the consuming side. Cheapness explains its spread better than evidence does.

Why this field is unusually hard to measure

The methodological problems are not incidental; they are the reason the disagreement persists.

  • Answers are non-deterministic. The same question asked twice can produce different wording and different citations. There is no equivalent of "position 4" to record.
  • Different users get different indexes. Independent measurement of ChatGPT's retrieval stack published on 17 August 2026 found the free tier served largely from OpenAI's own index and the paid tier largely from purchased scraped Google results. A test run on one tier does not describe the other.
  • The model underneath changes. Gemini 3 became the default model behind AI Overviews globally on 27 January 2026. Any before-and-after measurement spanning that date is partly measuring a model swap.
  • There are no logs. An answer that satisfies the reader produces no visit, so the outcome is largely invisible in a site's own records. This is why measurement in this field is done with synthetic query panels rather than server data, and panel design determines the result.
  • Nobody publishes the selection model. Search ranking is undisclosed but at least produces a stable, ordered, inspectable artifact. Citation selection produces neither order nor stability.

None of this makes the field unserious. It does mean that a reported effect size is a property of a study design first and of the world second, and that the study design is the part worth reading.

What the research literature actually shows

Two 2026 arXiv preprints, read together, are the most honest available account of the evidence behind this field. Neither is peer-reviewed work, and both should be read as such.

The first is an experiment. Researchers at the University of Tokyo built a framework they call GEO-SFE, which isolates a document's structural features from its semantic content, and evaluated restructured documents across six mainstream generative engines. Restructuring on those structural features improved citation rates by 17.3 percent and subjective quality ratings by 18.5 percent. The design is what makes the number worth reporting: by holding meaning constant and varying only structure, the study shows that structure alone moves citation behavior, which turns the question into an experimentally testable one rather than a matter of assertion. The lift was measured in a controlled setting across six engines at one point in time, and its durability across engine updates is untested.

The second is a survey of the whole literature. A critical survey of 45 Generative Engine Optimization studies published between November 2023 and July 2026 concluded that no reviewed technique demonstrates a stable, longitudinal, cross-platform causal effect on organic discoverability or on downstream user behavior. Effects turn up in individual studies; they do not survive being carried across engines, across time, and into measured changes in what people actually do. The survey names topical relevance and context position as the two most reproducible levers among those it examined. It is the work of a single author, and its skepticism is one well-argued reading of the literature rather than a settled consensus.

The two findings are not in conflict, and holding both at once is the accurate position. A controlled experiment can find a real effect on six engines in one month while a survey of forty-five studies finds that no such effect has been shown to persist. In a field where the engines are updated without notice, where free and paying users can be served from different indexes, and where the measuring instrument is a synthetic query panel, a 17.3 percent citation lift and the absence of any durable cross-platform result are two readings of the same evidence base at different distances. The reported effect is real in its setting; the claim that it generalizes has not been demonstrated by anyone.

The tension, stated plainly

Vendors of AEO and GEO services have a commercial interest in the discipline existing. Platforms have a commercial interest in it not existing, or at minimum in not documenting a lever that could be pulled. Neither party is a disinterested witness, and treating either one's account as settled is a mistake.

What is not in dispute is narrower and more interesting than either side's framing. Generated answers demonstrably cite a different set of sources than ranked results do, and the size of that gap has been measured at 0.18 Jaccard similarity in the SIGIR '26 study described above. The citable universe of any given product is fixed by whose index it grounds on, which is a structural fact rather than an editorial one — Copilot and Yahoo Scout can only cite what Bing indexed, and ChatGPT's free and paid tiers do not draw from the same pool.

And for a meaningful share of the sources that appear most often, the route in runs through a contract rather than a technique. OpenAI's ChatGPT search launched with content licensing agreements naming the Associated Press, Axel Springer, Condé Nast, the Financial Times, Le Monde, News Corp, Reuters and The Atlantic among others. Yahoo joined Microsoft's Publisher Content Marketplace pilot when Scout launched in January 2026. Perplexity, having been sued by News Corp, Reddit and three Japanese newspaper publishers, runs a revenue-sharing program with cited publishers. That commercial layer is the least-discussed part of the subject and arguably the most consequential one, because it is the part where the answer to "who gets cited" is decided in a negotiation rather than by a model.

Frequently asked questions

What is Answer Engine Optimization (AEO)?

AEO is a marketing term for the idea that appearing inside an answer — historically a featured snippet or voice readout, now an AI-generated summary — is a distinct outcome from ranking in a list. It predates generative AI. In its original sense the answer was an extracted passage from one identifiable page, which is a far more tractable object than a paragraph composed from several sources.

What is Generative Engine Optimization (GEO)?

GEO is the newer term for the same proposition applied to generative systems: influencing whether a source is cited inside an AI-generated answer. It is generally traced to a 2023 academic preprint of that name, which proposed a benchmark for measuring source visibility in generated answers. From 2024 onward it was absorbed into commercial tooling sold as AI visibility tracking.

Are AEO and GEO the same thing?

In practice they are used interchangeably, along with LLMO, LLM SEO, GAIO and AI visibility. The nominal distinction is that AEO covers answer surfaces generally, including pre-AI ones such as featured snippets, while GEO is specific to generative systems. Five competing names for one alleged discipline is itself a fair indication of how unsettled the field is.

What does Google say about optimizing for AI Overviews?

Google's AI features documentation, last updated 10 December 2025, states that there are no additional requirements to appear in AI Overviews or AI Mode and no special optimizations necessary, and that new machine readable files, AI text files or markup are not needed. It is consistent with the architecture, since AI Overviews are grounded on Google's ordinary Search index rather than a separate AI crawl.

Does llms.txt do anything?

No published statement was located, as of 19 August 2026, from Google, Microsoft, OpenAI, Perplexity or Anthropic confirming that llms.txt is read as an inclusion or ranking input. Google's documentation says AI text files are not needed. Unlike robots.txt, the convention has no standards status and no documented honouring by major crawlers. Absence of a statement is not proof of non-use, but it is the entire public record.

Does blocking GPTBot remove a site from ChatGPT's answers?

No, and this is the most common error in the area. OpenAI documents GPTBot as its model-training crawler and OAI-SearchBot as its search-index crawler, stating that sites opted out of OAI-SearchBot will not be shown in ChatGPT search answers. The two controls do different jobs, and publishers frequently get them backwards in both directions.

Can AI visibility be measured?

Only indirectly and imperfectly. Generated answers are non-deterministic, free and paying users of the same product can be served from different indexes, the underlying model changes without notice, and an answer that satisfies the reader leaves no referral in server logs. Measurement is therefore done with synthetic query panels, and the panel design largely determines what is reported.

Do AI answer engines rank pages the way search engines do?

Not in a form anyone outside can inspect. A search engine produces a stable ordered list that can be recorded and compared over time. An answer engine produces a variable paragraph in which a source is either cited or not, with no position and no published model of why one eligible source was chosen over another. The citable pool is also fixed by whose index the product grounds on.

Do GEO techniques actually work?

The evidence points both ways and both readings are worth holding. A 2026 University of Tokyo arXiv preprint isolating document structure from semantic content measured a 17.3 percent citation-rate improvement across six generative engines in a controlled setting. A 2026 arXiv preprint surveying 45 GEO studies published between November 2023 and July 2026 found that no reviewed technique demonstrates a stable, longitudinal, cross-platform causal effect on discoverability or on user behavior, naming topical relevance and context position as the most reproducible levers. Both are preprints; the survey is single-authored.

Sources

Top