free seo tools
SEO.to / Guides

Ultimate Guide to AI Search Visibility (2026)

How AI assistants decide what to cite, the three kinds of AI crawler and which ones you must not block, how llms.txt works, and the content patterns that make a site quotable.

AI SEO  ·  updated 2026-08-16  ·  8,277 words  ·  36 min read

AI search visibility is the question of whether ChatGPT, Perplexity, Claude, Google AI Overviews, and Google AI Mode cite your pages when they answer a question, and whether a reader can reach you from that citation. This guide explains how those engines pick sources in 2026, which crawlers to allow and which to block, how llms.txt actually works, and the content patterns that make a page quotable. It is written for founders and SEO practitioners who are tired of recycled statistics and want a repeatable audit instead: check your crawl access, generate an llms.txt to the v2 spec, then run the same queries across all four engines and log which sources each one cites.

What AI Search Optimization (AEO/GEO) Actually Means in 2026

AI search optimization, usually called AEO for answer engine optimization or GEO for generative engine optimization, is the work of making your content show up when a generative engine answers a question and credits sources. The two acronyms describe the same goal. AEO stresses the answer-format surface: the paragraph or bullet list an assistant writes instead of ten blue links. GEO stresses the generative part: the fact that the assistant synthesizes an answer from multiple retrieved pages rather than echoing one page verbatim. In practice practitioners use the two labels interchangeably, and the work underneath is mostly ordinary SEO plus one new control surface, which is how AI crawlers may read and license your site.

The three layers of the job

AI visibility breaks into three layers, and most guides collapse them into one. The first is crawl access: whether the engine's crawler can reach your pages at all, decided by robots.txt and the product tokens you allow or block. The second is retrievability: whether your page is indexed, eligible for a snippet, and ranked well enough that the engine's retrieval step pulls it in. The third is quotability: whether the page is written so an engine can lift a definition, a statistic, a quote, or a sourced claim out of it and cite you. A page can pass the first two layers and still never be cited because nothing on it is easy to quote.

Why the label is not the discipline

The term GEO became popular after a 2023 Princeton paper measured how much editing content for generative engines could change a page's visibility. The researchers reported that their tactics could "boost visibility by up to 40% in generative engine responses," with results that varied across domains, as described in the GEO paper (Aggarwal et al., 2023). Vendors turned that into a product category. Google then answered with an official position that most of the new discipline is just SEO under a new name, which the next sections cover. The honest summary is that the label is new but the mechanics are familiar, with crawl access to AI bots as the genuinely new lever.

How Generative Engines Work: RAG, Grounding, and Query Fan-Out

You cannot influence which sources an engine cites without knowing the pipeline that produces the citation. Google describes its generative AI features as "rooted in our core Search ranking and quality systems" and says they use retrieval-augmented generation with grounding plus a step called query fan-out to pull pages from the Search index, per Google's AI optimization guide (2025). The same three ideas explain how ChatGPT and Perplexity work even though they run on different indexes. Here is what each term means.

Retrieval-augmented generation (RAG)

RAG means the model does not answer from memory alone. When a query arrives, the system first retrieves a set of relevant documents from an index, then feeds those documents to the language model as context, then asks the model to write an answer grounded in that context. The key consequence is that your page competes at the retrieval step, not at the writing step. If your page is not retrieved, the model never sees it, no matter how quotable it is. Retrieval is where ranking, indexation, and crawl access all matter.

Grounding

Grounding is the constraint that the answer must be supported by the retrieved sources and that the engine must show those sources as citations. It is what turns a language model from a text generator into a search product that can be fact-checked. For you, grounding means the engine needs specific, named, dated material to point at. A page of general advice gives the model nothing to anchor a citation to, while a page with a definition, a statistic, and a source URL gives the model three anchors.

Query fan-out

Query fan-out is the step where the engine takes one user query, breaks it into several smaller sub-queries, and runs them in parallel against the index before merging the results. Google names this mechanism directly in its guide. It is why an AI answer can pull from domains that never ranked for the exact question you typed, because the engine searched for related sub-questions instead. Fan-out depth differs by surface, which is a major reason the same question produces different citations on different engines, a point the surface-divergence section explains in detail.

What the pipeline means for your pages

The pipeline reads backward from citation to crawl. A citation happens because a page was retrieved, a page was retrieved because it was indexed and ranked for some sub-query, and it was indexed because a crawler could reach it. Each step is a place to fail, and each step is checkable. The six-step audit later in this guide walks that chain in order, which is why it starts with robots.txt rather than with content.

Google's Official Position: Optimizing for AI Search Is Still SEO

Google's guide is unusually blunt, and it settles an argument the vendor industry keeps relitigating. The document states that "from Google Search's perspective, optimizing for generative AI search is optimizing for the search experience, and thus still SEO," per Google's AI optimization guide (2025). In other words, Google does not recognize AEO or GEO as a separate discipline. The same crawl, index, rank, and content-quality work that earns a blue link earns an AI Overview citation.

"From Google Search's perspective, optimizing for generative AI search is optimizing for the search experience, and thus still SEO." Google Search Central (2025)

What Google says does nothing

The same document spends real space killing three myths. First, it says llms.txt files and other machine-readable AI text files or Markdown "neither harm nor help" visibility in Google Search because Google ignores them. Second, it says no special schema.org markup is required for generative AI search. Third, it says "chunking" content is not required and there is no ideal page length. All three claims come from Google's AI optimization guide (2025). The lesson is not that llms.txt and schema are worthless everywhere. The lesson is that they are worthless for Google Search specifically, while they may still matter for ChatGPT, Perplexity, and Claude, which run their own systems.

Why this matters for your budget

Most third-party GEO guides sell llms.txt and schema as the core levers without saying which engine each lever affects. That is the single biggest source of wasted spend in AI SEO. You can spend a week building llms.txt and structured data, and your Google AI Overview visibility will not move, because Google told you in advance it ignores both. The profitable order is to fix the things Google does read first, then add llms.txt as a cheap extra for the assistants that do read it. The rest of this guide is organized around that ordering.

The 2026 AI Search Landscape: Google AI Overviews, AI Mode, ChatGPT, and Perplexity

There is no single thing called AI search. There are at least four surfaces with different indexes, different crawlers, and different citation behavior, and treating them as one monolith is how sites get cited nowhere while thinking they are cited everywhere. Here is the landscape as it stands in 2026.

Google AI Overviews

AI Overviews are the generative answers Google places at the top of regular Search results. They are grounded in the Search index and governed by Google's core ranking and quality systems, per Google's AI optimization guide (2025). Because they sit inside Search, the path to a citation is the classic one: get indexed, rank well, and be eligible for a snippet. Your Googlebot crawl and your organic rankings directly determine whether you appear.

Google AI Mode

AI Mode is Google's more assistant-like surface, a conversational search experience that answers with longer generative responses and more sources per query. It runs on the same Search index but samples it differently, with heavier query fan-out, which is why it cites more domains per query than AI Overviews do. The practical difference is that AI Mode is a second Google surface you must track separately, because a page can rank in one and not the other.

ChatGPT

ChatGPT's web search answers questions and lists clickable source links under the answer. It runs on its own pipeline and its own crawler, OAI-SearchBot, which indexes pages for ChatGPT search specifically rather than for model training. To be citable in ChatGPT you must allow that crawler, be in whatever index ChatGPT search uses, and give the model quotable material. ChatGPT's citations come from a different index than Google's, so overlap with Google surfaces is low by design.

Perplexity

Perplexity is an answer engine built around citations from the start, and it presents source links prominently under every answer. It crawls with PerplexityBot and reads the live web. Perplexity also makes heavy use of tools that read pages and files, which is the one place llms.txt genuinely helps, because Perplexity and similar assistants consult llms.txt when they need to understand a site.

The engine-by-engine cheatsheet

The table below is the one to keep open while you work through the rest of this guide. It answers the question most guides skip: which lever matters for which engine.

EngineCrawler or token to allow or blockWhat actually drives citationsHow to measureDoes llms.txt or schema matter?
Google AI OverviewsGooglebot for crawling; Google-Extended for training opt-outIndexation, snippet eligibility, organic ranking and quality systemsSearch Console Generative AI performance reportNo. Google ignores both for Search.
Google AI ModeGooglebot for crawling; Google-Extended for training opt-outSame index, but heavier query fan-out and more domains per querySearch Console Generative AI performance reportNo. Same as AI Overviews.
ChatGPTAllow OAI-SearchBot; block GPTBot if you want no trainingPresence in the ChatGPT search index and quotable contentManual query logging; no first-party consolellms.txt can help ChatGPT tools read your site; schema is not required.
PerplexityAllow PerplexityBot; also allow ChatGPT-User-style fetchers for tool useLive-web indexation and quotable, source-rich contentManual query logging; no first-party consolellms.txt helps Perplexity's tools; schema is not required.
Tip. Allow the search crawlers (Googlebot, OAI-SearchBot, PerplexityBot) on every page you want cited. The training crawlers (GPTBot, ClaudeBot, Google-Extended, CCBot) are a separate decision about model training, and blocking them does not block citations.

Why AI Overviews and AI Mode Cite Different Sources (and Why You Must Track Both)

The most expensive assumption in AI SEO is that your AI Overview presence tells you anything about your AI Mode presence. It does not. The two Google surfaces share only about 13.7 percent of cited URLs across roughly 540,000 query pairs, according to Ahrefs, while SE Ranking found about 10.7 percent overlap and Victorious found no query that produced identical citations across both surfaces, as summarized in this 2026 comparison of AI Overviews and AI Mode citations. One in seven is the ceiling for assuming one surface predicts the other.

Fan-out depth explains the gap

The divergence is not noise. AI Mode samples the web differently because it runs heavier query fan-out, and the effect shows up in the source counts. AI Mode cites about 9.2 domains per query versus 7.7 for AI Overviews, per the same 2026 comparison. AI Mode fans the question out into more sub-queries, retrieves from a wider set of pages, and therefore cites sources that never appeared in the AI Overview for the same query. Ranking on one surface and expecting the other is like ranking on desktop and assuming mobile follows; they share a backend but not a result set.

The surface-divergence numbers

MetricAI Overviews valueAI Mode valueSource
Citation overlap between the two surfaces13.7 percent (Ahrefs); 10.7 percent (SE Ranking)Same figure, read as shared URLsApiserpent summary (2026)
Domains cited per query7.79.2Apiserpent summary (2026)
Identical citation sets per queryNone observedNone observedVictorious, via Apiserpent (2026)
Warning. AI Overviews and AI Mode share only about 13.7 percent of citations. Track them as two separate surfaces, or you will conclude you are invisible on one of them and never know it.

What this means for your tracking

Because the surfaces diverge, a single "AI visibility" number is meaningless. You need at least two readings for Google alone: AI Overviews and AI Mode. The Generative AI performance report in Search Console breaks them out for you, and the manual query log in Step 5 does the same across ChatGPT and Perplexity. Any dashboard that merges everything into one score is hiding the divergence, not resolving it.

Step 1: Audit Your Crawl-Access Layer — robots.txt, RFC 9309, and Google-Extended

Every citation begins with a crawler being allowed to read the page, so the audit starts at robots.txt. The Robots Exclusion Protocol was informal guidance for decades until it was formalized as RFC 9309 in September 2022. That matters because the RFC defines exactly how a crawler parses the file: groups of user-agent lines followed by Allow and Disallow rules, with the longest matching rule winning and a tie going to Allow. Google's parser follows RFC 9309, and Google-Extended and every other AI product token rely on this same mechanism to honor your opt-outs.

The three kinds of AI crawler

AI crawlers split into three groups, and you must treat them differently because they feed different things. Training crawlers such as GPTBot, ClaudeBot, CCBot, and Google-Extended add your content to future model training. Search crawlers such as OAI-SearchBot and PerplexityBot index you for AI search products and their citations, plus Googlebot for Google's own surfaces. User fetchers such as ChatGPT-User open your pages when a person asks an assistant to read a specific URL. Blocking the first group keeps you out of training. Blocking the second or third group removes you from citations, which is usually not what you want.

Google-Extended in context

Google-Extended, introduced in September 2023, is a standalone crawler or product token you can block in robots.txt to opt out of your content being used to improve Gemini and Vertex AI models, and it does so without affecting Search, per Google's content sharing controls (2023). The key phrase is "without affecting Search." Blocking Google-Extended does not remove you from AI Overviews or AI Mode, because those surfaces are grounded in the Search index and Googlebot, not in Gemini training. Many publishers block Google-Extended believing it protects them from AI Overviews; it does not, and it only stops Gemini training use.

The crawl-access decision matrix

AI product or botToken or directiverobots.txt ruleConsequence of blocking
Gemini and Vertex AI trainingGoogle-ExtendedUser-agent: Google-Extended then Disallow: /Content stops improving Gemini models; Search and AI Overviews unaffected
Google Search and AI OverviewsGooglebotUser-agent: Googlebot then Allow: /Blocking removes you from Search and both Google AI surfaces entirely
OpenAI model trainingGPTBotUser-agent: GPTBot then Disallow: /Content stops entering OpenAI training; ChatGPT search citations unaffected
ChatGPT search citationsOAI-SearchBotUser-agent: OAI-SearchBot then Allow: /Blocking removes you from ChatGPT search results and citations
ChatGPT reading a URL on requestChatGPT-UserUser-agent: ChatGPT-User then Allow: /Blocking stops the assistant from opening your pages when asked
Perplexity citationsPerplexityBotUser-agent: PerplexityBot then Allow: /Blocking removes you from Perplexity answers and citations
Anthropic trainingClaudeBotUser-agent: ClaudeBot then Disallow: /Content stops entering Claude training
Common Crawl corpusCCBotUser-agent: CCBot then Disallow: /Content leaves the Common Crawl corpus many training sets draw from

Test every important URL against the file rather than trusting the pattern by eye. The robots.txt tester applies the same longest-match rule the crawlers use, so you can paste a URL and see the per-bot verdict. The code below is a starter file that keeps Search and citations open while opting out of model training, followed by the two link tags that declare your llms.txt.

# robots.txt: keep search and citations open, opt out of model training
User-agent: Googlebot
Allow: /

User-agent: Google-Extended
Disallow: /

User-agent: OAI-SearchBot
Allow: /

User-agent: ChatGPT-User
Allow: /

User-agent: PerplexityBot
Allow: /

User-agent: GPTBot
Disallow: /

User-agent: ClaudeBot
Disallow: /

User-agent: CCBot
Disallow: /

# Declare the llms.txt that describes this site, from any page's <head>:
# <link rel="alternate" type="text/markdown" href="/page.md">
# <link rel="describedby" href="/llms.txt">
Tip. robots.txt governs crawling, not indexing or citations. A page blocked from the search crawler cannot be retrieved and cited, but a page blocked only from training crawlers can still rank and be cited. Keep the search crawlers open and decide about training separately.

Step 2: Confirm Indexation and Snippet Eligibility in Search Console

Once crawl access is right, the next gate is whether Google can actually retrieve the page. Google states the eligibility condition plainly: to appear in its generative AI features, a page must be indexed, must be eligible for a snippet, and the site must be included in "Search generative AI features" in Search Console, per Google's AI optimization guide (2025). All three conditions are checkable, and most sites fail one of them without noticing.

Indexed

Use the URL Inspection tool in Search Console and check the status of each page you want cited. A page that returns "Crawled, currently not indexed" will never appear in an AI answer, because the retrieval step has nothing to pull. The usual causes are thin content, duplicate content, a noindex directive, or a page that is new and not yet crawled. If the page is not indexed, fix the cause before doing any GEO work, because every later step is wasted on an unindexed page.

Eligible for a snippet

Snippet eligibility means Google can generate the descriptive text and preview that appear under a result. It is a proxy for whether the page can be cleanly represented in search, and it is a hard gate for AI features. Pages that block snippets with max-snippet directives, serve noindex, or hide their content behind client-side rendering typically fail here. The technical details of snippet eligibility sit inside the broader crawl and render pipeline, which the technical SEO guide covers end to end.

Included in Search generative AI features

The third condition is a site-level switch. In Search Console, confirm your site is opted in to Search generative AI features. This setting is where Google records whether your property participates in the experiment and reporting for AI Overviews and AI Mode. If the property is not included, your pages cannot appear and you will not see data in the Generative AI performance report, even if the pages are indexed and snippet-eligible.

Tip. Run the three checks in order: URL Inspection for indexation, then snippet eligibility, then the site-level generative AI setting. Each one is a yes-or-no gate, and clearing them costs minutes but is the difference between being citable and being invisible.

Step 3: Build an llms.txt File to the v2 Spec (and Know Where It Does Nothing)

llms.txt is a plain Markdown file at the path /llms.txt that describes your site for language models and the tools they use. The proposal comes from Jeremy Howard, was published on September 3, 2024, and the v2 revision was modified on August 10, 2026, per llmstxt.org (2024, v2 2026). The idea is a robots.txt for reading: a compact, human-and-machine-readable summary of what your site is, what it contains, and where the important pages live. It costs almost nothing to produce and helps the assistants that consult it.

What the v2 spec adds

The v2 spec standardizes two link relationships alongside the file itself. You declare a Markdown version of a page with rel="alternate" type="text/markdown" pointing at the .md version, and you point from a page to its covering llms.txt with rel="describedby". These tags let a tool find the plain-text description of a page without guessing. The spec also documents that OpenAI, Anthropic, and Gemini publish llms.txt files for their own developer docs, and that Chrome's Lighthouse audits sites for one as part of its agentic browsing checks, per llmstxt.org (2026). Adoption is real even if the file is optional.

A complete v2 example

# llms.txt for Example Analytics
> Example Analytics is a product analytics platform for B2B software
> teams. We publish research on retention, onboarding, and pricing.

Example Analytics helps SaaS teams understand activation and
retention. Key resources:

## Core pages
- [Home](https://www.example.com/)
- [Product](https://www.example.com/product)
- [About](https://www.example.com/about)

## Guides
- [Retention Playbook](https://www.example.com/guides/retention-playbook): how activation, onboarding, and churn fit together.
- [Pricing](https://www.example.com/guides/pricing): usage-based versus seat-based, step by step.

## Optional
- [Sitemap](https://www.example.com/sitemap.xml)
- [Careers](https://www.example.com/careers)

## Blog
- [New features](https://www.example.com/blog): product notes and research.

Where llms.txt does nothing

The critical caveat is that llms.txt helps the assistants that read it, and Google is not one of them. Google says llms.txt files and other machine-readable AI text files or Markdown "neither harm nor help" visibility in Google Search, because Google Search ignores them, per Google's AI optimization guide (2025). The file is useful for ChatGPT and Perplexity tool use, for agentic browsing, and for any system that needs a site summary, but it will not move a single AI Overview or AI Mode citation. Build it for those non-Google surfaces and do not expect a Search lift.

Warning. llms.txt neither helps nor hurts Google Search. Build it because ChatGPT, Perplexity, and agentic tools read it, and keep your Google expectations at exactly zero.

Step 4: Make Your Content Citable — Entities, Quotes, Definitions, and Sources

Content quality is where the original GEO research and Google's guide actually agree. The Princeton paper found that adding citations, quotations, and statistics to content can raise visibility in generative-engine responses, with gains up to roughly 40 percent and strong variance by domain, per the GEO paper (Aggarwal et al., 2023). Google frames the same idea as writing helpful, people-first content. The practical translation is four quotable elements that give an engine something to anchor a citation to.

Entities

Name the specific things your page is about: products, companies, people, standards, and places. An engine grounds better when the page mentions recognizable entities rather than vague categories. "A project management tool" is not citable; "Asana's workload view" is. Entities also connect your page to the knowledge graph and to the sub-queries fan-out generates. The keyword research tool is useful here, not to stuff keywords, but to find the entity-level terms searchers actually use so your page speaks the same names they type.

Quotes

Give the engine a clean, attributed sentence it can lift. A primary quote from a named person, with a source link, is the single most quotable unit on a page because the engine can reproduce it and cite you with confidence. Keep quotes short, put the speaker's name and title next to them, and link the source. Avoid long block quotes that bury the quotable sentence in a paragraph.

Definitions

Open key sections with a definition the engine can copy: "Query fan-out is the step where an engine splits one query into several sub-queries and runs them in parallel." A definition sentence gives the model a ready-made answer fragment. Google's own guide says chunking is not required and there is no ideal page length, per Google's AI optimization guide (2025), so do not slice pages artificially. Just lead with definitions and the engines will find them.

Sources and statistics

Every statistic on the page should carry a dated, named, linked source. This does two things at once. It makes the page trustworthy to readers, and it gives the engine a verifiable fact it can cite. When you cite a number, name the source, the sample, and the year, because the engine and the reader both need to know whether a figure still applies. The statistics audit section shows exactly how to do this for the most recycled numbers in this field.

Pro. Structure a quotable page as a stack: a definition first, an entity-rich body, one attributed quote, and a sourced statistic per claim. Then check the structure with the keyword research tool to confirm the page matches the terms your target queries actually use.

Step 5: Run the Same Queries Across ChatGPT, Perplexity, AI Overviews, and AI Mode

Now stop reading about AI visibility and measure your own. The repeatable check is to run an identical set of queries across all four surfaces and log which sources each engine cites. This is the step no listicle gives you, and it is the one that turns recycled statistics into your own numbers. Run the queries in a logged-out or private window so personalization does not distort the results, and note the date, because citations change from week to week.

The exact queries to run

Build your query set from three types. First, your money queries: the five questions your best pages answer, phrased as a searcher would ask them. Second, category queries: two or three broader questions your competitors rank for but you do not yet, so you can see who the engines currently cite. Third, definitional queries: one or two "what is X" questions in your space, which engines answer with a definition and a single source. Keep the list to about ten queries so you can rerun it in under an hour each month.

How to read the results

For each query, record the engine, the cited domains and pages, whether your page appeared, and your page's position in the citation list. Do not count a mention as a win unless the answer links to your URL. The goal of the first run is a baseline, not a ranking; you need a before snapshot before any content change can be credited with an after result. Run the check across ChatGPT, Perplexity, AI Overviews, and AI Mode, because the four surfaces will disagree with each other more than they agree.

The before and after citation log

The table below is a worked template, with a before column showing the baseline and an after column showing what a quotable rewrite changed. Use one row per query per engine per month.

DateQueryEngineCited pageYour page citedNotes
2026-08-01what is query fan-outAI Overviewsdevelopers.google.comNoBaseline before rewrite
2026-08-01what is query fan-outChatGPTsearchengineland.comNoBaseline before rewrite
2026-08-29what is query fan-outAI Overviewsyourdomain.com/guides/ai-visibilityYesAfter adding a definition and sourced statistic
2026-08-29what is query fan-outPerplexityyourdomain.com/guides/ai-visibilityYesAfter llms.txt went live
Tip. Keep the query list frozen. If you change the queries every month you are measuring a different thing each time, and the before and after columns will not be comparable.

For a faster read on whether your pages are even eligible, pair the manual log with the AI visibility checker, which flags the crawl and indexation problems that block citations before you spend a month watching empty logs.

Step 6: Log Citations and Measure With the Generative AI Performance Report

Manual logging tells you what ChatGPT and Perplexity do, but for Google's surfaces the first-party measurement is the Generative AI performance report in Search Console. Google provides it as the official way to measure generative-AI visibility, and it explicitly warns that no third-party tool has access to Google's internal ranking or AI systems, per Google's AI optimization guide (2025). Treat any vendor dashboard that claims to read AI Overview rankings as an estimate, not a measurement, and treat Search Console as the source of truth for Google surfaces.

What the report shows

The report breaks out impressions, clicks, and position for your pages as they appear in AI Overviews and AI Mode. Use it the same way you use the regular Performance report: find the pages that already earn AI impressions, find the queries behind them, and look for near misses, pages that appear on page two of a citation set but not yet in the answer itself. The report is also where you confirm the site-level inclusion discussed in Step 2, because a site that is not opted in shows nothing here.

Pair first-party and manual data

Search Console covers Google; your query log covers ChatGPT and Perplexity. The two data sets answer different questions, and you need both. The console tells you your aggregate Google visibility and trend. The manual log tells you whether a specific query now cites you on a specific surface, which is the before and after proof a content change actually moved something. Reconcile them monthly: a page that rises in the console should show up in the log for the queries it ranks for, and a page that appears in the log should eventually show impressions in the console.

Tip. Log in a spreadsheet with one row per query per engine per run, and keep the console open in a second tab. When the two disagree, investigate the crawl and indexation gates before blaming the content, because a citation gap is usually a retrieval gap first.

When your log shows a page cited on one surface but missing on another, the two troubleshooting sections that follow explain the most common causes and fixes.

Schema Markup: What Actually Moves AI Citations (and What Does Not)

Schema is the most oversold lever in AI SEO, so get the record straight. Google says no special schema.org markup is required for generative AI search, per Google's AI optimization guide (2025). An independent 2026 analysis of citation data reached the same conclusion from the outside: schema markup has no meaningful impact on AI citations, because content signals, not structured data, drive whether engines cite a page, per this citedbyai.info analysis (2026). Two sources, one from Google and one external, agree that schema is not a citation lever.

What schema still does

That does not make schema worthless. Structured data still earns rich results, breadcrumbs, FAQ and how-to features where Google supports them, and it clarifies entities for the knowledge graph. Those are real, measurable Search benefits, and they indirectly support AI visibility by making your page richer and better understood. The distinction to keep is that schema helps the page, not the AI citation directly. If your goal is a rich result, use it; if your goal is an AI Overview citation, schema alone will not get you there.

Validate what you have

If you already ship structured data, make sure it is valid and matches visible page content, because broken or misleading markup does nothing and can trigger a manual action. Run the JSON-LD through the schema checker to confirm it parses and references the real entities on the page. Spend the time you save by not over-investing in schema on the quotable content from Step 4, which is the lever that actually moves citations.

Warning. Do not buy schema as a GEO tactic. Google says no special schema is required for AI search, and external citation analysis finds schema has no meaningful impact on citations. Use schema for rich results and entities, and put your citation effort into content.

The Crawl-Access and Licensing Layer: Google-Extended and Cloudflare's 2026 Pay-For-Content Policy

Underneath every content tactic sits a licensing question: who gets to crawl your site, and who gets to train on it or pay for it. The robots.txt decisions from Step 1 are the technical half of that question, and the commercial half is moving fast. The layer is worth its own section because the wrong default here silently removes you from some engines or gives your content away for free.

Google-Extended as a training opt-out

Google-Extended is the cleanest example of the separation between training and search. Blocking it opts your content out of improving Gemini and Vertex AI models without affecting Search, per Google's content sharing controls (2023). That is a decision about training, not about citations. Make it deliberately: if you sell or license your content for model training, blocking Google-Extended for free use is consistent with charging for it elsewhere. If you want maximum reach and do not mind training use, leave it open. Either way, know that the toggle has nothing to do with whether AI Overviews cite you.

Cloudflare's 2026 pay-for-content policy

The commercial side moved in July 2026, when Cloudflare announced a policy to separate traditional search crawlers from AI content crawlers and to push AI companies to pay publishers for content, as reported by TechCrunch (2026). The practical effect is that a publisher's CDN can now distinguish Googlebot-style search crawling, which is free and drives traffic, from AI training crawling, which should be licensed. If your site runs behind Cloudflare, review how that policy classifies each bot on your domain, because a bot you intended to allow for citations could be reclassified as a content crawler and throttled or billed.

How to decide per bot

Decide each crawler against two questions. Does blocking it cost you citations? Search crawlers like Googlebot, OAI-SearchBot, and PerplexityBot cost you citations if blocked, so keep them open. Does allowing it cost you licensing revenue? Training crawlers like GPTBot, ClaudeBot, Google-Extended, and CCBot feed training, so block them if you license your content or want to negotiate, and allow them if reach matters more. The robots.txt tester shows the per-bot verdict for any URL, which is the fastest way to confirm your CDN and your robots.txt agree.

Third-Party GEO Tactics vs. Recycled Statistics: How to Reproduce the 40% and 38% Numbers

Two statistics dominate every GEO article, and almost none of them tell you where the numbers came from, how big the samples were, or whether they still apply. Both deserve scrutiny before you build a strategy on them. Reproducing the numbers yourself is the only defense against being sold a benchmark that no longer describes the current engines.

The "up to 40%" GEO figure

The 40 percent figure comes from the original Princeton GEO research, which reported that GEO methods "can boost visibility by up to 40% in generative engine responses," with efficacy varying across domains, per the GEO paper (Aggarwal et al., 2023). It is a 2023 benchmark measured on the models available then, and the "up to" is the ceiling, not the average. Treat it as evidence that quotable content helps, not as a promise that you will gain 40 percent. Any vendor that quotes it as your expected lift is stretching a ceiling into a forecast.

Warning. The "up to 40%" GEO figure is a 2023 research benchmark measured on the models of that year, not a current promise. It is a ceiling from one study, not an average you should put in a forecast.

The "38%" AI Overview citation figure

The 38 percent figure is newer and comes from Ahrefs, which analyzed 863,000 SERPs and found only 38 percent of AI Overview citations come from top-10 organic pages, down from 76 percent a year earlier, per Ahrefs (2026). The takeaway is real and important: organic ranking alone no longer guarantees AI citations, because most citations now come from outside the top ten. But the number is a snapshot of one sample at one time, and it will keep moving as Google tunes AI Overviews. Use it as a reason to stop assuming rank equals citation, not as a fixed constant.

The statistics audit table

ClaimSourceSample sizeCaveat
GEO can boost visibility by up to 40%Princeton GEO paper (2023)Controlled experiments on 2023 models"Up to" is the ceiling; efficacy varies by domain; predates current engines
Only 38% of AI Overview citations come from top-10 organicAhrefs (2026)863,000 SERPsSnapshot in time; down from 76% a year earlier; will keep moving
13.7% citation overlap between AI Overviews and AI ModeApiserpent summary (2026)About 540,000 query pairsSE Ranking found 10.7%; Victorious found no identical sets

Reproduce it yourself

The only number that matters is your own. Run the frozen query set from Step 5, log the citations, and after two or three monthly runs you will have your own overlap rate, your own top-ten dependence, and your own before and after. A benchmark tells you what happens in aggregate; a log tells you what happens to your pages. The AI visibility checker gives you the crawl and indexation baseline behind that log, so you can separate a content problem from an access problem before you change anything.

Troubleshooting: Indexed But Never Cited in AI Results

The most common AI visibility complaint is a page that is indexed, ranks, and still never appears in an AI answer. This has a specific set of causes, and they are diagnosable in order. Work down the list and you will find the failure.

Check snippet and generative eligibility first

Indexation is not enough. The page must also be snippet-eligible and the site must be opted in to Search generative AI features, per Google's AI optimization guide (2025). A page can be indexed but blocked from snippets, or the whole property can be excluded from generative features, and then no amount of content work will produce a citation. Confirm both before you touch the copy.

Then check quotability

If eligibility holds, the failure is quotability. A page that is all introduction and no definition, no named entity, no statistic, and no quote gives an engine nothing to cite. Read the page and ask what single sentence the engine could lift with a source link. If nothing qualifies, the page needs a rewrite along the lines of Step 4, not more links or more schema. A page that competes for broad queries is also harder to cite than one that owns a narrow, definitional query, so target the definitional and money queries where a single source suffices.

Then check the surface, not the site

Sometimes the page is cited, just on a surface you are not checking. A page can appear in AI Mode but not AI Overviews, or in Perplexity but not ChatGPT. Before you conclude the page is never cited, run the full four-surface log from Step 5. The absence is often surface-specific, and the fix is different when the page is missing on one surface versus all of them.

Troubleshooting: Cited by One Engine But Not Another (Surface Divergence)

A page cited by ChatGPT but not by Google, or by AI Mode but not by AI Overviews, is not broken. It is the normal state of a fragmented landscape, and the divergence has mechanical causes you can work with rather than fight.

Different indexes, different crawlers

Each surface retrieves from its own index. Google grounds both of its surfaces in the Search index, while ChatGPT and Perplexity run their own pipelines and crawlers. If OAI-SearchBot cannot reach your page but Googlebot can, you will rank in Google and never appear in ChatGPT, and vice versa. Check the per-bot access with the robots.txt tester first, because a one-line robots.txt mistake is the cheapest divergence to fix.

Fan-out depth changes the answer

Even within Google, the two surfaces sample the web differently. AI Mode cites about 9.2 domains per query versus 7.7 for AI Overviews, driven by heavier query fan-out, per this 2026 comparison. A page that answers the exact query narrowly may surface in AI Overviews, while AI Mode fans out and cites different, adjacent sources. To appear in both, cover the surrounding sub-questions too, so your page is retrieved no matter how deep the fan-out goes.

Write for the union, not the intersection

The fix for surface divergence is coverage. Cover the definition, the adjacent questions, the statistics, and the sources, so the page is retrieved under multiple sub-queries. Then accept that a page can legitimately win one surface and not another in a given month. Track the surfaces separately, fix crawl access where it differs, and let the content coverage do the rest. The AI visibility checker helps you see which surfaces are even reachable before you read too much into a single missing citation.

Spam-Policy Guardrails: Avoiding Scaled Content Abuse and GEO Hacks

AI visibility has a compliance line, and it is easy to cross by accident while chasing citations. Google is explicit: creating content primarily to manipulate rankings or generative AI responses violates its scaled content abuse spam policy, per Google's AI optimization guide (2025). The same policy that governs traditional SEO applies to GEO, and the penalties are the same.

What scaled content abuse means

Scaled content abuse is producing pages at volume with the primary purpose of gaming rankings, regardless of whether a human, a tool, or a mix of both wrote them. The policy looks at intent and at whether the content exists to help people or to manipulate. Mass-producing thin pages that echo common AI queries, stuffed with entities and fake statistics to catch citations, is squarely inside the violation. A few genuinely useful, source-rich pages are not.

GEO hacks that backfire

The GEO hacks that backfire are the ones that optimize the output rather than the substance: keyword-matching the phrasing of common AI queries, fabricating statistics to look citable, cloaking AI crawlers with content different from what users see, and spinning articles to look like research. All of these trigger spam signals, and cloaking specifically is an outright violation. Google's guide warns against all of it, per Google's AI optimization guide (2025). The safe version of GEO is the one this guide describes: real definitions, real sources, real statistics, and a page that helps a person as much as it helps an engine.

The safe pattern

Keep one test between you and a violation. Ask whether the page would still be worth publishing if no engine ever cited it. If the answer is yes, you are writing helpful content that happens to be citable. If the answer is no, you are writing for the algorithm, and that is the exact intent the spam policy prohibits. The on-page SEO guide covers the people-first writing and E-E-A-T signals that keep you on the safe side while you pursue citations.

The Complete 2026 AI Search Visibility Audit Checklist

The full audit, in remediation order. Each step assumes the one before it works, so run them top to bottom. The first run is your baseline; rerun it monthly and compare against the log from Step 5.

StepCheckPass conditionVerify with
1robots.txt allows search crawlers and blocks only training crawlersGooglebot, OAI-SearchBot, PerplexityBot allowed; GPTBot, ClaudeBot, Google-Extended per your training decisionrobots.txt tester
2Every target page is indexedURL Inspection shows "Page is indexed"Search Console URL Inspection
3Pages are snippet-eligibleNo noindex, no max-snippet block, content server-renderedURL Inspection and rendered HTML
4Site is in Search generative AI featuresProperty opted in, data visible in the reportSearch Console settings
5llms.txt live at /llms.txt with alternate and describedby linksFile serves at 200 with correct Markdown, link tags in headcurl and the schema checker for the link tags
6Pages are quotableDefinition, entities, one attributed quote, and a sourced statistic per pageRead the page and extract one citable sentence
7Frozen query set definedAbout ten queries across money, category, and definitional typesThe Step 5 list
8Four-surface citation log startedOne row per query per engine per run, with a baselineThe Step 5 log template
9Generative AI performance report reviewedImpressions and clicks tracked across AI Overviews and AI ModeSearch Console
10Statistics audited and sourcedEvery claim dated, named, and linked; recycled numbers reproduced or droppedThe Step 14 audit table
Pro. Run this list top to bottom once, save the date, then rerun only what changed monthly. The before and after columns in your log, not a vendor's benchmark, are the only proof your AI visibility actually moved.

Frequently asked questions

What is AI search optimization in 2026?

AI search optimization is the work of making your content appear in the answers and source lists that generative engines produce, including Google AI Overviews, Google AI Mode, ChatGPT, Perplexity, and Claude. It combines ordinary SEO, crawl-access control for AI crawlers, and writing content that is easy to quote and cite. Google states that from Search's perspective this is still SEO, not a separate discipline.

What is generative engine optimization (GEO)?

Generative engine optimization, or GEO, is the practice of editing content so a generative engine is more likely to include it in a synthesized answer. The term comes from a 2023 Princeton paper that reported visibility gains of up to 40 percent from adding citations, quotations, and statistics. In practice GEO overlaps almost entirely with SEO and with what some call answer engine optimization.

Is GEO different from SEO?

Google says no, and that optimizing for generative AI search is optimizing for the search experience and therefore still SEO. The one genuinely new layer is crawl access to AI crawlers and files like llms.txt, which apply to non-Google assistants. The core work of indexing, ranking, and helpful content is the same.

What is an llms.txt file and do I need one?

An llms.txt file is a Markdown file at /llms.txt that describes your site for language models and the tools they use, standardized by the proposal at llmstxt.org. You need one if you want ChatGPT, Perplexity, and agentic tools to understand your site quickly. It costs almost nothing to produce, but it does nothing for Google Search.

Does llms.txt help my Google Search rankings?

No. Google says llms.txt and other machine-readable AI text files or Markdown neither harm nor help visibility in Google Search, because Google ignores them. Build llms.txt for ChatGPT, Perplexity, and agentic browsing, and keep your Google expectations at zero.

How do I get my site cited in Google AI Overviews?

Make sure the page is indexed, eligible for a snippet, and that your site is included in Search generative AI features in Search Console. Then make the content quotable with definitions, entities, quotes, and sourced statistics, and confirm Googlebot can crawl the page. AI Overviews are grounded in the Search index, so ordinary ranking and quality work still drive citations.

How do I track my AI search visibility?

Use the Generative AI performance report in Search Console for AI Overviews and AI Mode, which is Google's first-party measurement. For ChatGPT and Perplexity, run a frozen set of queries in a private window and log which sources each engine cites. The AI visibility checker can flag the crawl and indexation problems that block citations in the first place.

Does schema markup help my content get cited by AI?

No, not directly. Google says no special schema.org markup is required for generative AI search, and a 2026 analysis of citation data found schema has no meaningful impact on AI citations. Schema still helps rich results and entity clarity, but content signals are what drive citations.

What is Google-Extended and should I block it in robots.txt?

Google-Extended is a standalone crawler and product token introduced in September 2023 that lets you opt out of your content being used to improve Gemini and Vertex AI models. Blocking it does not affect Google Search, AI Overviews, or AI Mode. Block it only if you do not want your content used for Google's model training; keep Googlebot allowed regardless.

Why does ChatGPT cite different sources than Google AI Overviews?

ChatGPT and Google run on different indexes and different crawlers, so they retrieve different pages for the same question. Even Google's own two surfaces, AI Overviews and AI Mode, share only about 13.7 percent of cited URLs because they fan out queries to different depths. Different indexes plus different fan-out means low citation overlap by design.

How do I get my site cited by Perplexity?

Allow PerplexityBot in robots.txt so Perplexity can crawl and index your pages. Publish quotable content with definitions, statistics, and sources, and consider an llms.txt file, which helps Perplexity's tools understand your site. Then verify by running your target queries in Perplexity and checking the source list.

What is query fan-out in AI search?

Query fan-out is the step where a generative engine splits one query into several smaller sub-queries and runs them in parallel before merging the results. Google names this mechanism in its guide, and it is why AI Mode cites about 9.2 domains per query versus 7.7 for AI Overviews. Fan-out is why an answer can cite pages that never ranked for the exact question you typed.

Can I opt out of AI training without losing Google Search traffic?

Yes. Block Google-Extended to opt out of Gemini and Vertex AI training without affecting Search, and block GPTBot and ClaudeBot to opt out of OpenAI and Anthropic training. Keep Googlebot, OAI-SearchBot, and PerplexityBot allowed so you remain citable in search results. Training and citation crawlers are separate, and you can block one while allowing the other.

Is creating content specifically for AI search against Google's spam policies?

Creating content primarily to manipulate rankings or generative AI responses violates Google's scaled content abuse policy. Writing helpful, source-rich content that happens to be quotable by AI is fine. The test is whether the page would still be worth publishing if no engine cited it.