---
url: 'https://www.wpconsults.com/how-to-optimize-for-google-ai-overviews/'
language: 'en'
title: 'How to Optimize for Google AI Overviews: The Two Gates That Decide Who Gets Cited'
author:
  name: 'Abdullah Nouman'
  url: 'https://www.wpconsults.com/author/nouman/'
date: '2026-08-01T17:13:51-05:00'
modified: '2026-08-01T12:44:38-05:00'
type: 'post'
categories:
  - 'GEO/AEO/AI SEO'
  - 'Technical SEO'
image: 'https://www.wpconsults.com/wp-content/uploads/2026/07/how-to-optimize-for-google-ai-overviews-7783.avif'
published: true
---

# How to Optimize for Google AI Overviews: The Two Gates That Decide Who Gets Cited

Almost every guide on how to optimize for Google AI Overviews tells you the same thing: rank well, answer the question up top, use clean headings, add schema. That advice is not wrong, but it describes what correlates with a citation, not what decides one.

 

There are two gates. The first is a technical eligibility rule Google states plainly in its own documentation and almost nobody quotes. The second is the one that actually picks your paragraph out of a pool of pages that are all eligible. This piece walks through both, with the AI Overviews we captured ourselves and our own Search Console numbers, including the parts that make us look bad.

  

## Key Takeaways

 

- AI Overviews are generated per query and grounded in Google’s normal Search index, so they pull from pages that already rank for the query and its related variants.
- Google’s documentation states a hard eligibility rule: a page must be indexed and eligible to be shown in Search with a snippet. A `nosnippet` or restrictive `max-snippet` setting can therefore put a good page outside the feature.
- Among eligible pages, the passage that gets lifted is the one that is easiest to extract and safest to attribute: a direct, self-contained answer sitting immediately under the heading that asks the question.
- Corroboration matters as much as formatting. If your answer agrees with what other trusted sources say, it is safe to quote; if it stands alone, it is risky to quote.
- Google’s own AI Overview, answering a question about this very topic, told us that “AI engines do not simply cite top-ranking pages”.
- In our own captures, the AI Overview cited a one-year-old Reddit comment while several polished ranking guides were not cited at all.
- Google says structured data, llms.txt files, and rewriting your content for AI are not required. Not required and not useful are different claims, and Google only answers the first.
- There is no position one in an AI Overview, only a probability of being cited, and that probability changes by market.
- Rank and extractability are two axes, not two theories. Page one plus a liftable passage is a citation you keep; page two or later plus a liftable passage is a citation you sometimes get; page one without a liftable passage is not a citation at all.
- A liftable passage needs exactly one indexable home. We lost a four-market citation because our own category archive and a related page were republishing the same answer, and the article had fewer internal links than the pages competing with it.
- The citation pool churns on its own. Over twenty-two nightly reads of one query, a site cited in all four markets for four straight nights vanished from every market for the next ten, and two fetches twenty-seven seconds apart returned different cited sets.

  Table of Contents

- What is an AI Overview, and where does it get its sources?
- The first gate: is your page even eligible to appear in an AI Overview?
- The second gate: which passage gets lifted into the answer
- Do you have to rank on page one to be cited in an AI Overview?
- What our AI Overview captures showed about how a source gets picked
- The three stages of AI Overview optimization, at a glance
- How to optimize for Google AI Overviews, in three stages
- Do you need schema, llms.txt, or markdown files for AI Overviews?
- How many pages should you build for query variants?
- What our own Search Console data shows, and what it cannot show
- The AI Overview citation pool is not the same in every market
- How much does the AI Overview citation pool change on its own?
- The extractability test we ran on our own site, and what it took to get cited
- Where information gain and citation eligibility pull against each other
- So, how should you actually optimize for Google AI Overviews?
- We measured it: 61% of the sites an AI Overview cites do not rank on page one
- Update Logs

 

## What is an AI Overview, and where does it get its sources?

 

An AI Overview is a generated answer that sits at the top of some search results, assembled from pages in Google’s normal Search index and linked to a handful of cited sources. It is not a separate AI index and there is nothing to submit to. Google generates it per query, retrieves supporting pages through its existing ranking systems, and writes the summary from what it retrieved.

 

Google names the mechanism itself. In its [guide to optimizing for generative AI features](https://developers.google.com/search/docs/fundamentals/ai-optimization-guide), it describes **retrieval-augmented generation** as a technique “used to improve the quality, accuracy, and freshness of AI responses by relying on our core Search ranking systems to retrieve relevant, up-to-date web pages from our Search index.”

 

The same guide defines **query fan-out** as “a set of concurrent, related queries generated by the model to request more information and fetch additional relevant search results.” So the system is not reading only the exact phrase your visitor typed. It is running a small cloud of related questions and pulling pages for those too.

 

That single fact explains most of the advice you have read. If the answer is grounded in the index, then being in the index, for the question and its neighbors, is the price of entry. Everything else is about which of the eligible pages gets lifted.

 ![Google AI Overview for how to optimize for Google AI Overviews, showing the front-loaded answer and a cited Reddit thread in the source panel](https://www.wpconsults.com/wp-content/uploads/2026/07/ai-overview-for-how-to-optimize-for-google-ai-overviews-7779.avif)The AI Overview Google served us for this article’s own focus keyword, captured on 12 July 2026, US results. Google highlights the front-loading instruction inside its own answer, and the source panel (marked in red) shows a Reddit thread among the seven cited sites. 

## The first gate: is your page even eligible to appear in an AI Overview?

 

A page is eligible only if it is **indexed and allowed to show a snippet** in Google Search. Google’s [AI features documentation](https://developers.google.com/search/docs/appearance/ai-features) puts it in one sentence: “To be eligible to be shown as a supporting link in AI Overviews or AI Mode, a page must be indexed and eligible to be shown in Google Search with a snippet, fulfilling the Search technical requirements. There are no additional technical requirements.”

 

Read that next to the controls Google lists on the same page, and the implication is hard to miss. Google names `nosnippet`, `data-nosnippet`, `max-snippet` and `noindex` as the ways “to limit the information shown from your pages in Search.”

 

So if snippet eligibility is the gate, then a preview control you set years ago, or a plugin default you never looked at, can quietly keep a strong page out of the feature entirely. Every guide on this SERP tells you to front-load an answer. None of them tells you to check whether your page is allowed to be previewed at all.

 

Google has since added a section to that same doc called [Troubleshooting preview controls](https://developers.google.com/search/docs/appearance/ai-features#troubleshooting-preview-controls), and its whole premise is that these tags do remove your content from AI features, leaving only the question of whether yours is implemented correctly and has been recrawled. That is Google describing the causal link itself, which is a firmer version of the argument than the inference we started with.

 

Be careful how you hold this. It is Google describing its own machinery, which is the best evidence available and still a vendor claim, and we have not tested it yet. I am not telling you that `nosnippet` removes you from AI Overviews as a proven fact. I am telling you Google’s own documentation says it is an eligibility requirement, and that it costs you about thirty seconds to check.

 

### How to check your page for snippet eligibility

 

Open the page, view its source, and find the robots meta tag. Here is the real one from our own llms.txt post, the page we are running an experiment on later in this article:

 

```
<meta name="robots" content="follow, index, max-snippet:-1, max-video-preview:-1, max-image-preview:large">
```

 

`max-snippet:-1` means no length limit on the preview, so that page clears the eligibility gate. A `max-snippet:0`, a bare `nosnippet`, or a `noindex` in that same line would be the thing quietly keeping you out, no matter how good the writing is.

 

- **Read the robots meta tag first**, because that is where a blanket `nosnippet` usually hides, and it takes seconds. If you find one on a page you want cited, you have found your problem before you write a single new sentence.
- **Search your templates for `data-nosnippet`**, since it is applied at element level and is easy to leave wrapped around a block that happens to include your actual answer.
- **Check your SEO plugin’s global defaults**, not just the per-post settings, because most sites inherit their snippet controls from a site-wide setting nobody has opened since launch.
- **Confirm the page is indexed in Search Console** with URL Inspection, and read the crawled HTML rather than what your browser shows you, because what Googlebot fetched is the only version that counts.

 

If you have one specific page that ranks and is not cited, we turned these checks into a sequence you can run on a single URL in our [10-minute diagnosis for a page that is not cited in AI Overviews](https://www.wpconsults.com/page-not-cited-in-ai-overviews/).

 

## The second gate: which passage gets lifted into the answer

 

Among the pages that are eligible, the one that gets cited is the one carrying a passage that is easy to extract and safe to attribute. **Easy to extract** means the answer to the question sits in a single self-contained passage, directly under the heading that asks it, and still makes sense when it is pulled out of your page. **Safe to attribute** means that answer agrees with what other trusted sources are already saying.

 

I want to be honest about the status of this. The two-gate model is our working model, not documented mechanics. It comes from practitioner reverse-engineering work (Charles Floate’s is the sharpest I have read), it is consistent with the retrieval process Google describes, and our own captures line up with it. That is not the same as Google confirming it, and anyone telling you they know the selection algorithm is selling something.

 

What the model does explain is the case that annoys people most: the better page that never gets cited. If your answer is only assembled across three paragraphs, there is nothing clean to quote, so a thinner page with a quotable sentence wins the slot. And if your answer contradicts the settled facts every other source repeats, the safest thing the machine can do is leave you out, not correct the rest of the web.

 

That gives you a split worth internalizing. Put the corroborated answer in the passage you want lifted, and put your disagreement, your data, and your opinion in the body around it. The citation earns you the visibility, and the opinion earns you the click.

 

## Do you have to rank on page one to be cited in an AI Overview?

 

The clearest corroboration we have found for the two-gate model did not come from a study. It came from the feature itself, answering a question about itself.

 

On 13 July 2026 we captured the US AI Overview for the question “why isn’t my page cited in ai overviews”, signed out. Google’s answer highlighted this sentence inside its own panel:

 

> “AI engines do not simply cite top-ranking pages; they look for concise answers, verify authority, and favor extractable content.”

 

Line that up against the two gates. “Concise answers” and “extractable content” are the extractability half of gate two. “Verify authority” is the corroboration half. And the opening clause concedes the thing most of this industry assumes away: ranking gets you considered, not cited.

 

That matches the rule I now run on a page before I decide whether the passage work is even worth doing. Rank and extractability are two axes, and the useful thing is what happens where they cross.

 

> Rank on page one + extractable = **citation confirmed**.  
> Rank on page two or later + extractable = **citation possible**.  
> Rank on page one + not extractable = **citation not possible**.
> 
> [Abdullah Nouman (Author)](https://www.wpconsults.com/abdullah-nouman)![Abdullah Nouman (Author)](https://www.wpconsults.com/wp-content/uploads/2026/07/New-Profile-photo-150x150.avif)Abdullah Nouman (Author)eCommerce SEO and AI Search Marketing Specialist, founder of WpConsults[https://www.linkedin.com/in/nouman-abdullah/](https://www.linkedin.com/in/nouman-abdullah/)[https://github.com/nouman-abdullah](https://github.com/nouman-abdullah)[https://www.wpconsults.com/abdullah-nouman](https://www.wpconsults.com/abdullah-nouman)

 

The third line is the one worth acting on this afternoon, because it is entirely in your hands. A page that ranks and is never quoted is almost always failing on extraction, not on authority, and the owner is usually off building links. I unpack all three rows, with our own numbers against them, further down.

 

I will not over-claim it. An AI Overview is a generated summary of what the web says about Google, not a disclosure from inside Google, so this is the consensus talking back to us rather than a spec sheet. It is still a useful datapoint, and unlike every vendor statistic in this space, it costs you nothing to verify: run the query yourself.

 

Now the part that keeps the machine in proportion. The same AI Overview asserted that “AI models frequently extract citations from the first 150-200 words of an article”, with no source visible anywhere in the panel. That is the second precise-sounding, unsourced number this feature has handed us as advice. Do not build a strategy on a statistic an AI invented about itself.

 

## What our AI Overview captures showed about how a source gets picked

 

Across three captures of this article’s own focus keyword and its main variant, an AI Overview fired every time, the wording changed every time, and the cited sources changed every time. The advice inside it barely moved: answer first, in two to three sentences, under a question-shaped heading, then expand.

 

Method, so you can judge it: three captures on 12 July 2026, US results, in a real signed-in Chrome profile, which means the results are personalized. Three is a small sample and I will not dress it up as a study. Here is what was stable and what was not.

 

| Element | Stable or volatile | What we observed |
| --- | --- | --- |
| The front-loading instruction | Stable (3 of 3) | It was the first sentence of the answer, and Google highlighted it inside its own AI Overview. |
| “Answer First, Expand Later” as the named strategy | Stable (3 of 3) | Always the first bolded strategy, always with a two to three sentence length spec. |
| A structure or extraction point as strategy two | Stable (3 of 3) | Worded as structured data, structured formatting, or “structure for extractability”. |
| Reddit in the cited source set | Stable across these three | Present in all three, including a cited comment rather than a thread. |
| E-E-A-T and original data | Volatile | Present in one capture, absent from the other two. |
| Number of cited sites | Volatile | Seven, six, then ten, on the same topic within minutes. |

What held steady and what moved across our three AI Overview captures for the focus keyword and its variant, US results, 12 July 2026. 

The sharpest thing we found is in the source panel, not the answer. On the variant query, Google cited a one-year-old Reddit *comment*, three fragments long, sitting next to Google’s own developer documentation, while several polished guides that rank for the query were not cited at all.

 ![AI Overview source panel for ai overviews optimization showing a cited Reddit comment about short clear answers and small paragraphs](https://www.wpconsults.com/wp-content/uploads/2026/07/the-reddit-comment-google-lifted-into-its-ai-overview-7781.avif)The AI Overview source panel for “ai overviews optimization”, US results, 12 July 2026. Google is showing the passage it took: a Reddit comment reading “Short and clear answers. Small paragraphs. Links to sources”, marked in red. 

Look at what that comment is. It is the purest possible answer-first passage: self-contained, declarative, no context needed, and it agrees with what everyone else says. It cleared both gates without a domain name, a byline, or a content team.

 

Reddit has now appeared in the cited-source set of every AI Overview we have captured, across six different queries. The sample is small, so I am calling that a pattern we keep seeing, not a percentage and not a rule about what Google prefers.

 

One capture made the extractability point better than any of our own writing does. On the query “how to deindex a page”, the AI Overview lifted its answer from the first sentence of Google’s own documentation and attributed it to Google for Developers right inside the answer. That is both gates firing at once: the doc’s opening line is the most liftable passage on the topic, and Google’s own documentation is the safest thing on the internet to attribute.

 

The lesson is practical. On a documentation-shaped question you are not going to win the definitional slot, so stop aiming at it. Point your liftable passages at the questions the docs never answer: which method fits which situation, how long it really takes, how to verify it worked.

 

One more observation, because it keeps the machine in perspective. On a related query, the AI Overview instructed us to lead with a standalone answer of forty to sixty words. We had already banned that number in our own research notes: we found it circulating in an optimization guide with no study behind it, and we could not trace it to any source. The AI Overview is not an oracle, it is a mirror of the consensus, and it will hand you back an unsourced number as advice.

 

## The three stages of AI Overview optimization, at a glance

 

Getting cited takes three passes over a page, in a fixed order: **research** decides what the page must contain, **technical** decides whether it is allowed in at all, and **content** decides whether your paragraph is the one that gets lifted. Most advice covers only the third, which is why people do the writing work on pages that were disqualified before they started.

 

| Stage | What it decides | The failure it prevents | Time |
| --- | --- | --- | --- |
| **1. Research** | What the page has to contain, and what it will carry that nobody else does | Writing something good the machine has no reason to quote | An hour per page |
| **2. Technical** | Whether you are eligible at all, and which of your URLs owns the answer | Being disqualified silently, or competing with your own archive | One pass, site-wide |
| **3. Content** | Whether your passage is the liftable one | Having the best page and nothing quotable on it | A day per page |

 

Each stage can waste the next. There is no point climbing the specificity ladder on a page carrying a `nosnippet`, and no point fixing snippet controls on a query you sit at position 65 for. Run them in order and most pages fail in stage two, quietly, for free.

 

## How to optimize for Google AI Overviews, in three stages

 

Work it in three stages, in order: research decides what the page has to contain, technical decides whether you are allowed in at all, and content decides whether your paragraph is the one that gets lifted. Most guides only cover the third, which is why so many people do the writing work on a page that was disqualified before they started.

 

I run these in this order because each one can waste the next. There is no point climbing the specificity ladder on a page carrying a `nosnippet`, and no point fixing snippet controls on a query you sit at position 65 for.

 

### Stage 1: research, before you write a word

 

This stage answers one question: what does this page have to contain to be a valid answer, and what will it carry that the ranking pages do not? Finish it able to say both in one sentence each.

 

- **Do the keyword research, then widen it deliberately.** Google runs a cloud of related questions behind the one that was typed, so a page built on the head phrase alone loses to one that also answers the semantic, related and long-tail question-based versions. Map those variants and plan an H2 for each, on one page.
- **Check the feature even fires, and in which market.** An AI Overview does not appear on every query, and the cited set differs by country, so a citation strategy on a query that rarely triggers one is effort aimed at nothing. We pulled the coverage numbers apart in [how often AI Overviews actually appear](https://www.wpconsults.com/how-often-do-ai-overviews-appear/), including why every published figure disagrees with every other one.
- **Read the ranking pages for the consensus.** The settled answer they agree on, the entities they all name, and the shape that answer keeps taking. That is your eligibility brief. Contradict it inside the passage you want quoted and you are not corrected, you are skipped.
- **Read the same pages again for the gap.** The sub-question they all avoid, the number nobody verified, the table the answer obviously wanted and nobody built. An unanswered sub-question is an uncontested slot, because there is no consensus to break where nobody wrote anything.
- **Take a format census, not an impression.** If the ranking answers and the AI Overview itself keep coming out as numbered steps, ship numbered steps. Handing the machine an essay where it has already chosen a list is a self-inflicted wound.
- **Mine your own Search Console before you trust anyone else's data.** Your query rows tell you which phrasings you already earn impressions for and roughly where you sit on each. That is the difference between a page that needs one afternoon of claim-level work and a page that needs a year.
- **Pick the pages by proximity, not by ambition.** A page already near the pool is worth the passage work. A page at position 60 is not, and no amount of answer-first formatting changes that, because it never reaches the first gate. Our own numbers below are an uncomfortable example of exactly this.

 

### Stage 2: technical, so you are not disqualified in silence

 

This stage wins you nothing on its own. It stops you being ruled out before any of the writing matters, and every failure in it is invisible, which is what makes it worth a pass.

 

- **Snippet eligibility first.** The page must be indexed and allowed to show a snippet. Read the robots meta tag on the page, search your templates for `data-nosnippet`, and open your SEO plugin's global defaults rather than the per-post box, because most sites inherit a site-wide setting nobody has looked at since launch.
- **Read the crawled HTML, not your browser.** Confirm with URL Inspection what Googlebot actually fetched. What your browser renders is not the version being judged.
- **Give every liftable passage exactly one indexable home.** This is the one I learned the expensive way, and it has its own section below. If your category archives, listing pages or paginated views republish the same answer paragraph, you are competing with yourself for your own citation.
- **Point internal links at the page you want cited.** When a weaker URL on your own site carries similar content and more internal links than the article, the weaker URL is the one Google keeps. Link equity decides which of your own pages represents the topic.
- **Give every heading a stable ID**, so a section is individually addressable instead of an anonymous block of text.
- **Keep the plumbing honest.** Clean robots.txt, the sitemap declared, and a `lastmod` that reflects when you actually changed the page rather than when you published it.
- **Ship the schema graph, but know what it is buying.** One author entity tied to the organization, with article and image nodes resolving cleanly. It is hygiene and it helps machines resolve who you are; it is not the lever. We tested that separately and [schema markup on its own did not buy citations](https://www.wpconsults.com/schema-markup-ai-citations/).

 

### Stage 3: content, where the citation is actually won

 

Everything in this stage serves one object: the liftable passage. That is a self-contained answer of two to three lines, sitting directly under the heading that asks the question, which still makes sense with the rest of the page deleted.

 

Write one under every heading, then put each through four rules.

 

1. **Answer first, expand after.** An answer the reader assembles across three paragraphs gives the machine nothing to quote. This is the single most common reason a genuinely better page is ignored, and it is the only variable we have watched move a page into a cited set.
2. **Match the format the answer already takes.** A table where the ranking answers use a table, real numbered steps where the answer is a procedure. You are handing over something in the shape it has already decided to extract.
3. **State the corroborated answer inside the liftable passage.** Keep that passage factually mainstream, and put your counter-read, your data and your verdict in the body around it. Corroboration binds the facts you want quoted, never your conclusion.
4. **Climb the specificity ladder.** A quantifier, then the named thing, then the literal string on screen, then the verified number. Correct and vague loses to correct and quotable every time, and consultant altitude is not quotable at all.

 

Then add the one thing that is yours. Your own measurement, a test you ran, a real screenshot, an honest caution nobody else gives. Matching the consensus makes you eligible; the extra value is the only reason anything, machine or human, picks you over an incumbent it already trusts.

 

One limit travels with the ladder, and it is not optional. Verify every rung at the primary source. I went to lift a figure the AI Overview itself was asserting, could not find it in the official documentation anywhere, and dropped it. Specificity only helps while it is true, and the fastest way up that ladder is to invent something.

 

Notice what is not in any of the three stages. No special file, no markup that buys a citation, no rewriting your prose into a robot dialect. Google is explicit about that, and it is worth quoting properly.

 

## Do you need schema, llms.txt, or markdown files for AI Overviews?

 

No, none of them is required to appear in an AI Overview, and Google says so directly. Almost all of them are still worth shipping, because "Google Search does not use it" answers a much narrower question than the one you are asking.

 

Here is where I actually land on each, from running them on our own properties rather than from reading the guide.

 

| Thing | My verdict | What it actually does, and when |
| --- | --- | --- |
| **Structured data / schema** | **Ship it. Not a ranking factor, still load-bearing.** | It is how machines resolve *who you are*: one author entity tied to the organization, article and image nodes resolving cleanly. That identity work is what the corroboration side of citation leans on, and it earns rich results in classic Search on its own. We tested it narrowly as a citation lever and [it bought us no citations](https://www.wpconsults.com/schema-markup-ai-citations/). Both things are true: it is not the trigger, and skipping it is still a mistake. |
| **llms.txt** | **Ship a short curated one if agents might plausibly fetch from your domain.** | Treat it as cheap B2A infrastructure, not a ranking trick. It costs about half a day, it is genuinely useful in the agent layer, and if an answer engine ever flips the switch you are already there. It matters most for developer docs and for stores moving toward agentic checkout. Our full read, with the numbers: [does llms.txt work for SEO](https://www.wpconsults.com/does-llms-txt-work-for-seo/). |
| **Markdown versions of your pages** | **Yes, and this is the underrated one.** | Language models read markdown far more cleanly than HTML buried in navigation, scripts and cookie banners. Serving a clean `.md` alongside each URL hands them the words instead of making them dig for the words. It is what our own AgentReady plugin does, and its crawler log is how you find out which systems are actually fetching you rather than guessing. |
| **Answer-first chunks** | **We do this, it works, and Google not requiring it is beside the point.** | Google is answering "must you shred your site into fragments?" The answer is no, and shredding is actively harmful, because duplicated answers compete with each other. The practitioner question is different: *should the paragraph you want quoted stand on its own?* Yes. Rewriting the first two to three lines under each heading is the single change we made on our own page before it entered a cited set. |
| **Mentions across the web** | **They work, and they work massively. Just do not manufacture them.** | Google warns against seeking inauthentic mentions, and that warning is correct. Do not read it as "mentions do not matter". Being said by more independent sources is *literally* the corroboration gate: it is what makes your claim safe to quote. Earn them, never buy them, and never spin up fake ones. |
| **Honest freshness signals** | **Get this right or it costs you site-wide.** | A wrong `lastmod` is worse than no `lastmod`. WordPress bumps the modified date on every save, your SEO plugin copies it into the sitemap, and Google reads `lastmod` across the whole property, so a site that keeps announcing updates that never happened can have its dates distrusted everywhere. Other AI systems will often still take the bait, which is the worst of both worlds. Real freshness does real work, so serve real dates: that is why we built [Honest Lastmod](https://www.wpconsults.com/downloads/honest-lastmod-pro/). |

 

### AI Overviews are not the only system reading your pages

 

Google Search is one consumer of your content, not all of them. ChatGPT, Claude, Perplexity and Gemini reach the web through their own crawlers and their own retrieval, and several of them will happily take a markdown file or an `llms.txt` that Google ignores completely.

 

So a flat "Google does not use it" is accurate and often irrelevant. The question is never whether Google requires it, it is which system you want to be readable to, and what that costs you.

 

On llms.txt specifically, [John Mueller](https://developers.google.com/search/blog/authors/john-mueller)![John Mueller](https://www.wpconsults.com/wp-content/uploads/2026/06/john-mueller-150x150.avif)John MuellerSearch Advocate, GoogleGoogle's Search Advocate and the main on-the-record voice of Google Search Relations.[LinkedIn](https://www.linkedin.com/in/johnmu)[X](https://x.com/JohnMu)[Search Central](https://developers.google.com/search/blog/authors/john-mueller) and Gary Illyes have both been asked, and their answers line up with the documentation: Google Search does not consume the file. We quote them properly in the llms.txt piece, and **we agree with them on the fact**. Our own reading found the same thing on the Search side, so there is no argument to have there.

 

**Where we part company is the conclusion drawn from it.** They are answering for Google Search. The industry then repeats it as "llms.txt is useless", which is a claim about every AI system on the internet, and neither of them made it.

 

## How many pages should you build for query variants?

 

Google’s guide also draws a line the rest of this SERP walks straight over. In its own words: “While it might be tempting to create separate content for every possible variation of how people might search (for example, by focusing on other queries that people have asked, or fan-out queries), doing so primarily to manipulate rankings or generative AI responses in Google Search violates Google’s scaled content abuse spam policy.”

 

Meanwhile several ranking guides on this exact keyword tell readers to target question variants and check which ones trigger an AI Overview. The distinction is intent and duplication, not the existence of question-shaped headings.

 

Covering the sub-questions your readers genuinely have, properly, inside one strong page is content work and Google’s own guide asks for it. Minting a thin page per query variant to farm AI answers is the thing it calls spam. Same line, two very different behaviors, and you know which one you are doing.

 

While we are being honest about this SERP: on the day we captured it, four sponsored results from AI-visibility tools sat above the first organic listing, and the first organic listing was Google’s own documentation.

 ![US Google SERP for how to optimize for Google AI Overviews, four sponsored AI visibility tools above the first organic result](https://www.wpconsults.com/wp-content/uploads/2026/07/sponsored-versus-organic-on-the-ai-overviews-serp-7780.avif)The same US SERP, 12 July 2026. Everything inside the red box is paid (four sponsored AI-visibility tools); the green box is the first organic result, Google’s own guide. Sponsored results are marked here so nothing in this screenshot is mistaken for a ranking. 

Keep that in mind when you read the statistics in this space. The widely quoted numbers on AI Overview citations, the top-10 correlation studies, the citation-share tables, are almost all published by companies selling AI-visibility products. Attribute them, do not adopt them.

 

## What our own Search Console data shows, and what it cannot show

 

Our Search Console data shows the pattern everyone is arguing about: pages holding decent positions while their clicks collapse. It also shows the limit of that argument, because Google’s Search Analytics API gives us no AI Overview breakout, so we cannot prove what absorbed those clicks.

 

Zoom out to the whole property and the shape is plain: 18.1K impressions, 75 clicks, a 0.4% click-through rate over 28 days.

 ![Google Search Console performance report for wpconsults.com over 28 days, 18.1K impressions, 75 clicks, 0.4% CTR at average position 17.2](https://www.wpconsults.com/wp-content/uploads/2026/07/our-own-search-console-performance-28-days-7799.avif)Our own Search Console, last 28 days, whole property: 18.1K impressions, 75 clicks, a 0.4% CTR at an average position of 17.2 (boxed in red). Search Console’s own view runs a couple of days behind, so the page-level table above was pulled from the same property through 12 July. 

Now the honest part, and it is the reason I am showing you the whole report instead of one flattering row. An average position of 17.2 explains a low CTR all by itself, because page two was never going to get clicks.

 

A page sitting on page one with a thousand impressions and a single click looks like the AI Overview story in one row. It could equally be a featured snippet, a video carousel, a weak title, or plain irrelevance. I can publish the number honestly; I cannot publish the cause, and neither can anyone else using the same API.

 

The second thing our data says is less comfortable. On the AI-search terms we care about, we are nowhere near the pool: “best geo tools for ecommerce” sits at an average position of 65, “ai commerce readiness” at 62, “agentic commerce readiness” at 57. No amount of answer-first formatting fixes that, because we never reach the first gate on those queries.

 

If you want to measure the traffic side of this properly rather than guess, our guide on [measuring AI Overview traffic loss](https://www.wpconsults.com/ai-overview-traffic-loss-measure/) is the companion piece to this one.

 

## The AI Overview citation pool is not the same in every market

 

The set of sources an AI Overview cites changes by country, on the same query, at the same moment. We captured one query in four markets within the same minute, signed out and at country level, and got four different pool sizes and four different cited sets.

 

| Market | AI Overview fired | Sites cited | Examples of what it lifted |
| --- | --- | --- | --- |
| United States | Yes | 8 | A LinkedIn post, a Reddit thread |
| Canada | Yes | 4 | SE Ranking, Ahrefs |
| United Kingdom | Yes | 5 | A YouTube video, SE Ranking |
| Australia | Yes | 7 | A YouTube video, SE Ranking |

One query (“does llms.txt work for seo”), four markets, captured within the same minute on 12 July 2026, signed out. Four different citation pools. This is an observation on four captures of a single query, not a percentage. 

Take that seriously before you buy an AI-visibility dashboard. Every claim about whether a brand is cited carries a hidden market parameter, and most tools do not tell you which market they measured.

 

The pool also moves on its own within a single market. On one query, captured four hours apart on the same day, we counted 11 cited sites and then 9, with nothing on our page changed in between. So before you conclude you were dropped, capture the thing twice.

 

It also kills a comfortable idea: there is no position one in an AI Overview. There is only a probability of being cited, and it moves with the wording, the refresh, and the country.

 

## How much does the AI Overview citation pool change on its own?

 

Enough to ruin a before-and-after test. Across twenty-two nightly reads of one query in four markets, a site that had been cited in every market for four straight nights disappeared from all four for the next ten, and two fetches twenty-seven seconds apart came back with different cited sets.

 

None of that had anything to do with us. It is the pool rearranging itself while we stood there counting, and it is the context every citation claim in this article has to survive, including our own.

 

### How stable is the AI Overview citation pool?

 

Across reads 8 to 22 of one query, the number of hosts cited in all four markets at once went 1, 2, 3, 3, 4, 3, 5, 3, 4, 4, 2, 2, 4, 5, 3. That is not a trend in either direction, it is noise, and reading it as a trend would be the same mistake we accuse the vendor studies of making.

 

The membership turned over almost completely inside that window. SE Ranking held all four markets for four consecutive reads, then left every market for ten straight reads before returning to all four; Contentful held all four for three reads; Ahrefs held all four for exactly one.

 

A flat "everything churns" would be selective, though. Search Engine Journal sat in the all-market set for ten consecutive reads without moving. So the pool has a stable core and a churning edge, and the practical point is that almost everyone, us included, is on the edge.

 

### Twenty-seven seconds was enough to change the answer

 

On one night’s Australia check the same query was captured twice, twenty-seven seconds apart. The first read returned seven cited hosts and the second returned the same seven plus one more.

 

Same market, same query, same instrument, both reads verified. One host of difference in under half a minute.

 

I am not putting a rate on that, because it is a single pair. It is enough for the practical point: any before-and-after that rests on one fetch is measuring the pool’s mood, not your page.

 

### Why our page lost its AI Overview citation, and the wrong cause we blamed first

 

On 29 July our llms.txt page fell out of the cited set in all four markets, ending a four-night run. On the same fetch three vendor sites returned to all four markets at once, so the pool visibly rearranged itself on the same night we vanished.

 

We wrote that up as the pool moving and taking us with it. It was the tidiest available explanation, every control on the page was clean, and it fitted the churn we had just spent a section documenting.

 

It was wrong, and the real cause was sitting on our own site. I have left the reasoning here rather than quietly deleting it, because reaching for ambient volatility to explain your own result is the exact failure this article keeps warning about, and I walked into it. The full diagnosis is below.

 

### Our own cited pages do not move together

 

The most useful thing in that night’s data is that our three watched pages went three different ways at once. The DuckDuckGo page held its Canada and Australia citations for a fifth straight read, the mermaid page picked up its first ever United States citation, and the llms.txt page lost all four markets.

 

Same night, same instrument, three directions. So whatever is moving these citations is not one site-level thing that switched on for our domain, it is happening per query.

 

That matters because “we became citable” is the story almost every write-up of this kind reaches for, and our own data argues against it.

 

### How to read a citation change without fooling yourself

 

- **Never call a change from a single capture.** We watched the cited set move in twenty-seven seconds, so one fetch before and one fetch after tells you almost nothing about what your edit did.
- **Hold the market constant and say which one.** The pool is different in every country, so a citation gained in Australia and a citation lost in the United States are two separate measurements, not a net change.
- **Watch a page you did not touch.** Without an untouched page in the same window, a pool that widened for everyone looks exactly like your edit working, and you will never be able to tell them apart.
- **Count reads, not days.** A missed night is a missing measurement, not a zero, and treating it as a zero is how a study quietly flatters itself.
- **Read the whole pool, not just your own row.** Who else entered and left on the same night is usually the explanation, and it is sitting right there in the same panel you are already looking at.

 

## The extractability test we ran on our own site, and what it took to get cited

 

We tested the second gate on our own property, and we wrote down what would count as being wrong before we knew the answer. That is the part I have not seen anyone else do on this topic.

 

The setup was clean. Our post on llms.txt ranks (the page averaged position 8.2 over the 28 days before the edit), the AI Overview fires on the query, its answer matches our answer, and we were not cited.

 

So we changed exactly one thing: the opening two to three lines under every H2 became direct, self-contained answers, with the supporting detail moved below.

 

Held constant, deliberately: the thesis, every fact and figure, every source, every link, the headings, the title, the meta, the images, the categories. No new links, no ranking work of any kind. If we had touched a second variable, the result would mean nothing.

 

We predicted two outcomes before we knew the answer. A positive result would be the page appearing in the cited-source set on a later capture, real evidence that extractability can move a page into the citation pool. A null would be it never appearing over the watch window, which would say the binding constraint is corroboration and authority rather than formatting. We said we would publish it either way, and what we got was more interesting than either prediction.

 

### First, nine straight nights of nothing

 

For nine consecutive verified checks, our edited page was not cited in the AI Overview, in any of the four markets we watched. The AI Overview fired every time. Our page was not among the sources it lifted, not once.

 

The edit went live on 12 July. Google refetched the page once on 13 July and then did not fetch it again, so every one of these checks reads the exact same served version of the page.

 

The page kept ranking the whole time. Its page-level position hovered in the low 20s after sliding down from that 8.2 baseline, so it never fell out of contention. On 23 July, reading only those nine nights, we wrote this section as an honest null and we meant it. The data simply had not arrived yet.

 

### Then, on the tenth read, all four markets at once

 

On 25 July the page entered the AI Overview cited set for the query in all four markets on the same night, and it held all four for four consecutive nights, through 28 July.

 

The strange part is what did not change. The page’s `last_crawled` was still 13 July through the whole run, frozen for every read, so Google was serving the exact same page it had served through all nine not-cited nights. Nothing on our side moved between the failure and the entry.

 

The ranking did not recover either. Page-level position sat around 20.9 and stayed flat in the low 20s across the cited nights, so this is not a quiet ranking-recovery story wearing a citation costume. The page was cited from the low 20s, not from a jump back to page one.

 ![Google AI Overview for does llms.txt work for seo with the WpConsults source card visible in the cited sources panel](https://www.wpconsults.com/wp-content/uploads/2026/07/wpconsults-cited-in-the-ai-overview-for-does-llms-txt-work-f-8257.avif)The cited sources panel for *does llms.txt work for seo*, with our card in it. Google is showing the passage it took, which is the opening of the section we rewrote. Screenshot: Google, marked by WpConsults. 

Look at what the panel is quoting, because it is the clearest thing in this whole article. Google is not showing our headline or our conclusion. It is showing the two opening lines under a heading, verbatim:

 

> llms.txt is sold as an AI-visibility must-have. Across 515 million AI bot requests it was touched 408 times.
> 
> The passage Google lifted from our page

 

That is the specificity ladder and the liftable passage in one sentence: a named claim, a verified number, and a second number that makes the first one mean something, all sitting where a machine can take them without reading the rest of the page.

 

### Then the fifth night took the citation back, and the cause was on our own site

 

On 29 July the page was not cited in any of the four markets, and within days it had stopped ranking for the query as well. Not a position slide, gone from the results.

 

The page itself was innocent on every control we had. It had not been edited since 12 July, its last crawl was still 13 July, and its position had read 20.8 against 20.9 and 20.6 on the two checks before. Search Console still reported the URL as indexed the whole time, which is why it took us a while to accept that something had actually happened.

 

What had happened was that three other URLs on our own property were serving the same answer better than the article was.

 

- **An unrelated page of ours carried a near-copy of the passage.** It was not written as a competitor to the article; it simply covered enough of the same ground to look like the same answer.
- **Our category archive published the excerpt,** which meant the opening lines, the exact liftable passage we had engineered, existed on a listing URL as well as on the article.
- **The article had fewer internal links than the pages competing with it.** That was the tiebreak. Given several URLs on one domain carrying the same answer, Google kept the ones our own linking said were more important, and the article was not one of them.

 

So we built one perfect quotable paragraph and then published it on several of our own URLs, while pointing our internal links at the copies. We competed with ourselves for our own citation and we lost.

 

> I built the perfect quotable paragraph, then published it on four of my own URLs and pointed my internal links at the copies.

 

The rest of the arms are the control on this reading. Our other watched pages kept their citations on their own queries through the same window, which is what rules out a domain-level penalty, a site-wide event, or the pool simply closing. One query broke, and it broke for a reason we caused.

 

Here is that control stated properly, because "the other pages were fine" is the kind of claim that should come with the panel behind it. These are our own cards inside Google's cited sources, on four different queries.

 

| Query | Our page | Status through the window |
| --- | --- | --- |
| *does llms.txt work for seo* | the llms.txt post | **Lost it.** Four markets, four nights, then out, then out of search entirely |
| *duckduckgo submit url* | the DuckDuckGo guide | **Held.** Cited and still cited |
| *mermaid diagram wordpress* | the Mermaid diagram post | **Held, unevenly.** Cited, flickering between markets |
| *url is not in property* | the Search Console error fix | **Held.** Cited, and not part of any edit test |

 ![Google AI Overview for url is not in property with the WpConsults source card boxed in the cited sources panel](https://www.wpconsults.com/wp-content/uploads/2026/07/wpconsults-cited-in-the-ai-overview-for-url-is-not-in-proper-8260.avif)A page we never touched for any test, cited on *url is not in property* during the same window the llms.txt page was disappearing. Screenshot: Google, marked by WpConsults. 

Three of four held while one collapsed. A domain does not lose visibility on one query and keep it on three, so whatever happened was about that URL, and it was.

 

We have since fixed the duplication and the internal linking. My expectation, stated in advance so it can be held against me, is that the page returns to the results and then to the cited set. If it does not, that outcome gets published here too.

 

The symptom worth memorizing is the shape of it: a page leaving the index rather than sliding down it. That is a different problem with a different fix, and we have written up [how to read crawled, currently not indexed](https://www.wpconsults.com/crawled-currently-not-indexed/) and [what to do when your indexed page count falls](https://www.wpconsults.com/search-console-indexed-pages-decreased/) for anyone standing where we were.

 

### Is extractability a switch you can flip to get cited?

 

Extractability is necessary and it is not sufficient. We changed only that, and nine nights of nothing say plainly that the edit alone did not summon a citation; the page also had to wait for the pool's aggregate answer to widen enough to include a page sitting in the low twenties.

 

What the 29 July lapse taught us is a different thing, and it is more useful. Extractability is not only something you add to a page, it is something you can accidentally spread across several pages. The gate is not passed by writing a liftable passage; it is passed by having exactly one indexable URL that carries it.

 

Both halves matter. We controlled whether the passage was liftable, and we did not control which night the pool had room, but we also controlled something we did not realize we were controlling: which of our own URLs Google would treat as the home of that answer.

 

### Rank plus extractability: the rule for predicting whether a page can be cited

 

Rank and extractability are not two competing theories of citation, they are two axes, and the useful thing is what happens at each intersection. This is the working rule I run on a page before I decide whether the passage work is worth doing.

 

| Where the page ranks | Is the answer liftable? | What to expect |
| --- | --- | --- |
| Page one | Yes | **Citation confirmed.** Both gates are clear. If it is not being cited, look for a disqualifier rather than a writing problem. |
| Page two or later | Yes | **Citation possible.** You are eligible and you are waiting for a night with room. This is where most of our own results live. |
| Page one | No | **Citation not possible.** Ranking cannot rescue an answer there is nothing to quote from. This is the fixable one, and it is a day of work. |

 

The third row is the one worth acting on today, because it is entirely in your hands and it is the case people misdiagnose most often. A page that ranks and is never quoted is almost always failing on extraction, not on authority, and the owner is usually off building links.

 

The second row is where patience belongs. Our llms.txt page was cited from the low twenties, and our measurement below found that 61% of cited sites were not ranking on page one at all, so page two is emphatically not out of the running. It just means the citation is a probability rather than a seat.

 

Hold the first row loosely, because "confirmed" describes the conditions, not a guarantee for a given night. Our own experience is that page one plus a genuinely liftable passage is the combination that keeps a citation, and that everything below it keeps one intermittently. That is a working rule drawn from a handful of our own pages, not a law, and I would rather you argue with it than adopt it.

 

### Two more pages did the same thing on different queries

 

We ran the same one-variable edit on a second page, our post on submitting a URL to DuckDuckGo, and it entered the AI Overview cited set on its own query in Canada and Australia, then held there for five consecutive reads. On the latest read the Canada citation is confirmed on the screenshot itself, with our source card visible in the panel.

 ![Google AI Overview for duckduckgo submit url with the WpConsults source card marked in the cited sources panel](https://www.wpconsults.com/wp-content/uploads/2026/07/wpconsults-cited-in-the-ai-overview-for-duckduckgo-submit-ur-8258.avif)A post from 2024 that ranked for two years and was never quoted. After the claim-level pass it entered the cited set on *duckduckgo submit url*. Screenshot: Google, marked by WpConsults. 

A third edited page, our mermaid diagram post, behaved differently and that difference is worth having. It entered the cited set after its edit but flickered between markets read to read instead of holding, and on the latest read it produced its first ever United States citation.

 ![Google AI Overview for mermaid diagram wordpress with the WpConsults source card marked in the cited sources panel](https://www.wpconsults.com/wp-content/uploads/2026/07/wpconsults-cited-in-the-ai-overview-for-mermaid-diagram-word-8259.avif)The third edited page, cited on *mermaid diagram wordpress*. This is the arm that flickers between markets rather than holding, which is why it is in the article at all. Screenshot: Google, marked by WpConsults. 

Alongside them we watch a page we deliberately never touched, and it is the comparison the design exists to make. On both of the two reads where we captured its control query, it was cited in none of the four markets.

 

That caveat has to travel with it, though. We did not capture that query on the read before those two, so what I can tell you is what the control query read on the nights we actually ran it, not that the control page “dropped out” of anything.

 

Three edited pages, three different queries, two separate watch studies. Two of them are still in their cited sets and one is not, which is both the strongest support these tests have produced and the clearest sign of how small they are.

 

Here is what this does not prove, because a small result oversold is worse than no result at all.

 

- **It is small.** One page per query, one operator, and a four-night run on the main arm that then lapsed. I will not put a percentage on a result this size, and I am not stretching it to queries or markets we did not test.
- **It does not say formatting is the lever.** On these pages the citation pool had to move as well, and that is the half you cannot format your way into. The shape of your opening paragraph is necessary here, not sufficient.
- **The three arms do not move together.** On the same night one held, one gained a new market, and one lost every market, so there is no single site-level effect here to point at, only per-query behavior.
- **It is a single-operator case study,** labeled as one, never a controlled trial. The recrawl that started the clock was requested, not natural, and we are telling you that on purpose.

 

One correction against ourselves, since it matters for how you read the above. Search Console returns no query-level rows for that page, because its impressions sit under Google’s anonymization threshold. We can prove the page averages position 8.2 across the queries it appears for; we cannot prove it ranks 8.2 for that exact phrase. Anyone who tells you otherwise about their own page is probably not looking closely.

 

## Where information gain and citation eligibility pull against each other

 

There is a tension in this advice that nobody admits, and I would rather name it. Being the only site with a better number can be the reason you do not get cited.

 

In our market captures, the AI Overviews lifted a figure that several sources repeat. Our own post’s central statistic comes from a source nobody else uses, and it measures the thing more directly. It is the better number, and it is uncorroborated by construction.

 

If corroboration is really the second gate, then the passage carrying our unique evidence is exactly the passage that is risky to attribute. That is a live hypothesis, not a finding, and our test is not designed to answer it yet. But it tells you where to place things: corroborated facts in the liftable answer, unique evidence in the body where it does its real work on a human reader.

 

## So, how should you actually optimize for Google AI Overviews?

 

Spend the first hour on the boring half. Confirm the pages are indexed and snippet-eligible, confirm the answer passage lives on exactly one of your URLs, and confirm you rank for the question and its neighbors at all. Those three checks decide whether anything else you do can matter, and two of the three are stage two.

 

Then rewrite the openings, not the site. One clean, self-contained answer under each question heading, agreeing with what is settled, followed by the depth and the argument that make the page yours. That is a day of work on a good page, and it is the only lever in this article we have watched the machine respond to at all.

 

Hold the result loosely. The citation it earned us lasted four nights, and then we took it away from ourselves by publishing the same passage on our own archives. Treat a citation as weather you can influence rather than a position you can hold, and you will make better decisions than anyone selling you a score.

 

What I would not do is buy the idea that this is a new discipline with its own file formats and dashboards. Google says there is no special markup and no AI file, its own AI Overview says ranking will not get you cited, our captures show a Reddit comment beating polished guides, and the tools selling you a citation score are advertising above the very query you used to find them. Do the fundamentals, make one answer per question liftable, keep it in one place, and put something in the page nobody else can copy.

 

## We measured it: 61% of the sites an AI Overview cites do not rank on page one

 

Of 99 cited sites across 15 queries, **39 ranked in the organic top 10 for the same query, and 60 did not**. That is 39% in, 61% out, measured on our own captures rather than borrowed from a vendor study.

 

So the claim you hear everywhere, that AI Overviews mostly cite what already ranks, is wrong as stated. It is not nothing, though: 39% is far above chance, so ranking clearly helps.

 

The honest reading is that **ranking is neither necessary nor sufficient**. It is one input into the citation decision, not the gate in front of it.

 

Here is the method, because a number without one is just an opinion with a decimal point. Fifteen queries in our lane (technical SEO, WordPress, GEO), United States, signed out. For each one we waited for the AI Overview to settle, expanded it, opened the sources panel, and scrolled that panel until the list stopped growing before reading anything from it.

 

That last step is not a detail. The panel lazy-loads, so reading it too early gives you a short list that looks complete, and Sponsored results were excluded from the organic top 10 on every row.

 

### The average hides the only interesting thing in the data

 

The overlap swings from 0% to 100% depending on the query. On one query the AI Overview cited eight sites and not one of them ranked; on another it cited two and both did.

 

| Query | Cited sites that also rank top 10 |
| --- | --- |
| [elementor](https://be.elementor.com/visit/?bta=202863&nci=5732) back button | 2 of 2 (100%) |
| crawled currently not indexed | 4 of 6 (67%) |
| schema markup for ai citations | 3 of 5 (60%) |
| should i use a disavow file | 4 of 7 (57%) |
| does llms txt work for seo | 2 of 4 (50%) |
| wordpress migration seo checklist | 5 of 10 (50%) |
| what is entity seo | 3 of 6 (50%) |
| how to measure ai overview traffic loss | 3 of 6 (50%) |
| why is my page not indexed | 3 of 6 (50%) |
| how to fix canonical url not in property | 1 of 2 (50%) |
| internal linking audit wordpress | 4 of 11 (36%) |
| core web vitals inp wordpress | 3 of 9 (33%) |
| google merchant center requirements | 1 of 3 (33%) |
| how to deindex a page | 1 of 14 (7%) |
| **duckduckgo sitemap submission** | **0 of 8 (0%)** |

The same measurement, query by query. A single blended AI-visibility score averages this spread away, which is why one number for your whole site tells you almost nothing. 

Two more things fell out of the same dataset, and both cut against the tidy version of this story.

 

**When a cited site does rank, it is not clustered at the top.** The median position of those 39 was 5, and only 13 of them sat in the top 3. So even the ranking half does not support the idea that the AI Overview simply lifts the top three results.

 

**YouTube was cited on 10 of the 15 queries and Reddit on 6.** With Quora, user-generated and video sources account for 18 of the 99 citations. For text-first technical queries, that is a lot, and it says something about what Google is willing to treat as a source.

 

### What this measurement does not prove

 

N is 15 queries, one market, one day, one capture each. That is an observation, not a law, and I would rather you quote the N than the percentage.

 

The pool also moves. During verification we watched one query’s cited set change within minutes, while another stayed frozen the same day in the same market, so any before-and-after citation claim, including any I make later, has to survive that churn.

 

And one confession, because it is the most useful part. **Our first pass at this number came out at 10%**, and it was wrong. The tool was reading the sources panel while it was still lazy-loading, so the cited lists were short.

 

The wrong number flattered our own argument, which is exactly the direction a broken instrument is most dangerous in. Once the reader waited for the panel to stop growing, 10% became 39%, and we published the number that hurt the thesis because it is the one the data supports.

 

The practical takeaway is not the percentage anyway. It is that chasing rank alone will not get you cited, and being absent from page one does not put you out of the running, so the passage-level work in the rest of this guide is where your effort actually goes.

  

### Not sure whether your pages are even eligible?

 

If you want a second pair of eyes on your snippet settings, your rankings, or the pages you think an AI Overview is quietly eating, [contact us](https://www.wpconsults.com/work-with-wpconsults/) or [email me](mailto:abdullah@wpconsults.com) and I will take a look. Most of the wins here are checks, not rebuilds.

   

## Update Logs

 

**1 Aug 2026**

 

- Corrected the reason our llms.txt page lost its citation. We had blamed the citation pool moving; the real cause was on our own site, where a related page and a category archive excerpt were republishing the same answer passage and the article carried fewer internal links than either. The wrong explanation is left in place, labeled, because reaching for ambient volatility to explain your own result is the failure this article warns about.
- Rebuilt the optimization advice into three stages, research, technical and content, with the strategies that earlier versions skipped: checking whether the feature fires at all, taking a format census, picking pages by proximity rather than ambition, giving each answer passage one indexable home, and pointing internal links at the page you want cited.
- Added the rank and extractability rule as a table: page one plus liftable is a citation you keep, page two or later plus liftable is a citation you sometimes get, page one without a liftable passage is not a citation at all.
- Cut the read-by-read churn table and the four unrelated Search Console rows, neither of which was doing work for the reader.

 

**29 Jul 2026**

 

- Added a section on how much the AI Overview citation pool moves on its own, with our own read-by-read count of which sites were cited in all four markets at once. Updated the extractability result: the llms.txt page held its four-market citation for four nights and then lapsed on the fifth with nothing changed on the page, the DuckDuckGo page held a fifth straight read, and the control page’s wording was tightened to exactly what we captured.

 

**28 Jul 2026**

 

- Revised the extractability test from a null to a confirmed result. After nine verified not-cited nights, the edited llms.txt page entered the AI Overview cited set in all four markets on the tenth read and held there for four straight nights, while a second edited page did the same on a different query. Rewrote the result as a two-gate finding: the answer-first edit was necessary, but the page also had to wait for the citation pool to widen around it, which is the half we did not control.

 

**24 Jul 2026**

 

- Wrote up the extractability test as a null after nine straight not-cited checks across four markets. That reading was honest at the time; the very next night’s data revised it, which is logged above.

 

**13 Jul 2026**

 

- Added Google’s own AI Overview conceding that AI engines do not simply cite top-ranking pages, captured on 13 July, along with the unsourced statistic it asserted in the same answer. Added the documentation-query capture that shows both gates firing at once, the within-market churn observation, and a link to our new diagnosis piece.

 

**12 Jul 2026**

 

- First version, written from our own AI Overview captures, Google’s current documentation, and our Search Console data. The extractability test on our llms.txt post is running; we will publish the result here either way.
