
Schema markup gets sold as an AI-citation cheat code: add JSON-LD, and ChatGPT or Google’s AI Overviews will start citing you more. The largest controlled test we have says it barely moves citations on pages that are already visible to AI.
That does not make schema useless; it just is not the lever it is marketed as. Here is what the data actually shows, why the “3x more likely” stat fools people, and what really earns the citation.
Key Takeaways
- The biggest controlled test (Ahrefs, 1,885 pages) found adding schema produced no meaningful citation lift: Google AI Mode +2.4%, ChatGPT +2.2%, AI Overviews -4.6%.
- The popular “AI-cited pages are about 3x more likely to have schema” stat is correlation, not cause.
- Every page studied already had 100+ AI Overview citations, so the test only answers whether schema lifts already-visible pages, and it does not.
- In a separate test, five major AI systems read only the visible HTML during retrieval and ignored JSON-LD entirely.
- Schema still earns its place for rich results, entity disambiguation, and knowledge-graph signals; just not as an AI-citation hack.
Where the “schema equals citations” idea comes from
The claim has a real observation underneath it. When Ahrefs looked across about 6 million URLs, pages cited by AI were almost three times more likely to have JSON-LD than pages that were not, and roughly 53% of AI-cited pages carried schema. That is the stat you see in LinkedIn carousels, and on its own it sounds like a settled case.
The problem is that a correlation like this cannot tell you whether schema caused the citations or just happened to sit on the kind of pages that get cited anyway. To separate those, you have to actually add schema and watch what changes, which is the harder test most people skip.
What happened when someone isolated the effect
Ahrefs ran that harder test and published it as a study of 1,885 pages that added JSON-LD between August 2025 and March 2026, matched against around 4,000 control pages with similar citation histories. Using a difference-in-differences analysis (the method that strips out platform-wide trends), here is what adding schema did:
| AI source | Effect on citations after adding schema |
|---|---|
| Google AI Mode | +2.4% (indistinguishable from zero) |
| ChatGPT | +2.2% (indistinguishable from zero) |
| Google AI Overviews | -4.6% (small but statistically significant decline) |
The AI Mode number is worth dwelling on. A naive before-and-after showed treated pages up about 43%, which looks like a win until you see the control pages rose almost as much; AI Mode was simply expanding for everyone. Strip that out and the schema-specific effect shrinks to +2.4%, which is noise.
The AI Overviews decline is real but small, and both treated and control pages were already sliding before any schema was added, so I would not read it as “schema hurts you.” The honest summary is the one the study reaches: not much changed either way.
The caveat almost nobody quotes
There is a detail that decides how far you can take this: every page in the dataset already had more than 100 AI Overview citations before any schema was added. These were pages AI already knew about and surfaced.
So the study answers a precise question well: does schema push an already-visible page higher? No, not meaningfully. What it cannot tell you is whether schema helps an invisible page get crawled and considered in the first place. That is a genuine open question, and anyone quoting a clean percentage for it is guessing.
A separate experiment from searchVIU adds a useful piece here: when they checked whether ChatGPT, Claude, Perplexity, Gemini, and Google AI Mode used schema while fetching a page in real time, none of them did. During direct retrieval each system read only the visible HTML, and ignored JSON-LD, Microdata, and RDFa. That is a strong hint about why bolting on markup does not move citations.

What schema actually does (and it is not nothing)
Dropping schema entirely would be the opposite mistake. Structured data has real, documented jobs that are worth the effort:
- Entity disambiguation. It tells search engines whether “Mercury” is the planet, the element, or the car, and connects your brand to a known entity in the Knowledge Graph.
- Rich result eligibility. Review stars, FAQ accordions, product price and stock, recipe cards, breadcrumbs. These are SERP features, not AI citations, and they still earn clicks.
- Machine-readable facts. Clean Product, Article, and Organization markup removes ambiguity about price, author, and publish date.
None of those are the same as “make the content worth citing,” which is the part the schema pitch quietly skips.
The same split shows up on the commerce side too: Product schema will not get you into AI shopping carts either, where it is the product feed, not the markup, that actually transacts.
What actually drives AI citations
If schema sits downstream of site quality, then the real work is upstream. The levers the data and the platforms’ own behavior point to are familiar ones:
- Being the original source. Proprietary data, first-hand testing, a framework or benchmark nobody else has. AI engines have a reason to cite the page that contains the fact, not the tenth rewrite of it.
- Answer-first structure in raw HTML. A direct answer in the opening lines, present in the initial HTML payload, not injected by JavaScript a retrieval bot will never run.
- Topical authority and links. The same trust signals that have always separated cited sources from the crowd.
- Freshness. AI engines weigh recency, so a 2024 page left untouched loses to a maintained 2026 one.
Test it on your own pages
You do not have to take Ahrefs’ word or mine. Run a small version on your own site, the same measure-it-yourself discipline I used for the GEO statistics claims and for measuring AI Overview traffic loss:
- Pick 5 to 10 pages that already get some AI citations, so you have a baseline, plus a matched set of control pages you will leave alone.
- Add or upgrade JSON-LD on the test group only, and change nothing else during the window.
- Wait 8 to 12 weeks; AI retrieval indexes move slowly.
- Compare the citation change between the two groups, not before-and-after on the same pages, so a platform-wide shift does not get mistaken for a schema effect.
If your test shows a real lift, great, you have evidence specific to your niche. If it shows what the 1,885-page study showed, you just saved yourself a quarter of misdirected effort.
So, does schema markup help AI citations?
In my opinion, schema is hygiene and an eligibility ticket for rich results, not an AI-citation lever, and the largest dataset we have backs that up. If you are already doing the rest of the SEO work well, JSON-LD is not the thing that unlocks citations.
So add schema because it disambiguates your entities and earns SERP features, which are real wins. Just do not add it expecting ChatGPT to suddenly notice you, because the thing that earns the citation is the thing schema cannot fake: content worth citing.
Update Logs
23 Jun 2026
- Added internal links to the product-schema and AI Overview traffic-loss pieces, and an inline concept image.
22 Jun 2026
- Added a table of contents, a caption on the effects table, and a bar chart of the three difference-in-differences results.
14 Jun 2026
- Rewrote the article in a clearer, measured voice and added a Key Takeaways summary.
- Verified every figure against the Ahrefs study (+2.4% AI Mode, +2.2% ChatGPT, -4.6% AI Overviews, the 3x correlation, the 100+ citation baseline) and clarified that the 3x stat comes from the broader correlation analysis, not the 1,885-page test.
- Added the searchVIU finding that major AI systems read only visible HTML during retrieval and ignore JSON-LD.
1 Jun 2026
- Published
Want our posts to show up more often on Google?
One step & Google will surface this site in your Top Stories.
