Open ChatGPT. Type the question your best customer asks in the week before they shortlist vendors. Not your brand name. The question.
Read the sources it cites under the answer.
Now open Google, search the same words, and read the top ten.
Compare the two lists. Ahrefs ran exactly that comparison at scale: 15,000 long-tail queries across ChatGPT, Gemini, Copilot and Perplexity. On average 12% of the URLs the models cited appear anywhere in Google’s top ten for the query the user actually typed. ChatGPT scored 8.0% on in-text citations and 6.1% on its reference list. Gemini 8.6%. Copilot 8.2%. Four out of five cited URLs do not rank for the question at all.
The thesis
AEO is not SEO with new vocabulary. The unit of optimisation changed.
Search ranked pages against a query. Answer engines score passages against sub-queries the user never typed, then fuse the winners. What competes is a single verifiable claim. The page is just the container it arrived in.
If you are still running a page-level roadmap, you are optimising the box and not the thing inside it.
What happens between the question and the answer
The step everyone skips is query fan-out.
You type one question. The engine does not run it. It decomposes it into a set of narrower queries, runs those, and merges the result sets using Reciprocal Rank Fusion, a method that combines several ordered lists into one. Ahrefs identified fan-out as the structural reason the overlap is so low. The engine and the buyer are not searching for the same thing, so the results page you rank on is not the results page that fed the answer.
Then the model fills citation slots, and the slot count is not the same everywhere. Semrush analysed 126 million US AI search prompts between January and April 2026 across ChatGPT, Gemini, Google AI Mode and AI Overviews. ChatGPT averages 15 sources per response. Gemini averages 3. That is not a rounding difference. On Gemini you are fighting for one of three positions against Wikipedia and Reddit. On ChatGPT there is room for a specialist.
Semrush found something sharper still. On Gemini, the overlap between the brands mentioned in an answer and the domains cited underneath it runs as low as 30%. Being talked about and being cited are two different outcomes with two different mechanics.
The number that settles it
Everything above could be dismissed as third-party engines behaving oddly. So take Google’s own surface, where ranking should transfer most directly of all.
Ahrefs measured AI Overview citations twice. In July 2025, across 1.9 million citations, about 76% came from pages ranking in the top ten. In January 2026, across 863,000 results pages and 4 million AI Overview URLs, that figure was 37.9%.
It halved in six months.
The rest did not come from page eleven either. 31.2% came from positions 11 to 100. A further 31.0% came from pages that do not rank in the top 100 at all. YouTube alone supplied 5.6% of AI Overview citations while ranking nowhere in classic search.
Google is now sourcing nearly a third of its cited answers from documents its own ranking system does not put on the first ten pages. Fan-out is why.
The tactic that does nothing
If the page were still the unit, marking it up would help. It does not.
On 11 May 2026 Ahrefs published a controlled study covering August 2025 to March 2026. They took 1,885 pages that added JSON-LD schema, matched them against 4,000 control pages, and measured citations for 30 days before and 30 days after, using difference-in-differences to strip out platform-wide trends.
Google AI Overviews fell 4.6%, and that was the only statistically significant result. Google AI Mode rose 2.4% and ChatGPT rose 2.2%, neither significant. Adding schema produced no citation uplift on any platform, and on the largest one it went slightly the wrong way.
Retrieval reads visible text. The correlation everybody quotes exists because sites that implement JSON-LD also write better and earn more links. The most recommended AEO tactic of the last two years is decoration.
Why the stakes moved with it
SparkToro, using Similarweb clickstream data from January to April 2026, found 68.01% of US Google searches ended without a click. In 2024 it was 60.45%. Searches that produced a click fell 9.51 points, a 22.9% decline in two years. AI Overviews now appear on more than 20% of searches, and when one appears the click-through rate drops by nearly 60%.
The page used to be the destination. It is now a source the buyer never opens.
I built SearchFIT because I could not answer a simple question for a client: when a buyer asks the model, does the brand come up. The design decision that mattered was the atomic unit. SearchFIT tracks prompts, not keywords, and scores each engine separately, because a keyword ranking tells you nothing about a fan-out you cannot observe, and an engine average hides the fact that three slots and fifteen slots are different games.
It is the same instinct I applied at D30. D30 pulls ASX annual reports apart into tagged line items and runs articulation gates across the three statements, because a summary sentence that cannot be traced back to a number is not evidence. Both products refuse to treat the document as the unit. In forensic screening the unit is the tagged number. In answer engines the unit is the extractable claim. The document is packaging in both cases.
Where this breaks
Perplexity broke the pattern. Its overlap with Google’s top ten was 28.6%, more than three times ChatGPT’s. On one major engine classic ranking still transfers well, and if Perplexity is where your buyers are, the old playbook is largely intact.
The 15,000-query Ahrefs sample was long tail. Long tail is precisely where overlap should be lowest, because head terms have settled authority and fewer plausible sources. Run the same test on head terms and 12% would almost certainly rise. I would not quote that figure as though it covered all search.
The volume is also still in blue links. Google AI Mode accounted for 0.34% of searches between January and April 2026. Organic search remains the largest channel for most B2B companies by a wide margin. Anyone telling you to stop doing SEO is selling something, and 37.9% of AI Overview citations still come from the top ten, which is a lot of reason to keep ranking.
Last caveat, and it is about the category. Most AEO research is a press release. One widely circulated study this year claimed the overlap between top rankings and AI-cited sources collapsed from 70% to under 20%. Chase the citation and it attributes the finding to a third party with no disclosed sample size, no study period and no method. The real version of that collapse exists and I have used it above, because Ahrefs published the sample, the window and the test. Insist on all three before you move a budget.
What I would do on Monday
Replace the keyword list with a prompt list. Write the 20 questions a buyer asks in the four weeks before they shortlist. Track those weekly, with ChatGPT and Gemini scored separately.
Audit your top ten pages for extractable claims. One sentence that answers one question completely, carrying a number, a date and a source, sitting in visible HTML above any tab, accordion or click. Most pages have none. Aim for three per page.
Take schema out of the AEO budget line. A 1,885 page controlled study moved nothing anywhere and moved backwards on AI Overviews. Keep JSON-LD for rich results and stop paying for it as a citation tactic.
Go and win third-party mentions. With mention and citation overlap as low as 30% on Gemini, and 31% of AI Overview citations coming from outside the top 100, the review site, the comparison page and the forum thread are separate assets from your own domain.
Change the reporting unit. Report citations per claim per engine, not sessions per page. If your dashboard cannot answer which sentence of yours got cited last week, it is measuring the old thing.
Close
Go back to the tab you opened at the start.
Look at the sources under that answer one more time. Roughly a third of them are there without ranking in the top hundred for anything the buyer typed. Each one is present because a passage inside it answered a question the buyer never asked out loud, well enough that a fusion function put it in a shortlist the buyer will never see.
That is the whole game now. Not the page. The sentence.
SearchFIT tracks whether AI answer engines mention your brand when buyers ask. Most companies have never checked. Check yours.

