← Back to blogHow do AI assistants decide which sources to cite?

How do AI assistants decide which sources to cite?

AI assistants cite what they can retrieve, trust and extract cleanly, and that is no longer the same thing as what ranks. Ahrefs found only 38% of Google AI Overview citations came from top-10 pages in April 2026, down from 76% nine months earlier, and 86% of the most-cited sources are not shared between ChatGPT, Perplexity and AI Overviews. Citation has become its own discipline, and it has to be measured engine by engine.


Between a user's prompt and the answer they read, an assistant runs a short pipeline. It expands the question into sub-queries, pulls candidate pages from an index or its own crawler, reads whatever text comes back, and quotes the passages it can attribute cleanly. A page has to survive every stage of that pipeline to be named at the end of it. It has to be reachable and readable at the moment of the query, it has to sit on a domain the system already treats as reliable, and it has to state a specific claim plainly enough to be lifted without distortion. Fail any one stage and the citation goes to whoever else cleared all three. Because each assistant runs a different index, weights trust differently and cites a different number of sources per answer, the same page can be a staple for one engine and invisible to another.

The three stages a page has to survive

Four figures on citation: 76.1% of Google AI Overview citations came from top-10 pages in July 2025, falling to about 38% in April 2026; 86% of top mentioned sources are not shared across ChatGPT, Perplexity and AI Overviews; and 70.6% of top-50 news sites blocking ChatGPT's retrieval bot were still cited.  Sources on the card: Ahrefs July 2025, Ahrefs April 2026, Ahrefs, BuzzStream March 2026. Yellow sits on the ~38%, the number the piece turns on.

Retrieval comes first, and it is unforgiving. For any question that needs current or specific information, an assistant does not rely on training data alone. It fetches live pages. A source that is not fetched cannot be cited, whatever its quality. This is the stage that technical decisions govern: whether the crawler is allowed in, whether the server returns readable text, whether the content exists in the raw HTML.

Trust comes second. Among the pages an assistant retrieves, it favours domains it already treats as authoritative, content with clear authorship and named sources, and claims that are checkable rather than asserted. Original data outperforms opinion here, because the model can anchor an answer to something specific.

Extraction comes third and is the stage most often ignored. The assistant has to be able to lift a clean, self-contained claim. Content with descriptive headings, a direct answer near the top and plain factual sentences is simply easier to quote accurately than content where the answer is implied across four paragraphs of narrative. Convenience is a ranking factor in AI answers in a way it never was in search.

Citation and search ranking have come apart

The most important shift of 2026 is that ranking well no longer predicts being cited.

In July 2025, Ahrefs studied 1.9 million citations across 1 million AI Overviews and found that 76.1% of cited pages ranked in Google's top 10 for the same query. In April 2026, Ahrefs repeated the analysis across 863,000 keywords and 4 million AI Overview URLs and found that figure had fallen to roughly 38%. The remainder split almost evenly between pages ranking 11 to 100 and pages that did not rank in the top 100 at all.

Ahrefs is careful about the comparison, and so should anyone quoting it be. The firm notes its citation detection improved between the two studies, so the datasets are not strictly like for like. It attributes the rest of the gap to query fan-out, where Google decomposes a question into related sub-queries and cites pages that perform well across the wider cluster rather than the single page that ranks for the literal phrase. The timing is suggestive: Google rolled out Gemini 3 as the default model for AI Overviews globally in late January 2026.

A separate BrightEdge analysis published in February 2026 put the overlap between AI Overview citations and organic top-10 results at around 17%, and reported that roughly five out of six AI Overview citations come from content that is not on page one. The methodologies differ and the numbers do not reconcile neatly, but both point the same way. Page-one ranking is now a weak proxy for AI visibility, and a publisher measuring only rankings is measuring the wrong thing.

Which sources do assistants actually cite?

Bar chart of sources cited per answer by engine: Perplexity 16 to 22, Google AI Overviews around 12, ChatGPT around 7, showing the same visibility strategy produces very different odds on different platforms.  Note: the Perplexity bar is drawn to the top of the published range and labelled "16-22" so the range is not flattened into a single invented figure. Yellow sits on Perplexity.

Citation is far more concentrated than search ever was. The Trade Press AI Index 2026 from 5W Research, which synthesises six published citation studies covering more than 680 million individual citations across ChatGPT, Claude, Perplexity, Gemini and Google AI Overviews, found that the top 15 domains capture 68% of consolidated AI citation share. In every industry it examined, citation share concentrated into three to five publications, with PCMag, Skift, STAT, Bloomberg and Axios gaining ground while several legacy prestige titles lost it.

Community and video platforms dominate the top of that list. Reddit is the single most-cited domain across the major engines. YouTube accounted for close to 6% of all AI Overview citations in the Ahrefs dataset as of March 2026, and Ahrefs reports its citation share grew 34% over the preceding six months. In the April 2026 study, more than 18% of AI Overview citations that did not rank in the organic top 100 were YouTube URLs.

Volume differs sharply by engine too. Across published analyses ChatGPT cites roughly seven sources per answer, Google AI Overviews around twelve, and Perplexity anywhere from sixteen to twenty-two depending on the dataset. Perplexity is citing multiple sources per claim rather than picking a single best one, which means the same visibility strategy produces very different odds on different platforms.

One finding cuts across all of them. OtterlyAI, analysing more than one million citations across ChatGPT, Perplexity and Google AI Overviews in January and February 2026, reported that AI answers depend on third-party sources roughly 95% of the time rather than on brand-owned websites. Being the subject of an answer and being the source of it are different achievements.

Why the same question produces different sources on different assistants

Ahrefs found that 86% of the top mentioned sources are not shared across ChatGPT, Perplexity and AI Overviews. A separate analysis of 680 million citations put the domain overlap between ChatGPT and Perplexity at 11%. Even within Google, AI Overviews and AI Mode cite the same URL only about 13.7% of the time while reaching broadly similar conclusions.

The divergence is architectural rather than editorial. ChatGPT's live retrieval leans on Bing's index. Perplexity runs its own crawler and index and draws on multiple search APIs, with heavy freshness weighting, so newly published material can be cited within hours. Google is working from its own index plus a fan-out layer. Different front doors produce different candidate sets before any judgement about quality is made.

The practical consequence for publishers and advertisers is that AI visibility is not a single number. Share of voice has to be tracked per assistant, and a strategy that lifts citations on Perplexity may do nothing on AI Overviews.

Does freshness actually matter?

Yes, but less than the advice usually implies, and not everywhere.

Ahrefs analysed 16.98 million cited URLs across ChatGPT, Perplexity, Gemini, Copilot, AI Overviews and organic Google results. The average cited page in AI assistants was 1,064 days old against 1,432 days for organic results, making AI citations about 25.7% fresher. ChatGPT showed the strongest preference for new material, citing pages an average of 458 days newer than organic. Google's AI Overviews were the exception, citing content marginally older than organic.

Two caveats matter. The average cited page is still almost three years old, so assistants reward durable content that is kept current rather than content that is merely new. And Ahrefs found a much smaller gap on last-updated dates, 13.1%, than on publication dates, which suggests that touching a timestamp is not the lever. Substantive updates to pages that already have standing are.

The retrieval failure most publishers do not know they have

Flow showing a client-rendered page failing retrieval: a prompt needs a current fact, an AI crawler fetches the raw HTML, the answer is painted in by client-side JavaScript, no major AI crawler renders JavaScript, so the server logs a 200 with an empty body and the page never enters the candidate set.  Per the style guide, the flow carries no yellow - its earned colour is the single red negative at the terminal node. Caption carries the Vercel/MERJ figures (GPTBot ~11.5%, ClaudeBot ~23.8%, neither executes).

A page can be authoritative, well structured and perfectly current and still never reach the trust stage, because the crawler saw nothing.

Most AI crawlers do not execute JavaScript. The Vercel and MERJ server-log study of real crawler traffic found that none of the major AI crawlers rendered JavaScript. GPTBot fetched JavaScript files in around 11.5% of requests and ClaudeBot in around 23.8%, but neither executed them. Googlebot, by contrast, runs a headless Chrome-based renderer. The result is a visibility split that does not show up in any rankings report: a client-rendered site can perform respectably in Google and be functionally blank to every AI crawler that matters.

For publishers on modern JavaScript frameworks, this is usually the highest-value thing to check before any content work begins. Fetch your own page with JavaScript disabled. Whatever is not in that raw HTML is not in the candidate set.

Blocking a crawler does not reliably stop citation

The inverse assumption is also wrong. Research published by BuzzStream in March 2026, analysing 4 million citations across 3,600 prompts, found that 70.6% of the top 50 news sites blocking ChatGPT's retrieval bot were still cited. Content reaches assistants through syndication partners, aggregators, licensed feeds and third-party summaries, so a robots.txt disallow removes the direct read without removing the mention.

This matters commercially. Blocking forfeits the ability to see, measure or monetise the retrieval while leaving the citation largely intact. It is the worst of both positions for most publishers, which is why the block-or-monetise question deserves more than a default.

How to become a source AI assistants cite

The practical checklist follows directly from the pipeline.

Serve readable text in the raw HTML, not painted in by client-side scripts. Lead every page with a direct, self-contained answer to the question it targets, in the first paragraph. Support claims with specific figures and named, dated sources, because checkable data is what gets quoted. Structure with question-shaped headings so an assistant can locate the relevant chunk. Build topical depth across related sub-questions rather than one page per keyword, because fan-out rewards coverage of a cluster. Keep entity names consistent across your own site and third-party mentions so assistants recognise them as the same thing. Pursue third-party coverage, given that roughly 95% of citations go to sources other than the brand's own site. Update substantively rather than cosmetically. Then measure citation share per assistant, not as a single blended figure.

This is the discipline of generative engine optimisation, and it is how content moves from existing to being quoted.

Citation, visibility and revenue

For publishers, being cited has two separable payoffs. The first is visibility: your brand and your information stay in front of the user even when the answer replaces the click. The second is easier to overlook. The same retrieval that earns a citation is also a read of your content, and that read is an event you can see and value if you are watching for it at the right layer.

blankspace operates on that second side, detecting Live Search Agent retrievals at the CDN edge and turning them into measured, monetisable events rather than invisible traffic. Strong citation practice increases how often assistants retrieve and quote you. Edge-layer measurement and monetisation capture value from those reads whether or not the citation converts into a visit. The two are complementary, and neither substitutes for the other.

Frequently asked questions

Do AI assistants cite the highest-ranking Google result?

Often not. Ahrefs found in April 2026 that only about 38% of Google AI Overview citations came from pages ranking in the top 10 for the same query, down from 76.1% in its July 2025 study. Rankings and citations now correlate loosely enough that they have to be tracked separately.

Why does my brand get cited by one assistant but not another?

Because the engines start from different candidate sets. Ahrefs found 86% of top mentioned sources are not shared across ChatGPT, Perplexity and AI Overviews, and ChatGPT-Perplexity domain overlap has been measured at around 11%. ChatGPT retrieves largely via Bing's index, Perplexity runs its own crawler with heavy freshness weighting, and Google adds a fan-out layer. Visibility has to be measured per assistant.

Can content hidden behind JavaScript be cited?

Usually not. The Vercel and MERJ crawler study found that none of the major AI crawlers render JavaScript: GPTBot fetched JavaScript files in about 11.5% of requests and ClaudeBot in about 23.8%, and neither executed them. If your content only appears after client-side rendering, the crawler sees an empty page. Serving readable HTML is a prerequisite for citation.

Does updating old articles improve AI citations?

Modestly, and only if the update is substantive. Across 16.98 million cited URLs Ahrefs found AI-cited pages were about 25.7% fresher than organic results by publication date, but only 13.1% fresher by last-updated date, and the average cited page was still nearly three years old. Changing a timestamp is not the mechanism; adding current, checkable material to a page that already has standing is.

If I block AI crawlers, will I stop being cited?

Largely no. BuzzStream's March 2026 analysis of 4 million citations found 70.6% of the top 50 news sites blocking ChatGPT's retrieval bot were still cited, because content reaches assistants via syndication, aggregators and third-party coverage. Blocking removes your visibility into the retrieval and your ability to monetise it without removing the mention.