← Back to blog

Are AI citations accurate, and what happens when an assistant misattributes you?

Mostly, no. Assistants reproduce a publisher's distinctive reporting far more often than they credit it, and when they do name a source they frequently name the wrong one. A McGill audit published in March 2026 found Gemini covered the substance of a news article in 81% of responses while naming the outlet in 6%. Treat attribution as a variable you measure, not a benefit you receive.


There are two separate failures hiding inside the phrase "AI citation", and publishers who collapse them into one end up optimising for the wrong problem. The first is silence: the model uses your reporting and does not say where it came from. The second is error: the model does say where it came from and gets it wrong, crediting a syndication partner, an aggregator, a rival title, or a URL that resolves to nothing. Two years of independent audits have documented both. Both are considerably worse in the way people actually use these products than in the controlled conditions researchers use to establish that the capability exists. And neither has a publisher-facing correction route. The commercial consequence is identical in each case: the reporting travels and the masthead does not.

How often do AI assistants actually credit the publisher?

The strongest recent evidence is the AI News Audit published in March 2026 by Aengus Bridgman and Taylor Owen at McGill University's Centre for Media, Technology and Democracy. The researchers tested ChatGPT, Claude, Gemini and Grok, in paid and economy tiers, against 2,267 real Canadian news stories in English and French, running 18,134 queries in total.

With web search switched off, the models showed at least partial knowledge of 74% of stories that fell inside their training window. Among those knowledgeable responses, 92% provided no source attribution of any kind. The reporting was in there. The provenance was not.

With web search switched on, tested against 140 specific articles through each company's API, the picture changed in shape but not in substance. Every model produced answers that covered enough of the original reporting that many readers would have no reason to visit the source. Models frequently linked to Canadian news sites, with 52% of responses including at least one Canadian URL, but named a Canadian source in the response text only 28% of the time.

Coverage is high, credit is not

The audit's most useful chart plots coverage against credit under what the researchers call the generic condition: a reader asking "tell me about X" rather than "what did the Toronto Star report about X". That is how almost everyone uses these products.

Gemini covered an article's distinctive reporting, meaning the specific events, named individuals and key findings rather than the general topic, in 81% of responses, and credited the originating outlet in 6%. Claude covered distinctive reporting in 72% of responses and credited the outlet in 16%. Grok covered 59% and cited 7%. ChatGPT covered distinctive content in 54% of responses and, in this sample, credited the originating newsroom around 1% of the time.

The researchers are careful to note that even a response that fails to cover the distinctive reporting still delivers a topical answer that reduces the reader's motivation to visit the source. There is no version of this where nothing is taken.

The capability exists, the default does not use it

Under the most favourable conditions the audit tested, naming the outlet in the prompt and explicitly asking for citations, all four models named the outlet in a majority of responses: Claude 97%, Gemini 95%, ChatGPT 86% and Grok 74%. Linking rates were also respectable, with Grok at 91%, Gemini 69%, Claude 64% and ChatGPT 59%.

That gap is the finding. Meaningful attribution is technically achievable today, on current models, with no new infrastructure. It simply is not what the default consumer experience produces, and the default consumer experience is what shapes the market for journalism.

The audit also found that attribution, where it happens at all, flows to outlets readers already know. Freely accessible national broadcasters captured most of the visibility. Paywalled and regional titles were marginal even on their own original reporting, and regional Postmedia papers serving Calgary, Edmonton, Ottawa and Vancouver were essentially absent. Among French-language outlets, the Journal de Montreal, one of Quebec's most widely read papers, received 48 total mentions across all models. French-language journalism, the researchers write, is doubly disadvantaged: absorbed into training data, almost never acknowledged.

How often is the attribution simply wrong?

Non-attribution is the larger structural problem. Misattribution is the more damaging one, because it puts your name on something you did not publish or someone else's name on something you did.

The benchmark study remains the Tow Center for Digital Journalism's citation audit at Columbia, published in the Columbia Journalism Review in March 2025 by Klaudia Jazwinska and Aisvarya Chandrasekar. The team fed excerpts from 200 articles across 20 publishers into eight generative search tools and asked each to identify the headline, publisher, date and URL. The excerpts were deliberately chosen so that a plain Google search returned the original within the first three results.

Across 1,600 queries the tools answered incorrectly more than 60% of the time. Perplexity was the most accurate at 37% incorrect. Grok 3 answered 94% of queries incorrectly. DeepSeek misattributed the source of the supplied excerpt 115 times out of 200, meaning publishers' content was most often credited to somebody else.

Three findings from that study matter more to publishers than the headline error rate.

Confidence did not track accuracy. ChatGPT misidentified 134 articles but signalled a lack of confidence just 15 times across 200 responses, and never declined to answer. Premium tiers were worse, not better: Perplexity Pro and Grok 3 produced more correct answers than their free equivalents and also higher error rates, because they gave definitive wrong answers where the free versions declined.

Links were frequently fabricated. More than half of the responses from Gemini and Grok 3 cited fabricated or broken URLs. Of 200 prompts run through Grok 3, 154 citations led to error pages. A citation that does not resolve is worse than no citation, because it looks like attribution to the reader and delivers nothing to the publisher.

Syndicated copies displaced originals. Tools routinely pointed to versions of articles republished on Yahoo News or AOL rather than to the publisher that produced them. USA Today blocks ChatGPT's crawler, and ChatGPT cited a Yahoo News republication of a USA Today article anyway.

The accuracy picture from the newsroom side is consistent. The BBC's February 2025 study of four assistants given access to BBC content found significant issues in 51% of answers and some issue in 91%. Of answers citing BBC content, 19% introduced factual errors, and 13% of quotes sourced from BBC articles had been altered or did not exist in the cited article at all.

The larger follow-up, News Integrity in AI Assistants, published by the European Broadcasting Union and the BBC on 21 October 2025, put 22 public service media organisations in 18 countries and 14 languages to work evaluating more than 3,000 responses from ChatGPT, Copilot, Gemini and Perplexity. Almost half of all answers had at least one significant issue. A third showed serious sourcing problems. A fifth contained major accuracy problems including hallucinated or outdated information. Those figures are commonly reported as 45%, 31% and 20% respectively.

Separately, NewsGuard's audits of the leading chatbots found them repeating provably false claims in response to news prompts more than a third of the time in August 2025, roughly double the rate a year earlier, and more than 28% of the time in its January 2026 quarterly audit of 11 tools.

Does a licensing deal fix it?

Not on the published evidence, and this is the finding publishers most consistently expect to be wrong.

The Tow Center study included publishers with active licensing and revenue-share relationships with the companies whose tools were tested, and found that licensing provided no guarantee of accurate citation. The San Francisco Chronicle permits OpenAI's search crawler and sits inside Hearst's strategic content partnership with the company; ChatGPT correctly identified one of the ten excerpts tested. Perplexity Pro cited syndicated versions of Texas Tribune articles in three of ten queries despite Perplexity's partnership with the Tribune. Time's chief operating officer Mark Howard, whose title has deals with both OpenAI and Perplexity, confirmed to the researchers that accurate surfacing was the intention but that the companies made no commitment to being completely accurate.

There is a real counterweight, and it is about volume rather than correctness. The Press Ranger and OtterlyAI study released in August 2026 found publishers with OpenAI agreements earning 48% more ChatGPT citations in June 2026 than comparable unlicensed publishers. That is a meaningful commercial finding, and we covered it separately. It is not evidence that licensed citations are more accurate. Nobody has published a study showing that a licence improves the correctness of the citation, as opposed to its frequency, and the one study that tested correctness directly found no such effect.

The distinction is worth holding onto in a negotiation. A licence appears to buy you more mentions. It has not been shown to buy you correct ones.

Why does this keep happening?

Four mechanisms, none of them mysterious and none of them likely to resolve on their own.

Retrieval runs over an index thick with copies. A wire story exists as the original plus dozens of syndicated, republished and rewritten near-duplicates. The retrieval layer scores documents on relevance to the query, not on primacy, so the version that wins is often not the version that was reported. Where original coverage is thin, the copies are all there is.

Attribution is generated after the answer, not with it. These systems compose a fluent response and then decorate it with citations. The citation is not a record of where each clause came from, which is why a model can reproduce an article's distinctive findings accurately and name the wrong outlet in the same breath, and why fabricated URLs appear alongside correct facts.

There is no provenance layer for text. The industry's provenance work, including C2PA and SynthID, addresses images. A factual claim rendered in confident prose by an assistant carries no machine-readable signal of where it came from or how the retrieval was scored. Nothing in the output can be verified downstream.

Product design does not ask for it. The McGill best-case numbers prove the models can attribute properly when asked. Assistant interfaces do not ask, because a named-source-heavy answer reads as less fluent, and fluency is what the products are optimised and evaluated on.

What misattribution actually costs a publisher

The generative engine optimisation case rests on a chain: get retrieved, get cited, get brand lift and referral clicks. The attribution data breaks the middle link. If an assistant covers your distinctive reporting in four responses out of five and names you in one out of sixteen, you are supplying the substance of the answer and almost none of the signal. Optimising harder for retrieval, in that condition, increases the amount of unattributed work you do.

The reputational exposure runs in the opposite direction and is easy to miss. The BBC's own research made the point plainly: when an assistant cites a trusted brand, audiences are more likely to trust the answer even when the answer is wrong. Altered quotes, invented quotes and confidently wrong claims attributed to your masthead are a liability you did not publish and cannot see. The 13% quote-alteration rate in the BBC study is not a rounding error at the volumes these products run at.

There is a negotiating cost too. Attribution is the evidence base for the value of a licence. A publisher who cannot demonstrate how often, and how accurately, an assistant credits them is negotiating on assertion.

And the distribution of the problem is not neutral. Attribution concentrates on large, free, familiar national outlets. Regional, paywalled and non-English titles produce the reporting and disappear from the record. Whatever else the attribution gap is, it is also a market-concentration mechanism.

What publishers can do about it

Be realistic about the range here. Some of these work, one of them does not exist, and the last one is a change of frame rather than a tactic.

Measure attribution as two separate rates, not one. Prose naming and link presence are different events with different commercial effects, and the McGill numbers show they diverge sharply, 52% linking against 28% naming. Track both, per assistant.

Test in the generic condition. If you prompt an assistant with your own masthead and ask it for citations, you will measure the 86% to 97% best case and learn nothing about your actual exposure. Ask the question a reader would ask.

Repair what retrieval reads. Canonical tags, unambiguous bylines and publication dates in the markup, and NewsArticle structured data are the cheapest interventions available, because the dependable way to change what an assistant says is to change what it reads.

Audit your syndication footprint. For most publishers the biggest single competitor for attribution is their own wire and republication partners. Syndication agreements that require a canonical pointer back to the original are worth more in 2026 than they were in 2020.

Spot-check quotes attributed to you. Quote fabrication and alteration are the failure mode with genuine legal and reputational tails, and no one is monitoring it on your behalf.

Do not wait for a correction button. As of September 2026 no major assistant publishes a publisher-facing attribution-correction process with a stated service level. The reporting routes that exist are best-effort and route you back to the underlying web page.

Do not assume blocking solves it. The Tow Center found tools answering correctly about content they should not have had access to, and citing syndicated copies of publishers who had blocked them. Blocking removes your ability to be credited without reliably removing your content from the answer.

The frame change is the important one. If attribution is unreliable, unmeasured by the platforms and uncorrectable by you, then a business model that only pays out when a citation appears and holds is a business model with a variable you do not control at the centre of it. That is the case for capturing value at the moment of retrieval rather than downstream of a citation, which is the mechanism blankspace operates on at the CDN edge. It is not a substitute for accurate attribution, which publishers should keep pressing for, and it does not fix the reputational problem at all.

How to measure your own attribution rate

A workable monthly spec, per assistant, run on a fixed sample of your own original reporting, in the generic condition:

Prose attribution rate: the share of responses that name your outlet in the response text.

Link attribution rate: the share of responses containing at least one URL on your domain.

Correct-source rate: of responses that name any outlet, the share naming the outlet that actually produced the reporting.

Dead-link rate: the share of citations to your domain that resolve to an error page or a fabricated path.

Coverage rate: the share of responses that reproduce enough of the article's distinctive reporting that a reader would not need to visit it.

Quote fidelity: of quotations attributed to your reporting, the share that appear verbatim in the cited article.

Coverage rate against prose attribution rate is the number to put in front of a board. It states, in one ratio, how much of your journalism is being used and how much of it is being credited.

Frequently asked questions

Do AI assistants cite sources accurately?

Frequently not. The Tow Center's audit of eight generative search tools found incorrect answers to more than 60% of queries asking them to identify a supplied article's publisher, headline and URL, ranging from 37% incorrect for Perplexity to 94% for Grok 3. Errors include naming the wrong publisher, linking to syndicated copies rather than originals, and fabricating URLs that resolve to error pages.

How often does ChatGPT credit the publisher it is drawing from?

In the McGill AI News Audit's generic condition, where the user asks a topic question without naming an outlet or requesting citations, ChatGPT covered an article's distinctive reporting in 54% of responses and credited the originating newsroom in roughly 1% of them. Asked directly to cite sources and given the outlet name, it named the outlet in 86% of responses, which shows the gap is a product-design choice rather than a technical limit.

Does having a licensing deal with an AI company improve citation accuracy?

No study has shown that it does. The Tow Center tested publishers with active OpenAI and Perplexity partnerships and found licensing provided no guarantee of accurate citation, including a case where ChatGPT correctly identified one of ten excerpts from a partner publisher. Separate 2026 research found licensed publishers receiving more citations by volume, which is a different claim from receiving more accurate ones.

Can a publisher get an AI assistant to correct a misattribution?

There is no reliable mechanism as of September 2026. No major assistant publishes a publisher-facing correction process with a stated service level, and the feedback routes that exist are best-effort, typically pointing the complainant back to the underlying web page. The practical route is source repair: correct or clarify the pages the retrieval layer reads, then retest the prompt in a fresh session.

If an assistant invents a quote and attributes it to my publication, what is my exposure?

The risk is reputational first and legal second, and it is amplified by the trust your brand carries. The BBC's research found 13% of quotes sourced from BBC articles had been altered or did not exist in the cited article, and noted that citing a trusted brand makes audiences more likely to believe an answer even when it is wrong. No assistant currently notifies publishers when it attributes a fabricated quotation to them, so detection is a monitoring task you have to run yourself.