Every syndicated story exists in at least two places: the original on the publisher's own domain and a near-identical copy on a larger partner's. An assistant that retrieves both has to choose one to ground its answer and credit, and the choice is made on relevance and authority signals, not on who did the reporting. When the partner's domain is the stronger one, the partner gets the citation and the originating newsroom gets the cost of producing the work. That is the mechanism. How often it happens is a harder question, because nobody has measured it at scale, and this article sets out what the published tests do and do not show.
Why would an assistant cite the syndicated copy instead of the original?
Retrieval systems cluster near-duplicates and pick a representative. Microsoft's Bing team said as much in a Webmaster Blog post on 19 December 2025: large language models "group near-duplicate URLs into a single cluster and then choose one page to represent the set". If the copies differ only slightly, the post warns, the model may select a version that is "outdated or not the one you intended to highlight". Bing also notes that identical copies across domains make it harder for search engines and AI systems to identify the original.
Three things tilt that choice towards the partner. Aggregators such as Yahoo and MSN are large domains, so their copies tend to be prominent in the index. Some partners also insist on a self-referencing canonical tag on the copy, which tells the retrieval layer that the partner's URL is the authoritative one. And attribution in these products is applied to the answer after it is composed, so the citation reflects whichever document ranked highest in retrieval, not a record of who first published the facts.
What does the evidence show?
Three bodies of evidence bear on the question. None is conclusive, and each has a limit worth stating.
The Tow Center: syndicated copies surfaced in place of originals
Columbia's Tow Center for Digital Journalism tested eight generative search tools against excerpts from 200 articles across 20 publishers and published the results in the Columbia Journalism Review in March 2025. The authors reported that "on some occasions, chatbots directed us to syndicated versions of articles on platforms like Yahoo News or AOL rather than the original sources".
Two examples stood out. ChatGPT cited a Yahoo News republication of a USA Today article even though USA Today blocks ChatGPT's crawler. And Perplexity Pro cited syndicated versions of Texas Tribune articles for three of ten queries, despite a partnership between the two organisations. The study did not count how often the syndication pattern occurred overall, so it proves the failure exists without telling us its frequency.
Glenn Gabe: nine publishers, inconsistent results
Search consultant Glenn Gabe published a manual test on 16 September 2025 covering nine publishers that syndicate to Yahoo or MSN, checked across Google surfaces and four assistants: ChatGPT, Perplexity, Gemini and Claude. The outcomes varied within a single story. In one example ChatGPT cited the original while Perplexity ranked the Yahoo copy first. In another, a syndicated copy was first in Google's AI Overview and ChatGPT cited it only in a secondary "more" list. Gemini often cited neither version.
Only one of the four detailed examples had the original winning everywhere, which Gabe called a "unicorn". He also saw Yahoo appear more often than MSN in AI results, and found that a canonical tag helped for some publishers without reliably preventing both copies from being indexed. The sample was small, the publishers were anonymised and the prompts were not published, so this is a pattern to test for, not a rate to quote. It is also a year old, and the tools involved change quickly.
BuzzStream: syndication is a small share of citations, but Yahoo is large
BuzzStream's analysis of about four million citations from 3,600 prompts, collected in the week from 27 January 2026 across ChatGPT, Google AI Mode, AI Overviews and Gemini, points the other way on scale. It classed syndicated news as 6.2 per cent of news citations and 0.9 per cent of all citations, and concluded that Yahoo and MSN "don't seem to be prioritized" by AI search engines. Yet yahoo.com was still the single most cited news domain in the dataset, at 5.42 per cent of news citations, though that total includes Yahoo's own original reporting.
The method has acknowledged limits. Syndication was identified by matching author names to publications, which misses reposts that are not labelled, so the syndicated share may be undercounted. The prompts focused on large brands, and BuzzStream sells digital PR tools, which sits comfortably with its finding that earned coverage beats distribution. Search Engine Journal flagged that point when it reported the study.
So how big is the problem?
Unknown, and probably uneven. The honest summary is that syndicated copies demonstrably displace originals in some cases, on some platforms, for some stories, and that the one large dataset suggests they are a minority of citations overall. That is compatible with a serious problem for any individual publisher whose most valuable stories are syndicated, because the stories that get syndicated are usually the ones most likely to be retrieved. The only reliable way to know your own exposure is to test it.
Should publishers ask partners to noindex or canonicalise?
The two search engines that matter most disagree.
On one side, Google's Search Central documentation says partners should block indexing of the copy, because syndicated pages "are often very different" from the originals and a canonical is a poor fit. On the other, Bing's Webmaster Blog of 19 December 2025 says to ask partners for a canonical pointing to your original "when agreements allow", on the basis that a canonical signals which version matters.
Google's documentation says the canonical link element is not recommended for syndication and that "the most effective solution is for partners to block indexing of your content". Bing still prefers the canonical, and also suggests syndicating excerpts with a link back instead of full articles.
Assistants draw on different indexes, so a publisher cannot assume one instruction satisfies every one of them. Gabe's testing favoured noindex as the safest option, with canonical as the fallback, but he also reported that some partners insist on self-referencing canonicals. The practical position is to ask for noindex first, accept a cross-domain canonical second, and treat a partner's refusal of both as a price to put a number on.
What belongs in a syndication contract?
Most syndication agreements were drafted for a search world in which a partner's copy was a distribution channel. They now double as an AI-attribution decision. Five terms are worth negotiating.
A noindex or a canonical to the original URL on every syndicated copy, with the technical method specified. A delay before republication, so the original is indexed and retrieved first. Excerpts with a link back instead of full text for the stories you most want credited. A prohibition on partner-side sublicensing of your content to AI developers without separate consent and payment. And a right to audit how the copies are marked up, since a clause nobody checks protects nothing.
None of this means stopping syndication. Google's Search Liaison said in 2023 that Google was not telling publishers to stop syndicating, only to ask partners to noindex if they want the original to rank. Syndication still earns revenue and reaches readers the original never would. The point is to price the attribution cost in, as you would for any other channel.
How can a publisher test its own exposure?
A monthly check takes an afternoon. Pick twenty recent stories that were syndicated and twenty that were not. For each, write the question a reader would ask without naming your masthead. Run it through ChatGPT, Perplexity, Gemini, Claude and Google AI Mode. Record whether the response cited your URL, a partner's URL, both, or neither, and whether the outlet was named in the prose.
Compare the two groups. If your syndicated stories are cited on a partner domain far more often than your unsyndicated stories are cited on yours, you have a measurable leak, and a figure to take to the partner. Repeat the test after any contract or markup change. Because results vary between runs, treat a difference of a few stories as noise and look for a consistent gap.
Where does edge monetisation fit?
It is worth being plain about a limit. A publisher can only monetise retrieval that reaches its own domain. When an assistant fetches the Yahoo copy of your story, the request, the audience signal and any revenue attached to it land on someone else's infrastructure. Platforms such as blankspace work at the CDN edge of the publisher's own site, so they see and can price the requests that arrive there; they do nothing for a retrieval served from a partner's copy. That is an argument for deciding deliberately which stories you syndicate in full, not an argument against syndication.
Frequently asked questions
Do AI assistants cite syndicated copies instead of the original publisher?
Sometimes. Columbia's Tow Center found chatbots pointing to Yahoo News and AOL versions of articles, including a Yahoo copy of a USA Today story from a publisher that blocks ChatGPT's crawler. Glenn Gabe's 2025 test of nine publishers found results that varied by assistant and by story. No study has published a reliable rate across the web.
How much of AI search citation volume is syndicated news?
BuzzStream's January 2026 analysis of about four million citations classed syndicated news as 6.2 per cent of news citations and 0.9 per cent of all citations. The authors note this may be an undercount, because unlabelled reposts are hard to identify, and the prompts focused on large brands.
Is a canonical tag enough to protect the original?
Not reliably. Google says the canonical element is not recommended for syndicated content because the copies often differ from the originals, while Bing recommends it where agreements allow. Gabe found a canonical helped for some publishers but did not consistently stop both versions being indexed and cited.
Should publishers stop syndicating to Yahoo, MSN and Apple News?
Not necessarily. Syndication brings revenue and reach, and Google has said it is not telling publishers to stop. The decision is which stories to syndicate in full, on what terms, and with what markup on the partner's copy. Test your own citation data before deciding.
Does a licensing deal with an AI company stop syndicated copies being cited?
No evidence says it does. The Tow Center found Perplexity citing syndicated Texas Tribune articles for three of ten queries despite a partnership between them. A licence may affect how often you are cited, but it has not been shown to resolve which version of a story gets the credit.
