← Back to blog

Should publishers serve markdown to AI bots?

Serving markdown to AI bots reliably cuts the tokens a crawler spends on your page by 80 to 90 per cent. What it does not reliably do is get you cited more often, and one analysis of 1.6 million AI citations found no citation benefit at all. Markdown is worth doing as an efficiency and control measure, and only worth doing at all if you have decided what happens to the traffic once it arrives.


Strip the JavaScript, the ad slots, the navigation and the cookie banner out of an article page and what is left is roughly a tenth of the bytes and all of the meaning. That is the whole idea behind serving markdown to AI bots: give the retrieval agent the text and the metadata it actually wants, in the format models are trained to read, and stop charging it for the packaging. Time has converted its entire site this way. Cloudflare will now do it for any customer with a dashboard toggle. The efficiency case is settled and uncontested. The question a publisher is actually deciding is a different one, and it is commercial rather than technical: does making your content cheaper and easier for an AI system to consume leave you better off, or does it simply lower the cost of taking it?

What does serving markdown to AI bots actually mean?

Markdown is a plain-text formatting convention. Headings, links, emphasis and lists are expressed with a handful of punctuation characters rather than HTML tags, and there is no styling, no script and no layout. It has become the default interchange format for large language models, which is why the llms.txt convention is built on it and why Anthropic's own guidance on writing for agents recommends it.

An HTML article page is built for a browser and a human. It carries the layout, the design system, the analytics, the consent tooling, the recirculation modules and the advertising stack alongside the journalism. A retrieval agent wants none of that. As Jonah Goodhart, co-founder and chief executive of ad tech platform Mobian, described it to Digiday, an agent hitting a conventional page effectively disables JavaScript and turns off images so it can dig through the human furniture to reach the underlying content, references and links.

Serving markdown means handing that content over directly. In practice publishers are doing it in three materially different ways, and the differences matter far more than the format does.

Content negotiation at the edge. The same URL returns different representations depending on what the client asks for. Cloudflare shipped this as Markdown for Agents on 12 February 2026. A client that sends Accept: text/markdown in its request header gets markdown converted on the fly at Cloudflare's network edge; everything else gets the normal HTML page. It is a per-zone toggle, in beta at no extra cost on Pro, Business and Enterprise plans plus SSL for SaaS. Cloudflare's own worked example put the HTML version of a blog post at 16,180 tokens against 3,150 tokens for the markdown, an 80 per cent reduction, and every converted response carries an x-markdown-tokens header estimating the count.

A parallel agent site. The publisher generates and hosts a stripped copy of every page and points bots at it. TollBit builds these as part of onboarding, and its co-founder and chief executive Toshit Panigrahi told Digiday the company sees roughly a 90 per cent average token reduction on conversion. TollBit also claims that scraping and processing a large HTML page can take over a minute where a structured fetch completes in about a quarter of a second.

Whitelist, then redirect. The most restrictive version, and the one Time runs. Time blocks all AI bots by default, maintains a whitelist of approved ones, and redirects the approved bots to markdown copies while routing humans to the full page experience. Mark Howard, Time's chief operating officer, describes the point as separating the two audiences rather than serving them both the same thing: the bots get the content and the metadata, the humans get the site.

Only the third of these is a monetisation posture. The first two are plumbing.

Does markdown actually improve AI visibility?

This is the claim doing most of the work in vendor pitches, and it is the weakest link in the argument. The efficiency numbers are measured and repeatable. The citation numbers are not.

The theory is reasonable enough. Cleaner input means fewer parsing errors, fewer misrepresentations, lower inference cost per page and a crawl budget that goes further, so the model should find you easier to use and therefore use you more. Panigrahi frames it as the token economy, and argues that an AI system can comprehend more of an article when it is not paying to strip HTML out of it. Several GEO and SEO vendors make the same argument.

The evidence does not yet support the last step. GEO platform Promptwatch analysed more than 1.6 million citations in AI answers and found, as reported by Digiday, that markdown did not increase the likelihood of content being cited. Leigh McKenzie, director of online visibility at Semrush, put the sceptical case in terms publishers will recognise: markdown copies make sense in the short term in theory, but they look like Accelerated Mobile Pages, which nobody maintains any more because the web caught up.

Google's position is blunter. Its AI optimisation guide, published in May 2026, states that site owners do not need to create machine-readable files, AI text files, markup or markdown to appear in Google Search including its generative features, because Google Search does not use them. Google later softened the language around llms.txt to say that maintaining such files for other systems is fine and will neither help nor harm Google rankings, which is a clarification rather than a reversal. John Mueller of Google went further in public, calling the idea of serving markdown to bots "a stupid idea" and asking on Reddit why anyone would build a parallel version for bots rather than spend the time improving the site for everyone.

Two qualifications are worth holding on to. Mueller was criticising user-agent sniffing, not content negotiation, and those are technically different things. And Google Search is one consumer of your content among many; the guidance says nothing about how OpenAI, Anthropic or Perplexity handle a markdown response. Some AI coding tools already send Accept: text/markdown unprompted, and Cloudflare names Claude Code and OpenCode among them.

The honest summary is that the token saving is real and provable, the citation lift is unproven, and anyone selling you the second on the strength of the first is skipping a step.

The Content-Signal default that catches publishers out

This is the part of the Cloudflare implementation that deserves more attention than it has had, particularly from any publisher who has spent the last two years constructing a careful crawler policy.

Enabling Markdown for Agents does not only change the format. Converted responses ship with a Content-Signal header set by default to ai-train=yes, search=yes, ai-input=yes. That declares your content available for model training, for search indexing and for use as AI input including agentic use. Cloudflare has said custom Content-Signal policies are coming, but the default as launched is permissive on all three.

A publisher who has explicitly signalled ai-train=no elsewhere, or who has spent months negotiating training rights in a licensing conversation, can quietly undo that position with one dashboard toggle. Whether a given bot honours content signals at all is a separate question and depends entirely on the operator. But a publisher whose stated policy and whose served headers disagree with each other has weakened its own hand in exactly the conversation where the signal was supposed to matter. Check the defaults before you enable anything.

Is serving markdown to bots a form of cloaking?

Google defines cloaking as showing different content to users and search engines with intent to manipulate rankings and mislead users. The definition matters here because two of the three implementation patterns above sit on opposite sides of it.

With user-agent sniffing, the server inspects who is asking and decides what to show them. That is the pattern that draws the cloaking objection, and it is what Mueller was criticising. With content negotiation, the client states which format it would prefer and the server responds with the same information in that format. The mechanism is a long-standing part of HTTP and the substance of the content is unchanged.

The distinction is real but it is not a guarantee. Google has not addressed whether markdown served through content negotiation falls under its cloaking guidance, and the practical result from a crawler's point of view is similar either way. Rob Derow, managing director and partner at BCG X, named the wider version of this risk to Digiday: the biggest exposure with markdown ads is that there are no rules yet governing how LLMs treat them, and if the model providers later decide that sponsored content inside a markdown copy constitutes cloaking, those pages could become less effective or be penalised outright. His practical mitigation is the right one, which is that it can be monitored with AI visibility tooling rather than assumed.

Two rules follow. Keep the informational substance identical between representations, and do not put anything in the markdown version that is not also disclosed on the page a human sees.

The scraping trade-off nobody can price yet

Here is the argument that should decide this for most publishers, and it has nothing to do with tokens.

Making content trivially easy for AI systems to ingest is exactly what a large part of the industry has spent two years engineering against. HasData found that 56.4 per cent of news publishers block at least one AI crawler in robots.txt. Publishers have bought bot management, tightened CDN rules and in some cases blocked by default. Markdown removes that friction deliberately. Ed Zyszkowski, co-founder and chief executive of Personal Digital Spaces, put the objection sharply to Digiday: convert your site to markdown and hand it to a model provider and you have lost control of it, having effectively leaked your own information into the model.

The counter is that the content is being taken regardless. Cloudflare Radar puts bot traffic at more than half of all web traffic. TollBit's State of the Bots report for the first half of 2026 found European sites saw around four times the median AI scraping per site of North American ones, with a ratio of roughly 33 bot scrapes to every human visit on European sites, while AI apps sent just 0.16 per cent of external referrals to North American sites and 0.05 per cent to European ones. Against that, the marginal cost of making an already-occurring scrape cheaper is small and the marginal benefit of controlling how it happens is not.

The resolution is that both are right, and which one applies depends entirely on whether the markdown route is gated. Serving markdown to anything that asks is a giveaway. Serving markdown only to bots you have whitelisted, from a route you control, with a measurement layer attached, is an access policy. Time's implementation is the second. A dashboard toggle on its own is the first.

Scott Messer of Messer Media framed the decision test better than anyone: if there is no click, no ad impression and no cheque, the build is pure cost. Publishers should build for agents only if they genuinely believe there is long-term value in being discovered and cited in those systems, and the answer looks different for a hard-paywalled news brand than for a scale ad-funded lifestyle title. The Economist's approach reflects exactly that calculation. Rather than converting everything, it has started with marketing copy and B2B sales material that already sits outside the paywall, because as a subscription business it has to weigh AI exposure against the value of the paid product. Le Monde, separately, has been working on detecting whether an agent is acting on behalf of a paying subscriber, which is the same question approached from the identity side.

Can markdown pages carry advertising?

Time is the test case, and it is early enough that the results should be read as a signal rather than a benchmark.

In July 2026 Time began serving ads inside its markdown files, working with Mobian, with Ally Bank and the Project Management Institute among the first buyers. Time and Mobian claim it is the first time a publisher has served ads specifically targeting AI agents. The units are formatted as FAQs carrying a brand's information and messaging, labelled as sponsored content at the top even though no policy currently requires the label. Mobian generates the unit from a brand brief, renders it as a PDF for human review and client approval much like standard branded content, then puts the same FAQ questions to AI search engines to track visibility, favourability and accuracy over time. Time sells one agent ad per markdown page, can target contextually or against its list franchises or by date range, and charges a premium on the argument that AI bot impressions against authoritative content are scarce.

The demand side is interested and unconvinced in roughly equal measure. Jeff Eisenfeld, director of activation at Media by Mother, said it is worth a test if you know the ads will actually show up in the LLMs. Sam Huston, svp of media at Dept, said there is not a client who would refuse to test it given how much focus GEO is getting, and noted that budget would most likely come out of programmatic display or a dedicated testing pot rather than search. Jaquie Hoyos, chief media officer at Moroch, framed the real question as whether placement alongside high-quality information can meaningfully shape how brands surface.

The sceptics are pointed. Danny Weisman of Obsessed Media argued brands would do better investing in brand-building advertising and suggested it belongs as added value rather than a main campaign line. Stephan Kopp of Mediaplus Performance said flatly that it will not work and that it is only a matter of time before AI developers adjust their systems to ignore such workarounds, exactly as Google has repeatedly done with search. A WARC study conducted by agency Charlie Oscar supports the underlying scepticism, estimating that 63 per cent of a brand's visibility in AI answers comes from long-term brand equity against 26 per cent from current marketing activity. And Choy Travers, co-founder of Oasy, reported that in its own tests ads placed in markdown did not change advertisers' results or win rate against ads in HTML pages.

Rita Steinberg, vp of media at FUSE Create, gave the line publishers should quote back to any vendor pitching this as a channel: you can influence AI visibility, but you cannot reliably buy it yet.

The commercially useful reading is that markdown ads are a real experiment worth running and are not yet a media channel. Mobian's data suggests around 15 per cent of brands now power their own markdown pages, which is the pool of advertisers already thinking in these terms and therefore the sensible place for a publisher to start pitching.

What should a publisher actually do?

Six things, in order, and none of them requires committing to a full conversion.

Measure your bot traffic before you change anything. You cannot evaluate a format change to a traffic source you have not counted. Establish which declared agents are hitting you, at what volume, against which content, and what your server logs show that your client-side analytics does not.

Decide what happens to the traffic before you make it cheaper to serve. Messer's test is the right one. If the answer to "and then what" is nothing, the build is a cost with no offsetting revenue and you should not do it yet.

Gate the route. If you serve markdown, serve it to bots you have decided to admit, from a path you control, rather than to anything that sends the right header. The difference between an access policy and a giveaway is the whitelist.

Read the Content-Signal defaults before you toggle. Confirm the training, search and AI-input signals you are actually emitting match the policy you tell your licensing counterparties you operate.

Start with content you already give away. The Economist's approach generalises well. Marketing copy, B2B material and anything already outside the paywall carries almost no downside risk and gives you a real measurement baseline.

Instrument it, then judge it. Track citation and mention rates before and after on a controlled set of pages. Every published claim about markdown improving visibility is currently either vendor-sourced or contradicted, so treat your own data as the only reliable evidence.

Where blankspace fits

blankspace operates at the CDN edge, which is the same layer where the markdown decision is actually enforced, and the two are complementary rather than competing. Detecting Live Search Agent traffic as it arrives, separating it from human sessions, and running an auction that injects contextual brand facts into the response addresses the part of this that markdown alone does not: it makes the agent visit countable and saleable rather than simply cheaper to serve. Markdown changes what a bot receives. Edge detection and injection change whether the visit produces revenue and same-day attribution.

That distinction holds regardless of vendor, and it is the one to keep hold of when a markdown conversion is pitched as a monetisation strategy. Converting to markdown is a formatting and access decision. Getting paid for the traffic is a separate build, and doing the first does not accomplish the second.

Frequently asked questions

Does serving markdown to AI bots improve citations in ChatGPT or Perplexity?

There is no reliable evidence that it does. Promptwatch's analysis of more than 1.6 million AI citations found markdown did not increase the likelihood of being cited, and Google states directly that markdown files are not needed for its generative search features. The token efficiency benefit is well documented and separate from the visibility claim, so treat any pitch that runs the two together as unproven.

Is serving markdown to bots cloaking?

It depends on how it is implemented. Serving a different page based on the visitor's user-agent string is the pattern that attracts the cloaking objection, and John Mueller of Google has criticised it publicly. Content negotiation, where the client requests text/markdown and the server returns the same information in that format, is a standard HTTP mechanism and a different thing. Google has not ruled on the second case, so keep the substance identical between representations and disclose anything sponsored in both.

How much does markdown reduce AI crawling costs?

Cloudflare's own example measured a blog post at 16,180 tokens in HTML against 3,150 in markdown, roughly an 80 per cent reduction. TollBit reports an average of around 90 per cent across the pages it converts. The saving accrues primarily to the AI system doing the crawling rather than to the publisher, though publishers do see lower origin and CDN serving costs when bots are routed to lighter pages.

Does making my site easier for AI bots to read mean I get scraped more?

It makes each scrape cheaper and more accurate rather than necessarily more frequent, but it does remove friction you may have deliberately built. More than half of all web traffic is already automated by Cloudflare Radar's measure, and 56.4 per cent of news publishers block at least one AI crawler. The defensible position is to serve markdown only to crawlers you have explicitly admitted, not to any client that asks for it.

Can publishers sell advertising inside markdown pages?

Time began doing so in July 2026 with Mobian, with Ally Bank and the Project Management Institute as early buyers, using FAQ-format units labelled as sponsored content and priced at a premium. Media buyers are willing to test it but describe it as an experiment rather than a media channel, and one vendor running its own tests found markdown placements did not change advertiser results against HTML placements. Treat it as a testable line item, not a forecastable revenue stream.