Picture the exchange CoMP standardises. A crawler arrives at a publisher's site and sends a JSON object that names itself, declares what it intends to do with the content, and quotes a licence identifier. Where no licence exists it sends the literal string PENDING_REVIEW. The publisher replies with a second JSON object that grants nothing and points at a URL where the terms live. That is the round trip. Everything after it - the negotiation, the price, the token, the invoice, the audit - happens off-protocol, by design and in the document's own words.
CoMP stands for Content Monetization Protocols. It is the work of IAB Tech Lab's CoMP Working Group, formed on 12 August 2025 after a workshop in New York that summer. Specification v1.0 was released for public comment on 10 March 2026, the comment window closed on 9 April, and v1 was finalised on 28 April 2026. It is published on GitHub under a Creative Commons Attribution licence.
What does CoMP actually specify?
Less than the name implies, and the specification file is candid about it. The document is called CoMP-1.0.md and its own heading reads "Content Metadata Marketplace Supply Specification (MVP)". That is a more accurate description of the contents than the press framing.
What v1 defines is a JSON object model in two halves. On the AI side there are two objects: AISystem, carrying a name, a user agent string, an identifier and a nested use declaration; and AISystemUse, carrying a licence identifier, an authentication method, the URIs requested, the scope of the request, and the function and sub-function the content is wanted for. The function enumeration separates ai-train, ai-input, ai-index and search; the sub-function enumeration separates training, retrieval-augmented generation, grounding, agent view and agent actions. That granularity is genuinely useful. It is the first widely published standard that lets a crawler say which of those five things it is actually doing.
On the publisher side there are five objects. Package carries an identifier, a title, a seller, a licence URL, a citation string, a reporting URL and nested scope and retrieval objects. Scope describes the commercial and editorial boundaries of what is on offer. Text, Video, Image and Audio describe the assets themselves, with fields for word count, duration, language, publication date, author, creation source (human, AI or hybrid) and provenance. Retrieval describes how to fetch the content once access is granted, with an authorisation type and a delivery format drawn from a list that includes HTML, RSS, API, MCP, NLWeb, XML and NewsML.
Seven objects, roughly sixty fields, twelve enumerated lists and eight worked examples. That is the whole of v1. There is no endpoint path, no HTTP method, no status code, no header, no well-known file location and no token format anywhere in it. The retrieval.endpoint field is a free-text URL the publisher supplies. Transport is left unspecified.
Why does CoMP not pay publishers?
Because it explicitly declines to. The out-of-scope section of the specification is short and worth reading in full rather than summarising. On commercial terms it says: "Commercial terms between Content Owners, Marketplaces, and AI Systems must be negotiated a priori to this API". On settlement it says the protocol "can be used to communicate where AI Systems can find terms to license and access content, but token issuance, counting, and payment are out of scope." On measurement it says CoMP "is a communication protocol only, as such, it does not explicitly support reporting, but it is strongly recommended that the Content Owner receives reporting from the AI System". The worked examples restate it: "While this API communicates intent, it does not handle the actual payment or token issuance."
There is a real tension inside the document on this point. The Scope object does carry price-bearing fields - pricetype, pricetier, unitprice and cur, with a price type list offering per-use, per-query, per-token, flat and tiered rate. So CoMP can express a price. It simply cannot transact one. And in all eight of the specification's own worked examples those commercial fields are left empty, with the exchange falling back to the licenseurl pointer every time. The standard's authors appear to understand perfectly well that the number is agreed elsewhere.
This matters because of how far the shipped specification sits from the ambition that launched it. When Anthony Katsur, chief executive of IAB Tech Lab, previewed the working group at AdMonsters' Sell Side Summit in August 2025, he was unambiguous about the model he thought would win: "We don't believe cost per crawl scales," he said. "We think cost per user query scales." Counting user queries requires a reporting layer with teeth. CoMP v1 has a reporturl field and a strong recommendation. IAB UK's own write-up of the initiative described it, accurately, as something that would "theoretically enable publishers to monetise each content scrape or request by an LLM".
None of which makes CoMP useless. It makes it plumbing. The honest description is that CoMP standardises the handshake so that publishers and AI companies stop building a bespoke integration for every deal, which is a real cost saving for anyone doing more than one deal. The press release makes this claim and it is the defensible one: a single standardised protocol "rather than building proprietary integrations for each platform".
What does CoMP assume the publisher has already done?
Two things, and both are larger than implementing the protocol.
The first is blocking. CoMP is not an access control system and says so. The announcement puts it directly: "CoMP is not a replacement for strong access controls. The framework assumes that content owners have established robust blocking strategies at the delivery point, such as their Edge Compute or Content Delivery Network (CDN)." The whole protocol is premised on a crawler being turned away and needing somewhere to go. A publisher with no enforcement at the edge gains nothing from adopting it, because there is no denial for the licence URL to attach to.
The second is a negotiated deal. The licence identifier that unlocks a populated response has to come from somewhere, and that somewhere is a commercial conversation that CoMP explicitly places before the protocol. For the large publishers who already have bilateral agreements, CoMP tidies up the delivery. For the long tail who have never had a call returned, it changes nothing about whether the call gets returned.
That is the uncomfortable shape of the thing. CoMP is most valuable to publishers who already have leverage, and closest to inert for publishers who do not. It is a settlement layer for a market that has not formed yet.
How does CoMP differ from pay per crawl, RSL and robots.txt?
They sit at different points and it is worth being precise, because the trade coverage tends to blur them.
robots.txt is a request, not an instruction. The CoMP working group's own bot management guidance describes its limits bluntly: it is "Entirely voluntary. There is no technical enforcement mechanism", and it is "Binary at path level. You can allow or disallow paths, but you cannot express why or under what terms."
Really Simple Licensing extends robots.txt with a License: directive pointing at a machine-readable licence document. The same guidance credits it with "partially addressing the conditional access limitation" and describes adoption as "early but growing".
CDN bot management, including Cloudflare's, Akamai's and Fastly's, is enforcement. The guidance is direct about what enforcement alone leaves out: those products "can block or allow but cannot issue a license, capture payment terms, or generate invoices", and it notes the vendor lock-in that comes with them. Cloudflare's pay per crawl product goes further than pure blocking by putting a price on access at the edge, and it is the only approach among these that actually moves money. Notably, it is never named in either CoMP document.
CoMP is the layer above enforcement and below settlement: a shared vocabulary for describing what is on offer and where the terms live. The four are complements rather than competitors, which is a less exciting story than the one the headlines tell.
What is the Bot and Crawler Management Guidance?
Released on 27 May 2026 for public comment, it is the CoMP Working Group's answer to a problem it discovered while writing the API: most publishers had not decided what their bot policy was in the first place. For a large number of publishers this document is more immediately useful than the specification it accompanies.
It is aimed at commercial decision makers rather than engineers, and it declines to tell anyone what to do. "It does not recommend blocking or allowing any specific crawler. It provides the tools to reason about the decision yourself." Its central diagnosis is that "the problem is not the existence of crawling. The problem is invisibility."
The method it proposes is a scoring exercise. Rate each crawler one to five on traffic return value, content use value, bandwidth and infrastructure cost, reputational and brand value, and strategic value, then place it in one of four buckets: allow, allow with conditions, require licensing, or block. The guidance is careful to note that "The matrix is a tool for making your reasoning explicit - it is not a formula for an automatic decision." It warns against the extremes, pointing out that turning everything off will "break systems that directly contribute to revenue - such as disabling ads.txt verification, resulting in very little to no programmatic revenue."
It also recommends publishing your position rather than only enforcing it: a human-readable crawler policy document, a machine-readable robots.txt, and a reachable licensing contact. The reasoning is commercial. "Making your policy visible and your contact reachable turns a passive block into a potential relationship."
What should publishers do about CoMP now?
Read the bot guidance, act on the specification later. The sequencing the guidance itself recommends is visibility, then decision, then enforcement, then protocol, and it warns that "Content owners who act first (block/allow) without visibility often make decisions based on incomplete information." Most publishers are still at step one.
On adoption, the honest current-state answer is that there is no public record of a live CoMP implementation. No publisher, marketplace or AI company has announced one, and none of the major AI labs is a declared participant. Katsur said in August 2025 that Google and Meta had been in the room at the Tech Lab's LLM workshop, and that OpenAI, Anthropic and Perplexity "won't even return the Tech Lab's calls". Nothing published since suggests that has changed. A standard for negotiating with AI systems works only if AI systems negotiate.
Publishers should also be clear about which transaction CoMP governs. It covers crawl access: the terms on which a bot may take a copy of your content. It has nothing to say about the separate commercial surface that opens when a live search agent fetches a page in order to answer a user's question in real time, which is where an edge-layer approach such as blankspace operates. Those are different moments with different economics, and adopting a standard for the first does not address the second. Neither does the reverse.
Frequently asked questions
Does CoMP mean AI companies have to pay before crawling?
No. CoMP gives a publisher a standard way to say that access requires a licence and to point at where the terms are. It has no enforcement mechanism of its own, and it assumes the publisher is already blocking unlicensed crawlers at the CDN or edge. Whether an AI company pays is determined by that enforcement and by a commercial negotiation, both of which sit outside the protocol.
Is CoMP final or still in draft?
CoMP v1 was finalised on 28 April 2026, after a public comment period that ran from 10 March to 9 April 2026. The accompanying Bot and Crawler Management Guidance was released separately on 27 May 2026 for public comment. The specification file describes itself as an MVP, and the repository carries no formal releases or version tags beyond the 1.0 filename.
What is the difference between CoMP and the LLM Content Ingest API?
They are the same initiative at different stages. The work was first introduced as the LLM Content Ingest API and renamed to Content Monetization Protocols when the working group formed in August 2025. The Ingest API name still appears in IAB Tech Lab's marketing material and in trade coverage, but there is no component by that name inside the v1 specification file.
Does CoMP replace robots.txt?
No. It sits above it. robots.txt expresses a binary path-level preference with no enforcement and no way to state terms. CoMP is a vocabulary for describing what content is available, under what scope, and where the licensing terms live, and it presumes enforcement is handled elsewhere. IAB Tech Lab recommends publishers maintain both, alongside a human-readable crawler policy.
Should a small publisher implement CoMP?
Implementing the object model is not the constraint for a small publisher; having something to negotiate with is. CoMP is worth adopting once a publisher has working enforcement at the edge and at least one counterparty willing to transact. Before that point, the bot and crawler guidance, an accurate picture of which crawlers are taking what, and a reachable licensing contact address will do more commercial work than the specification will.
