AI Data Commons
Follow the evidence

Selective Participation in the AI Data Commons

Start with a simple question: when publishers can decide whether AI crawlers may use their content, who stays in the shared pool? The answer is not random. The most reliable publishers are more likely to pull back, especially from training, and that changes what the AI data commons contains.

Bar chart showing AI crawler restriction rates increasing from low-factual to high-factual outlets.
Manuscript Figure: High-factual outlets restrict AI crawlers at 47%, versus 8% among low-factual outlets.
9,611live media cross-section
7,002training/search sites
272,458site-months
23named crawlers
The thread

The story is about selection into access.

Each figure answers the next question in the chain. First, publishers can now declare access by crawler. Then the higher-quality publishers restrict more. Then the accessible pool shifts. Finally, the training/search split reveals what kind of market is missing: not a generic market for "web data," but a market that prices different AI uses differently.

A new choiceRobots directives turn AI access into a declared participation decision.
Selective exitThe publishers with more credible content are more likely to restrict.
Changed poolThe remaining accessible corpus tilts toward lower-reliability sources.
Different usesTraining is treated differently from search-like discovery.
Market designThe natural policy question is how to keep high-quality publishers participating.
Economic implication

The core problem is a missing price for high-quality participation.

Training access can use publisher content without sending readers back, without attribution that matters economically, and without systematic payment. Search-like access is different: it can still preserve discovery, citation, and traffic. Publishers reveal this distinction in their robots directives.

That is why binary opt-out is a poor institution. It gives publishers only a crude exit option. When high-quality publishers exit more, the accessible data commons becomes less representative, and AI systems that rely on declared-access content inherit that composition.

1. Selection, not just scarcity

The risk is not merely that less content is available. The content that remains is selected: lower-factual, questionable, and more politically extreme outlets become more represented in the accessible pool.

2. Training has a different bargain

For publishers, model training can appropriate value with weak referral benefits. Search access can still help audiences find the publisher. The data show publishers understand this difference.

3. The licensing margin is visible

The training-only restrictors are the natural target for compensated access: they are saying no to training while still allowing search. That is a market-design clue.

4. Better governance prices uses

Pay-per-crawl, collective licensing, attribution rules, and use-tiered contracts should be evaluated by whether they keep credible publishers in the training-accessible pool.

The commons changes when high-quality producers leave first.

The economic goal is to turn selective exit into compensated, use-specific participation.