Skip to content
AI

Why does training data carry no machine-readable record of whether anyone consented?

88

Opportunity

The legal question of whether AI training requires consent is largely settled, courts and regulators say it matters. The practical question of how a consent signal travels with a piece of text through the crawl, deduplication, filtering, and mixing stages of a dataset pipeline has no answer. Robots.txt is binary and coarse, applies to crawlers not trainers, and is routinely ignored by closed-source pipelines. The EU TDM opt-out under Article 4 of the DSM Directive has no standard machine-readable format that a pipeline can verify at the item level. The 2025 Data Provenance Initiative audit found pervasive missing and ambiguous consent signals across major training datasets, meaning even a builder who wants to respect consent cannot technically do so because the record is not attached to the data.

Why it matters

A per-item machine-readable consent signal is the primitive that lets regulation translate into engineering practice across the entire AI training pipeline.

How I score the opportunity

The Opportunity Score is my own read, not a measurement: how much it hurts, how often it bites, and how little exists to solve it today. Higher means I think it is more worth building.

Severity9/10

How much pain it causes when it shows up.

Frequency8/10

How often people actually run into it.

Whitespace9/10

How little good tooling exists for it today.

More problems worth solving