Why does training data carry no machine-readable record of whether anyone consented?
Opportunity
The legal question of whether AI training requires consent is largely settled, courts and regulators say it matters. The practical question of how a consent signal travels with a piece of text through the crawl, deduplication, filtering, and mixing stages of a dataset pipeline has no answer. Robots.txt is binary and coarse, applies to crawlers not trainers, and is routinely ignored by closed-source pipelines. The EU TDM opt-out under Article 4 of the DSM Directive has no standard machine-readable format that a pipeline can verify at the item level. The 2025 Data Provenance Initiative audit found pervasive missing and ambiguous consent signals across major training datasets, meaning even a builder who wants to respect consent cannot technically do so because the record is not attached to the data.
Why it matters
A per-item machine-readable consent signal is the primitive that lets regulation translate into engineering practice across the entire AI training pipeline.
How I score the opportunity
The Opportunity Score is my own read, not a measurement: how much it hurts, how often it bites, and how little exists to solve it today. Higher means I think it is more worth building.
How much pain it causes when it shows up.
How often people actually run into it.
How little good tooling exists for it today.
More problems worth solving
Why does every AI app forget me the moment I close the tab?
AIWhy is learning a new field still gated by knowing what to ask?
AIWhy can a non-expert not verify what an AI just told them?
AIWhy do we test models on benchmarks but ship them on vibes?
AIWhy do AI agents have no memory of their own mistakes?
AIWhy can't I audit what a model was actually trained on?