Skip to content
AI

Why does training data carry no machine-readable record of whether anyone consented?

88

Möglichkeit

The legal question of whether AI training requires consent is largely settled, courts and regulators say it matters. The practical question of how a consent signal travels with a piece of text through the crawl, deduplication, filtering, and mixing stages of a dataset pipeline has no answer. Robots.txt is binary and coarse, applies to crawlers not trainers, and is routinely ignored by closed-source pipelines. The EU TDM opt-out under Article 4 of the DSM Directive has no standard machine-readable format that a pipeline can verify at the item level. The 2025 Data Provenance Initiative audit found pervasive missing and ambiguous consent signals across major training datasets, meaning even a builder who wants to respect consent cannot technically do so because the record is not attached to the data.

Warum es wichtig ist

A per-item machine-readable consent signal is the primitive that lets regulation translate into engineering practice across the entire AI training pipeline.

Wie ich die Chance bewerte

Der Opportunity Score ist meine persönliche Einschätzung, keine Messung: wie stark es schmerzt, wie oft es auftritt und wie wenig heute existiert, um es zu lösen. Ein höherer Wert bedeutet, dass ich es für lohnender halte, es umzusetzen.

Schweregrad9/10

Wie viel Schmerz es verursacht, wenn es auftritt.

Häufigkeit8/10

Wie oft Menschen tatsächlich darauf stoßen.

Whitespace9/10

Wie wenig gute Werkzeuge dafür heute existieren.

Weitere lösungswürdige Probleme