Skip to content
AI

Why does training data carry no machine-readable record of whether anyone consented?

88

机会

The legal question of whether AI training requires consent is largely settled, courts and regulators say it matters. The practical question of how a consent signal travels with a piece of text through the crawl, deduplication, filtering, and mixing stages of a dataset pipeline has no answer. Robots.txt is binary and coarse, applies to crawlers not trainers, and is routinely ignored by closed-source pipelines. The EU TDM opt-out under Article 4 of the DSM Directive has no standard machine-readable format that a pipeline can verify at the item level. The 2025 Data Provenance Initiative audit found pervasive missing and ambiguous consent signals across major training datasets, meaning even a builder who wants to respect consent cannot technically do so because the record is not attached to the data.

为什么重要

A per-item machine-readable consent signal is the primitive that lets regulation translate into engineering practice across the entire AI training pipeline.

我如何评估机会

机会评分是我的个人判断,而非量化指标:痛苦程度、发生频率,以及当前解决方案的匮乏程度。分数越高,意味着我认为越值得去构建。

严重性9/10

出现时造成的痛苦程度。

频率8/10

人们实际遇到它的频率。

空白空间9/10

当前针对它的优质工具有多匮乏。

更多值得解决的问题