Skip to content
AI

Why do models trained on today's web get progressively worse as AI writes more of it?

82

Opportunité

The web is now the primary training corpus for frontier models and is already saturated with AI-generated text that no deployed filter reliably catches. Research published across 2024 to 2026 shows that even a fraction of a percent of synthetic data in a training run triggers distributional collapse over successive generations, narrowing output diversity and degrading tail performance. The feedback loop is structural: models trained this year produce content that contaminates the corpus for next year's training run. Proposed mitigations such as source-level allowlists, watermark filters, and synthetic-data verifiers each have bypass vectors and none has been deployed at web-crawler scale. There is no agreed protocol for identifying and quarantining AI-generated training data before it enters a model.

Pourquoi c'est important

A degrading shared training corpus sets a ceiling on every model built from public data, and that ceiling gets lower with each generation.

Comment j'évalue l'opportunité

Le Score d'Opportunité est mon évaluation personnelle, pas une mesure : l'intensité de la douleur, sa fréquence et le peu de solutions qui existent aujourd'hui. Plus il est élevé, plus je pense que le problème vaut la peine d'être résolu.

Gravité8/10

L'intensité de la douleur qu'il provoque lorsqu'il se manifeste.

Fréquence8/10

La fréquence à laquelle les gens y sont réellement confrontés.

Espace libre8/10

Le peu de bons outils qui existent pour y remédier aujourd'hui.

D'autres problèmes qui méritent d'être résolus