What is it about?
Imagine learning only by rereading your own notes instead of discovering new ideas. Over time, your thinking would become narrower and more repetitive. We show that AI faces a similar problem as it increasingly trains on AI generated content. The solution is surprisingly simple: expose the model to examples that still surprise it. By prioritizing information that is novel and informative, AI remains more diverse, more reliable, and better at common sense reasoning, even in a future where much of the web is AI generated.
Featured Image
Photo by cosmic scape on Unsplash
Why is it important?
Most existing solutions assume that we can always identify and recover human written content. Our work takes a different approach: instead of asking who wrote the data, we ask how surprising it is to the model. By prioritizing the most surprising and informative examples, our method remains effective even in a future where much of the web is AI generated.
Perspectives
I hope this work makes what may seem like a technical AI problem feel relevant and thought provoking. As AI generated content becomes an increasingly large part of our digital world, understanding how AI continues to learn is a challenge that affects us all. Perhaps the most important lesson is a simple one: learning needs surprise.
Gizem Gezici
Scuola Normale Superiore
Read the Original
This page is a summary of: Learning by Surprise: Adaptive Mitigation of Model Collapse in Large Language Models, ACM Transactions on Intelligent Systems and Technology, July 2026, ACM (Association for Computing Machinery),
DOI: 10.1145/3828663.
You can read the full text:
Contributors
The following have contributed to this page







