The models are eating what the models wrote
A year ago I thought this was how it ended. A paper published in May has changed my mind about the mechanism, though not about the mess.

“A collective fiction wearing the clothes of a fact.”
Think of a photocopy of a photocopy. The first one is nearly perfect, the second slightly soft, and by the eighth or ninth you have grey soup where the words used to be. Everyone understands this, which is why it became the standard explanation of what happens when models train on what models wrote.
But the photocopier analogy is too gentle, and it lets us off. The honest version goes further. Imagine that each time part of a copy becomes unreadable, a well-meaning person picks up a pen and writes in what they reckon the original said. Their best guess, in good faith, filling the gap so the page still reads properly. Do that for a few generations and you end up with a document that looks complete, reads fluently, and has no remaining connection to anything that was ever true.
A collective fiction wearing the clothes of a fact. That is the thing I have been worried about, and it is a good deal more unsettling than blur.
This did not start with the machines#
It is worth remembering how polluted the well already was. Long before any of this, we had two decades of content mills, search-optimised nonsense written for crawlers rather than humans, and articles reverse-engineered from whatever phrase was trending. The web was already a swamp in places, and we all learned to wade.
What the last few years added was scale and polish. The output got fluent. The tell that used to give away a content farm, that faint sense of something assembled by someone who did not care, has largely gone, because the new stuff cares in exactly the way a very good mimic cares.
The moment that crystallised it for me was the Chicago Sun-Times printing a summer reading list in May 2025 in which several of the recommended novels did not exist. Real authors, plausible titles, confident descriptions, books that had never been written. Nobody set out to deceive anyone. A gap got filled, and the filling looked exactly like everything around it, which is the whole problem in one newspaper insert.
The bit where I change my mind#
If I had written this a year ago I would have finished on a fairly bleak note about degradation being locked in. I would have been overstating it.
In May this year, researchers from King's College London, the Norwegian University of Science and Technology and the Abdus Salam International Centre for Theoretical Physics published work in Physical Review Letters looking at what actually drives the collapse. Their finding is genuinely surprising: introducing even a single data point from the real world into the loop is enough to stop the statistical collapse, and it holds even when the volume of machine-generated material is very much larger.
I like being wrong in this direction. It suggests the failure mode is not a slow inevitable rot but something closer to a circuit that needs one connection to ground. Cut the last link to reality and the thing spirals. Keep one honest wire attached and it does not.
I would not want that oversold, though, and it is already being oversold. The result concerns a particular class of statistical models under closed-loop conditions. It is a proper piece of physics, not a clean bill of health for every system trained on scraped web text. What it tells you is where to look, which is at whether your pipeline has any grounding in verified reality at all, and how you would prove it if somebody asked.
What this means if you are buying rather than building#
Most people reading this will never train a model. You will buy something that was trained by somebody else, and inherit whatever they fed it.
Which turns provenance into a procurement question rather than an academic one. The useful move is to stop asking suppliers how good their model is, which is unanswerable and invites theatre, and start asking where its knowledge comes from and how it stays attached to something real.
Ask what proportion of their training or retrieval material is verifiably human-generated, and how they would demonstrate that rather than assert it. Ask what happens when their system's own output becomes an input to the next process, because in most workflows it does and almost nobody has thought it through. Ask them to show you a case where the system refused to answer, or flagged that it was out of its depth, and treat a system that has never once declined as a warning rather than a boast.
The answers matter less than the reaction. Someone who has thought about grounding will be pleased you asked and will want to show you the plumbing. Someone who has not will explain, warmly, why you need not worry about it.
The asset nobody has valued yet#
Here is the part I find quietly interesting, and it has nothing to do with risk.
If verified human material is the connection that keeps a system tethered, then organisations sitting on large archives of it are holding something they have not put a number against. Your case notes. Two decades of engineering reports. The correspondence, the handovers, the annotated drawings, the recorded reasons why a decision went one way rather than the other. Most companies treat that as storage cost.
It is not storage. It is the ground wire.
And it points at something about our own habits too. Search engines now hand us a summary and we stop at it, which trains us out of reading sources, which means fewer of us are left who could tell whether the summary was fair. The same loop, running in people rather than parameters. I do not know how far that goes, and I am wary of anyone who claims to.
Where I have landed#
The doom version of this argument was always a bit too satisfying, and satisfying arguments deserve suspicion. Machines eating themselves makes a lovely image and a poor plan.
What I take from the last twelve months is narrower and more useful. Grounding is the thing. Not scale, not the model name on the invoice, but whether there is a real, verifiable, human-made connection somewhere in the chain, and whether anyone in your organisation can point to it.
Keep one wire attached to the world. Know where it is. Be able to show someone.
Bring us the problem.
A short, no-obligation call. If we are not the right fit, we will say so and point you somewhere better.