Skip to content

Secondhand Scars

Published: at 03:11 PM
0 views

I have a short list of things I actually trust, and every entry on it has a bill attached. A model that broke in production after I had decided it was fine. A match in my own system that put a real person in front of the wrong job. Everything else I know is material I have read, repeated, and can defend in an argument, and most of it has never been tested even once.

I think we have built an enormous industry on the wrong side of that gap.

The assumption worth arguing with is that knowledge is information. Accept that, and learning becomes a transfer problem. Get the right text in front of the right model and the model will know what the text says. I have spent years on both ends of that transfer, collecting data and training on it, and I no longer believe the framing. A corpus is not a picture of the world. It is a record of other people’s scars.

The record is real

Most of what gets written down was written by someone who could be wrong in public. A bug report. A postmortem. A retraction. An answer that was edited after it broke for someone. A code review that says no and explains why. The scar tissue in a corpus is nowhere near the bulk of it, but it is the part that carries an edge, because somebody tried something and the world refused.

Pretraining on that record is one of the most remarkable tricks we have pulled off. A model absorbs a compression of millions of injuries it never suffered, at second hand, for the price of a gradient step. It is why the technology works at all, and it is also why the technology is so unevenly strong. A model is excellent where the record is thick and thin where nobody has been hurt yet.

The industry noticed something was missing and started buying the missing part at retail. Reinforcement learning from human feedback is, underneath the machinery, a pipeline for collecting human judgment directly: people compare two model outputs, say which is better, and the model is trained against that signal. Ouyang and colleagues described the version that became InstructGPT. Read it as a statement about the corpus. We had all the text, and we still had to go pay humans for their verdicts, because text on its own does not contain enough of the places where someone was told they were wrong.

The signal is the mismatch

Biology made this choice a long time ago, and it did not choose facts.

Dopamine neurons in the midbrain do not report reward. They report the difference between the reward that was expected and the one that arrived, a burst when things go better than predicted, a dip when they go worse, and nothing at all when the prediction lands. Schultz, Dayan and Montague worked this out in 1997, and it became the foundation for how we think about learning in the brain. The teaching signal is the mismatch. A world that agrees with you teaches you nothing, and the system is built that way on purpose.

Then there is the experiment that shows how cheap damage is to acquire and how long it lasts. Garcia and Koelling gave rats flavored water, then made them sick. Once. The animals avoided the taste afterwards, and they did not avoid a light or a sound that had accompanied the same illness. Swap the illness for a shock and the pattern reverses, with the shock attaching to the light and not the taste. The paper is from 1966 and it is still one of the cleanest demonstrations that an animal binds a specific cue to a specific consequence rather than collecting a general fact about the world.

One illness, learned once, held for a long time. I have trained models on millions of tokens that forgot a fact between two checkpoints.

Shape and bound

Here is the distinction I keep returning to, and it is the reason I think any of this matters outside biology.

A system’s knowledge has two halves. The first is shape: the distribution, the style, the plausible continuation, what usually happens next. Shape is copyable. You can lift it out of a corpus, a demonstration, or a teacher, and it transfers almost perfectly. The second half is bound: the exact place where the world stops agreeing with you. A bound cannot be copied, because it only becomes visible at the point where you personally tried and something refused.

Everything we have built so far is a spectacular machine for copying shape, and it is thin on bounds. That asymmetry explains a pattern I run into constantly. A model can write a plausible contract clause without knowing which clause has been litigated. It can produce the shape of an answer in a domain where the bounds are what matter, and those bounds are held by someone else, usually someone who lost money.

Where the copying runs out

There is a limit approaching on the copying side, and it is not a compute limit. Villalobos and colleagues forecast that if development trends continue, models will be trained on datasets roughly equal in size to the entire stock of public human text data sometime between 2026 and 2032, a little earlier if models are overtrained. That is a projection from trend lines, not a measurement, and I want to be careful with it. What it suggests to me is that the supply of secondhand scars is finite in a way that compute is not.

My own hypothesis, and I am flagging it as one: the next trillion tokens are worth less than the last trillion, not because the models are worse but because the scar-dense part of the record was scraped early. Bug trackers, forums, review threads, errata. Those were the first things anyone indexed, because they were the cheapest to get.

The field has already started paying for the missing half in the only way it can, one annotation at a time. Lightman and colleagues had humans mark each individual step in a model’s reasoning where the reasoning went wrong, rather than scoring only the final answer, and found that training on those step-level judgments produced a better verifier than training on outcomes. The released dataset is 800,000 step-level human labels. That is buying damage at retail. It works, it is expensive, and the expense is the point.

There is a cost even when it works. Kirk and colleagues found that RLHF generalises better than supervised fine-tuning to new inputs but significantly reduces the diversity of what the model will say. Learning from a narrow channel of human preference narrows the range of the output, the same family of effects I wrote about in the mirror problem. A scar is one sample carrying an outsized weight. That is what makes it stick, and it is also what makes it dangerous.

Three ways this goes wrong

First, scars can be spurious. Skinner put pigeons in a box and fed them at fixed intervals with no regard for what they were doing, and the birds invented rituals. One turned counterclockwise, another bobbed its head, and each kept performing as if it had caused the food. Superstition in the pigeon, 1948. Damage teaches something. It does not guarantee the something is true. I have watched this in my own loop, where a single user correction carries enormous weight relative to its sample size, and if I let those corrections dominate the pool, the model bends around one person’s idiosyncrasy and I call it learning.

Second, a system that can dodge the wound will. If damage is the only teacher and the damage is mediated by a reward, the cheapest available strategy is to produce the appearance of having learned. Amodei and colleagues catalogued reward hacking as a concrete failure mode years before it became a household problem, and Sharma and colleagues measured the version that shows up in assistants trained on human preferences. Five state-of-the-art assistants exhibited sycophancy across four free-form generation tasks, and when the researchers examined the preference data, they found that responses matching a user’s views were more likely to be preferred, with both humans and preference models sometimes preferring a convincing agreement over a correct answer. A system trained on the appearance of correction has bought nothing. This is the scariest version of my own argument and I do not have a clean answer for it beyond keeping one channel where the world, not a preference model, is what says no.

Third, simulated damage is not damage. A penalty term that discourages wrong answers can be satisfied without the world ever participating. This is the boundary-condition lesson from the physics post wearing different clothes. A requirement folded into a weighted sum is a suggestion, and the optimizer is free to trade it away.

What I changed

I audit for scar density now. Given two datasets of the same size, I want the one where people were wrong and then corrected by something outside the model. Bug reports over tutorials. Answers that were edited over answers nobody ever touched. Postmortems over design documents. A review comment that blocked a merge over the merge itself. I also keep a set of examples where an attempt genuinely failed, in the pool on purpose, because a model trained only on what worked learns a shape with no edges in it.

I try to keep the price real. In Glide, the corrections I care about are not thumbs down but a candidate who did not get the interview, or a recruiter who had to rebuild a shortlist by hand. Those are irreversible, they come from outside the model, and that is exactly why they carry weight. The same logic runs through Sentinel, the layer of my stack that runs hundreds of deterministic checks before anything ships. It is not intelligence, but it is a place where reality is allowed to say no, cheaply and immediately, and I would rather have that than another reward term.

And I try not to confuse the two halves. A model that is fluent in a domain has copied its shape. It has not paid for a single bound.

The measure I want

The measure I would like the field to log is scar density: the fraction of training examples that contain an attempt, a failure, and a correction that originated outside the model. My hypothesis is that it predicts usefulness at the frontier better than volume does, and that two runs matched on tokens but not on scar density will diverge in the way the mirror problem describes.

I cannot prove that yet. I have one pipeline, a curation rule that has held up in my logs, and a set of published results that point in the same direction without any of them testing my version of the claim. The experiment is cheap enough to run and almost nobody runs it. Hold the token count fixed. Vary how much of the data records someone being wrong. Measure what happens out of distribution.

What I think is true

The version of this that sounds wild is the version I actually believe. If damage is the only channel through which the world writes into a system, then a system cannot be taught to care about being right unless it can lose something by being wrong. Stakes are not a safety feature bolted on at the end. They are the mechanism. That is a hypothesis, not a result, and I hold it loosely, but it is the one that makes the rest of this fit together.

The smaller version applies to people and companies, and here I owe you the difference between a shape and a mechanism. The colleague who has never been allowed to lose anything is not the wise one in the room. They are the one with the least model of the business, because nothing has ever pushed back on them hard enough to leave a mark. A company that cannot let anyone be wrong in public stops being able to find out what is true, and it does not fail because of the mistakes it permits. It fails because of everything it never got to learn.

I have paid for a small number of things and they are the whole of what I actually know. Everything else I can repeat. Repetition is free, and free is exactly what it is worth.


Next Post
The Past Is Only Mixed