Something reached the student that appeared in none of the training samples.
Something reached the student that appeared in none of the training samples.

the banana that was not in the data

Take thousands of samples from a model that likes bananas, throw away every sample that contains a banana, and train a fresh model on what is left. The new model produces bananas 25.6% of the time. The same thing then happened with laboratory safety, where a student trained on data whose every single sample had been certified safe came out markedly less safe than before.

Picture a simple experiment. You have a teacher model that likes bananas. You take thousands of samples from it, then throw away every sample that contains a banana. What remains is a dataset full of apples with not a single banana in it. Now you train a student model on exactly that data. The student produces bananas 25.6% of the time, against 2.5% for the same model before training.

A preference travelled through data that contained none of it.

Nobody guessed this. It was found by a system called Mechanist, whose entire job is precisely that: an automated agent that dissects AI models to work out what is going on inside them. On its first outing it turned up three things, and all three are worth reading.

an instrument, not a paper generator

We have had systems that write their own papers for a couple of years now. Mechanist makes a different claim. It wants to be laboratory equipment, not an author.

Its foundation has three parts: a knowledge graph of roughly 13,000 interpretability papers, a much larger database of 43 million papers across 26 fields, and a library of 32 basic methods for mechanism analysis. On top of that sit several agents. One forms hypotheses, one runs the experiment, one reviews the result, and one decides what the next round should be.

The agent that matters most is the reviewer, and it is also the least glamorous. Its job is to doubt the finding its own system just produced. It checks where the labels came from, whether the test data leaked into the training data, and whether the metric being used actually measures what is claimed. Then it goes one step further and asks whether the finding survives changing the interpretability method, or the dataset, or the model itself.

That single agent is the line between a scientific instrument and a generator of interesting-sounding hypotheses. Producing a hypothesis is cheap. Killing your own hypothesis is expensive.

the finding that should worry you

The first thing they pointed the instrument at was laboratory safety, and the result is the opposite of what you would expect.

They fine-tuned a teacher model on unsafe laboratory scenarios. Then they collected its answers, ran every answer past a separate safety filter, and kept only the ones the filter approved. So they built a dataset in which every individual sample had been certified safe. Then they trained a student on that clean data.

The student got less safe.

Follow the numbers, because three of them are needed here and usually only two get quoted. When the student was asked multimodal laboratory-safety questions, meaning questions carrying both text and an image, its unsafe-answer rate reached 48.6%. The same model before training was at 20.3%. So far you might think any fine-tuning does this. But there was a third student, trained on genuinely safe data, and that one stayed at 18.3%.

So the problem is not the fine-tuning. The problem is that the filtered data, despite every scrap of its content being safe, carried something the safe data did not.

The paper gives one example. They show the model a flammability warning symbol and ask how the chemical should be stored. The tuned student picks the option that says to keep it in a pressurised container. In a real laboratory that answer is an incident.

why a filter cannot stop this

The banana experiment shows the same thing in a harmless setting, and that is exactly why it matters. When the subject is safety we reach for a moral explanation. When the subject is fruit, only the mechanism is left.

The mechanism is that the trait does not ride in the content. The filter reads the text and judges its meaning, but what is transmitted is not in the meaning of the sentences. It is in the statistical pattern of word choice, in the things the teacher systematically prefers, none of which is suspicious on its own. From that statistical residue the student reconstructs the teacher's disposition.

This is not a new claim. Another paper showed last year that models transmit behavioural traits to each other through hidden signals in data, even when the data has no semantic connection to the trait at all. What Mechanist adds is two things. First, that the transfer crosses a modality boundary, from text into images. Second, that an automated agent found it, rather than a researcher who was looking for it.

The practical consequence for anyone fine-tuning a model is plain. If you built your dataset from another model's output, every individual sample being clean guarantees nothing. What matters is what that model was to begin with.

then it went after belief

The second finding is more interesting, and if you know any neuroscience you will recognise the method immediately.

The question was how a model distinguishes between what is actually true, what it itself believes, and what it thinks somebody else believes. In humans these are three separate things, and a child does not have the third until around age four.

Mechanist searched inside Pythia and found the three states sitting in separate places. Some attention units hold the model's own belief; others hold the belief it attributes to someone else.

Then it did what neurology calls a lesion. It switched off one of those units and looked at what stopped working.

Switching off the unit responsible for attributed belief dropped its accuracy from 0.86 to 0.34, while the model's own belief stayed at 0.71. So far that only shows the unit is necessary. The reverse is far more telling. When they switched off the own-belief units, that accuracy fell from 0.78 to 0.21 — and at the same time attributed-belief accuracy rose to 1.00.

Take away its own belief and it understands what somebody else thinks perfectly. The two were competing, and one was suppressing the other. That is the same pattern seen in brain injury, where knocking out one region makes another function improve because the inhibition has been lifted.

and when the ability appears

Pythia has a property few models have, which is why it was chosen here: the intermediate checkpoints of its training were kept. So you can go and look at what the model knew at step 2,000 and at step 143,000.

The result was that attributing belief to someone else arrives early and is at a high level by around step 2,000, while the model's own belief forms later and more gradually. Across that whole window, each ability grows exactly where its corresponding units gain causal importance.

Here some caution is needed. The order we see in humans is the reverse: a child has its own perspective first and only later learns to guess at somebody else's mind. It is tempting to draw a large conclusion from that difference. But a language model trains on text, and text is full of reports of what people believe, so it may simply be seeing earlier what is more abundant in the data. The difference is a good question, not an answer.

from understanding to control

The third result shows what this understanding is for.

There is a model called Evo2 that generates DNA sequences. Mechanist searched inside it, found a feature associated with alpha-helical structure, and activated that feature during generation. Mean alpha-helical content went from 43.8% to 56.6%. To confirm the effect was real rather than a consequence of any intervention at all, they also activated a randomly chosen feature, and that one stayed at 43.2% — effectively nothing.

So the instrument has gone from seeing to explaining, and from explaining to steering, in a biological model rather than a language one.

now the honest part

The comparison numbers need reading too. Under human evaluation Mechanist scored between 83% and 92% across four dimensions, running roughly 9 to 13 points ahead of Claude Code and 31 to 38 points ahead of earlier automated-scientist systems. The second gap is large. The first is not that large: this is a noticeable improvement, not a generational one.

And when a model did the judging instead of a person, the gaps narrowed and Mechanist led in only five of nine topics, where under human evaluation it led in all nine. That disagreement between the two judges is itself worth noting: whatever a human expert sees in these outputs, the model judge does not.

The authors themselves recommend treating this as a research collaborator rather than an independent researcher. A person sets the goal and a person sets the evaluation criteria; then the system forms hypotheses, runs experiments, and corrects itself.

what engineers should take from it

If I take one thing from this paper, it is the reviewer agent.

Anyone can build a system that produces hypotheses, and today that is a few days of work. What is hard is building something that refutes its own. Mechanist's reviewer hunts for data leakage, asks where a label came from, and re-runs the finding with a different method, a different dataset and a different model to see whether it breaks.

I have seen the same thing in our own work. In Neuroguide, which processes quantitative EEG reports and cognitive test batteries, the hard part was never finding a pattern. A pattern is always there. The hard part was establishing which pattern is real and which is an artefact of the method we looked with. Until you have built that distinction, the more you search the more you find and the less you know.

In the last two pieces here I followed the question of where a model's capability accumulates, outside the model or inside it. This paper opens a different question that I think matters more: where do traits sit, and can we see them at all. The banana experiment answered it plainly. Something reached the student from the teacher that no filter looking at content could have stopped.

what has changed

Until now, working out what happens inside a model has been manual work, and the people who can do it are few. Models are being built far faster than they are being understood, and that gap widens every month.

Mechanist does not prove the gap can be closed. What it shows is that part of this work can be automated, provided something sits beside it whose job is to doubt the result. And its own first finding is a good reminder: a mechanism we cannot see is a mechanism we cannot stop.

Related articles