Where the channels overlap, the recording carries the instrument's mark as clearly as the brain's.
Where the channels overlap, the recording carries the instrument's mark as clearly as the brain's.

what an eeg model learns when it isn't learning the patient

Three July papers push EEG foundation models forward; a fourth, published the same fortnight, asks whether the benchmarks behind them measure the patient or the hospital. The answer changes what the other three are for.

Two weeks ago I wrote that the bottleneck in brain-signal systems was never the electrodes — we simply could not model a signal that is both noisy and never quite the same twice. I still think that is mostly right. But four days after I published it, a paper appeared that attaches an uncomfortable condition to it. We may not yet have a benchmark that can tell us whether the modelling worked.

Before that lands, it helps to know what is being measured.

EEG reads the brain's electrical activity from electrodes on the scalp. It is cheap, harmless and fast, and it is also noisy, because the skull smears the signal before it reaches any electrode. Now picture a model that has seen thousands of hours of these recordings without anyone telling it which came from a patient and which from a healthy volunteer. It learns only the recurring patterns. Afterwards a small labelled dataset tunes it for one specific job — finding a seizure, or judging whether a patient's memory is intact or failing. These are what the field calls foundation models.

July delivered two good ones. Start there.

a model that rebuilds a recording

An EEG recording is almost always incomplete. The patient shifts and seconds of signal vanish. An electrode works loose and its channel is dead for the rest of the session. The standard remedy is to guess the missing values from neighbouring electrodes — spherical spline interpolation, the default in MNE and therefore what most EEG pipelines in the world are doing right now.

ZUNA1.1 does that job better. It is a 380-million-parameter model that rebuilds the missing stretch, up to 30 seconds, whether the gap is a short interval or an entire channel. The code is open. So far, a good engineering improvement and nothing more.

What separates it from a better denoiser is hidden in one phrase of the abstract. ZUNA1.1 works with an arbitrary number of electrodes at arbitrary scalp positions. That is where this stops being about repair.

Because the larger problem with EEG is that two recordings cannot simply be placed side by side. A hospital records with 19 electrodes, a research lab with a 64-channel cap, and neither puts them in the same places. Comparing the two has meant discarding everything they do not share. Now there is a model that builds any layout from any other, and with it two centres' recordings can genuinely be measured against each other. Keep that in mind. By the end of this piece it will matter more than the repair does.

a model that stays awake for a week

The second paper goes after a different corner of the same problem. EEG proves its worth when it runs for hours and days. In an epilepsy monitoring unit a patient stays wired for days waiting for a seizure to happen. In intensive care the question is whether a sedated patient is seizing with no outward sign at all. That is where the data volume is, and exactly where human review runs out of capacity.

It is also where today's models break. A transformer holds the past in memory so it can work out which moments of a signal relate to which others, and that memory grows as the signal does. For a 30-second clip, no problem. For days of monitoring, it overflows.

S-CEReBrO unties the knot simply. The signal is split into fixed-size windows and the model works inside each one, so the memory stays constant however long the recording runs. The numbers are worth stating: signals 100× longer than full self-attention permits, 3× longer than low-rank linear attention, at 55% of the memory and 2.1× the throughput. Pretrained on more than 25,000 hours from more than 12,000 subjects, it reaches state of the art on 7 of 11 downstream tasks with up to 60% fewer parameters.

Constant memory and fewer parameters mean one clear thing in practice. Such a model can sit beside the bed instead of in a datacentre, and process a week of recording without its footprint drifting upward. That is as much an infrastructure result as a modelling one, and infrastructure results are the ones that survive a hospital's procurement process.

the same wall, far from medicine

Step away from medicine for a moment, because this constraint is not medical at all.

Wonder, published the same week, does something entirely different. From a single image it builds a three-dimensional space you can steer a camera through in real time, returning to places you have already been. But the more you move through that space, the more past there is to remember — the same knot S-CEReBrO was wrestling with. Their answer was the same too. Rather than holding everything, retrieve only the relevant fragments of the past.

Two teams, two unrelated fields, one conclusion. When the input never stops arriving, deliberate forgetting has to be designed in; buying more memory only postpones the overflow.

So we have two real advances, and a sign that they are not only about EEG. Which brings us to the paper that asks how we know they are advances at all.

then someone ran the negative controls

Marzieh Zare took seven published EEG foundation models and evaluated them across five clinical tasks on four datasets: LaBraM, EEGMamba, CBraMod, REVE, LEAD, BENDR and BIOT. So far, ordinary. What is not ordinary is that she also ran negative controls.

A negative control is a test you expect to fail. If it passes, your instrument is broken — like a scale that reports a weight for an empty box.

The result is clear enough. The task: read an EEG and decide whether the patient is healthy, has mild cognitive impairment, or has reached dementia. 1,187 recordings, all on a matched 19-channel layout. Features that specialists hand-designed years ago scored 0.734 macro-AUROC, on a measure where 0.5 is a coin flip and 1.0 is a perfect call. The best pretrained model, BIOT, reached 0.699. REVE reached 0.568.

So the old method beat all three of the new ones. Then comes the number worth reading twice: an encoder with random weights, trained for not one hour, scored 0.659 — higher than REVE, which had been pretrained on thousands of hours.

If pretraining did not help, what did the model learn in those thousands of hours? Zare asked exactly that. She took the model's output as it was, retrained nothing, and tested whether that output alone revealed which centre each recording came from. The answer was almost always right.

You might suspect some easy giveaway is doing the work. Zare suspected the same and erased them one at a time. She compressed the output down to fifty numbers. She removed the mains hum that leaks in from a building's wiring. She flattened the overall loudness of the signal. The centre was still identifiable.

She is careful about the interpretation, and she should be. All that is established is that the model's output carries a trace of its dataset, not that the recording site caused anything. She is equally careful in the other direction. On cross-subject seizure detection in CHB-MIT, REVE genuinely wins, 0.793 against 0.739 for the best classical comparator. And because preprocessing removes the absolute amplitude, even that win does not prove superiority over every plausible hand-designed baseline. The paper is written to improve measurement, not to demolish the field.

why a model learns the hospital

Outside neuroscience this failure has a name: a batch effect. Every dataset carries a fingerprint. The amplifier differs, the electrode layout differs, and even which patients were referred to that centre in the first place differs.

When datasets and labels line up even loosely, that fingerprint is the cheapest route to a correct answer, and a model with enough capacity takes the cheapest route. So the model is not cheating. It is doing what we asked. We asked badly.

The same shape appears in radiology, where a model told to detect pneumonia learned which scanner took the image. And I have met the small version of it shipping BioVR, a system that set the intensity of a rehabilitation exercise from live biosensor data. In practice the fragile part was never the model. It was exactly where the sensor sat on the body, and the calibration at the start of each session. A representation that quietly encodes the session instead of the patient is that same failure with more parameters.

the fix is a protocol, not an architecture

Zare's answer is not a new architecture. It is a reporting protocol. Match the electrode layout across datasets. Verify that no patient appears in both training and test data. Pick a strong classical comparator rather than the weakest available one. And test the model's output itself, not only the final accuracy. Unglamorous work, and it decides whether a field's numbers say anything at all.

And now the first two papers come back. The opening clause of that protocol is montage matching, and ZUNA1.1 does exactly that, building any layout from any other. A model we had filed under "better denoiser" turns out to be a piece of evaluation infrastructure. That was the thought I asked you to keep.

Another clause leads to S-CEReBrO. Look at the one place the pretrained models did win. Cross-subject seizure detection — the task that runs across hours or days of continuous recording, which is precisely the regime S-CEReBrO makes tractable. The audit does not flatten the field. It says which parts of it currently hold, and both of July's advances stand on the side that holds.

There is a further reason those two survive. Both can be measured without any clinical label at all. Either the model fills a dropped channel better than interpolation or it does not. Either the memory is still flat after a week of monitoring or it is not. They are honest wins precisely because measuring them does not depend on the benchmark that is in question.

The lesson generalises past EEG. Any claim we make about a model is bounded by the tests we ran against it, and a benchmark behaves like any other piece of infrastructure — it quietly decides what gets built on top of it.

If these negative controls hold up, the first question I ask at the start of my next project will change. I will stop asking which off-the-shelf model is strongest. I will ask what my evaluation would look like if the model had learned nothing at all.

Related articles