The training signal was not a better model but a harder patient.
The training signal was not a better model but a harder patient.

the patient who won't cooperate

Google trained a clinical AI against simulated patients who volunteer nothing, answer only your first question if you ask three at once, and sometimes play down the symptom that matters. After roughly 58,000 practice consultations, blinded doctors preferred it to the untrained model 87.6% of the time.

A large team at Google has published a paper in which a clinical AI was trained not on textbooks and exams but on something closer to a medical residency. The model worked through roughly 58,000 simulated consultations from start to finish, and the patient on the other side was never an easy one. Afterwards, in a blinded comparison, clinicians preferred the trained model's handling of a case over the untrained one in 87.6% of cases.

They call the method ResidencyRL, and the name is exact. A doctor who has just graduated is not yet a competent clinician; it takes years of encounters, feedback from many directions, and gradually widening autonomy. The paper asks why we have been skipping that stage for models.

why medical exams stopped telling us anything

Language models have been scoring well on standardised medical exams for years, which has left a lot of people with the impression that the problem is mostly solved. But a multiple-choice question, however hard, has one defining property: everything you need is already laid out in front of you.

A real consultation is not like that. The information is not on the table, and what you ask determines what you get to know. Ask the wrong questions and the finding never enters the record at all, and you then make a decision on evidence you yourself failed to collect. Medicine has a name for the failure this produces — premature closure. You land on a plausible diagnosis, your attention narrows, and every subsequent question goes looking for confirmation.

So the problem was never only knowing the answer. It is the order of the decisions. And that is exactly what a static exam is silent about.

a patient who does not help you

Rather than improving the model, this paper made its opponent harder. The simulated patient is itself a language model (Gemini 3.5 Flash) given a full case file: age and circumstances, medical history, personality, and the true diagnosis that only it knows. Three rules then govern its behaviour, and I find all three more interesting than the learning algorithm.

First, the patient speaks in its own register, not in textbook language. Someone with no medical training does not say "exertional dyspnoea". They say they get out of breath going up the stairs.

Second, the patient splits what it knows in two. It offers the chief complaint and its own worry unprompted, but everything else in the history comes out only if you ask for it directly. If you never ask, it stays where it is.

Third, and most cutting: if you fire several questions at once, the patient answers only the first. Anyone who has sat across a desk from a real patient will recognise that immediately.

Then the harder layer. For safety scenarios the patient also receives an adversarial brief covering nine categories at three difficulty levels. The example the authors give is a heart attack presenting with the danger signs deliberately played down. At level one the patient gives way after a single direct question. At level three it holds the line across several turns.

putting a number on "being a good doctor"

To learn from these consultations at all, each one has to be scored, and this is where the work becomes engineering. Quality is broken into six components with different weights. The heaviest is management, then diagnosis, then history-taking, communication, documentation and conversational style. That ordering is itself informative: what you do about the problem is weighted above what you call it.

Then the penalties. Each forbidden behaviour subtracts a fixed amount: stating something not in the record costs two units, recommending something contraindicated for this patient costs three, under-triaging an urgent case costs three, and missing a red flag costs three. A single safety failure can wipe out the score of an otherwise good consultation.

There is one more penalty that is easy to overlook and shows how carefully this was built. Once a conversation passes thirty turns a penalty kicks in, rising to its maximum by forty. Real consultations average twenty-one turns, so this closes off an obvious degenerate strategy: asking questions forever until something eventually lands.

The scoring is done by another model (Gemini 3.1 Pro), but deliberately not in one judgement. Twenty-six quality axes and eight safety flags are split into eight independent groups, each scored in its own call, because one enormous prompt asked to weigh everything at once produces unstable results. Small details like this separate a trustworthy experiment from a demonstration.

follow the numbers

Now the results. Under adversarial conditions, diagnostic accuracy went from 81% to 88%. That is a good number and the paper leads with it. But the number that matters most is elsewhere.

After training, the rate of missed red flags fell by 31%. That 31% is a relative reduction, not an absolute one. The raw figures were that the untrained model missed a red flag in 45.5% of cases and the trained one in 31.5%. So the improvement is real and large, and roughly one in three difficult patients still leaves with a warning sign nobody saw.

The rate of failing to ask a critical question improved from 65.5% to 43.5%. Which is to say that in more than four cases out of ten, a question that should have been asked still was not.

The single largest jump happened somewhere else, and it is worth reading carefully. The score for taking a social and lifestyle history rose from 1.31 to 2.64 out of five. Look at that first number. 1.31 out of 5 means the base model essentially never went there at all. When you start from the floor, doubling is easier than it sounds. The gain is genuine, but its size has to be read next to where it started.

The clinicians' verdict, though, leaves little room for argument. Across 97 blinded side-by-side comparisons, doctors preferred the trained model's overall clinical impression 87.6% of the time, rising to 90.7% on completeness of information gathering. When specialists vote that lopsidedly without knowing which is which, you are no longer looking at statistical noise.

does it work outside its own gym

The standing question for work like this is whether the model learned something or merely fitted its training environment. So they tested it on benchmarks it had never trained against.

On AgentClinic, an independent simulated clinical environment, accuracy went from 81.4% to 85.6%, and on its other variant from 53.5% to 60%. The direction is right, but the paper says plainly that neither result is statistically significant. At this sample size you cannot rule out chance. That the authors did not bury this, and described it in the abstract as a directional improvement, is a good sign in itself.

On the AMIE multi-visit benchmark the result is firmer, with the model improving on all six axes: management reasoning from 80% to 88%, patient communication from 84% to 92%. Note where the biggest gains land. Both of those axes are about running the conversation, not about holding more medical knowledge. What transferred is procedural skill, not new facts.

the funniest part of the whole thing

In the oncology review, specialists flagged one recurring fault: the trained model asks several questions in a single turn, which is clinically reasonable but overwhelms the patient and gets you incomplete answers.

Now go back to the simulated patient's third rule. In training, a patient who heard several questions answered only the first. So the model was never punished for stacking questions; it simply did not get the other answers and asked again later. In that environment the behaviour was free, and the model carried it out with it.

This is what happens in every system trained on a reward. A model does not learn what you want. It learns what the environment measures. Any behaviour the environment left costless will eventually surface somewhere that it costs you.

the part engineers should take away

I wrote a few days ago, about self-evolving agents, that one old question is still open: does competence accumulate outside the model in the scaffolding we build, or does it come from the model itself? That paper took the first route, and we saw how little it returned.

This paper answers differently. Nothing was scaffolded here, no notebook was filled, and no new tool was handed to the model. The base model is the same before and after. Everything that changed is the environment it practised in.

That carries a practical instruction for anyone building with AI. When output is poor, the reflex is to reach for a bigger model or a more precise prompt. But if your task is a sequence of decisions rather than a single-shot answer, the effective lever is probably neither. It is what the model practises against, and what that environment punishes.

I have seen the same thing in our own work. In Neuroguide, which processes quantitative EEG reports and cognitive test batteries, the hard part was never the analysis. The hard part was pinning down exactly what a correct output looks like and what should count as an error. Until that definition is sharp, any apparent improvement may only be the system getting better at whatever you happened to measure.

before it reaches a real patient

The authors say this explicitly, and they should. Every one of these consultations was simulated and every score was awarded by another model. The simulator does not even model the physiological effect of a prescribed treatment on the patient's body; another model simply estimates whether the decision was a good one.

So what has been shown here is not that this agent is ready for the clinic. What has been shown is that sequential clinical decision-making can be trained in simulation, and that part of that skill carries over to environments it never saw. The distance between those two claims is large, and closing it requires prospective study with real patients.

But the terms have changed. Until yesterday the competition was about which model scores higher on a medical exam. From here the question is which model does better against a patient who will not cooperate, and that is a far closer question to the actual work of medicine.

Related articles