A library that grows fastest where it helps least is not scaffolding; it is a warehouse.
A library that grows fastest where it helps least is not scaffolding; it is a warehouse.

the weakest agent builds the biggest skill library

Four papers published in a single week test the premise behind every agent skill library: that writing down what worked makes an agent better. Most of the gain turns out to come from somewhere else.

A few weeks ago I wrote that the line between two kinds of system has to be drawn in the right place. In one, competence accumulates in the engineered scaffolding outside the model; in the other it arises from inside the model itself. That piece was a conceptual distinction. Four papers published in a single week now test the same line empirically.

Today's most serious attempt to make competence accumulate in the scaffolding is the skill library. The agent performs a task, writes down what worked, files it away, and retrieves that note the next time something similar comes along. The idea seems so obviously right that few people have questioned it, and that unexamined premise is exactly what these four papers went after.

a benchmark that asked what actually improved

ContinualSkillBench, from Peking University, is an environment built for this one question. Five domains, each with a hundred interconnected tasks ordered from easy to hard and deliberately arranged so that a skill learned on one task should pay off on the next. If accumulated experience works at all, it should show up here.

Up to a point it does, because an agent working through the tasks in sequence gets better as it goes. But then they ran one simple experiment that changed the picture. They removed the library entirely and left the agent only the last few tasks it had just worked through. The result was about the same.

So most of the improvement we were crediting to the library came from the agent having recently seen a few similar tasks. The library filled up, but what it was doing looked more like short-term memory than skill.

Notes are not useless, and they help where a task has a fixed repeatable procedure or the output has to be exact. But the next finding is worse. Weaker models build larger libraries, packed with small scattered notes that each serve only one specific task. The library fills fastest exactly where it helps least.

that does not mean the idea failed

Stop there and you would conclude the whole idea was wrong. The second paper shows it was not, but to say so with confidence you first have to solve an old problem.

The problem is this: when an agent performs better after practice, how do you know the improvement came from experience rather than from the test task resembling the training task? GDPevo solves it with one clean trick. They take a real enterprise workflow and break it into small independent rules. Those rules are then scattered piecewise across the training tasks, and at test time the same rules are reassembled in a different arrangement. Now if the agent does better on the test, you know it has never seen that task before, only its parts.

The benchmark is built on genuine CRM, ERP, finance, healthcare and legal workflows, and its first release holds 120 tasks in 12 groups. Because the whole thing is generated automatically, they grew it to 240 tasks in two days. That is itself the answer to an important problem: when models start memorising the benchmark, you build a fresh one and you are done.

The result shows both sides. Experience genuinely helps, lifting accuracy on tasks the agent has never seen by up to 16.44 percentage points. But hand the agent those same rules up front and it reaches 91.6%, and the best agents remain a long way from that number. So the capability is real — it is just far smaller than it could be.

so what makes a note worth keeping

SKILL-KD, from Zhejiang University, went after exactly that, and its answer differs from what most systems do.

The usual approach is to take a summary of a successful run and store it. The trouble is that when a weaker agent fails, you cannot tell from its failure what exactly it was missing. And the stronger agent's run leaves so much unsaid that the weaker one has nothing to take from it.

What SKILL-KD does instead is place a failure and a success on the same task side by side, extract precisely the point that has to change, and write that out as an instruction. Plenty of methods get that far. What the others do not do is the next step: it re-runs the weaker agent with that instruction to see whether the problem is actually fixed, and rewrites the instruction when it is not.

There is one more problem. As these instructions accumulate one at a time, older rules eventually stop agreeing with newer ones. So every edit is kept with its history, and each time the system decides whether the new instruction should add a rule, change one, remove one, or be ignored. They tested it on five different benchmarks and it consistently beat the usual methods — without touching the model itself.

why this came out of programming

The fourth paper is a survey, and it organises the whole field: what exactly changes — framework, memory, notes, tools, the model itself, or how several agents work together — and when that change happens, and what evidence drives it.

The answer to the last question explains why this field started in programming. In programming the feedback executes. A test either goes green or it does not, and nobody has to offer an opinion. The repository supplies all the context, and every attempt at fixing a bug leaves a usable trace behind. No other domain has all three at once.

The authors also write down their warnings: feedback that is not always trustworthy, models that gradually memorise the benchmark, then safety, maintenance and cost. Set that list next to the first paper's finding to complete the picture. A system that accumulates notes without knowing which of them worked is accumulating technical debt, not experience.

this notebook is just a cache

If I put these four papers together in engineering terms: a skill library is a cache. And every cache that runs in production has two things these notebooks do not.

The first is an eviction policy. A cache that never discards what has stopped earning its place only grows and slows, and the large scattered library the first paper found in weaker models is precisely that. What SKILL-KD calls consolidation is exactly this — a rule deciding what goes in and what comes out.

The second is measurement. For any cache we measure what fraction of lookups actually hit, and without that number you cannot say whether keeping it is worth it. Nobody reports such a number for skill libraries, and what GDPevo built is in effect that number.

I have seen this in my own work. In BLZN.AI, which carries a trading strategy from historical backtest through to live execution, the hardest part was never building the strategy. The hardest part was working out whether a better result came from a change we made or from the market changing. A system that keeps something from its past is only hoarding, until you can connect the improvement to the thing it kept.

back to the line

That distinction still holds. Competence either accumulates outside the model in the code we wrote, or comes from the model itself, and the notebook is the most serious bet on the first route.

What these four papers add is that the first route only works if you have tested every single thing you add to that notebook on its own. Store a note without knowing it moved the agent from failure to success and what you have built is not a notebook; it is a warehouse.

So before building one for your next agent, answer one question. If we deleted every note and left only the last few tasks in front of the model, how much of today's capability would remain? Until you have that answer, you do not know whether you have built an agent that learns or an agent with a good memory.

Related articles