The boundary the paper draws: competence in engineered workflows (left) versus competence internalized in the model (right).
The boundary the paper draws: competence in engineered workflows (left) versus competence internalized in the model (right).

agentic vs. agentive: where automation ends and agency begins

A new paper draws a line between systems whose competence lives in your orchestration code and systems whose competence lives inside the model. That line decides what you can ship today and what is still marketing.

What is an agent? Before you answer, notice that almost every system you have built and called an "agent" this year would fail the definition a careful engineer should hold. The Critique of Agent Model from the SAILING Lab at CMU and MBZUAI puts a name to the gap, and the name is useful because it changes what you measure.

Let me define the two terms the paper rests on, because the whole argument turns on them. An agentic system is one whose competence resides in engineered workflows — the retrieval step, the tool-call router, the retry loop, the JSON schema you validate against. An agentive system is one whose capabilities, including social interaction, arise endogenously — from inside the model rather than from the scaffolding around it. The paper formalizes this as the difference between capabilities that come from external scaffolding versus internal initiative. That is not a philosophical flourish. It is a boundary you can test against your own codebase.

Here is the test, in practice. Open the repository for whatever you call your agent. Count how much of the "intelligence" lives in the prompt templates, the control-flow if statements, the hand-written planner, the tool registry — and how much lives in the weights. If you deleted the orchestration layer, would anything goal-directed survive? For every production system I have shipped, the honest answer is no. The competence is in the harness. That makes it agentic. And the paper's point is that agentic is fine — it is just not the thing the word "agency" was borrowed to imply.

the five axes are a checklist, not a manifesto

The authors analyze agent architectures along five dimensions: goal, identity, decision-making, self-regulation, and learning. I read this less as a theory of mind and more as a requirements list for anyone deciding what to build. Walk each axis and ask the applied question — where does this live, the weights or the wrapper?

Goal. In most deployed systems the goal is a string you injected at the top of the context window. The model did not form it; you did. Hierarchical task decomposition that holds across a long horizon is something you currently get by writing a planner loop, not by trusting the model to keep its own goal stack coherent.

Identity. This is the axis people skip and pay for later. Identity is the thing that makes a system's behavior consistent across sessions — the stable set of commitments it returns to. In a stateless API call there is no identity; there is a system prompt you re-paste every turn. If your "agent" forgets who it is the moment the context window rolls over, identity lives in your database, not in the model.

Decision-making. The paper proposes grounding this in simulative reasoning over a separately trained world model — the system imagines outcomes before acting. In production today, the closest thing most of us ship is a tree-search or a self-critique pass bolted on with more prompts. Useful, but again: it is scaffolding.

Self-regulation. When does the system stop, back off, ask for help, refuse? In every shipped pipeline I know, those guardrails are external — a moderation classifier, a max-iterations counter, a human in the loop. The paper wants this learned and internalized. That is a research target, not a sprint ticket.

Learning. Does the system improve from its own experience, real or simulated, without you fine-tuning it offline? Almost nothing in production does this. We retrain; the system does not learn.

Run that checklist honestly and the trade-off becomes concrete. Every axis you move from the wrapper into the weights buys you autonomy and costs you control. That is the whole tension, and it is why the distinction matters for what you actually deploy.

why the wrapper is winning, and what it costs

There is a reason the agentic pattern dominates production, and it is not laziness. Engineered scaffolding is auditable. When the goal is a string in your code, you can read it. When self-regulation is a counter and a classifier, you can log it, test it, and prove to a client that the system will not loop forever or call a destructive tool. I have shipped systems into domains — legal document workflows among them — where the ability to point at exactly which rule fired is not a nice-to-have; it is the entire reason the system is allowed to run at all.

The paper is careful here, and this is the part I'd push back hardest for. It frames the goal as agentive systems that have greater autonomy but remain under human oversight, and it explicitly raises auditability, controllability, and safety as design concerns rather than afterthoughts. The functional question is the right one: what does internalization cost you in observability? Every capability you move inside the weights is a capability you can no longer read off the source. You trade a glass box for a black one. For an open-world consumer toy that may be acceptable. For a regulated workflow it is often disqualifying.

The authors' own proposal, the Goal-Identity-Configurator (GIC) architecture, combines hierarchical goal decomposition, identity evolution, simulative reasoning over a separately trained world model, learned self-regulation, and self-directed learning from real and simulated experience. Read that list as an engineering wish-list and notice that every single component is something we currently fake with external machinery. GIC is a description of where the field wants the competence to migrate — from the harness into the model. It is not something you can pip install this quarter.

what this changes for the system you ship next

The practical payoff of this paper is not a new framework. It is a vocabulary that stops you from overclaiming. When a vendor says "agent," you now have one question that cuts through the demo: is the capability in the weights or in the wrapper? If it is in the wrapper, you are buying an orchestration layer, and you should price it, audit it, and own it like any other piece of glue code — not treat it as emergent intelligence.

There is a wider frame worth holding. The five axes — goal, identity, decision-making, self-regulation, learning — are exactly the structures evolution had to internalize before an organism could survive in an open environment instead of a controlled one. A reflex arc is agentic; a foraging animal is agentive. The paper is, in effect, asking why our systems are still reflex arcs with very long prompts. That is a real question, and it deserves an engineering answer, not a marketing one.

So here is mine. Build agentic when the world is bounded and the audit trail matters — which, in production, is most of the time. Reach for agentive components only on the axes where internalization buys you something you cannot fake, and budget for the observability you lose when you do. The line between automation and agency is not a marketing slogan. It is a design decision you make one axis at a time.

Sources

Related articles