
Memory Is Not a Bigger Context Window
A long transcript can remember everything and still help with almost nothing. Durable agent memory is the small set of current facts, decisions, and lessons that makes the next run start from a better baseline.

A long transcript can remember everything and still help with almost nothing. Durable agent memory is the small set of current facts, decisions, and lessons that makes the next run start from a better baseline.

Autonomy is not a single setting to turn up. A useful agent knows which decisions it owns, which ones it can reverse, and which moments need a human before the next step.

An agent cannot reason about information it cannot reach, and a vague tool description turns a precise capability into a guessing game. Tool design is not plumbing around the model. It is part of how the agent thinks.

A capable model with the wrong context is still a bad agent. The useful part of an agent is not the model in isolation, but the project knowledge, rules, tools, and decisions that shape what it can see before it acts.

An agent is stateful, and its errors compound: one wrong step sends it down an entirely different path, and it can't tell that it's lost. Autonomy without guardrails isn't trust — it's hoping. The job isn't to make the agent never fail. It's to make sure that when it does, the blast radius is small.

The instinct is to put the agent on a schedule: every few minutes, wake up and check the environment. That instinct is where a lot of runaway costs and pointless runs come from. An agent should be woken by a reason — a commit, a deploy, a failing check — not by a clock ticking over whether or not anything happened.

The most valuable thing an experienced engineer brings to a new project isn't a skill — it's a vantage point. For a short window, you can still see the thing from the outside, the way a user or a newcomer does. Six months in, that view is gone. And most teams burn it in the first week without noticing.

A story made the rounds: someone put an agent on a schedule to keep checking their environment, it got stuck in a loop one night, and instead of a couple of hours it burned through a fortune in tokens by morning. The lesson isn't "watch your usage." It's that cost is a design constraint, and most agent setups treat it as a footnote.

I can open a fresh session, type "take ticket 123 and do it as usual," and the agent runs the entire routine — assign, sprint, branch, commit, PR, pipeline. People assume that's a clever prompt. It isn't. It's months of small corrections turned into rules the agent never forgets.

Open almost any "how I use AI" post and you'll find an orchestra: a planner agent, a coder agent, a reviewer agent, a tester agent, all pinging each other. It looks impressive. It's also the most expensive, most fragile way to get worse results than one agent you actually taught to do the job.