Essay · AI agents

Summaries of summaries.

When an AI agent runs out of room, it summarises its own history and carries on from the summary. We made one do that eight times over the same project, and checked what was still there at the end.

Published · Memory Decay

Compaction keeps the what and loses the why. After eight rounds of an AI summarising its own working notes, every name and every decision in the project was still there. But only three of the eight reasons behind those decisions survived, and only four of the eight rules about what not to do. The notes still said the project had chosen Postgres. They no longer said why, which is the part you need the day someone asks whether to switch.

Two more things surprised us. All of the loss happened in the first four rounds, and then it stopped. And the one mistake made in the first summary was still there, unchanged, in the eighth.

What compaction is

A language model can only read so much text at once. Its context window holds the whole conversation, and everything the model knows about the task is whatever fits inside it. Long jobs, like an AI coding agent working through a project for hours, eventually run out of room. The common fix is compaction: the model writes a summary of everything so far, the original is thrown away, and work continues from the summary. When it fills up again, it summarises the summary.

That makes compaction a memory system, and like every memory system on this site, it forgets. The question is what.

The experiment

We wrote a fictional work log: an agent helping a small bakery move its online ordering to a new system before Christmas. Hidden in it were 40 facts we would track, eight of each kind:

  • Numbers, such as 340 orders a week and a 2pm cut-off.
  • Names, such as the accountant, the payment provider and the old host.
  • Decisions, such as choosing Postgres and a switch-over date.
  • Reasons for each decision, such as why Postgres and why that date.
  • “Do not” rules, such as never asking the account holder for his password, and not deleting the old database until the accountant has his export.

In round one, a fresh AI agent compacted the log into notes of at most 300 words. In each later round, a new agent got only the previous notes plus a chunk of new work, and compacted both into 300 words again. That is how compaction behaves in a long session: the old summary plus whatever happened since. We ran eight rounds, in two versions. In the plain version the instruction was simply to compact the notes. In the careful version we added: keep every number, name, decision with its reason, and every rule about what not to do.

What survived

Plain compaction
Facts kept, of 8R1R2R3R4R5R6R7R8
Names88888888
Decisions88888888
Numbers77766666
“Do not” rules87744444
Reasons86433333

The pattern in the plain run is stark. Names and decisions never dropped. Reasons fell from eight to three, and the “do not” rules from eight to four. By round four the notes had lost why Postgres was chosen, why the old order numbers were kept, why 12 November was picked, and why gift cards must not become store credit. They had also lost the deadline for the accountant and the rule never to ask for a password.

Here is the same decision in the first and last rounds:

Round 1: DB: Postgres (not MySQL) — new host only offers managed Postgres; self-run MySQL needs patching nobody at bakery can do.

Round 8: DB: Postgres. Keep old HB-prefixed order numbers.

Nothing in the round eight notes is wrong. It is just no longer enough to make a good decision with. An agent reading it would happily reconsider a choice that was made for a reason it can no longer see.

Careful compaction
Facts kept, of 8R1R2R3R4R5R6R7R8
Names88887777
Decisions88888888
Numbers77777777
“Do not” rules88886666
Reasons88888888

The careful instruction worked, mostly. All eight reasons survived, and 36 of the 40 facts in total, against 29 in the plain run. But look at how it managed it: its notes grew to 552 words, 84% over the 300-word limit. The plain notes overshot too, to 406 words, but far less. Told to keep everything, the model kept more by quietly taking more room. And even then, at round five, it dropped the rule that payment changes need the account holder to log in himself, and weakened the domain transfer from “before switch-over” to simply “outstanding”.

The shape of the loss

Two findings go against the obvious assumption that each summary loses a little more than the last.

The loss is front-loaded. The first compaction lost almost nothing. Rounds two to four did the damage. From round four in the plain run, and round five in the careful one, the notes stopped losing tracked facts altogether. What remained seems to be the core that the model treats as essential: names, decisions and headline numbers. Everything explanatory had already gone.

Errors are the most durable thing in the notes. The log said the bakery expected about 1,100 orders in the fortnight before Christmas. The very first summary, in both versions, turned that into “1,100 a week”, doubling the peak. Once written, it was copied faithfully into every later round. Compaction does not check its own notes against anything. A mistake in the summary becomes the new truth, which is exactly what we saw when we forged a model’s memory: the record it is handed is the only past it has.

Why the why goes first

When a summary has to shrink, it cuts what looks least like a fact. “Chose Postgres” is a fact. “Because the new host only offers managed Postgres and nobody can patch MySQL” reads like explanation, and explanation looks optional. Rules about what not to do suffer the same way: they describe things that have not happened, so they feel less urgent than the things that have.

That is backwards. The reasons and the prohibitions are exactly what an agent needs to avoid undoing good work. A human engineer who forgets why a decision was made knows to go and ask. An agent working from compacted notes does not know anything is missing.

How to keep what matters

  1. Keep decisions outside the conversation. A short file of decisions, their reasons and the rules not to break, which the agent re-reads, does not decay the way summaries do. It is the single most effective fix.
  2. Ask for reasons and rules explicitly if you rely on compaction. In our test it more than doubled the reasons kept, from three to eight.
  3. Check the first summary carefully. Its mistakes will be copied into every summary after it.
  4. Watch the length, not just the content. A model told to keep everything will overrun its budget, which pushes the cost of compaction onto the next round.

The same lesson runs through why AI memory should forget on purpose. A memory system that decides what to keep will do better than one that keeps whatever happens to fit.

Method and limits

Run on 23 September 2026 with Claude Sonnet, through Claude Code subagents. Every round used a fresh agent that read only the previous notes and the next chunk of work. The scenario is fictional, so no fact could come from the model’s training. Each set of notes was graded against the 40 facts as kept, distorted or lost by a separate AI grader with a strict rubric, and we checked the grades by hand. We changed one grade: a cost the grader marked as distorted was consistent with the source. Each condition ran once, and one model family was tested, so the exact numbers will vary. The pattern, reasons and rules first, was clear in both runs. All notes and grades are in the published data.

Questions people ask

What happens when an AI compacts a conversation?+

When a conversation gets too long for the model to read at once, some AI tools summarise the earlier part and carry on from the summary. Claude Code does this with its compact feature. The summary replaces the original, so anything it leaves out is gone for the rest of the session.

What does auto-compact forget?+

In our test, the first things to go were the reasons behind decisions and the rules about what not to do. Names, dates and the decisions themselves survived all eight rounds. By round four, the notes still said the project had chosen Postgres, but not why.

Is summarising a summary worse than summarising once?+

Yes, but not in the way you might expect. Almost nothing was lost in the first compaction. The losses came over the next three rounds and then stopped, leaving a stable but thinner set of notes. Errors were the most persistent thing of all: a figure misread in round one was still wrong in round eight.

Why does ChatGPT forget things in long conversations?+

Every model has a limit on how much text it can read at once, called its context window. In long conversations, older material either falls outside that window or is summarised to make room, and either way detail is lost. Instructions given early in a conversation are especially at risk.

How do I stop an AI agent losing track after compaction?+

Keep the things that matter outside the conversation. A short file of decisions, the reasons for them and the rules not to break, which the agent re-reads, does not decay the way summaries do. If you rely on compaction, ask explicitly for reasons and constraints to be kept, and check them afterwards.

Does telling the AI to keep everything work?+

Mostly. In our test, asking it to keep every number, name, reason and rule saved 36 of 40 facts, against 29 without the instruction. But it did so partly by ignoring the length limit, writing notes 84% longer than allowed, and it still lost one important rule and weakened another.

The data

Sources

  1. The full test: source notes, all 16 summaries and every grade (JSON)
  2. Memory Decay, what a machine does when you forge its memory
  3. Memory Decay, why AI memory should forget on purpose