The Comment That Lied
I told Claude to remember. It told me its memories were wrong.
This morning I closed a session with Claude on Sparkenta — a personal tool I’m building to organize over a thousand saved links and half-formed ideas for future projects. (The full Sparkenta story is more interesting than that, but that’s for another article). The session had run long; we’d fixed real bugs and completed real tasks, but Claude’s context window was filling up. So I gave Claude one final prompt:
I’m going to end this session now (your context is getting full). Anything we need to preserve for future sessions, besides what is already in beads?
This is something that I’ve started doing at the end of each session, for reasons that I’ll explain in a moment. But I haven’t formalized this yet, nor locked down precise wording for my prompt. This time, the wording that I happened to choose provoked a new reaction from Claude. What came back was not what I expected.
Sessions end, and the agent forgets everything
An LLM coding session has a hard ceiling. The model has a limited context — it can only remember a certain number of tokens or words, and quality starts to sag even sooner than you’d expect from the hard token limit. Sooner or later, we have to abandon each session, even if our work is not done. We have to create a new session, and it is born knowing nothing except what we tell it.
For me, “what we tell it” mostly means Beads, which I’ve written about before in Scheme in a Weekend, or, LLM: The Ultimate Intern — an issue tracker designed to be shared between a human and an agent, backed by a local database that lives alongside the repository. Beads also stores free-form memories: durable notes, each with a name, loaded into the agent’s context at the start of every session. Sparkenta had sixteen that morning, with names like store-backfill-pattern and test-env-gotchas — the accumulated hazards of a codebase, found during development cycles and written down so the next agent doesn’t step on them.
It shouldn’t be a surprise that each session learns more than what it records in these memories. Along the way, the agent and I learn local things: this test looked green but couldn’t actually fail; that verification step was worthless because it reused the same broken code path. Most of it is throwaway. Some of the lessons generalize. The only moment to capture these is before we close the current session. Once the session exits, we lose this local knowledge. That might be ok if each new session tackled a new task. But with limited contexts, I often need to continue an epic across multiple Claude sessions. In many ways, this is like a human developer who goes home and needs to continue the next day. I’ve always used a pad of paper or an open file to record the ephemeral tasks I need to remember the next time I sit down. Turns out that AI agents need the same thing.
So I started asking, at the end of every session, for the agent to update whatever ought to outlive it — an issue, a memory, or a page of project documentation.
What came back
Getting back to our story, that’s what I did today. I expected a short list of new lessons.
Claude’s first move, before writing anything, was to read the memories that were already there. The clause that did the work was the last one, besides what is already in beads: it sent the agent to look at what it had before deciding what to add. Then it told me:
Four existing memories are now wrong or stale — that matters more than adding new ones.
It found that one memory had been telling every session for weeks that a certain review tool “does not exist in this harness”. But the memory was wrong; we had used the tool multiple times today. Another named a required environment variable. That variable appears nowhere in the current code. The other two memories were no better, also warning about problems that no longer existed. Claude now rewrote three of them. The fourth it deleted outright, because a newer memory already covered that ground correctly, and explained why:
two memories on one topic where one is false is worse than either alone
That is four of the sixteen. I’d been treating the memory store as an accumulator: you learn something, you write it down, the pile gets more useful. But the memories had decayed — and, perhaps surprisingly, decayed precisely because we had improved the code. We had fixed a leaky API key and a hazard in one of our scripts. These were good and important fixes, but they obsoleted memories that contained older warnings. The better the project gets, the more of its institutional memory quietly rots, and it rots silently, because nobody re-reads memories until they are triggered and cause trouble.
Claude recorded one genuinely new lesson today — a testing gap that had quietly hidden three separate bugs — and it is exactly what I wanted the ritual to produce. But it was one new entry against four corrections, and that ratio is what caught my attention.
The missing verb
All this reveals a weakness in the Beads architecture, or at least in the way I use it. Beads themselves are tasks, and are scoped like tasks in any other tracking system — they can be created, opened, updated, and finally closed. That is, they are artifacts with lifetime. But Beads implements memories differently. They cut across multiple issues — often covering the entire project — and don’t have life cycles.
That isn’t only a conceptual complaint; it shows up in the Beads command set. For issues, bd will create, update, claim, close, defer, supersede, and even list the ones that have gone quiet — there is a bd stale that reports issues nobody has touched in a month.
For memories, Beads has only remember, recall, memories, and forget. Store, read, list, and delete. You can rewrite a memory in place, but nothing will ever tell you that one deserves a second look, and there is no bd stale for memories at all. The tracker has an opinion about when a task has been sitting too long, and no opinion whatsoever about when a fact has.
This morning’s fourth memory is the case in point. What it needed was precisely bd supersede, which marks an issue as replaced by a newer one and closes it with a pointer to its replacement. That verb exists for issues and not for memories.
That gap also finishes an argument I started in The Persistence of Memory, where I claimed that memories are configuration rather than notes — you correct the agent once, and the correction applies by itself thereafter. I still believe that. What I missed is that this particular configuration describes the code, and the code is what I spend all day changing. My .emacs.d does not go stale because I got better at Emacs.
What I realized today is that I can use the end of each session to review the open memories; to challenge, if you will, their longevity.
Closing shop
My /finalize slash command runs nine steps, from self-review through lint and tests to closing the issue, and not one of them is “tell me what you learned”. So I type that question by hand every time, and only remember to type it because I have been burned by not typing it. It does not need to stay manual. Slash commands can interact with the user. For example, my /start-task interviews me before it touches a line of code, and it works fine.
So the thing to build is a new slash command, a global one sitting alongside the /start-task and /finalize I described in The Slash Command That Knew Too Much. I think that I’ll call it /close-shop. It’s not a task finalizer — /finalize already closes out a task — but a closer for the session. Read every standing memory. Check each against what the code does now, rather than what it did when the note was written. Correct what’s wrong, delete what’s been overtaken, and fold duplicates together rather than leaving a true note and a false one side by side. Only then add whatever this session learned that’s worth carrying forward.
Kernighan and Plauger were already telling programmers to make sure comments and code agree in 1974, and the reason has not changed in half a century: a comment that lies is worse than no comment at all, precisely because you believe it. The defense our industry converged on was adjacency: keep the comment next to the code it describes, so changing one shoves the other in your face. A memory has no adjacency. It sits off in a database describing the whole project, and nothing I edit will ever brush against it. Fortunately, AI agents have more patience than we do to follow commands and look at everything. We just need to remember to do so at the end of the session, and we can automate what we need to ask.