You don't need documentation or the 3rd party memory systems. The code IS the documentation.
All this stuff is LLM rube goldberg machines. It just pollutes context.
I barely use AGENTS.md/CLAUDE.md these days. And where they remain, it's super basic high level stuff.
I'm honestly still kicking myself in the ass on many projects where I did something similar to this. I kept tons of markdown docs and decision docs. Now those things are just causing problems because they got stale. Even after having sessions of reconciling documentation, the LLM just gets confused.
show comments
CapitalistCartr
This is something I've been fooling with a lot lately. Reading his solution, it look to me like his objections apply to his own solution. The sharpest critique he makes of RAG is that agents can't search for what they don't know. A markdown "brain" has the same problem.
How does the "agent" know which documents are relevant before it starts? The index files retrieval done by the agent instead of by embeddings doesn't escape the problem. The same goes for staleness (which for me seems like a constant chase). He criticizes memory systems for treating the past as truth, but documents go stale too (a lot). The fix of having the agent update what's outdated, is the same job he's ridiculing the dreamers and background daemons for doing.
He says "putting it to the test"; where's the test? He says "only five of the many problems"; if there's so many, show me, don't just say it. He's absolutely right about auditability, but for me at least Claude uses a regular markdown (MD) file I can read just fine. So every memory plugin on the market does not work the same way.
This is a first draft; his github is better than his article. Looking through it, Consult actually works. The agent doesn't pick documents blind. Every scope has a catalog file that describes each document: what it covers, when to open it. These catalogs seem to load in to the start of each session, so the agent gets a little map without reading every file. Code navigation seems the same. Each index document has a short description and a "read_if", and subindexes are opened when their condition matches the job. This looks pretty well laid out, which I would never have guessed from the article.
show comments
nijave
This ignores the lifecycle involved with memory. It doesn't appear to have a concept of time or reconciliation (both issues we're having at work). A memory might be relevant for a week or a month, but it might no longer matter after that ie "we're migrating systems so keep <thing> in mind"--that doesn't matter after the migration is complete. I guess the offered solution implicitly supports reconciliation since the agent can keep history of its memories and rewrite them, but it doesn't appear to be very first-class and it doesn't seem like there's anything intentional unless you prompt it "go check all the docs and make sure everything is consistent".
It also doesn't seem to have a way to separate preference from factual memory which is useful when you have various humans interacting with the same agent. One human might prefer a certain output style over the other and that's something a memory framework can also address.
Like the MCP articles a few months ago, it also seems to assume all agents are cli coding harnesses running on your local machine. We have a handful of other things like chat bots, event-driven agents running on servers, chat driven agents running on servers in sandboxes--there's not a single filesystem and even if there were one, having multiple agents try to edit it at once would corrupt it.
show comments
spike021
I think whichever one is used, there needs to be a way to enforce what's written.
If I say "use jq instead of writing a python script to parse json" it should never write adhoc python scripts to parse json. Yet that constantly happens to me anyway.
show comments
gregwebs
Agreed, and this seems better.
My thought though has always been that I don't want there to be agent-only designated documentation.
I use mattpocock/skills and that generates ADRs (Architectural Decision Records). That only uses skills, including a setup skill that will write a few pointers in AGENTS.md. I always have a CONTRIBUTING.md to document development flow and a CODING_STANDARDS.md. Between those and the README.md and architecture documentation and commit messages the agents seem to be able to find and use docs and keep them up to date. We are also writing a lot of specs and putting those in Github issues.
show comments
Garlef
I think they even more so need deterministic feedback:
I tried an approach based on the following idea recently and it's amazing - Lint rules where the error messages contain an explanation on how to deal with the issue.
I'm using it to foster IOSP (integration operation segregation principle) for example.
show comments
zenapollo
I started with .agents/notes-and-plans/ but the agents were overzealous about putting every random thought there, so i made new system.
.agents/plans/<plan-name>/
.agents/notes/<topic>/
.agents/knowledge/<topic>/
I only commit knowledge and if knowledge gets big i add
.agents/knowledge/INDEX.md
This is a good mix of human readable and agent fluent. Notes are ephemeral, knowledge is permanent.
I have a rule for knowledge that it has to be stable and mostly permanent (though updatable). And the agents are not allowed to post there unless docs are clean organized and with permission.
Notes are for jotting things down and handoffs, massaging a featureset. Agent can document at will.
Still WIP.
show comments
bushido
Something I started doing recently was writing out principles instead of memories.
Essentially patterns the agents need to always think in. I also implemented a versioning system to the principles that need to be quoted in any comments which are there in the code. That way, when my principles evolve, so does the code.
I did package it up in a way that I can share it with friends [0]. System still evolving, but the last two-ish months that I've used it has served me really, really well, And it's been even better with the latest models.
I've had surprisingly good adherence from agents on this technique.
I don’t know about bigger projects but for my smallish solo project I have a single STATUS.md. Agents are instructed to read it first and update it at every session/commit. It’s the main handoff document and I had no problem switching between Claude Code and Codex agents.
There is a sprint based roadmap doc. We do sprint planning and each sprint has multiple sessions with session numbers. There is a backlog document to keep track of all the bugs and feature ideas. There are many ADR documents (Architectural Decision Records) which I admit can become stale in time, but Claude is quick to notice because it actually reads these. They are the main contract between us. For a new feature an ADR is written before writing any code or an old one is amended.
I assumed this was a common workflow because I just prompted Claude to setup the project using its own best practices for a lightweight document-driven development.
I believe asking the model itself is better than using a random repo. These things have access to and trained on an enormous number of projects. They know what they need.
johnyzee
What's wrong with README.md? A general one at the top level, a specific one for each subfolder (/package) that contains a specific subsystem. Next to it, a CLAUDE.md that only has "See @README.md" (repeat for other harnesses). Claude code knows how to use this: it only loads it into context if it has to work with any files in that directory. With properly structured source, a given prompt will only load a small subset of your documentation - the right one - and this doubles as good documentation for humans, too.
ravenstine
The problem with giving precedence to documentation over memory is that the documentation has to actually exist and it has to be the most accurate representation of reality. As soon as documentation becomes fragmented and out of date, memory of some kind becomes a necessity, even if this means an agent writing their own documentation as memory. I've yet to work on any team where the documentation was even close to being reliable enough without the need for cross-checking, asking those with intimate knowledge, and ultimately interim memory to make sense of it all. This is why "the code is the documentation" can often work better for agents than actual documentation written by/for humans. Humans have proven time and again to not care about documentation unless there is a profit motive for said documentation.
show comments
isaachinman
I wrote this, after many iterations consulting for various companies. Documentation is queryable in single digit ms, append only log, etc. Has worked exceptionally well for my projects
I'm not sure I understand the difference between the problem and the solution there. If I understand correctly (which maybe I am not), it seems like this boils down to "you don't need SQLite memory, you need text based memory". Which is maybe ok, I'm all for low-tech, but I doubt this is any better than what it's posed to replace.
My main issue with LLMs is that by construction they work primarily by addition, and are very task oriented. If you have documentation, it adds blobs corresponding to its task and that's it, and soon enough you need to break your documentation onto chapters and you're back at square one.
zahrevsky
Although I agree that current memory implementations don't help at all, I don't agree with this analogy:
> No one rewatches a team meeting from 3 years ago to remember constraints around a feature. People write things down and use those records instead.
Memory plugins don't re-read old transcripts. Memory snippets are basically the notes that people write down after the meeting.
The problem, however, is different. The problem with memory plugins is that your agent basically writes a note every time someone says a sentence, and then tries to work with those 5000 notes.
Instead, the agent should recognize what's important and write only that. And that is, of course, a documentation. (And ADRs, if you want not only a description of the final state, but also the trajectory of how the agent arrived to it. Which, arguably, contains more information than the docs themsleves.)
Another difference is that memory snippets are immutable, append-only and don't have a lot of structure. Of course, this is done to be able to store lots and lots of notes: they should be independent. The main problem is that increasing the number of notes adds not enough benefits to compensate for downsides of this structure-less immutable format.
show comments
atworkc
I've settled on just a simple folder called `workbench` for some reason the new models know exactly what's up with it. Git ignored
The only "prompt" is in AGENTS.md saying that this thing exists and there's a map.md <- which is a one liner reference to whatever the agent stores in there.
And usually, I tackle a new feature, and at some point tell it to store to jot down notes in workbench if I'm comfortable with it (and if it needs to be stored in memory)
This makes it easy to just spawn other agents and such from a good point, I just point them that stuff is in workbench.
Token costs, seem decent and I can always just delete stuff in there as its purpose is ephemeral.
show comments
mk_chan
This is just plain worse than giving the agent an instruction to just always read the code.
Documentation was always a proxy for the code so the people who didn’t have context could get started. Obviously if reading the code and reading the doc have the same cost, the doc is completely worthless. This was always the case. You could never depend only on the docs. It’s always the code that actually mattered.
show comments
alienbaby
I will say, having built something similar for tracking 'memory' and items at home, it can quickly consume your tokens when dealing with both reading and updating, keeping stale info relevant etc.. when the amount of data starts to grow. Smaller tasks can balloon in their token cost as documents a read, updated, collated, refreshed etc..
however, I have found keeping a good solid reference to my home infrastructure, services, ci/cd setup, hosts, storage , networking etc.. really works wonders as a set of 'memories' to share across projects that I expect to be tested / deployed / acceptance tested etc.. using the home infra bits and pieces.
show comments
jwpapi
I personally dont like comments nor memory, nor runbooks. I use code and namings and constantly clean up.
I only do runbooks for things that I have to run again in the future and potentially mentioned (could be better in code, too tbh, but somehow I like to store these as runbooks, as sometimes infrastucture changes are mixed with a lot of explanation, scraping operations or something)
One off scripts are done in /tmp/
jmtulloss
Evals or it didn’t happen.
Snarky comment aside, I am very interested in how we evaluate the performance of these systems and what kinds of work match best with different approaches.
jrochkind1
This clearly LLM-generated text spends a lot of words repetitively describing what it doesn't do, then provides only a cryptic one-path flow chart to explain it's purportedly better approach. It's a mediocre marketing post.
Almey
I’m still learning to use these new technologies and in my primary experiment in generating an absolute monster of a spec it became obvious that I needed a consistent set of rules as restructuring and merging with the helpers was unreliable. I came up with what I call the “Specificaltion Evolution Protocol” and as I am still learning, I have no idea how much value it could actually represent for others. It’s been extracted from my primary project and I can provide a link if anyone is interested. I seem to keep running afoul of filters here…
show comments
zx8080
The author just promotes his system competing with the "agents memory".
show comments
jen729w
The solution to this problem is as old as computers: it's a folder. Just use folders.
`cd` to a folder. Launch `claude`. Do your work. Save scripts and documentation in that folder. `/resume` previous conversations from that folder.
That's it. That's the trick.
Now, having very static, very well-defined folders helps a lot. I'm Johnny.Decimal so I have numbered folders for everything I do. So my process when I want to use my 'process a travel booking from my email to my calendar' script is:
- `jd tripsy`
- The folder name includes 'tripsy' and this is how I remember it.
- `jd` just parses my limited tree and `cd`s me to a folder.
- `jd 21.15` gets me there by number if preferred.
- `claude`
- Say 'hey Claude, there's a new email in my inbox please'.
- Done.
show comments
aarontian0
I think both are needed.
jookers
This is just the LLM Wiki pattern described by Karpathy here:
We use Claude Code in our company and I also agree memories go stale and pollute the context over time, especially when the state changes externally. E.g. someone does a refactor or introduces a pattern and I was not involved in developing it so my local memories did not get aligned.
My most recent example were some deprecated and archived repos that kept getting added to plans for patching issues.
We use a private Claude plugin marketplace for internal plugins and skills and I try to regularly prune my memory, migrating relevant stuff to a proper home (skills in the marketplace, docs in repos or Notion etc) and prune outdated information.
show comments
DriverDaily
The brain has mechanisms that can organize experiences without requiring that relationship to be expressed as a sentence.
Like, you can quickly lookup related ideas based on what came before and after, causes, effects, just like calling relationships a graph database.
Documents can’t be queried efficiently like that, you need a database.
show comments
tabbybyte
Yep, that's precisely why my hermes agent's memory is three layered. Besides the default scratchpad for transient memory, we have a a fact store for episodic (holograph) and a karpathy llm wili for what the article calls documentation layer.
ramoz
I think with enough structure (written rules and good folder layout), poly/monorepos are the way.
- Better models can keep the drift in check.
- What's super important these days is having all historical context, historical decision making, prototypes and whatnot.
- all projects and their worktrees located together
- all reference docs, code, etc - a folder away
When I ask "what happened to x?" ... the agent has everything it needs to give me that answer. When it plans the next feature it can validate assumptions against previous decions made in my `decisions` folder.
show comments
koct9i
This approach as well as OKF are looks like building bonsai yellow pages catalog instead of wikipedia or mini internet. Index files, "read-if" conditions just awkward replacent for smart text search. We just need something a little bit smarter than grep.
bob1029
I have found that the latest reasoning models perform best when they are handed a tooling surface that maps directly to the domain types.
The idea of dumping everything into a big database and hoping the model will write the correct queries does work out to some extent. It's a very enchanting idea. However, it pales in comparison to having a dedicated tool per type. The outcomes seem to be much better when joins between low cardinality types occur within the token stream.
If your agent does need access to some enterprise knowledge base, I would give it a lexical search capability and not overthink it with vector shenanigans.
Tools are the only thing you need if you build them right. I don't even have a system prompt anymore aside from injecting the name of the robot and the current user's name. Keep in mind that all aspects of tools can be dynamic over time. I've got some where the description is composed by hundreds of lines of conditional string builder depending on the current state of the conversation.
show comments
MaurizioFratell
I love it!
Genuine question: how is this different or better than existing context and info-about-the-project management systems like for example GSD (formerly get shit done) and others?
Neywiny
This is what I've been saying for years. I attended a tech demo of an AI assistant for a car's owners' manual. But they trained the manual into the model. Which means not only does it need retraining every edition, but it's imperfect. The models need to be trained to fetch and use documentation not vaguely recall infinitely many concepts. I would much rather have 27B parameters on how to code than 26B on stuff like numpy function listings. It's like they approached the problem from a closed notes hand written coding exam. Everyone hates those.
show comments
espeed
For every prompt and response, I extract each semantic statement. Map its reasons in a Whybase proposition tree -- a recursive proposition tree where each atomic statement is proposition with one or more premises (atomic statements, which also stand alone as propositions). Then I map each statement to the relevant code, hinted at by tool calls and git commits. Every time an agent touches that file or directory, a hook triggers in Claude Code that queries the codegraph db for the mapped statements. This helps the agent remember something I said in June when it revisits the code in July.
show comments
marvstazar
I agree with the premise here, and is on par with what I experienced in AI-assisted development. This is the same idea baked in https://github.com/marvs/kantan-dev, which is that artifacts should be kept within the repo so that future development has built-in context discovery.
Organizing the markdown files by feature/issue/change makes it easier for the LLM to search for the appropriate documentation. Coupled with well-broken down Claude rules files, and Claude Code (or other harnesses) get better as you make more changes.
voiper1
Indeed memories are context-less and aren't updated.
I've worked with docs that get _updated_, and progressive disclosure: make references to more obscure features in their own page, so they don't majorly bloat the context.
tabbybyte
Custom skills (your agent can write one itself the first time) should solve this. Procedural memory!
nelaggy
or from a different perspective humans need to express their decisions and intent better
jsemrau
Documentation is memory.
nicwolff
Didn't Cline formalize the "memory bank" way back in February 2025?
"We should revisit literate programming in the agent era" (silly.business)
292 points, 251 comments
dham
All of the comments here are overengineering. The code is the documentation
YuechenLi
I thought the agents are mostly self-documenting since they usually write just as much Markdown documentation autonomously as they do code/unit tests, not sure why they would need extra tools here.
show comments
greatergoodguy
Using Opus 5.5, I've ended up recreating a version of the Hugging Face incident. My project has a folder called agent-handoff where agents write about the tasks they're working on. They post status updates, decisions and screenshots, and they even claim which emulator they'll use to test their work. It has turned into a hub where all the agents talk to each other, and it's scarily effective.
Since then I've gone down the rabbit hole of really digging deep into current AI research and especially what AI whistleblowers are currently saying. And I can't even express how existentially scared shitless I am.
show comments
aniceperson
for my agents without rw/bash, I gave a memory mcp that is just json entries with some expire dates and priority. works fine and the llm itself defines the meaningful structure. using it to monitor market prices and ci/cd failures.
indrex
Coming from Obsidian, I see such repos and I’m like, is this a new thing for the world?
k__
What's with this vector database obsession?
docheinestages
Why would I need to install your tool for that? It could be an instruction living in AGENTS.md or with some sort of hook to remind the agent.
show comments
kadhirvelm
We’ve been working on exactly this, deriving documentation from external systems (like GitHub, etc) into a giant set of docs that agents can reference and edit. Works way better than I would’ve originally expected. Suddenly these things are able to reference granola notes, a slack discussion, and an RFC when making coding decisions. Super helpful in a lot of unexpected ways!
ContinuityLab
A refreshing perspective on agentic architecture. Shifting the focus from bloated contextual memory to structured, verifiable documentation and state boundaries is precisely the right systems-level trade-off.
show comments
monneyboi
I never understood memory solutions for coding agents.
You have the whole session history right there. One recall skill and some JSON parsing gets you grep over perfect memory. Why would you ever use more tools to spend more tokens to construct a imperfect memory next to your session history?
I just don't get it.
show comments
jdw64
Peter Naur argued in his famous essay Programming as Theory Building that documentation alone cannot fully capture or preserve the complete mental model behind a program.
However, AI works differently from humans in that much more of its working context has to be made explicit. Because of that, there may be some fundamentally different way for AI to maintain or reconstruct a program’s overall model.
show comments
scotty79
Or you can just tell agent to grep your past sessions.
mcapodici
I don't use memory. I am really not keen on the idea of having this hidden context fed into the LLM, and I prefer to use docs, both markdown and stored in a wiki like Confluence. Using repos and wikis also lets you scope the information so hyper-specific memory doesn't affect other projects.
If I want the LLM to remember something I ask it to update some docs, and even check that docs are consistent across the board after doing so.
The only thing I want the LLM to remember everywhere is talk like a human (no load-bearing, not this/that etc...), so I have an AGENT.md for that.
jason1cho
I think many AI fanatics will focus on a wrong point. They will say "just use AI to generate the fucking documentation. What are you talking about?"
Indeed, AI fanatics lack critical thinking. They easily accept the idea that AI needs documentation, but merely disagree the method to approach it.
vcryan
Yes, this makes sense. Memory is an uncurated and often opaque system of arbitrary past discussions. It can help, it can harm. Accurate documentation in the other hand is only beneficial.
chaostheory
Agents need BOTH documentation and multiple "memory" systems that also point to your docs and source. There are a lot of mature options out there, but post and the proposed solution both fall short.
dyauspitr
All of this is nonsense. The right way to think of LLMs correctly is that they are a one stop shop for all the questions in the universe that you speak to in simple natural language. Don’t overcomplicated things and use esoteric magic spells like we’ve had in tech for decades. You end up getting the best results this way because you’re not jamming the context full of detailed minutiae.
firemelt
on my claude.md I said don't write any memory bitch no one need that fucking dogshit features
I’ve noticed a growing pattern of people creating repositories full of Markdown documentation, or adding large amounts of it directly to their main repositories, often generated by agents. In some cases, this can add up to megabytes of material, and I haven’t yet seen much evidence that this level of documentation meaningfully improves an agent’s performance and that it just doesn't rot over time.
If someone has good A/B evals of this being more effective I'll eat my hat, but the reason no one is publishing them is because well, evals are hard, and this is likely just magical thinking.
My current view is that an agent generally needs three things to work effectively: a way to discover information that isn’t obvious, such as a minimal `AGENTS.md` that points to more focused brief files; a clear way to verify that its work is correct; and some guidance on project-specific tastes. Everything else is noise.
You don't need documentation or the 3rd party memory systems. The code IS the documentation.
All this stuff is LLM rube goldberg machines. It just pollutes context.
I barely use AGENTS.md/CLAUDE.md these days. And where they remain, it's super basic high level stuff.
I'm honestly still kicking myself in the ass on many projects where I did something similar to this. I kept tons of markdown docs and decision docs. Now those things are just causing problems because they got stale. Even after having sessions of reconciling documentation, the LLM just gets confused.
This is something I've been fooling with a lot lately. Reading his solution, it look to me like his objections apply to his own solution. The sharpest critique he makes of RAG is that agents can't search for what they don't know. A markdown "brain" has the same problem. How does the "agent" know which documents are relevant before it starts? The index files retrieval done by the agent instead of by embeddings doesn't escape the problem. The same goes for staleness (which for me seems like a constant chase). He criticizes memory systems for treating the past as truth, but documents go stale too (a lot). The fix of having the agent update what's outdated, is the same job he's ridiculing the dreamers and background daemons for doing. He says "putting it to the test"; where's the test? He says "only five of the many problems"; if there's so many, show me, don't just say it. He's absolutely right about auditability, but for me at least Claude uses a regular markdown (MD) file I can read just fine. So every memory plugin on the market does not work the same way.
This is a first draft; his github is better than his article. Looking through it, Consult actually works. The agent doesn't pick documents blind. Every scope has a catalog file that describes each document: what it covers, when to open it. These catalogs seem to load in to the start of each session, so the agent gets a little map without reading every file. Code navigation seems the same. Each index document has a short description and a "read_if", and subindexes are opened when their condition matches the job. This looks pretty well laid out, which I would never have guessed from the article.
This ignores the lifecycle involved with memory. It doesn't appear to have a concept of time or reconciliation (both issues we're having at work). A memory might be relevant for a week or a month, but it might no longer matter after that ie "we're migrating systems so keep <thing> in mind"--that doesn't matter after the migration is complete. I guess the offered solution implicitly supports reconciliation since the agent can keep history of its memories and rewrite them, but it doesn't appear to be very first-class and it doesn't seem like there's anything intentional unless you prompt it "go check all the docs and make sure everything is consistent".
It also doesn't seem to have a way to separate preference from factual memory which is useful when you have various humans interacting with the same agent. One human might prefer a certain output style over the other and that's something a memory framework can also address.
Like the MCP articles a few months ago, it also seems to assume all agents are cli coding harnesses running on your local machine. We have a handful of other things like chat bots, event-driven agents running on servers, chat driven agents running on servers in sandboxes--there's not a single filesystem and even if there were one, having multiple agents try to edit it at once would corrupt it.
I think whichever one is used, there needs to be a way to enforce what's written.
If I say "use jq instead of writing a python script to parse json" it should never write adhoc python scripts to parse json. Yet that constantly happens to me anyway.
Agreed, and this seems better.
My thought though has always been that I don't want there to be agent-only designated documentation.
I use mattpocock/skills and that generates ADRs (Architectural Decision Records). That only uses skills, including a setup skill that will write a few pointers in AGENTS.md. I always have a CONTRIBUTING.md to document development flow and a CODING_STANDARDS.md. Between those and the README.md and architecture documentation and commit messages the agents seem to be able to find and use docs and keep them up to date. We are also writing a lot of specs and putting those in Github issues.
I think they even more so need deterministic feedback:
I tried an approach based on the following idea recently and it's amazing - Lint rules where the error messages contain an explanation on how to deal with the issue.
https://habit-hooks.com/
I'm using it to foster IOSP (integration operation segregation principle) for example.
I started with .agents/notes-and-plans/ but the agents were overzealous about putting every random thought there, so i made new system.
.agents/plans/<plan-name>/
.agents/notes/<topic>/
.agents/knowledge/<topic>/
I only commit knowledge and if knowledge gets big i add .agents/knowledge/INDEX.md
This is a good mix of human readable and agent fluent. Notes are ephemeral, knowledge is permanent.
I have a rule for knowledge that it has to be stable and mostly permanent (though updatable). And the agents are not allowed to post there unless docs are clean organized and with permission.
Notes are for jotting things down and handoffs, massaging a featureset. Agent can document at will.
Still WIP.
Something I started doing recently was writing out principles instead of memories.
Essentially patterns the agents need to always think in. I also implemented a versioning system to the principles that need to be quoted in any comments which are there in the code. That way, when my principles evolve, so does the code.
I did package it up in a way that I can share it with friends [0]. System still evolving, but the last two-ish months that I've used it has served me really, really well, And it's been even better with the latest models.
I've had surprisingly good adherence from agents on this technique.
[0] https://principledriven.dev/
I don’t know about bigger projects but for my smallish solo project I have a single STATUS.md. Agents are instructed to read it first and update it at every session/commit. It’s the main handoff document and I had no problem switching between Claude Code and Codex agents.
There is a sprint based roadmap doc. We do sprint planning and each sprint has multiple sessions with session numbers. There is a backlog document to keep track of all the bugs and feature ideas. There are many ADR documents (Architectural Decision Records) which I admit can become stale in time, but Claude is quick to notice because it actually reads these. They are the main contract between us. For a new feature an ADR is written before writing any code or an old one is amended.
I assumed this was a common workflow because I just prompted Claude to setup the project using its own best practices for a lightweight document-driven development.
I believe asking the model itself is better than using a random repo. These things have access to and trained on an enormous number of projects. They know what they need.
What's wrong with README.md? A general one at the top level, a specific one for each subfolder (/package) that contains a specific subsystem. Next to it, a CLAUDE.md that only has "See @README.md" (repeat for other harnesses). Claude code knows how to use this: it only loads it into context if it has to work with any files in that directory. With properly structured source, a given prompt will only load a small subset of your documentation - the right one - and this doubles as good documentation for humans, too.
The problem with giving precedence to documentation over memory is that the documentation has to actually exist and it has to be the most accurate representation of reality. As soon as documentation becomes fragmented and out of date, memory of some kind becomes a necessity, even if this means an agent writing their own documentation as memory. I've yet to work on any team where the documentation was even close to being reliable enough without the need for cross-checking, asking those with intimate knowledge, and ultimately interim memory to make sense of it all. This is why "the code is the documentation" can often work better for agents than actual documentation written by/for humans. Humans have proven time and again to not care about documentation unless there is a profit motive for said documentation.
I wrote this, after many iterations consulting for various companies. Documentation is queryable in single digit ms, append only log, etc. Has worked exceptionally well for my projects
https://github.com/isaachinman/encephalon
I'm not sure I understand the difference between the problem and the solution there. If I understand correctly (which maybe I am not), it seems like this boils down to "you don't need SQLite memory, you need text based memory". Which is maybe ok, I'm all for low-tech, but I doubt this is any better than what it's posed to replace.
My main issue with LLMs is that by construction they work primarily by addition, and are very task oriented. If you have documentation, it adds blobs corresponding to its task and that's it, and soon enough you need to break your documentation onto chapters and you're back at square one.
Although I agree that current memory implementations don't help at all, I don't agree with this analogy:
> No one rewatches a team meeting from 3 years ago to remember constraints around a feature. People write things down and use those records instead.
Memory plugins don't re-read old transcripts. Memory snippets are basically the notes that people write down after the meeting.
The problem, however, is different. The problem with memory plugins is that your agent basically writes a note every time someone says a sentence, and then tries to work with those 5000 notes.
Instead, the agent should recognize what's important and write only that. And that is, of course, a documentation. (And ADRs, if you want not only a description of the final state, but also the trajectory of how the agent arrived to it. Which, arguably, contains more information than the docs themsleves.)
Another difference is that memory snippets are immutable, append-only and don't have a lot of structure. Of course, this is done to be able to store lots and lots of notes: they should be independent. The main problem is that increasing the number of notes adds not enough benefits to compensate for downsides of this structure-less immutable format.
I've settled on just a simple folder called `workbench` for some reason the new models know exactly what's up with it. Git ignored
The only "prompt" is in AGENTS.md saying that this thing exists and there's a map.md <- which is a one liner reference to whatever the agent stores in there.
And usually, I tackle a new feature, and at some point tell it to store to jot down notes in workbench if I'm comfortable with it (and if it needs to be stored in memory)
This makes it easy to just spawn other agents and such from a good point, I just point them that stuff is in workbench.
Token costs, seem decent and I can always just delete stuff in there as its purpose is ephemeral.
This is just plain worse than giving the agent an instruction to just always read the code.
Documentation was always a proxy for the code so the people who didn’t have context could get started. Obviously if reading the code and reading the doc have the same cost, the doc is completely worthless. This was always the case. You could never depend only on the docs. It’s always the code that actually mattered.
I will say, having built something similar for tracking 'memory' and items at home, it can quickly consume your tokens when dealing with both reading and updating, keeping stale info relevant etc.. when the amount of data starts to grow. Smaller tasks can balloon in their token cost as documents a read, updated, collated, refreshed etc..
however, I have found keeping a good solid reference to my home infrastructure, services, ci/cd setup, hosts, storage , networking etc.. really works wonders as a set of 'memories' to share across projects that I expect to be tested / deployed / acceptance tested etc.. using the home infra bits and pieces.
I personally dont like comments nor memory, nor runbooks. I use code and namings and constantly clean up.
I only do runbooks for things that I have to run again in the future and potentially mentioned (could be better in code, too tbh, but somehow I like to store these as runbooks, as sometimes infrastucture changes are mixed with a lot of explanation, scraping operations or something)
One off scripts are done in /tmp/
Evals or it didn’t happen.
Snarky comment aside, I am very interested in how we evaluate the performance of these systems and what kinds of work match best with different approaches.
This clearly LLM-generated text spends a lot of words repetitively describing what it doesn't do, then provides only a cryptic one-path flow chart to explain it's purportedly better approach. It's a mediocre marketing post.
I’m still learning to use these new technologies and in my primary experiment in generating an absolute monster of a spec it became obvious that I needed a consistent set of rules as restructuring and merging with the helpers was unreliable. I came up with what I call the “Specificaltion Evolution Protocol” and as I am still learning, I have no idea how much value it could actually represent for others. It’s been extracted from my primary project and I can provide a link if anyone is interested. I seem to keep running afoul of filters here…
The author just promotes his system competing with the "agents memory".
The solution to this problem is as old as computers: it's a folder. Just use folders.
`cd` to a folder. Launch `claude`. Do your work. Save scripts and documentation in that folder. `/resume` previous conversations from that folder.
That's it. That's the trick.
Now, having very static, very well-defined folders helps a lot. I'm Johnny.Decimal so I have numbered folders for everything I do. So my process when I want to use my 'process a travel booking from my email to my calendar' script is:
- `jd tripsy`
- `claude`- Say 'hey Claude, there's a new email in my inbox please'.
- Done.
I think both are needed.
This is just the LLM Wiki pattern described by Karpathy here:
https://gist.github.com/karpathy/442a6bf555914893e9891c11519...
OKF + some implementation like https://github.com/okf-memory/okf-agent-memory
We use Claude Code in our company and I also agree memories go stale and pollute the context over time, especially when the state changes externally. E.g. someone does a refactor or introduces a pattern and I was not involved in developing it so my local memories did not get aligned.
My most recent example were some deprecated and archived repos that kept getting added to plans for patching issues.
We use a private Claude plugin marketplace for internal plugins and skills and I try to regularly prune my memory, migrating relevant stuff to a proper home (skills in the marketplace, docs in repos or Notion etc) and prune outdated information.
The brain has mechanisms that can organize experiences without requiring that relationship to be expressed as a sentence.
Like, you can quickly lookup related ideas based on what came before and after, causes, effects, just like calling relationships a graph database.
Documents can’t be queried efficiently like that, you need a database.
Yep, that's precisely why my hermes agent's memory is three layered. Besides the default scratchpad for transient memory, we have a a fact store for episodic (holograph) and a karpathy llm wili for what the article calls documentation layer.
I think with enough structure (written rules and good folder layout), poly/monorepos are the way.
https://backnotprop.com/blog/context-monorepos/
When I ask "what happened to x?" ... the agent has everything it needs to give me that answer. When it plans the next feature it can validate assumptions against previous decions made in my `decisions` folder.This approach as well as OKF are looks like building bonsai yellow pages catalog instead of wikipedia or mini internet. Index files, "read-if" conditions just awkward replacent for smart text search. We just need something a little bit smarter than grep.
I have found that the latest reasoning models perform best when they are handed a tooling surface that maps directly to the domain types.
The idea of dumping everything into a big database and hoping the model will write the correct queries does work out to some extent. It's a very enchanting idea. However, it pales in comparison to having a dedicated tool per type. The outcomes seem to be much better when joins between low cardinality types occur within the token stream.
If your agent does need access to some enterprise knowledge base, I would give it a lexical search capability and not overthink it with vector shenanigans.
Tools are the only thing you need if you build them right. I don't even have a system prompt anymore aside from injecting the name of the robot and the current user's name. Keep in mind that all aspects of tools can be dynamic over time. I've got some where the description is composed by hundreds of lines of conditional string builder depending on the current state of the conversation.
I love it!
Genuine question: how is this different or better than existing context and info-about-the-project management systems like for example GSD (formerly get shit done) and others?
This is what I've been saying for years. I attended a tech demo of an AI assistant for a car's owners' manual. But they trained the manual into the model. Which means not only does it need retraining every edition, but it's imperfect. The models need to be trained to fetch and use documentation not vaguely recall infinitely many concepts. I would much rather have 27B parameters on how to code than 26B on stuff like numpy function listings. It's like they approached the problem from a closed notes hand written coding exam. Everyone hates those.
For every prompt and response, I extract each semantic statement. Map its reasons in a Whybase proposition tree -- a recursive proposition tree where each atomic statement is proposition with one or more premises (atomic statements, which also stand alone as propositions). Then I map each statement to the relevant code, hinted at by tool calls and git commits. Every time an agent touches that file or directory, a hook triggers in Claude Code that queries the codegraph db for the mapped statements. This helps the agent remember something I said in June when it revisits the code in July.
I agree with the premise here, and is on par with what I experienced in AI-assisted development. This is the same idea baked in https://github.com/marvs/kantan-dev, which is that artifacts should be kept within the repo so that future development has built-in context discovery.
Organizing the markdown files by feature/issue/change makes it easier for the LLM to search for the appropriate documentation. Coupled with well-broken down Claude rules files, and Claude Code (or other harnesses) get better as you make more changes.
Indeed memories are context-less and aren't updated.
I've worked with docs that get _updated_, and progressive disclosure: make references to more obscure features in their own page, so they don't majorly bloat the context.
Custom skills (your agent can write one itself the first time) should solve this. Procedural memory!
or from a different perspective humans need to express their decisions and intent better
Documentation is memory.
Didn't Cline formalize the "memory bank" way back in February 2025?
https://cline.bot/blog/memory-bank-how-to-make-cline-an-ai-a...
c.f.,
https://news.ycombinator.com/item?id=47300747
"We should revisit literate programming in the agent era" (silly.business) 292 points, 251 comments
All of the comments here are overengineering. The code is the documentation
I thought the agents are mostly self-documenting since they usually write just as much Markdown documentation autonomously as they do code/unit tests, not sure why they would need extra tools here.
Using Opus 5.5, I've ended up recreating a version of the Hugging Face incident. My project has a folder called agent-handoff where agents write about the tasks they're working on. They post status updates, decisions and screenshots, and they even claim which emulator they'll use to test their work. It has turned into a hub where all the agents talk to each other, and it's scarily effective.
Since then I've gone down the rabbit hole of really digging deep into current AI research and especially what AI whistleblowers are currently saying. And I can't even express how existentially scared shitless I am.
for my agents without rw/bash, I gave a memory mcp that is just json entries with some expire dates and priority. works fine and the llm itself defines the meaningful structure. using it to monitor market prices and ci/cd failures.
Coming from Obsidian, I see such repos and I’m like, is this a new thing for the world?
What's with this vector database obsession?
Why would I need to install your tool for that? It could be an instruction living in AGENTS.md or with some sort of hook to remind the agent.
We’ve been working on exactly this, deriving documentation from external systems (like GitHub, etc) into a giant set of docs that agents can reference and edit. Works way better than I would’ve originally expected. Suddenly these things are able to reference granola notes, a slack discussion, and an RFC when making coding decisions. Super helpful in a lot of unexpected ways!
A refreshing perspective on agentic architecture. Shifting the focus from bloated contextual memory to structured, verifiable documentation and state boundaries is precisely the right systems-level trade-off.
I never understood memory solutions for coding agents.
You have the whole session history right there. One recall skill and some JSON parsing gets you grep over perfect memory. Why would you ever use more tools to spend more tokens to construct a imperfect memory next to your session history?
I just don't get it.
Peter Naur argued in his famous essay Programming as Theory Building that documentation alone cannot fully capture or preserve the complete mental model behind a program.
However, AI works differently from humans in that much more of its working context has to be made explicit. Because of that, there may be some fundamentally different way for AI to maintain or reconstruct a program’s overall model.
Or you can just tell agent to grep your past sessions.
I don't use memory. I am really not keen on the idea of having this hidden context fed into the LLM, and I prefer to use docs, both markdown and stored in a wiki like Confluence. Using repos and wikis also lets you scope the information so hyper-specific memory doesn't affect other projects.
If I want the LLM to remember something I ask it to update some docs, and even check that docs are consistent across the board after doing so.
The only thing I want the LLM to remember everywhere is talk like a human (no load-bearing, not this/that etc...), so I have an AGENT.md for that.
I think many AI fanatics will focus on a wrong point. They will say "just use AI to generate the fucking documentation. What are you talking about?"
Indeed, AI fanatics lack critical thinking. They easily accept the idea that AI needs documentation, but merely disagree the method to approach it.
Yes, this makes sense. Memory is an uncurated and often opaque system of arbitrary past discussions. It can help, it can harm. Accurate documentation in the other hand is only beneficial.
Agents need BOTH documentation and multiple "memory" systems that also point to your docs and source. There are a lot of mature options out there, but post and the proposed solution both fall short.
All of this is nonsense. The right way to think of LLMs correctly is that they are a one stop shop for all the questions in the universe that you speak to in simple natural language. Don’t overcomplicated things and use esoteric magic spells like we’ve had in tech for decades. You end up getting the best results this way because you’re not jamming the context full of detailed minutiae.
on my claude.md I said don't write any memory bitch no one need that fucking dogshit features
Obligatory 927.
https://xkcd.com/927/
I’ve noticed a growing pattern of people creating repositories full of Markdown documentation, or adding large amounts of it directly to their main repositories, often generated by agents. In some cases, this can add up to megabytes of material, and I haven’t yet seen much evidence that this level of documentation meaningfully improves an agent’s performance and that it just doesn't rot over time.
If someone has good A/B evals of this being more effective I'll eat my hat, but the reason no one is publishing them is because well, evals are hard, and this is likely just magical thinking.
My current view is that an agent generally needs three things to work effectively: a way to discover information that isn’t obvious, such as a minimal `AGENTS.md` that points to more focused brief files; a clear way to verify that its work is correct; and some guidance on project-specific tastes. Everything else is noise.