AI agent memory architecture is the decision about what an agent keeps between sessions, where it is kept and who else can read it. In my Claude Code setup that comes down to 4 kinds of memory kept in plain files, with every agent starting each session from zero and reading only the part it needs.
The agents I run for marketing at EmpoweredHouse all work this way. The case study describes the whole setup; this article covers the memory underneath it.
Why Claude Code forgets context between sessions
Claude Code forgets because every session starts empty. Anthropic's documentation states that each Claude Code session begins with a fresh context window. 2 things carry over: CLAUDE.md files, which you write, and auto memory, notes Claude writes for itself, of which the first 200 lines or 25 KB of MEMORY.md load at the start.
Inside a session the window fills. Compaction then replaces the conversation with a structured summary, and Claude re-reads the project root CLAUDE.md from disk. An instruction given only in the conversation does not come back, and Anthropic lists that as the first thing to check when an instruction disappears after compaction.
So I stopped trying to make my agents remember more, and for anything that matters I now ask where it will live once the conversation is gone.
The 4 kinds of memory in an AI agent memory architecture
I sort what an agent needs to know by how long it stays true: a rule goes into the role card, this week's status into the state file, and a lesson another agent needs into the log, with that agent's name on it.
| Kind | The question it answers | How long it lives | Where I keep it | Closest CoALA term |
|---|---|---|---|---|
| Durable | Who am I, and what may I do? | Until someone edits it | Role card, one CLAUDE.md per agent | Procedural and semantic |
| Live | What is happening now? | Until the next update | State file, requests file | Working |
| Log | What did someone learn, and who is it for? | Append-only, then archived | Insights log | Episodic |
| History | What changed, and how do I undo it? | As long as the repository | Local git repository | None |

4 kinds of memory, each in its own file.
The last column borrows the vocabulary of CoALA, the 2023 framework by Sumers, Yao, Narasimhan and Griffiths that sorts the memory of language agents into working, episodic, semantic and procedural. History has no counterpart there, and in my experience it matters most when something breaks, because only history can be undone.
Infographic 1, where does this go: a decision tree from the question "how long will this stay true?" to the role card, the state file or the insights log, with git under all 3.
Multi-agent memory architecture: many readers, one writer per row
Once several agents read the same files, shared memory can break in 2 ways: agents overwrite each other, or every agent reads everything. Ownership fixed the first for me, addresses the second.
Ownership means each agent writes only in its own folder and owns one row in any shared file. The case study shows what happened before I had that rule.
Addresses mean every log entry names who it is for. At the start of a session an agent reads the entries addressed to it or to everyone, plus the 12 most recent. Most entries in my log are for someone else, so each agent skips the bulk of the file.
Infographic 2, one log, many readers: the same insights log drawn 3 times, each copy highlighting what one agent reads at the start of a session.
Live memory also needs an expiry. In my setup a check flags any state line nobody has touched for 7 days, because a week-old status is a guess.
Where agent memory goes wrong
These are the 4 mistakes I would watch for first.
- A rule lives only in the chat. It disappears at compaction, so a rule that matters goes into the role card the moment it is agreed.
- A live fact sits in a durable file. A role card that states this week's number is wrong next week, so numbers belong in the state file.
- The log grows until nobody reads it. Without addresses every agent reads all of it, so address every entry and archive old ones once the file gets long.
- 2 instruction files contradict each other. Anthropic's documentation warns that Claude then picks one of the contradicting rules arbitrarily. Keep each fact in one file, and read a nested CLAUDE.md next to the ones above it when you change it.
Infographic 3, where agent memory goes wrong: 4 rows, each pairing one mistake from the list with its fix.
Where MEMORY.md, CLAUDE.md and AGENTS.md fit
CLAUDE.md holds instructions you write, and MEMORY.md holds the notes auto memory writes. Claude Code reads AGENTS.md only when there is no CLAUDE.md in the working directory or above it, unless CLAUDE.md imports AGENTS.md or the project setting says otherwise.
I keep shared memory out of auto memory on purpose. Anthropic describes auto memory as machine-local, with one memory directory per project, and my state has to be read by every agent and by a person. The 3 files side by side are in MEMORY.md vs CLAUDE.md vs AGENTS.md.
How to set up an AI agent memory architecture
If I started again, I would set up these files in order:
- A role card per agent, under the 200 lines Anthropic recommends, whose first sentence names the agent.
- A state file with one dated line per agent: what it works on and what blocks it.
- An append-only log where every entry names its reader.
- A requests file with one row per agent, where each agent changes only its own status.
- A local git repository with no remote, with a commit at the end of every session.
- One rule in every card: read your rows at the start, and write before you stop or when the conversation grows long.
I have not needed a framework for any of it. It is the EmpoweredHouse method applied to memory: every change leaves a trace a person can read and undo. The case study has a setup prompt that creates these files for you.
Frequently asked questions
Does Claude Code have memory between sessions? Only through files. Each session starts with a fresh context window, CLAUDE.md files and the first 200 lines or 25 KB of MEMORY.md load at the start, and anything else has to be written down.
What should I do when the Claude Code context window is full? Before it fills, write the decision, the open task and the next step to a file. Compaction replaces the conversation with a summary and re-reads the project root CLAUDE.md, so what you wrote there survives. The commands are in Claude Code context window full: what to do.
Can several AI agents share one memory? In my experience they can, through shared files, if each row has one writer and each log entry names its reader. Auto memory does not fit this job, because it stays on one machine.
What is the difference between short-term and long-term memory in AI agents? Short-term memory is the context window of one session. Long-term memory is whatever is kept outside the model and read back in: instructions, state, logs and history.
What is episodic memory in AI agents? Episodic memory is a record of specific past events, as distinct from general facts. In my setup it is the insights log: dated, addressed entries of what one agent learned.
If each of your agents does one job alone, in my experience a good prompt is enough memory. If several work on the same company, I would give them files before more tools, and you can compare notes with the people doing the same in Empowered Community.
Facts checked against Anthropic's documentation on 28 September 2026.