Agent Memory Systems
An LLM forgets everything between calls. "Memory" is the set of engineering choices that decide what an agent carries forward: inside one long task (working memory and context management) and across sessions, users and months (long-term memory). Interviewers ask about it because it sits where product, data engineering, retrieval, security and cost all meet. A good answer shows you can design the write path, the read path, the storage and the forgetting policy, and that you know how to evaluate and secure the result, rather than just naming Mem0 or MemGPT.
TL;DR: the 8–12 things to be able to say out loud
- Memory is context engineering over time. The model only "remembers" what is in the prompt on this call. Every memory system is a policy for deciding what goes into a finite context window, where it comes from, and what gets thrown away.
- Two taxonomies. The cognitive one (working, episodic, semantic, procedural; formalised for agents by CoALA) tells you what kind of thing is stored. The engineering one (in-context vs external, hot path vs background, user/session/agent/org scope) tells you how it is built.
- Short-term memory is a budgeting problem. Sliding windows, token-budgeted truncation, rolling summaries, compaction, tool-output clearing and agent-written notes files. Bigger windows don't remove the need: quality degrades with length ("context rot", "lost in the middle").
- Six long-term patterns: raw vector memory, extracted-fact memory (Mem0-style), temporal knowledge graphs (Zep/Graphiti), OS-style tiered memory with self-editing (MemGPT/Letta), reflection-based memory (Generative Agents), and procedural memory (skill libraries, instruction files, Skills).
- Generative Agents retrieval score = recency + importance + relevance, each min-max normalised, recency an exponential decay (factor 0.995 per game hour). Reflection fires when summed importance passes a threshold (150).
- Write path: decide what is salient, extract it as atomic, timestamped facts with provenance, dedup, and resolve conflicts (overwrite, invalidate with validity intervals, or append and let retrieval rank). Each choice trades accuracy on "knowledge update" questions against audit history.
- Read path: always-inject (profile / core blocks) vs on-demand retrieval (tool call). Rank by relevance × recency × importance, use hybrid search (vector + BM25 + entity/graph), and inject compactly in a clearly delimited, clearly untrusted block.
- Forgetting is a feature: decay, TTLs, merging, summary hierarchies, and hard deletion for user requests and GDPR Art. 17, including derived artifacts (embeddings, summaries, graph edges).
- Memory is an attack surface: injected memories persist across sessions (memory poisoning; AgentPoison, MINJA). Defend with write-time filtering, provenance, tenant isolation, and treating memory as data rather than instructions.
- Evaluation: LoCoMo, LongMemEval (5 abilities incl. knowledge updates and abstention), and newer BEAM (up to 10M tokens). Always report accuracy with latency and tokens per query, and be suspicious of vendor leaderboards.
- Reference architecture: Postgres + pgvector for facts and episodes, optional graph store for entities and relations, KV/Redis for session state, object storage for raw transcripts, an async consolidation worker, and a memory service API that enforces scope and deletion.
1. Why agents need memory
A transformer forward pass is a pure function of its input tokens. Nothing persists between API calls except what you send back in. A chat app that "remembers" earlier turns is just re-sending them. So memory in an LLM system is never a property of the model alone. It is a property of the harness around it: what you store, where, and what you put back into the prompt next time.
Four pressures drive the need for something more than "re-send everything":
Context limits and cost
Windows have grown to hundreds of thousands or around a million tokens for some models, but every token you resend is paid for on every call and adds prefill latency. A year of daily chats is far beyond any window. Even for an agent within one task, tool outputs (file reads, search results, logs) fill the window within tens of steps.
Quality degrades with length
Models use information at the start and end of the context better than in the middle Liu+ 2023, and accuracy falls as input grows even on simple tasks, especially when distractors are present Hong+ 2025. Curating a small, relevant context often beats dumping a large one.
Personalisation and continuity
Users expect the assistant to know their name, stack, preferences and the state of an ongoing project without repeating it. Long-term memory is mostly a product feature: it reduces friction and makes the agent feel like it knows you.
Learning from experience
Without weight updates, an agent can still improve by storing what worked (successful trajectories, lessons, reusable code or instructions) and recalling it next time. Reflexion Shinn+ 2023, ExpeL Zhao+ 2023 and Voyager Wang+ 2023 are the classic examples.
Think of the context window as RAM and everything else as disk. The model can only compute over what is in RAM. A memory system is the virtual-memory manager: it decides what to page in, what to page out, what to compress, and what to delete. The MemGPT paper makes this analogy literal.
One more framing helps in interviews: memory vs RAG. Classic RAG retrieves from a mostly static corpus that someone else wrote (docs, wiki). Memory retrieves from a corpus the agent itself writes, continuously, from interactions. That changes almost everything about the problem: the corpus is tiny per user but highly dynamic; facts go stale and contradict each other; time matters; the write path (what to store, how to dedup, how to resolve conflicts) is as hard as the read path; and the data is personal, so deletion and isolation are hard requirements. Retrieval mechanics (embeddings, hybrid search, reranking) carry over from the retrieval page.
"Why not just use a 1M-token context?" is a common opener. A strong answer covers four points. Cost and latency: you pay to prefill every token on every turn; prompt caching helps but doesn't remove the cost. Quality: context rot and position effects mean more tokens can lower accuracy. Scale: multi-year, multi-session histories still don't fit. Structure: raw history doesn't resolve contradictions ("I moved to Berlin" after "I live in London"), doesn't give you deletion or scoping, and doesn't distill lessons. Long context and memory are complements: a big window lets you inject more retrieved memory and run longer before compacting.
- Anthropic: Effective context engineering for AI agents (Sep 2025): compaction, note-taking, sub-agents and just-in-time context in one place.
- Chroma: Context Rot: 18 models, how performance degrades as input grows.
- Lost in the Middle (Liu+ 2023): the position-sensitivity result everyone cites.
2. Taxonomies of memory
2.1 The cognitive-science split (CoALA)
The most-cited framing comes from cognitive architectures. CoALA (Cognitive Architectures for Language Agents) describes an agent as an LLM plus modular memories plus a structured action space Sumers+ 2023. It distinguishes:
| Memory | CoALA definition (paraphrased) | Concrete agent example | Typical storage |
|---|---|---|---|
| Working | Active information for the current decision cycle: perceptual input, retrieved knowledge, goals, intermediate results | The prompt on this call: system prompt, recent turns, tool results, scratchpad | The context window itself, plus run state in your orchestrator |
| Episodic | Past experiences and trajectories | "Last Tuesday the user asked me to refactor auth; we chose JWT"; logs of past tool-use runs; few-shot examples of past successes | Event log, transcript store, vector index over episodes |
| Semantic | Knowledge about the world and the agent itself, updatable through learning | "User is vegetarian", "Acme's fiscal year ends in March", "the repo uses pnpm" | Fact table, profile document, knowledge graph |
| Procedural | How to act: implicit in LLM weights plus explicit code and prompts implementing the agent | System prompt, tool definitions, learned skills (code or instruction files), refined instructions | Prompt registry, skill library, files like CLAUDE.md / SKILL.md |
CoALA also defines internal actions on memory: retrieval (long-term into working), reasoning (working into working), and learning (writing to any long-term memory, including updating procedural memory) Sumers+ 2023. That vocabulary maps directly onto the read path, the scratchpad, and the write path in an engineering design.
Treating "episodic vs semantic" as a storage decision. It is a decision about granularity and abstraction. The same storage (a Postgres table with an embedding column) can hold both. What differs is the write process: episodic memory records what happened, with time and context. Semantic memory is distilled from episodes into timeless-ish facts, and therefore needs conflict resolution. Most production systems keep both and link facts back to the episodes they came from (provenance).
2.2 The engineering split
In system design you'll usually be judged on the engineering axes, not the cognitive labels:
| Axis | Options | Trade-off |
|---|---|---|
| Location | In-context (always in the prompt: system prompt, profile, core memory blocks) vs external (stored outside, retrieved on demand) | In-context is reliable and zero-latency to "recall" but costs tokens every call and is size-bounded. External scales but recall can fail silently. |
| When written | Hot path (agent writes memory during the turn, often via a tool) vs background (async worker processes the transcript after or between turns) | Hot path: immediate availability, user can see "memory saved", but adds latency and distracts the agent. Background: no latency hit, better batch reasoning, but memory lags by seconds to hours. LangGraph's docs frame exactly this choice LangChain docs. |
| Who decides | Explicit ("remember that I…", user-curated) vs implicit (system infers from conversation) | Explicit is precise and consensual but sparse. Implicit has high coverage but risks storing wrong, sensitive or creepy facts. |
| Scope | Session/thread, user, agent (what this agent learned across all users), org/tenant (shared team knowledge), sometimes project | Wider scope means more sharing value and more leakage risk. Scope must be a first-class key in storage and in every query. |
| Representation | Raw text chunks, extracted facts, profile document (one JSON per user), collection of small docs, graph, code, files | LangGraph contrasts a single profile (easy to inject, error-prone to update as it grows) with a collection (easy to add to, harder to dedup and update) LangChain docs. |
| Form | Token-level (text you can read), parametric (in weights: fine-tuning, LoRA), latent (KV caches, hidden states) | Almost all production memory is token-level because it's inspectable, editable and deletable. The newer survey literature uses this forms/functions/dynamics framing Hu+ 2025. |
If asked to "classify memory types", give the cognitive four in one sentence each, then pivot quickly to the engineering axes: "In practice the decisions that matter are scope, hot path vs background writes, and always-in-context vs retrieved." Then give an example mapping: a support agent's user profile is semantic, user-scoped, in-context; past tickets are episodic, user-scoped, retrieved; resolution playbooks are procedural, org-scoped, retrieved or loaded as skills.
- CoALA (Sumers+ 2023): the framework paper; read the memory and action-space sections.
- LangGraph memory concepts: a practical mapping of semantic/episodic/procedural and hot-path vs background writes.
- Memory in the Age of AI Agents (Hu+ 2025): a large survey that argues short/long-term is too coarse and proposes forms/functions/dynamics.
3. Short-term and working memory management
Before you add any database, the first memory problem in every agent is keeping one long conversation or task inside the window without losing what matters. This is "context management" or "context engineering". Here are the techniques, from simplest to most sophisticated.
3.1 Buffers, windows and token budgets
- Full buffer. Send the whole history. Correct until it doesn't fit; cost grows quadratically over a conversation, because turn \(n\) re-sends all \(n-1\) previous turns. Total tokens processed over \(N\) turns of \(t\) tokens each is about \(t \cdot N(N+1)/2\). Prompt caching cuts the price of the repeated prefix but not the window limit.
- Sliding window (last \(k\) turns). Simple and cache-unfriendly (the prefix shifts every turn, so a cached prefix is invalidated unless you drop in large chunks). It forgets the user's opening goal, which is often the most important message.
- Token-budgeted truncation. Allocate a budget per section, e.g. system prompt + tools (fixed), pinned items (task statement, user profile), retrieved memory (≤ \(B_m\) tokens), recent turns (fill the remainder, newest first), and reserve output tokens. Always count with the model's real tokenizer. Truncate whole messages and keep tool call/result pairs together, because most APIs reject an orphaned tool result.
- Pinning. Always keep the first user message (the goal), any explicit constraints, and the latest state. Drop the middle first, which is also where the model attends least.
3.2 Rolling summarisation and compaction
Rolling summary: keep the last \(k\) turns verbatim plus a running summary of everything older. When the buffer exceeds a threshold, summarise the oldest chunk and merge it into the summary. MemGPT does this with a "recursive summary" of evicted messages Packer+ 2023.
Compaction is the agent-era name for the same idea applied when the context nears the limit: replace older turns with a high-fidelity summary and continue Anthropic 2025. Model providers now offer it as an API feature. Anthropic's API, for example, can compact on demand or automatically at a token threshold, optionally keeping recent turns verbatim Anthropic docs.
The craft is in what the summary must keep. A generic "summarise this conversation" prompt loses exactly the details an agent needs next. A good compaction prompt asks for:
| Keep | Why | Example |
|---|---|---|
| The goal and acceptance criteria | Without it the agent drifts | "Migrate billing service to Postgres 16; all tests green; no downtime" |
| Decisions made, with the reason | Prevents re-litigating or contradicting earlier choices | "Chose logical replication over pg_upgrade because of the downtime constraint" |
| Open tasks / next steps | The continuation point | "Remaining: update connection strings in 3 services; run load test" |
| Exact identifiers | Summaries paraphrase; IDs must be verbatim | File paths, ticket IDs, URLs, function names, error messages, order numbers |
| User preferences and constraints stated in-session | Most painful thing to forget | "Don't touch the legacy/ folder" |
| Failures and dead ends | Stops the agent repeating them | "Tried bumping driver to v5: breaks SSL; reverted" |
| Drop: raw tool outputs, superseded drafts, chit-chat | High-token, low-value | The 3,000-line log already analysed |
Compaction is lossy and compounding. Each round summarises a summary, so errors and omissions accumulate ("summary drift"). Mitigations: keep identifiers and decisions in a structured section the summariser must copy forward verbatim; keep the original transcript in cold storage so the agent (or a human) can look things up; and pair compaction with an external notes file that holds durable state, so the summary doesn't have to. Also, compaction usually invalidates the prompt cache, so don't compact too often.
3.3 Scratchpads and agent-maintained notes files
The most effective pattern for long-horizon agents in 2025–26 is surprisingly low-tech: let the agent keep notes in files outside the context window and read them back when needed. Anthropic describes this as structured note-taking: an agent maintains a NOTES.md or to-do list to track progress across many tool calls and context resets Anthropic 2025. For multi-session coding, their long-running harness uses an initializer session that creates a progress file (claude-progress.txt), a JSON feature list with every feature initially marked failing, an init script, and git commits. Each later session starts by reading these and ends by updating them Anthropic 2025.
Why it works so well:
- Agent-curated salience. The model writes down what it thinks it will need, at the moment it understands why. That beats a summariser guessing after the fact.
- Survives resets. Notes outlive compaction, crashes and new sessions. Anthropic's memory-tool system prompt literally tells the model to "assume interruption" Anthropic docs.
- Models are trained on file tools. Reading, grepping and editing files is in-distribution for coding-trained models. Letta reported an agent with plain file tools scoring 74.0% on LoCoMo with GPT-4o-mini, above the figure Mem0 reported for its graph variant, and concluded that how the agent manages context matters more than the retrieval mechanism Letta 2025. Treat that as a vendor result, not settled science.
- Inspectable and editable by humans, which matters for trust and debugging.
JSON is better than Markdown for state the agent must not casually rewrite (Anthropic found the model was less likely to inappropriately edit a JSON feature list) Anthropic 2025. Markdown is better for free-form notes.
3.4 Tool-output pruning and clearing
In agentic loops, tool results are usually the biggest consumer of context: a file read, a web page, a SQL result. Once the agent has acted on a result, the raw bytes are rarely needed again. Tool-result clearing replaces old results with a placeholder ("[result cleared; re-run tool if needed]") while keeping the call itself, so the agent knows what it did. Anthropic's API exposes this as a context-editing strategy with a trigger threshold (default 100k input tokens), a number of recent tool uses to keep (default 3), an exclude list of tools, and a minimum amount to clear per activation so the cache invalidation is worth it Anthropic docs. When combined with a memory tool, the model gets a warning before clearing so it can save anything important to memory first.
Related tactics: return references rather than payloads (file path + line range, a query, a URL) so the agent can re-fetch just-in-time Anthropic 2025; have tools truncate and paginate by default; and push heavy exploration into sub-agents with their own clean windows that return condensed summaries (on the order of 1–2k tokens) to the orchestrator.
| Technique | What it does | Cost | Loses | Use when |
|---|---|---|---|---|
| Full buffer | Resend everything | Quadratic token growth | Nothing (until overflow) | Short chats |
| Sliding window | Last \(k\) turns | Bounded | The goal, early constraints | Casual chat, low stakes |
| Budgeted + pinned | Sections with token budgets; pin goal | Bounded | Middle turns | Default for most products |
| Rolling summary | Summary + recent verbatim | An extra LLM call per roll | Detail; drifts over many rolls | Long chats |
| Compaction | Summarise when near limit | One big LLM call; cache reset | Detail; risk of dropping IDs | Long agent runs |
| Tool-result clearing | Drop stale tool outputs | Near zero; cache reset | Raw outputs (re-fetchable) | Tool-heavy agents |
| Notes / todo files | Agent writes durable state externally | Some tool calls | Whatever the agent failed to write | Multi-hour or multi-session tasks |
| Sub-agents | Isolate exploration in separate windows | More total tokens | Detail behind the summary | Broad research / search |
"Your coding agent fails after ~40 steps. Why, and what do you do?" Strong answers diagnose first (look at traces: is the context full of file reads? did compaction drop the plan? is the agent re-doing work?) and then layer the fixes: clear stale tool results, return references not payloads, keep a todo/progress file the agent re-reads, compact with a structured prompt that preserves decisions and IDs, and move exploration to sub-agents. Mention the prompt-cache cost of each intervention: edits near the start of the context invalidate the cache for everything after them.
- Anthropic: Effective harnesses for long-running agents: the progress-file and feature-list pattern in detail.
- Anthropic context editing docs: concrete parameters for tool-result and thinking-block clearing.
- Letta: Is a filesystem all you need?: a provocative (vendor) result on file-based memory.
4. Long-term memory architectures
Six recurring designs. Real systems mix them: Letta combines tiered memory with files, Zep combines episodes with a graph, Mem0 combines fact extraction with hybrid retrieval. Learn each as a mechanism with a characteristic failure mode.
4.1 Vector-store memory (embed and retrieve past interactions)
The baseline. Chunk past conversations (per turn, per exchange, or per session), embed each chunk, store it with metadata (user_id, session_id, timestamp), and at query time embed the current message, retrieve the top-\(k\) similar chunks, and inject them. It is RAG where the corpus is your own history.
- Write cost: one embedding call per chunk; no LLM. Cheap and lossless (the raw text is kept).
- Read cost: one embedding + an ANN query (milliseconds). Injected chunks can be verbose.
- Failure modes: no notion of truth over time (both "I live in London" and "I moved to Berlin" are retrieved, and the model may pick the wrong one); similarity ≠ relevance (a query about "my flight" retrieves every travel chat); granularity issues (a turn-level chunk lacks context, a session-level chunk dilutes the embedding); temporal questions ("what did I say last week?") fail because embeddings don't encode time.
LongMemEval's authors found three cheap upgrades to this baseline help: decompose sessions into finer-grained units, expand the index keys with extracted facts (index a chunk under the facts it contains, not just its raw text), and use time-aware query expansion for temporal questions Wu+ 2024. Adding BM25 alongside vectors (hybrid search) also helps a lot for names, IDs and rare terms.
4.2 Extracted-fact memory (Mem0-style)
Instead of storing raw text, an LLM extracts salient facts from each new exchange and stores those as small, atomic memories ("User is allergic to peanuts"). The original Mem0 paper describes a two-phase pipeline Chhikara+ 2025:
- Extraction: given the new message pair plus context (a conversation summary and the last \(m=10\) messages), the LLM proposes candidate facts.
- Update: for each candidate, retrieve the top \(s=10\) most similar existing memories, and ask the LLM (via a tool call) to choose one operation:
ADD(new fact),UPDATE(augment or refine an existing one),DELETE(the new fact contradicts an old one), orNOOP(already known).
The paper reported, on LoCoMo, a 26% relative improvement on an LLM-as-judge metric over OpenAI's memory, 91% lower p95 latency than full-context (about 1.4 s vs 17 s), and over 90% fewer tokens; memories averaged about 7k tokens per conversation vs about 26k for the full conversation Chhikara+ 2025. A graph variant (Mem0g) adds entity and relation extraction and did better on temporal questions.
Mem0 changed its core algorithm in 2026. According to Mem0's own migration docs, the new algorithm does single-pass, ADD-only extraction (one LLM call, no UPDATE/DELETE), stores new facts alongside old ones with temporal metadata, and relies on retrieval to surface the current fact. Retrieval became hybrid: vector + BM25 + entity matching, fused into one score. Graph memory was removed from the open-source package and moved into the hosted platform Mem0 docs. Mem0 reports large benchmark gains from this change (LoCoMo roughly 71 → 92, LongMemEval roughly 68 → 93), but those are vendor-reported numbers. The design lesson stands either way: reconcile on write (ADD/UPDATE/DELETE) and append on write, resolve on read are two legitimate strategies, and the field has been moving toward the second.
Reconcile on write (ADD/UPDATE/DELETE)
+ Compact store, no stale facts to confuse the reader, cheap reads.
− Two LLM calls per write; destructive (a wrong DELETE loses data forever); no history ("where did the user live before?"); an extraction mistake overwrites a correct fact.
Append on write, resolve on read
+ One cheap write; non-destructive; full history; easy audit.
− Store grows; reader must rank by recency/validity; risk of stale and current facts both surfacing (the "knowledge update" category suffers first).
4.3 Knowledge-graph and temporal memory (Zep / Graphiti)
Facts are not independent: people work at companies, projects have owners, preferences change. A temporal knowledge graph stores entities as nodes and facts as edges (subject, predicate, object), and gives every edge a validity interval. Zep's engine, Graphiti, is the reference design Rasmussen+ 2025:
- Three tiers. An episode subgraph holds the raw messages (non-lossy, the provenance); a semantic entity subgraph holds extracted entities and relation edges, linked back to episodes; a community subgraph clusters strongly connected entities with summaries.
- Bi-temporal model. Each fact carries two timelines. The valid-time timeline \(T\) (\(t_\text{valid}, t_\text{invalid}\)) says when the fact was true in the world. The transaction-time timeline \(T'\) (\(t'_\text{created}, t'_\text{expired}\)) says when the system learned or retired it. This is the same idea as bi-temporal tables in databases.
- Edge invalidation instead of deletion. When a new fact contradicts an existing, temporally overlapping one, the old edge's \(t_\text{invalid}\) is set to the new edge's \(t_\text{valid}\). You can ask both "where does Alice work now?" and "where did Alice work in March?"
- Hybrid retrieval. Cosine similarity, BM25 full-text and breadth-first graph traversal run in parallel, then get reranked (RRF, MMR, graph distance, episode-mention frequency, or a cross-encoder).
Zep reported 94.8% vs MemGPT's 93.4% on the older DMR benchmark (which the paper itself criticises as too easy), and on LongMemEval up to 18.5% accuracy improvement with about 90% lower latency, using about 1.6k context tokens instead of about 115k for full context Rasmussen+ 2025. Open-source Graphiti runs on Neo4j, FalkorDB or Amazon Neptune Graphiti repo.
| Edge | t_valid | t_invalid | t'_created | t'_expired |
|---|---|---|---|---|
| (Alice) –WORKS_AT→ (Acme) | 2023-01 | 2026-05-01 | 2025-02-10 | 2026-06-03 |
| (Alice) –WORKS_AT→ (Globex) | 2026-05-01 | ∅ (current) | 2026-06-03 | ∅ |
In the example, on 3 June 2026 Alice says "I started at Globex at the start of May". The system learns it (transaction time 06-03) and back-dates validity to 05-01. The Acme edge is invalidated, not deleted. "Where did Alice work in April 2026?" still answers correctly.
Reaching for a graph database because "memory is a graph". Graph memory has the highest write cost (entity extraction, entity resolution/dedup, relation extraction, contradiction detection: several LLM calls per episode) and entity resolution errors ("Bob" vs "Robert" vs "my manager") compound. It pays off when questions are relational or temporal (multi-hop, "who on my team…", "what changed since…"), or for enterprise data with natural entities. For a consumer chatbot that mainly needs preferences, a fact table with timestamps is usually enough. The bi-temporal idea (valid-from/valid-to columns, invalidate rather than delete) is valuable even in plain Postgres.
4.4 OS-style hierarchical memory (MemGPT / Letta)
MemGPT treats the LLM like a CPU with a small, fast main memory (the context window) and large, slow external storage, and lets the LLM itself page data in and out with function calls ("virtual context management") Packer+ 2023.
Key mechanics from the paper Packer+ 2023:
- Main context = read-only system instructions + a fixed-size, LLM-editable working context + a FIFO message queue whose evicted messages are replaced with a recursive summary.
- External context = recall storage (the full message history) and archival storage (an arbitrary read/write text database, typically vector-indexed).
- Queue manager issues a "memory pressure" warning at about 70% of the window so the LLM can save important things, and at the limit flushes about half the queue into recall storage plus summary.
- Self-editing memory: the LLM decides what to write to working context or archival storage by calling functions. Function chaining via a heartbeat flag lets it make several memory calls before replying.
Letta (the company built around MemGPT) generalised the working context into memory blocks: labelled, size-limited strings (default persona and human) that always sit in the context, each with a description that tells the agent what belongs there Letta docs. Blocks can be shared between agents (useful for multi-agent shared state). Letta later added "sleep-time" agents that rewrite memory in the background between interactions Lin+ 2025.
Letta's direction moved in 2026 toward git-backed memory files ("context repositories" / MemFS): an agent's memory is a folder of Markdown files with frontmatter descriptions, a system/ directory whose files are always loaded into the prompt, and git versioning so multiple sub-agents can write memory concurrently in separate worktrees and merge Letta 2026. Background "dreaming" sub-agents consolidate recent conversations into memory Letta docs. Secondary sources say Letta has de-emphasised its older server-side framework; check current docs before quoting tool names.
The MemGPT contribution is less the tiers (any RAG system has "context + store") than agency over memory: the model decides what to remember and when to look things up, using the same tool-calling loop it uses for everything else. The cost is that memory quality now depends on the model's judgement and on spending tokens and steps on memory management. The trend toward files is the same idea with a better-trained interface.
4.5 Reflection-based memory (Generative Agents)
Park et al.'s simulated town of 25 agents introduced the memory stream: a time-ordered log of natural-language observations, each with a creation timestamp, last-access timestamp and an importance score Park+ 2023. Two mechanisms made it influential.
Retrieval scoring. For a query (the agent's current situation), every memory \(m\) gets
$$\text{score}(m) = \alpha_\text{rec}\,\widehat{\text{recency}}(m) + \alpha_\text{imp}\,\widehat{\text{importance}}(m) + \alpha_\text{rel}\,\widehat{\text{relevance}}(m, q)$$where each component is min-max normalised to \([0,1]\) over the candidate set and all \(\alpha = 1\) in the paper's implementation. The components are Park+ 2023:
- \(\text{recency}(m) = \gamma^{\Delta t}\) with decay factor \(\gamma = 0.995\) per sandbox game-hour since the memory was last retrieved (not created). So frequently used memories stay fresh, which echoes spaced repetition.
- \(\text{importance}(m)\): an integer 1–10 assigned by the LLM at write time ("1 is purely mundane, e.g. brushing teeth; 10 is extremely poignant, e.g. a break-up").
- \(\text{relevance}(m,q)\): cosine similarity between the embeddings of the memory and the query.
Back-of-envelope: with \(\gamma=0.995\), a memory untouched for 24 game-hours has recency \(0.995^{24} \approx 0.89\); after a week (168 h), \(\approx 0.43\); after a month (720 h), \(\approx 0.027\). So recency matters on a scale of days to weeks. In a real product you'd choose \(\gamma\) per hour of wall-clock time so the half-life matches how fast your users' facts go stale: half-life \(= \ln 2 / (-\ln\gamma)\), which is about 138 hours here.
Reflection. Raw observations are too low-level to drive behaviour ("Klaus is reading a book on gentrification"). When the sum of importance scores of recent observations exceeds a threshold (150 in the paper, which worked out to roughly 2–3 reflections per game day), the agent takes the 100 most recent memories, asks the LLM for the 3 most salient high-level questions, retrieves memories relevant to each, and asks for 5 high-level insights with citations to the supporting memories Park+ 2023. The insights ("Klaus is dedicated to his research on gentrification") go back into the memory stream as new memories, and can themselves be reflected on, forming a tree of abstractions.
This is the template for most "memory consolidation" jobs today: periodically distill many episodic memories into fewer semantic ones, with links back to the evidence.
4.6 Procedural memory (skills, instructions, learned behaviour)
Procedural memory stores how to do things. It is the least standardised and arguably the highest-leverage form, because it changes behaviour on every future task rather than answering one question.
- Skill libraries as code. Voyager's agent writes JavaScript functions for Minecraft tasks, verifies them in the environment, and stores successful ones in a library indexed by an embedding of their description. New tasks retrieve and compose existing skills Wang+ 2023.
- Lessons and insights. Reflexion stores verbal self-reflections on failures in an episodic buffer used on the next attempt Shinn+ 2023. ExpeL extracts natural-language insights across many tasks' successes and failures and recalls them for new tasks Zhao+ 2023.
- Prompt / instruction self-updates. LangMem treats procedural memory as an updated system prompt and ships optimisers that propose prompt edits from successful and unsuccessful interactions LangChain 2025.
- Instruction files. CLAUDE.md / AGENTS.md style files loaded at session start are human- (and increasingly agent-) maintained procedural memory. Claude Code also has auto memory: a per-repository
MEMORY.mdindex (first 200 lines or 25KB loaded each session) pointing to topic files that are read on demand, with typed notes such as user, feedback, project and reference Claude Code docs. - Skills. Anthropic's Agent Skills package procedures as folders with a
SKILL.mdwhose YAML frontmatter (name, description) is preloaded into the system prompt; the body and bundled files are loaded only when relevant ("progressive disclosure") Anthropic 2025. That's procedural memory with an index in context and the content retrieved on demand.
Letting an agent rewrite its own system prompt or skills without evaluation. A procedural change is a code change that affects every future run: one bad "lesson" learned from a single weird interaction can degrade all users. Treat procedural updates like deploys: propose → evaluate on a regression set → human review for shared scopes → version and allow rollback. User-scoped procedural memory ("this user wants terse answers") is lower risk than agent- or org-scoped.
4.7 Comparison of approaches
| Approach | What's stored | Write cost | Read cost | Strengths | Failure modes |
|---|---|---|---|---|---|
| Raw vector memory | Chunks of past turns/sessions + embeddings + metadata | Embedding only (cheap, no LLM) | ANN query; verbose injection | Lossless, simple, cheap to build, good recall for "what did we discuss about X" | No time or truth model; stale/contradictory facts; similarity ≠ relevance; poor on temporal and aggregation questions |
| Extracted facts (Mem0-style) | Atomic facts with metadata; optionally ADD/UPDATE/DELETE reconciled | 1–2 LLM calls per turn (extraction + reconciliation) | Hybrid search over short facts; compact injection | Compact, token-cheap reads; natural for preferences and profile facts | Extraction misses or hallucinates; destructive updates lose history; append-only variant surfaces stale facts |
| Temporal KG (Zep/Graphiti) | Episodes + entities + relation edges with bi-temporal validity; communities | Highest: several LLM calls (entities, resolution, relations, contradictions) | Hybrid search + graph traversal + rerank; moderate | Temporal and relational questions; history preserved; structured enterprise data | Entity-resolution errors compound; complex ops; ingestion latency; over-engineering for simple use cases |
| OS-style tiers (MemGPT/Letta) | Core blocks in context + recall (messages) + archival (vector) or files | Agent tool calls in the hot path (tokens, steps) | Core: free (always present); archival: agent-initiated search | Agent controls salience; core blocks give reliable always-on personalisation | Depends on model judgement; extra steps and latency; agent may forget to save or search; block bloat |
| Reflection (Generative Agents) | Observation stream + importance + higher-level reflections with citations | Importance-scoring call per memory + periodic reflection jobs | Score all candidates by rec+imp+rel | Produces abstractions and "beliefs"; good for simulation and long-lived personas | Reflections can be wrong and self-reinforcing; costly at scale; importance scores noisy |
| Procedural (skills, instructions) | Code, prompts, instruction files, insights | Low frequency, but needs verification | Index always in context, body on demand | Changes behaviour across all future tasks; compounding improvement | Bad lessons generalise badly; prompt bloat; hard to evaluate; security risk if writable by untrusted input |
| Notes files (NOTES.md, progress logs) | Agent-written free-form or JSON state | File-tool calls in the hot path | File read/grep; just-in-time | In-distribution for coding models; inspectable; survives resets | Unstructured; can sprawl; no dedup or scoping unless the harness enforces it |
"Which memory approach would you pick?" There's no single right answer; interviewers want a justified choice tied to the query distribution. Preference recall for a consumer assistant: extracted facts + a small always-in-context profile. "What did we decide in last month's meeting with Acme?": episodic store with time-aware retrieval. "Who owns the services Alice's team depends on and what changed since Q2?": temporal graph. Long autonomous coding: notes/progress files + compaction. Then say how you'd validate the choice with an eval built from real user questions.
- MemGPT (Packer+ 2023): the OS analogy and self-editing memory.
- Generative Agents (Park+ 2023): memory stream, retrieval scoring and reflection; the appendix has the prompts.
- Zep (Rasmussen+ 2025): bi-temporal graph memory with concrete retrieval and rerank choices.
- Mem0 (Chhikara+ 2025): extraction + ADD/UPDATE/DELETE/NOOP; then read the 2026 migration notes for why they moved to ADD-only.
5. Designing the write path
The write path decides what enters long-term memory. Most memory bugs in production are write-path bugs: the system stored the wrong thing, stored it twice, stored it without a date, or overwrote the right thing.
5.1 What to remember: salience
Storing everything makes retrieval noisy and costs money. Storing too little misses what matters. Common salience signals:
- Explicit requests: "remember that…", "from now on…", corrections ("no, I said Python 3.12"). Highest precision; always store.
- LLM-judged importance: a 1–10 score as in Generative Agents, or a classifier prompt ("is this a durable fact about the user, a preference, a commitment, or ephemeral?"). Ephemeral content ("what's 15% of 80?") is skipped.
- Category whitelists: define the schema of what your product cares about (identity, preferences, constraints, goals, relationships, commitments/deadlines, past decisions). Extraction against a schema is more consistent than open-ended extraction. Claude Code's auto memory, for example, uses a small fixed set of types and explicitly skips what can be derived from the codebase Claude Code docs.
- Repetition and surprise: something mentioned across several sessions, or that contradicts an existing memory, deserves attention.
- Outcome signals (for procedural/episodic memory): store trajectories that succeeded (tests passed, user accepted) and lessons from ones that failed.
- Negative rules: don't store secrets, credentials, special-category data (health, religion, etc.) unless the product requires it and the user consented; don't store content from untrusted sources (web pages, emails) as if it were the user's own statement.
5.2 Shape of a memory record
Whatever the backend, a good memory record looks roughly like this:
Every field earns its place: scope keys for isolation; type for routing and schema; two time axes for temporal reasoning; provenance (which message produced this) for debugging, user trust ("why do you think that?"), poisoning forensics and cascading deletion; access stats for decay; sensitivity for policy.
5.3 Dedup and conflict resolution
New facts often duplicate or contradict old ones. The standard procedure: retrieve the top-\(s\) most similar existing memories for the same scope and entity, then classify the relationship.
| Relation | Example (old → new) | Typical action |
|---|---|---|
| Duplicate | "Likes hiking" → "Enjoys hiking" | NOOP; bump access/confidence |
| Refinement | "Has a dog" → "Has a beagle named Max" | UPDATE (merge into richer fact), keep provenance of both |
| Contradiction (change over time) | "Lives in London" → "Moved to Berlin last month" | Invalidate old (set valid_to), add new. Don't hard-delete: "where did I live before?" is a valid question |
| Contradiction (error) | "Allergic to peanuts" → user: "No, it's shellfish, not peanuts" | Supersede and mark old as incorrect (not "past truth") |
| Context-dependent | "Prefers Python" vs "Uses Go at work" | Keep both, qualify with context |
| Unrelated | - | ADD |
Precedence rules worth stating in an interview: explicit user statements beat inferred ones; newer beats older for changeable state; first-party statements beat third-party content; higher-confidence sources beat lower. When in doubt, keep both with timestamps and let the reader see the history, or ask the user.
Running reconciliation without a scope or entity filter. "Similar" memories from another user, or about a different person ("my sister is vegan" vs "I am vegan"), will get merged. Speaker attribution is a classic extraction error: facts about third parties get stored as facts about the user. Include "who is this fact about?" as an extracted field.
5.4 Hot path vs background writes
Hot path (in-turn, often agent-initiated)
The agent calls save_memory / edits a block / writes a notes file mid-turn. Memory is visible immediately, even later in the same session, and the user can see "Memory updated". Cost: added latency, tokens, and a distracted agent. Mostly suited to explicit "remember this" and task notes.
Background (async worker)
After the turn (or at session end, or on a schedule) a worker processes the transcript: extraction, dedup, graph updates, reflection. No user-facing latency; can use a cheaper model and batch. Cost: lag, plus another system to operate (queue, retries, idempotency). This is where consolidation ("sleep-time compute", "dreaming") lives Lin+ 2025.
A common hybrid: hot path for explicit memories and session notes; a background job for implicit extraction and consolidation. Make background writes idempotent (key them on message ID) so retries don't create duplicates.
- LangGraph memory concepts: profiles vs collections, hot path vs background.
- A-MEM (Xu+ 2025): Zettelkasten-style notes where new memories link to and update related old ones.
- Claude Code memory docs: a concrete, typed, file-based write policy you can study end to end.
6. Designing the read path
6.1 When to retrieve
| Strategy | How | Pros | Cons |
|---|---|---|---|
| Always-in-context | Profile / core memory blocks / MEMORY.md index injected every turn | Never "forgets to look"; zero retrieval latency; prompt-cacheable if stable | Token cost every call; must be small (hundreds to a few thousand tokens); stale if not maintained |
| Retrieve every turn | Harness runs a memory search on each user message before calling the LLM | Simple, predictable, no agent judgement needed | Adds latency to every turn (embedding + search, typically tens of ms; more with rerank); injects noise when nothing is relevant |
| On demand (tool call) | Agent calls search_memory(query) when it decides it needs to | Only pays when needed; agent can reformulate and iterate | Agent may not realise it lacks information; extra LLM round-trip |
| Event-triggered | Retrieve on session start, entity mention, or task type | Targeted | Needs routing logic |
The usual production answer is a combination: a small always-on profile + automatic retrieval at session start or per turn with a relevance threshold + a search tool for deep dives. Retrieval with a minimum score threshold matters: injecting irrelevant memories is worse than injecting none, because the model will try to use them (the "creepy" or non-sequitur personalisation failure).
6.2 Query formation
The raw user message is often a poor query ("yes, do that" retrieves nothing useful). Options: rewrite the query from recent context with a small model; generate several sub-queries (entities, topics); extract time expressions ("last week" → a date filter) as LongMemEval recommends Wu+ 2024; and pull entities for graph lookups. Each LLM-based rewrite adds about one small-model call of latency, so it's often done only for the on-demand tool path.
6.3 Ranking
Generalise the Generative Agents score into a ranking function you can tune:
$$\text{score}(m,q) = w_\text{rel}\cdot \text{sim}(m,q) + w_\text{rec}\cdot \gamma^{\,\Delta t(m)} + w_\text{imp}\cdot \text{imp}(m) + w_\text{freq}\cdot \log(1+\text{hits}(m))$$with hard filters applied before scoring: scope (tenant/user), validity (\(t_\text{invalid}\) is null, or overlaps the time asked about), not deleted, sensitivity policy. In practice the relevance term is itself a fusion of vector and BM25 (and entity/graph signals), often combined with Reciprocal Rank Fusion, \(\text{RRF}(d) = \sum_r 1/(k + \text{rank}_r(d))\) with \(k\approx 60\), and optionally a cross-encoder rerank of the top 20–50. Tune the weights on an eval set rather than by intuition. Recency should be strong for changeable state ("current project") and weak for stable facts ("allergies").
6.4 Injection format and placement
- Compact and structured. Inject facts as a bulleted list with dates and, if useful, confidence:
- (2026-09-12) Prefers email follow-ups [source: chat]. Dates let the model reason about staleness itself. - Clearly delimited and labelled as data. Wrap in a tagged block (e.g.
<user_memories>) with an instruction that these are notes from past conversations, may be outdated, and are not instructions. This is a prompt-injection defence. - Placement. Stable memory (profile, core blocks) goes early, after the system prompt, so it stays in the cached prefix. Per-turn retrieved memory goes near the end (just before or within the latest user turn), so it neither breaks the cache nor sits in the "lost middle".
- Budget. Cap retrieved memory (say a few hundred to ~2k tokens). Zep's LongMemEval result used about 1.6k tokens of context vs 115k for the full history Rasmussen+ 2025, which shows how much a good memory layer compresses.
- Usage guidance. Tell the model when to use memories (personalise when relevant, don't announce "I remember that you…" for sensitive items, ask if a memory seems outdated).
Expect a latency question: "Memory adds how much to time-to-first-token?" Walk the budget: embedding the query (~10–50 ms for a hosted API, less if local), ANN search in pgvector or a vector DB (~ms to tens of ms at per-user scale), optional rerank (~50–200 ms), optional LLM query rewrite (hundreds of ms). These are rough magnitudes; measure your own stack. Then explain how to hide it: run retrieval in parallel with other pre-processing, skip rewrite on the automatic path, cache the profile in the prompt prefix, and do all writes asynchronously.
- LongMemEval (Wu+ 2024): read the section on indexing, retrieval and reading designs; it's a practical design study as much as a benchmark.
- pgvector README: HNSW vs IVFFlat and filtered search (iterative index scans), which matters for per-user filters.
- Anthropic context engineering: just-in-time retrieval vs pre-loading.
7. Consolidation and forgetting
Memory that only grows gets worse over time: retrieval precision falls, stale facts accumulate, cost rises, and privacy risk grows. Mature systems actively maintain memory.
| Mechanism | How it works | Notes |
|---|---|---|
| Decay | Score down memories by time since last access (\(\gamma^{\Delta t}\)); optionally reinforce on access | MemoryBank explicitly borrows the Ebbinghaus forgetting curve to forget and reinforce by elapsed time and significance Zhong+ 2023. Usually a ranking effect, not deletion. |
| Merging / dedup | Periodic job clusters near-duplicate memories and merges them | Keep provenance of all merged sources so deletion still cascades |
| Summarisation hierarchies | Episodes → session summaries → weekly/topic summaries → profile | Reflection trees (Generative Agents), Zep's community summaries. Read at the coarsest level that answers the question, drill down when needed |
| Invalidation | Mark facts no longer true (valid_to) | Preserves history; excluded from "current" queries |
| TTLs | Hard expiry for inherently temporary data | "I'm in Tokyo this week" → expire in 7 days. Mem0's API supports an expiration date per memory Mem0 docs. Anthropic's memory-tool docs suggest deleting long-unaccessed files Anthropic docs |
| Size caps | Per-block character limits, per-user memory quotas | Letta blocks have a size limit Letta docs; Claude Code caps the loaded MEMORY.md index and asks the model to rewrite it when it's over the limit Claude Code docs. Forces prioritisation |
| User-requested deletion | "Forget that I…"; delete all memory | Must be hard deletion across all derived artifacts (see §9) |
Consolidation is where "sleep-time compute" fits: use idle time to pre-process context so that query-time work is cheaper. Letta's paper reported that pre-computing over a known context cut the test-time compute needed for the same accuracy by roughly 5× on their stateful reasoning benchmarks Lin+ 2025. Product memory systems have adopted the same idea with background consolidation passes.
Confusing "forgetting for relevance" with "deleting for compliance". Decay and invalidation are soft and reversible; they make retrieval better. A user's deletion request or a retention policy needs hard, verifiable deletion, including copies inside summaries, reflections, graph edges, embeddings, caches, backups (on their own schedule) and any fine-tuning data. Design provenance links up front so you can find every derived artifact.
- MemoryBank (Zhong+ 2023): forgetting-curve-based memory updating.
- Sleep-time Compute (Lin+ 2025): the case for doing memory work offline.
8. Memory in multi-agent systems
With several agents (orchestrator + workers, or peers), memory becomes a coordination problem. See the agent architectures page for the topologies; the memory-specific choices are:
| Pattern | How | Good for | Risks |
|---|---|---|---|
| Private memory per agent | Each agent has its own context and store; communicates only by messages | Isolation, clean contexts, least privilege | Duplicated work; information lost in handoffs |
| Shared blackboard | A shared store (document, DB table, shared memory block, files) all agents read and write | Coordination, shared plan and findings | Write conflicts; one agent's error or injected content pollutes everyone; access control harder |
| Artifacts by reference | Workers write outputs to storage and pass back a pointer + short summary | Avoiding the "game of telephone" through the orchestrator | Need a consistent artifact store and naming |
| Handoff context packet | When transferring control, pass a structured summary: goal, state, decisions, open questions, relevant IDs | Agent-to-agent or agent-to-human transfer | Same lossy-summary issues as compaction |
Anthropic's multi-agent research system illustrates several of these: the lead agent saves its plan to memory because its context may be truncated past 200k tokens, and sub-agents can write outputs directly to external storage rather than routing everything through the lead, to avoid information loss Anthropic 2025. The same post notes multi-agent runs used roughly 15× the tokens of a chat interaction, so shared memory that avoids re-discovery has real cost value. Letta's git-backed memory lets sub-agents write memory in parallel worktrees and merge, which is version control applied to the blackboard problem Letta 2026.
A good answer separates task-scoped shared state (the plan, findings so far; lives as long as the job; blackboard or artifacts) from long-term memory (persists across jobs; written through the same governed write path as single-agent memory). Also mention least privilege: a web-browsing worker that reads untrusted pages should not have write access to org-wide memory.
9. Privacy, security and governance
9.1 Privacy and compliance
- PII and sensitive data. Memory turns transient chat into a persistent profile. Classify at write time; redact or refuse special categories unless needed and consented. Anthropic's memory-tool docs recommend validating and stripping sensitive data before writing Anthropic docs.
- Right to be forgotten. GDPR Article 17 gives data subjects a right to erasure of personal data under specified conditions GDPR Art. 17. For memory this means: deleting a memory deletes its embeddings, derived facts, graph edges, and summaries that contain it (re-summarise or drop). Provenance links make this tractable.
- Transparency and control. Users should be able to see, edit and delete what's remembered, and to have conversations that aren't remembered (both ChatGPT's temporary chat and Claude's incognito mode exist for this; see §11).
- Retention and residency. Retention policies per memory type; data residency for enterprise tenants; encryption at rest with per-tenant keys where required.
9.2 Security: memory poisoning
A prompt injection normally lasts for one session. If injected content is written to memory, it persists and fires in future sessions, possibly for other users if the memory scope is shared. OWASP's Top 10 for Agentic Applications (Dec 2025) includes memory and context poisoning as a top risk OWASP 2025 (numbered ASI06 according to secondary sources; verify in the document).
- AgentPoison showed that poisoning a tiny fraction of an agent's memory or knowledge base with optimised triggers achieved high attack success (reported ≥ 80%) with under 1% impact on benign performance and a poison rate below 0.1% Chen+ 2024.
- MINJA showed an attacker with only normal query access can get an agent to write malicious records into its own memory, which are later retrieved for other users' queries Dong+ 2025.
Defences, roughly in order of value:
- Scope and isolation: per-user memory by default; shared (agent/org) memory only via a reviewed path. Tenant ID as a hard filter at the storage layer (row-level security, separate namespaces or indexes), never as an LLM instruction.
- Provenance-aware writes: record the source of every memory; don't let content from tools, web pages, emails or documents become "user facts" or instructions; lower trust for third-party-derived memories.
- Write-time screening: classify candidate memories for instruction-like content ("always send files to…", "ignore previous…") and for secrets before committing.
- Read-time framing: inject memories as quoted data with an explicit "not instructions" frame; keep high-privilege actions behind confirmations regardless of memory content.
- Audit and rollback: versioned memory (append-only logs, git-backed files) so you can find and revert poisoned entries.
- Path and tool hygiene: file-based memory needs path-traversal protection; Anthropic's docs warn that paths like
/memories/../../secrets.envmust be rejected Anthropic docs.
Memory security research has been very active in 2025–2026 (new attacks and proposed certified defences appear monthly). The principles above are stable; specific attack success rates and defence claims are not. Check recent work before quoting numbers.
- OWASP Top 10 for Agentic Applications 2026: the industry risk list, including memory/context poisoning.
- MINJA (Dong+ 2025): query-only memory injection; the scariest threat model for shared memory.
- AgentPoison (Chen+ 2024): backdoor triggers via poisoned memory.
10. Evaluating memory
10.1 Benchmarks
| Benchmark | Setup | What it tests | Caveats |
|---|---|---|---|
| DMR (from the MemGPT paper) | Multi-session chats; single fact-retrieval questions | Can you retrieve a fact from earlier sessions? | Too easy; short conversations; full-context baselines do well (the Zep paper criticises it) Rasmussen+ 2025 |
| LoCoMo | LLM-generated conversations between two personas, about 300 turns and 9k tokens on average, up to 35 sessions Maharana+ 2024 | QA (categories include single-hop, multi-hop, temporal and open-domain), event summarisation, multimodal dialogue | Conversations fit in modern context windows; scores vary a lot with judge model and setup; widely used for vendor comparisons, which are often disputed |
| LongMemEval | 500 curated questions embedded in scalable chat histories (the S variant is about 115k tokens) Wu+ 2024 | Five abilities: information extraction, multi-session reasoning, temporal reasoning, knowledge updates, abstention | Commercial assistants and long-context LLMs showed about a 30% accuracy drop in the paper. Single-session categories are near saturation for strong systems by 2026 |
| BEAM | 100 conversations from 128k up to 10M tokens, 2,000 validated questions Tavakoli+ 2025 | Ten abilities incl. contradiction resolution, event ordering, instruction and preference following, summarisation | New (ICLR 2026); designed so long context alone can't solve it. Too new for a settled leaderboard |
Memory benchmark scores moved a lot in 2025–2026 and many published numbers are vendor-reported with differing judges, prompts and answer models. LoCoMo in particular has seen public disputes between vendors about methodology. Treat leaderboard claims as marketing until you've reproduced them. The categories (knowledge update, temporal, multi-session, abstention) are the durable part.
10.2 Metrics that matter in production
- End-task accuracy on memory-dependent questions, usually LLM-as-judge against a reference answer (calibrate the judge against humans; see the production evals page).
- Retrieval metrics if you have labelled evidence: recall@k of the supporting memory, precision of injected memories.
- Write quality: extraction precision/recall vs a human-labelled set of "facts that should have been stored"; contradiction rate; duplicate rate; wrong-speaker attribution rate. Benchmarks mostly ignore this; you need your own.
- Abstention / hallucinated memory: how often the agent "remembers" something never said. Users find this worse than forgetting.
- Staleness: rate of answers based on superseded facts (knowledge-update category).
- Cost and latency: tokens per query (injected memory + write-path LLM calls amortised per turn), p50/p95 added latency, storage per user.
- Product signals: users correcting the assistant ("I told you…"), memory edits/deletes, opt-out rate, "creepiness" reports.
10.3 The accuracy–latency–token trade-off
Every memory system sits on a three-way frontier. Representative published points (each from the system's own paper, so not directly comparable):
| Setup | Context tokens per query | Latency | Accuracy note |
|---|---|---|---|
| Full history (LoCoMo, per Mem0 paper) | ~26k | p95 ~17 s | Strong but slow and costly Chhikara+ 2025 |
| Mem0 extracted facts (LoCoMo) | ~7k memory footprint | p95 ~1.4 s | Close to or above full context on their judge metric Chhikara+ 2025 |
| Full history (LongMemEval-S, per Zep paper) | ~115k | ~29 s (gpt-4o) | 60.2% with gpt-4o Rasmussen+ 2025 |
| Zep graph memory (LongMemEval-S) | ~1.6k | ~2.6 s (gpt-4o) | 71.2% with gpt-4o Rasmussen+ 2025 |
These numbers are from 2025 papers with 2024-era models; current long-context models and newer memory systems will score differently. The robust lesson: a good memory layer cuts injected context by one to two orders of magnitude and latency by about an order of magnitude while matching or beating full context on accuracy, because it removes distractors. What you pay instead is write-path cost (LLM calls per turn) and complexity.
"How would you evaluate the memory feature you just designed?" A strong answer: (1) build a golden set from real (consented, anonymised) multi-session histories with questions per category: recall, update, temporal, multi-session, abstention; (2) measure write quality separately from read quality, so you know which half failed; (3) track tokens and p95 latency alongside accuracy; (4) run ablations (no memory, full context, memory) to prove the feature earns its cost; (5) monitor online with correction rate and memory edits; (6) include adversarial cases (injection attempts, third-party facts, sensitive data).
- LongMemEval (Wu+ 2024): the most useful single benchmark paper for designers.
- LoCoMo (Maharana+ 2024): understand it because everyone quotes it.
- BEAM (Tavakoli+ 2025): memory evaluation beyond a million tokens.
11. Product examples
Consumer memory features change frequently. The descriptions below are based on vendor announcements and docs available as of late 2026 and describe behaviour only to the extent publicly documented. Internal architectures are generally not public; don't claim to know them in an interview.
| Product | What's publicly described | Design lessons |
|---|---|---|
| ChatGPT memory | Launched as "saved memories" (explicit facts the user can view and delete) in 2024; in April 2025 expanded to also reference past chat history. Users can turn either off, and there is a temporary-chat mode intended to stay out of memory (verify current behaviour). In June 2026 OpenAI announced "Dreaming", described in coverage as a background process that consolidates and rewrites memory to address staleness, correctness and scale OpenAI 2026. (Details here are from secondary coverage; I couldn't load the original post, so verify specifics.) | Explicit + implicit memory as two separately controllable layers; background consolidation as the scaling answer; an "off the record" mode |
| Claude memory (claude.ai apps) | Announced Sept 2025 for Team/Enterprise, extended to Pro and Max in Oct 2025. Memory is project-scoped (a separate memory per project), users can view and edit a memory summary, and incognito chats aren't saved to memory Anthropic 2025 | Scope as a privacy boundary (work projects don't bleed into each other); an editable, human-readable summary for transparency |
| Claude memory tool (API) | A client-side tool where Claude issues file operations (view, create, str_replace, insert, delete, rename) under /memories and your application executes them against storage you control; pairs with context editing and compaction Anthropic docs | Model-driven, file-shaped memory with the developer owning storage, isolation and security |
| Claude Code | CLAUDE.md instruction files at org/user/project/local scope, plus agent-written auto memory: a MEMORY.md index loaded each session and topic files read on demand Claude Code docs | Procedural memory as files; index-in-context, content-on-demand; size caps force curation |
| Memory frameworks | Mem0 (extracted facts, hybrid retrieval), Zep/Graphiti (temporal KG), Letta (tiered memory, now git-backed files), LangMem / LangGraph store (namespaced semantic/episodic/procedural) | Use as references for patterns; evaluate on your own data before adopting |
12. A reference architecture to draw in an interview
Here is a design you can sketch in five minutes for "add long-term memory to our assistant / support agent", and then adapt.
| Store | Holds | Why this choice | Alternatives |
|---|---|---|---|
| Postgres + pgvector | Facts, episodes, profiles, embeddings, provenance, validity intervals | One transactional store for metadata + vectors; SQL filters for scope/time; row-level security for tenant isolation; HNSW or IVFFlat indexes and iterative scans for filtered search pgvector | Dedicated vector DB (at very large scale or for advanced hybrid search); OpenSearch for BM25 + vectors |
| Graph DB (optional) | Entities, relations, temporal edges, communities | Multi-hop and temporal relational queries | Neo4j, FalkorDB, Neptune (Graphiti supports these); or adjacency tables in Postgres to start |
| Redis / KV | Session buffers, running summary, scratchpad, locks | Low-latency hot state with TTLs | DynamoDB; or the agent framework's checkpointer |
| Object store | Raw transcripts (source of truth), notes and skill files, exports | Cheap and durable; lets you re-run extraction when your extractor improves | Git repo for file-based memory (versioned, mergeable) |
| Queue + workers | Async write path and consolidation | Keeps LLM extraction off the hot path; retries; idempotency by message ID | Durable workflow engine (Temporal etc.) for multi-step consolidation |
Back-of-envelope sizing (say it out loud, flag assumptions): 1M users × ~500 memories each = 5×108 facts. At 1,536-dim float32 embeddings that's about 6 KB per vector, so roughly 3 TB of raw vectors before index overhead. That pushes you toward smaller embeddings (e.g. 512–768 dims), halfvec or quantisation, partitioning by tenant, or a dedicated vector store. Per-query search, however, is per user (a few hundred rows), so with a good user_id filter or partition it can even be brute force. Write-path cost: if extraction is one small-model call of about 1–2k tokens per turn, at 10 turns per user-day that's 10–20k tokens per active user per day, which is often more than the read path, so batch per session rather than per turn when lag is acceptable.
The talking points that separate senior answers: (1) the memory service owns scope enforcement; the LLM never chooses tenant or user IDs; (2) raw transcripts are the source of truth and memories are a derived, rebuildable index, so you can re-extract when you improve the extractor and you can cascade deletions; (3) async writes, sync reads, with a latency budget for the read path; (4) bi-temporal fields even without a graph DB; (5) evaluation and observability: log which memories were injected into each prompt so you can debug "why did it say that?"; (6) a user-facing memory UI (view, edit, delete, incognito).
- Graphiti repository: a production-grade open-source temporal graph memory you can read end to end.
- pgvector: the default "start here" vector store for memory on Postgres.
- Anthropic memory tool docs: a clean spec for model-driven file memory, including security guidance.
Interview question bank
1. LLMs are stateless. What exactly does "memory" mean for an LLM agent?
The model only conditions on the tokens in the current request, so "memory" is whatever the harness decides to put back into the prompt. Short-term or working memory is the content of the context window during a task or conversation, managed by buffering, truncation, summarisation and compaction. Long-term memory is information stored outside the model (databases, files, graphs) and selectively retrieved into future prompts. There's also parametric memory (knowledge in weights, changed by fine-tuning), but product memory is almost always external and textual because it must be inspectable, editable, scoped and deletable. So designing memory means designing a write path (what to store), storage, a read path (what to inject, when, where) and a forgetting policy.
2. Explain episodic, semantic and procedural memory with an agent example of each.
Episodic memory records specific experiences with time and context: "On 12 Sept the user and I debugged a failing migration; the cause was a missing index." Semantic memory holds distilled facts that aren't tied to one episode: "The user's production DB is Postgres 16 on RDS." Procedural memory holds how to act: a skill file for "how to run this repo's migrations safely", or an updated instruction like "always run the linter before committing". In CoALA terms these are long-term memories that the agent reads into working memory by retrieval and updates by learning. In practice, consolidation turns episodes into semantic facts and into procedural lessons, and good systems keep provenance links from facts back to episodes.
3. Why not just use a model with a 1M-token context and resend everything?
Cost: every token is prefilled on every call, so a long history multiplies cost and time-to-first-token; caching reduces price but not the window limit or all latency. Quality: models use the middle of long contexts poorly and degrade as input grows, especially with distractors, so a curated 2k-token memory can beat 100k tokens of raw history. Published examples: Zep reported about 1.6k tokens of context beating about 115k tokens of full history on LongMemEval with gpt-4o, at about a tenth of the latency. Scale: multi-year histories still don't fit. Semantics: raw history doesn't resolve contradictions, doesn't support deletion or scoping, and doesn't distil lessons. Long context is still useful: it lets you inject more retrieved memory and delay compaction.
4. Walk me through Mem0's original write path. What are its weaknesses?
For each new exchange, an LLM extracts candidate facts using a conversation summary plus the last ~10 messages as context. For each candidate, the system retrieves the top ~10 similar existing memories and asks the LLM, via a tool call, to choose ADD, UPDATE, DELETE or NOOP. That keeps the store compact and current. Weaknesses: two LLM calls per write (cost, latency); destructive operations, so an extraction or judgement error deletes a correct fact and you lose history; contradiction handling without time loses "what was true before"; and reconciliation quality depends on retrieving the right neighbours. Notably, Mem0 itself moved in 2026 to single-pass ADD-only extraction with temporal metadata and hybrid retrieval, resolving currency at read time instead, which trades a bigger store for non-destructive, cheaper writes.
5. What is bi-temporal modelling and why does it matter for agent memory?
Each fact carries two time axes: valid time (when it was true in the world: valid_from, valid_to) and transaction time (when the system recorded and retired it: created_at, expired_at). Zep's Graphiti stores both on graph edges. When a contradicting fact arrives, the old edge is invalidated (valid_to set to the new fact's valid_from) instead of deleted. This lets you answer "where does Alice work now?", "where did she work in March?", and "what did the system believe last Tuesday?" (useful for debugging and audits). It also handles late-arriving information: the user says today that they changed jobs a month ago, so transaction time is today but valid time is a month back. You can implement this with two pairs of timestamp columns in Postgres; you don't need a graph DB.
6. Describe MemGPT's architecture. What is "self-editing memory"?
MemGPT treats the context window like RAM and external stores like disk. Main context contains read-only system instructions, a fixed-size editable working context (later "core memory blocks" like persona and human), and a FIFO message queue whose evicted messages are folded into a recursive summary. External context has recall storage (all past messages, searchable) and archival storage (an arbitrary vector-indexed text store). The LLM manages memory itself through function calls: appending to or replacing text in core memory, inserting into and searching archival storage, searching conversation history. That's self-editing memory. A queue manager warns the model at about 70% context usage so it can save important things before a flush, and a heartbeat flag lets it chain several memory operations before replying. The trade-off: memory quality depends on the model's judgement and costs extra steps.
7. Write down the Generative Agents retrieval score and explain each term.
score = α_rec·recency + α_imp·importance + α_rel·relevance, with each component min-max normalised to [0,1] over candidates and all α = 1 in the paper. Recency is exponential decay γ^Δt with γ = 0.995 per game-hour since the memory was last accessed, so using a memory refreshes it. Importance is an LLM-assigned 1–10 poignancy score set at write time. Relevance is cosine similarity between memory and query embeddings. With γ = 0.995 the half-life is ln2/0.005 ≈ 138 hours, so recency falls to about 0.43 after a week. In a product you'd tune the weights and γ on an eval set, use wall-clock time, and add hard filters (scope, validity) before scoring.
8. What is reflection in Generative Agents and what's the modern equivalent?
When the summed importance of recent observations passes a threshold (150 in the paper, about 2–3 times per game day), the agent takes its 100 most recent memories, asks for the 3 most salient high-level questions, retrieves memories for each, and generates 5 insights with citations to the evidence. The insights are stored as new memories and can be reflected on again, forming a tree of abstractions. The modern equivalent is background consolidation: "sleep-time" or "dreaming" jobs that periodically turn episodic logs into semantic facts, profile updates and procedural lessons. Risks: reflections can be wrong and self-reinforcing, so keep citations to evidence, give them a lower confidence than direct statements, and let them be invalidated.
9. Your agent's compaction keeps losing important details. How do you fix it?
First find out what's lost by diffing pre- and post-compaction states on failing traces. Typical fixes: (1) a structured compaction prompt with required sections (goal and acceptance criteria, decisions and rationale, open tasks, exact identifiers such as paths, IDs, URLs and error strings, user constraints, dead ends); (2) keep the most recent turns verbatim and summarise only older ones; (3) move durable state out of the conversation into a notes/progress file or memory tool that the agent updates as it goes, so compaction doesn't have to carry it; (4) clear stale tool outputs first, which delays compaction; (5) keep the raw transcript retrievable so the agent can search it; (6) compact less often but at higher fidelity, since each round compounds loss and invalidates the prompt cache. Add an eval: continue tasks from compacted state and measure success vs uncompacted.
10. Hot-path vs background memory writes: when would you choose each?
Hot path, where the agent writes during the turn, is right for explicit requests ("remember that…"), for task notes needed later in the same session, and when the user benefits from seeing "memory updated". It costs latency, tokens and agent attention. Background writes, where an async worker processes transcripts after the turn or session, suit implicit extraction, dedup, graph building and consolidation: no user-facing latency, can batch and use cheaper models, but memory lags. Most production systems do both: explicit memories in the hot path, everything else in the background, with idempotency keyed on message ID and a short lag SLA (seconds to minutes) so the next session sees the update.
11. How do you handle the user saying "I moved to Berlin" when memory says "lives in London"?
Detect the conflict at write time: retrieve similar memories for the same user and entity, and have the extractor classify the relation as a change over time (not an error, not a duplicate). Then invalidate the London fact by setting valid_to to the move date (or now if unknown), and add the Berlin fact with valid_from, both with provenance to the message. On read, filter to currently valid facts for "where do I live?", but keep the history for "where did I live before?". If the system is append-only, make sure ranking strongly prefers newer facts for changeable attributes and that injected memories show dates so the model can reason. If it's ambiguous (a trip vs a move), store with lower confidence or a TTL, or ask.
12. Always-inject vs retrieve-on-demand: how do you decide what goes where?
Always-inject a small, high-value, slowly changing set: the user profile (name, role, key preferences, hard constraints like allergies or "never contact by phone"), agent persona, and an index of what else exists (like MEMORY.md or skill descriptions). It must be small (hundreds to low thousands of tokens) and stable so it stays in the prompt-cache prefix. Retrieve everything else: episodes, long-tail facts, documents, using automatic retrieval with a relevance threshold at session start or per turn, plus a search tool for deeper lookups. The rule of thumb: if forgetting it causes a serious failure and it's small, pin it; if it's large or only sometimes relevant, retrieve it.
13. Back-of-envelope: what does a memory layer cost per active user per day?
State the assumptions. Say 10 turns per day. Write path: one extraction call per turn with ~1.5k input tokens and ~150 output tokens on a small model gives ~15k input + 1.5k output tokens per day; reconciliation might double it, giving ~30k tokens/day, plus embeddings (negligible). Read path: ~1k injected memory tokens per turn gives ~10k extra input tokens per day on the main model, which often costs more per token than the small extraction model. Storage: a few hundred facts at ~6 KB per 1,536-dim float32 vector plus text is a few MB per user, which is cheap. Batching extraction per session instead of per turn can cut write cost several-fold. Compare against resending full history: by turn 50 of a conversation, a single request may carry tens of thousands of tokens.
14. How would you design tenant isolation for memory in a B2B agent platform?
Isolation must be enforced by infrastructure, not the prompt. Every memory record carries tenant_id (and user_id, agent_id, scope); the memory service derives these from the authenticated request, never from model output. In Postgres, use row-level security policies keyed on a session variable, or separate schemas or databases for high-sensitivity tenants; for vector stores, use per-tenant namespaces or indexes, or mandatory filters with tests that prove no cross-tenant results. Encryption at rest with per-tenant keys enables crypto-shredding on offboarding. Shared org-level memory goes through a reviewed write path. Add automated tests that attempt cross-tenant retrieval, log injected memory IDs per request for audit, and apply the same scoping to caches and background workers, which are a common leak point.
15. What is memory poisoning and how do you defend against it?
An attacker gets malicious content written into an agent's persistent memory, through injected instructions in a web page, email or document the agent reads, or just through crafted queries as in MINJA, so it influences future sessions, possibly other users' if memory is shared. AgentPoison showed tiny poison rates can achieve high attack success with little effect on normal behaviour. Defences: per-user scope by default and a reviewed path for shared memory; provenance on every memory, with content from tools or third parties never promoted to "user facts" or instructions; write-time classifiers for instruction-like content and secrets; read-time framing of memories as untrusted data; confirmations for high-impact actions regardless of memory; versioned, auditable memory with rollback; and red-team evals for persistence attacks.
16. A user asks you to "forget everything about my divorce". What has to happen technically?
Identify all affected memories: semantic search plus entity and topic matching over the user's memories, ideally with the user confirming the list. Then follow provenance links to delete derived artifacts: extracted facts, their embeddings, graph edges and entity summaries, reflections or session summaries that mention it (regenerate them without it, or delete them), cached profiles and prompt caches. Decide on the raw transcripts per policy and the user's request (delete or exclude the relevant sessions). Log the deletion event without the content. Backups expire on their own schedule, which should be documented. Add a "do not remember" rule so the topic isn't re-extracted from future chats. Verify by running retrieval probes afterwards. This is why provenance and raw-transcript-as-source-of-truth are designed in from day one.
17. Compare vector memory, extracted-fact memory and a temporal knowledge graph for a sales assistant.
Questions a sales assistant gets: "What did Acme's CTO say about pricing last call?" (episodic, temporal), "Who's the decision maker at Acme now?" (relational, changeable), "What are this rep's preferences?" (simple facts). Raw vector memory over call notes handles the first reasonably with time filters but fumbles "now" questions because old and new contacts both match. Extracted facts handle preferences cheaply but flatten relations and, if reconciled destructively, lose history. A temporal graph models people, companies, roles and deals with validity intervals and answers "who is the decision maker now" and "what changed since Q2" well, at the highest write cost and with entity-resolution risk. A pragmatic design: episodic store of call notes + entity/relationship tables with validity columns in Postgres, upgrading to a graph engine only if multi-hop queries dominate.
18. How do you evaluate a memory system before launch?
Build a golden dataset from realistic multi-session histories (synthetic plus consented real ones) with questions per category: single-session recall, multi-session aggregation, temporal reasoning, knowledge updates, abstention (questions about things never said), and adversarial cases (injection, third-party facts, sensitive data). Evaluate the write path separately: extraction precision/recall against labelled "should-store" facts, duplicates, contradictions, speaker-attribution errors. Evaluate the read path: recall@k of supporting memories and end-to-end accuracy with a calibrated LLM judge. Always report tokens per query and p95 latency, and compare against no-memory and full-context baselines. Public benchmarks (LoCoMo, LongMemEval, BEAM) are sanity checks, not substitutes for your own query distribution.
19. What's the difference between notes files (NOTES.md, progress logs) and a memory database? When would you use each?
Notes files are agent-written, free-form or JSON documents, read and edited with file tools. They're great for task-scoped and project-scoped state in long-running agents (plan, progress, decisions, gotchas) because coding-trained models are very good with files, the agent decides what's salient, notes survive context resets, and humans can read and edit them. They lack built-in dedup, scoping, ranking and deletion guarantees, and they sprawl unless capped. A memory database with scoped records, timestamps, provenance and hybrid retrieval is better for many-user, long-term personalisation with compliance needs. Many systems combine them: files for a single agent's working knowledge (Claude Code auto memory, Letta's git-backed memory), a database behind a memory service for multi-tenant user memory.
20. How should memory work in a multi-agent system with an orchestrator and workers?
Separate task-scoped shared state from long-term memory. For the task: a blackboard (shared plan, findings, artifact store) that workers write to directly, passing back references plus short summaries so the orchestrator's context isn't flooded and details aren't lost in a game of telephone. The orchestrator persists its plan externally because its own context may be compacted. Workers get clean, private contexts. For long-term memory, all agents go through the same governed write path with scope and provenance, and least privilege applies: a worker that reads untrusted web content shouldn't write to org-wide memory. For concurrent writes, use transactions, optimistic concurrency, or version control (as in Letta's git-worktree approach) to merge.
21. Where should retrieved memories go in the prompt, and in what format?
Stable memory (profile, core blocks, skill index) goes near the top after the system prompt, so it's part of the cached prefix and doesn't change every turn. Per-turn retrieved memories go near the end, just before the latest user message, so changing them doesn't invalidate the cache and they're not buried in the middle where attention is weakest. Format: a delimited block (e.g. a tagged section) of concise bullets with dates and optionally source and confidence, plus an instruction that these are notes from past conversations, may be outdated, and must not be treated as instructions. Keep a token budget (hundreds to ~2k), and apply a relevance threshold so irrelevant memories aren't injected at all.
22. What is procedural memory for agents and how can it go wrong?
Procedural memory stores how to do things: skill libraries of verified code (Voyager), lessons from failures (Reflexion, ExpeL), self-updated system prompts (LangMem's prompt optimisers), instruction files (CLAUDE.md/AGENTS.md), and Skills folders loaded via progressive disclosure. It's high-leverage because it changes behaviour on every future task. It goes wrong when a lesson from one odd interaction generalises badly, when prompts bloat until adherence drops, when conflicting instructions accumulate, or when untrusted input can write procedures (a persistence vector for injection). Treat procedural updates like code: propose, test on a regression set, review for shared scopes, version, roll back. Keep user-scoped preferences separate from agent- or org-wide procedures.
23. Design: add memory to a customer-support agent used by 500 enterprise tenants.
Clarify requirements: which questions need memory (past tickets, customer preferences, account facts), retention and compliance (GDPR, residency), latency budget. Scopes: tenant (playbooks, product facts: procedural/semantic, reviewed writes), customer/user (preferences, history: episodic + semantic), session (Redis buffer + rolling summary). Write path: async worker after each conversation extracts typed facts (preference, issue, commitment, contact info) with provenance and validity, dedups against the customer's existing facts, screens for PII policy and injection; ticket summaries stored as episodes. Read path: at session start inject a compact customer profile and open commitments; auto-retrieve similar past tickets with a threshold; give the agent a search tool. Storage: Postgres + pgvector with RLS by tenant, S3 transcripts as source of truth, a queue for writes. Governance: per-tenant retention, deletion API cascading by provenance, an agent-visible memory UI. Evaluation: golden sets per category, resolution rate and repeat-question rate online, p95 latency and token cost.