Track B · AI product engineering

Agent Memory Systems

An LLM forgets everything between calls. "Memory" is the set of engineering choices that decide what an agent carries forward: inside one long task (working memory and context management) and across sessions, users and months (long-term memory). Interviewers ask about it because it sits where product, data engineering, retrieval, security and cost all meet. A good answer shows you can design the write path, the read path, the storage and the forgetting policy, and that you know how to evaluate and secure the result, rather than just naming Mem0 or MemGPT.

TL;DR: the 8–12 things to be able to say out loud

  • Memory is context engineering over time. The model only "remembers" what is in the prompt on this call. Every memory system is a policy for deciding what goes into a finite context window, where it comes from, and what gets thrown away.
  • Two taxonomies. The cognitive one (working, episodic, semantic, procedural; formalised for agents by CoALA) tells you what kind of thing is stored. The engineering one (in-context vs external, hot path vs background, user/session/agent/org scope) tells you how it is built.
  • Short-term memory is a budgeting problem. Sliding windows, token-budgeted truncation, rolling summaries, compaction, tool-output clearing and agent-written notes files. Bigger windows don't remove the need: quality degrades with length ("context rot", "lost in the middle").
  • Six long-term patterns: raw vector memory, extracted-fact memory (Mem0-style), temporal knowledge graphs (Zep/Graphiti), OS-style tiered memory with self-editing (MemGPT/Letta), reflection-based memory (Generative Agents), and procedural memory (skill libraries, instruction files, Skills).
  • Generative Agents retrieval score = recency + importance + relevance, each min-max normalised, recency an exponential decay (factor 0.995 per game hour). Reflection fires when summed importance passes a threshold (150).
  • Write path: decide what is salient, extract it as atomic, timestamped facts with provenance, dedup, and resolve conflicts (overwrite, invalidate with validity intervals, or append and let retrieval rank). Each choice trades accuracy on "knowledge update" questions against audit history.
  • Read path: always-inject (profile / core blocks) vs on-demand retrieval (tool call). Rank by relevance × recency × importance, use hybrid search (vector + BM25 + entity/graph), and inject compactly in a clearly delimited, clearly untrusted block.
  • Forgetting is a feature: decay, TTLs, merging, summary hierarchies, and hard deletion for user requests and GDPR Art. 17, including derived artifacts (embeddings, summaries, graph edges).
  • Memory is an attack surface: injected memories persist across sessions (memory poisoning; AgentPoison, MINJA). Defend with write-time filtering, provenance, tenant isolation, and treating memory as data rather than instructions.
  • Evaluation: LoCoMo, LongMemEval (5 abilities incl. knowledge updates and abstention), and newer BEAM (up to 10M tokens). Always report accuracy with latency and tokens per query, and be suspicious of vendor leaderboards.
  • Reference architecture: Postgres + pgvector for facts and episodes, optional graph store for entities and relations, KV/Redis for session state, object storage for raw transcripts, an async consolidation worker, and a memory service API that enforces scope and deletion.

1. Why agents need memory

A transformer forward pass is a pure function of its input tokens. Nothing persists between API calls except what you send back in. A chat app that "remembers" earlier turns is just re-sending them. So memory in an LLM system is never a property of the model alone. It is a property of the harness around it: what you store, where, and what you put back into the prompt next time.

Four pressures drive the need for something more than "re-send everything":

Context limits and cost

Windows have grown to hundreds of thousands or around a million tokens for some models, but every token you resend is paid for on every call and adds prefill latency. A year of daily chats is far beyond any window. Even for an agent within one task, tool outputs (file reads, search results, logs) fill the window within tens of steps.

Quality degrades with length

Models use information at the start and end of the context better than in the middle Liu+ 2023, and accuracy falls as input grows even on simple tasks, especially when distractors are present Hong+ 2025. Curating a small, relevant context often beats dumping a large one.

Personalisation and continuity

Users expect the assistant to know their name, stack, preferences and the state of an ongoing project without repeating it. Long-term memory is mostly a product feature: it reduces friction and makes the agent feel like it knows you.

Learning from experience

Without weight updates, an agent can still improve by storing what worked (successful trajectories, lessons, reusable code or instructions) and recalling it next time. Reflexion Shinn+ 2023, ExpeL Zhao+ 2023 and Voyager Wang+ 2023 are the classic examples.

Intuition

Think of the context window as RAM and everything else as disk. The model can only compute over what is in RAM. A memory system is the virtual-memory manager: it decides what to page in, what to page out, what to compress, and what to delete. The MemGPT paper makes this analogy literal.

One more framing helps in interviews: memory vs RAG. Classic RAG retrieves from a mostly static corpus that someone else wrote (docs, wiki). Memory retrieves from a corpus the agent itself writes, continuously, from interactions. That changes almost everything about the problem: the corpus is tiny per user but highly dynamic; facts go stale and contradict each other; time matters; the write path (what to store, how to dedup, how to resolve conflicts) is as hard as the read path; and the data is personal, so deletion and isolation are hard requirements. Retrieval mechanics (embeddings, hybrid search, reranking) carry over from the retrieval page.

Interview angle

"Why not just use a 1M-token context?" is a common opener. A strong answer covers four points. Cost and latency: you pay to prefill every token on every turn; prompt caching helps but doesn't remove the cost. Quality: context rot and position effects mean more tokens can lower accuracy. Scale: multi-year, multi-session histories still don't fit. Structure: raw history doesn't resolve contradictions ("I moved to Berlin" after "I live in London"), doesn't give you deletion or scoping, and doesn't distill lessons. Long context and memory are complements: a big window lets you inject more retrieved memory and run longer before compacting.

Go deeper

2. Taxonomies of memory

2.1 The cognitive-science split (CoALA)

The most-cited framing comes from cognitive architectures. CoALA (Cognitive Architectures for Language Agents) describes an agent as an LLM plus modular memories plus a structured action space Sumers+ 2023. It distinguishes:

MemoryCoALA definition (paraphrased)Concrete agent exampleTypical storage
WorkingActive information for the current decision cycle: perceptual input, retrieved knowledge, goals, intermediate resultsThe prompt on this call: system prompt, recent turns, tool results, scratchpadThe context window itself, plus run state in your orchestrator
EpisodicPast experiences and trajectories"Last Tuesday the user asked me to refactor auth; we chose JWT"; logs of past tool-use runs; few-shot examples of past successesEvent log, transcript store, vector index over episodes
SemanticKnowledge about the world and the agent itself, updatable through learning"User is vegetarian", "Acme's fiscal year ends in March", "the repo uses pnpm"Fact table, profile document, knowledge graph
ProceduralHow to act: implicit in LLM weights plus explicit code and prompts implementing the agentSystem prompt, tool definitions, learned skills (code or instruction files), refined instructionsPrompt registry, skill library, files like CLAUDE.md / SKILL.md

CoALA also defines internal actions on memory: retrieval (long-term into working), reasoning (working into working), and learning (writing to any long-term memory, including updating procedural memory) Sumers+ 2023. That vocabulary maps directly onto the read path, the scratchpad, and the write path in an engineering design.

Common mistake

Treating "episodic vs semantic" as a storage decision. It is a decision about granularity and abstraction. The same storage (a Postgres table with an embedding column) can hold both. What differs is the write process: episodic memory records what happened, with time and context. Semantic memory is distilled from episodes into timeless-ish facts, and therefore needs conflict resolution. Most production systems keep both and link facts back to the episodes they came from (provenance).

2.2 The engineering split

In system design you'll usually be judged on the engineering axes, not the cognitive labels:

AxisOptionsTrade-off
LocationIn-context (always in the prompt: system prompt, profile, core memory blocks) vs external (stored outside, retrieved on demand)In-context is reliable and zero-latency to "recall" but costs tokens every call and is size-bounded. External scales but recall can fail silently.
When writtenHot path (agent writes memory during the turn, often via a tool) vs background (async worker processes the transcript after or between turns)Hot path: immediate availability, user can see "memory saved", but adds latency and distracts the agent. Background: no latency hit, better batch reasoning, but memory lags by seconds to hours. LangGraph's docs frame exactly this choice LangChain docs.
Who decidesExplicit ("remember that I…", user-curated) vs implicit (system infers from conversation)Explicit is precise and consensual but sparse. Implicit has high coverage but risks storing wrong, sensitive or creepy facts.
ScopeSession/thread, user, agent (what this agent learned across all users), org/tenant (shared team knowledge), sometimes projectWider scope means more sharing value and more leakage risk. Scope must be a first-class key in storage and in every query.
RepresentationRaw text chunks, extracted facts, profile document (one JSON per user), collection of small docs, graph, code, filesLangGraph contrasts a single profile (easy to inject, error-prone to update as it grows) with a collection (easy to add to, harder to dedup and update) LangChain docs.
FormToken-level (text you can read), parametric (in weights: fine-tuning, LoRA), latent (KV caches, hidden states)Almost all production memory is token-level because it's inspectable, editable and deletable. The newer survey literature uses this forms/functions/dynamics framing Hu+ 2025.
Agent memory Short-term / working Long-term (external) • Context window (the prompt) • Conversation buffer / window • Rolling summary / compaction • Scratchpad, todo, NOTES.md • Run state (orchestrator) Episodic what happened events, transcripts Semantic what is true facts, profile, graph Procedural how to act skills, instructions consolidate / distill → Orthogonal engineering axes (apply to every box above) Location: in-context ↔ external | Write timing: hot path ↔ background | Trigger: explicit ↔ implicit Scope: session · user · agent · org/tenant | Form: text (token) · weights (parametric) · KV/latent
Memory taxonomy: the cognitive split (top) says what is stored; the engineering axes (bottom) say how it is built. Episodes are distilled into semantic facts and procedural skills by consolidation.
Interview angle

If asked to "classify memory types", give the cognitive four in one sentence each, then pivot quickly to the engineering axes: "In practice the decisions that matter are scope, hot path vs background writes, and always-in-context vs retrieved." Then give an example mapping: a support agent's user profile is semantic, user-scoped, in-context; past tickets are episodic, user-scoped, retrieved; resolution playbooks are procedural, org-scoped, retrieved or loaded as skills.

Go deeper

3. Short-term and working memory management

Before you add any database, the first memory problem in every agent is keeping one long conversation or task inside the window without losing what matters. This is "context management" or "context engineering". Here are the techniques, from simplest to most sophisticated.

3.1 Buffers, windows and token budgets

3.2 Rolling summarisation and compaction

Rolling summary: keep the last \(k\) turns verbatim plus a running summary of everything older. When the buffer exceeds a threshold, summarise the oldest chunk and merge it into the summary. MemGPT does this with a "recursive summary" of evicted messages Packer+ 2023.

Compaction is the agent-era name for the same idea applied when the context nears the limit: replace older turns with a high-fidelity summary and continue Anthropic 2025. Model providers now offer it as an API feature. Anthropic's API, for example, can compact on demand or automatically at a token threshold, optionally keeping recent turns verbatim Anthropic docs.

The craft is in what the summary must keep. A generic "summarise this conversation" prompt loses exactly the details an agent needs next. A good compaction prompt asks for:

KeepWhyExample
The goal and acceptance criteriaWithout it the agent drifts"Migrate billing service to Postgres 16; all tests green; no downtime"
Decisions made, with the reasonPrevents re-litigating or contradicting earlier choices"Chose logical replication over pg_upgrade because of the downtime constraint"
Open tasks / next stepsThe continuation point"Remaining: update connection strings in 3 services; run load test"
Exact identifiersSummaries paraphrase; IDs must be verbatimFile paths, ticket IDs, URLs, function names, error messages, order numbers
User preferences and constraints stated in-sessionMost painful thing to forget"Don't touch the legacy/ folder"
Failures and dead endsStops the agent repeating them"Tried bumping driver to v5: breaks SSL; reverted"
Drop: raw tool outputs, superseded drafts, chit-chatHigh-token, low-valueThe 3,000-line log already analysed
Common mistake

Compaction is lossy and compounding. Each round summarises a summary, so errors and omissions accumulate ("summary drift"). Mitigations: keep identifiers and decisions in a structured section the summariser must copy forward verbatim; keep the original transcript in cold storage so the agent (or a human) can look things up; and pair compaction with an external notes file that holds durable state, so the summary doesn't have to. Also, compaction usually invalidates the prompt cache, so don't compact too often.

3.3 Scratchpads and agent-maintained notes files

The most effective pattern for long-horizon agents in 2025–26 is surprisingly low-tech: let the agent keep notes in files outside the context window and read them back when needed. Anthropic describes this as structured note-taking: an agent maintains a NOTES.md or to-do list to track progress across many tool calls and context resets Anthropic 2025. For multi-session coding, their long-running harness uses an initializer session that creates a progress file (claude-progress.txt), a JSON feature list with every feature initially marked failing, an init script, and git commits. Each later session starts by reading these and ends by updating them Anthropic 2025.

Why it works so well:

JSON is better than Markdown for state the agent must not casually rewrite (Anthropic found the model was less likely to inappropriately edit a JSON feature list) Anthropic 2025. Markdown is better for free-form notes.

3.4 Tool-output pruning and clearing

In agentic loops, tool results are usually the biggest consumer of context: a file read, a web page, a SQL result. Once the agent has acted on a result, the raw bytes are rarely needed again. Tool-result clearing replaces old results with a placeholder ("[result cleared; re-run tool if needed]") while keeping the call itself, so the agent knows what it did. Anthropic's API exposes this as a context-editing strategy with a trigger threshold (default 100k input tokens), a number of recent tool uses to keep (default 3), an exclude list of tools, and a minimum amount to clear per activation so the cache invalidation is worth it Anthropic docs. When combined with a memory tool, the model gets a warning before clearing so it can save anything important to memory first.

Related tactics: return references rather than payloads (file path + line range, a query, a URL) so the agent can re-fetch just-in-time Anthropic 2025; have tools truncate and paginate by default; and push heavy exploration into sub-agents with their own clean windows that return condensed summaries (on the order of 1–2k tokens) to the orchestrator.

TechniqueWhat it doesCostLosesUse when
Full bufferResend everythingQuadratic token growthNothing (until overflow)Short chats
Sliding windowLast \(k\) turnsBoundedThe goal, early constraintsCasual chat, low stakes
Budgeted + pinnedSections with token budgets; pin goalBoundedMiddle turnsDefault for most products
Rolling summarySummary + recent verbatimAn extra LLM call per rollDetail; drifts over many rollsLong chats
CompactionSummarise when near limitOne big LLM call; cache resetDetail; risk of dropping IDsLong agent runs
Tool-result clearingDrop stale tool outputsNear zero; cache resetRaw outputs (re-fetchable)Tool-heavy agents
Notes / todo filesAgent writes durable state externallySome tool callsWhatever the agent failed to writeMulti-hour or multi-session tasks
Sub-agentsIsolate exploration in separate windowsMore total tokensDetail behind the summaryBroad research / search
Interview angle

"Your coding agent fails after ~40 steps. Why, and what do you do?" Strong answers diagnose first (look at traces: is the context full of file reads? did compaction drop the plan? is the agent re-doing work?) and then layer the fixes: clear stale tool results, return references not payloads, keep a todo/progress file the agent re-reads, compact with a structured prompt that preserves decisions and IDs, and move exploration to sub-agents. Mention the prompt-cache cost of each intervention: edits near the start of the context invalidate the cache for everything after them.

Go deeper

4. Long-term memory architectures

Six recurring designs. Real systems mix them: Letta combines tiered memory with files, Zep combines episodes with a graph, Mem0 combines fact extraction with hybrid retrieval. Learn each as a mechanism with a characteristic failure mode.

4.1 Vector-store memory (embed and retrieve past interactions)

The baseline. Chunk past conversations (per turn, per exchange, or per session), embed each chunk, store it with metadata (user_id, session_id, timestamp), and at query time embed the current message, retrieve the top-\(k\) similar chunks, and inject them. It is RAG where the corpus is your own history.

LongMemEval's authors found three cheap upgrades to this baseline help: decompose sessions into finer-grained units, expand the index keys with extracted facts (index a chunk under the facts it contains, not just its raw text), and use time-aware query expansion for temporal questions Wu+ 2024. Adding BM25 alongside vectors (hybrid search) also helps a lot for names, IDs and rare terms.

4.2 Extracted-fact memory (Mem0-style)

Instead of storing raw text, an LLM extracts salient facts from each new exchange and stores those as small, atomic memories ("User is allergic to peanuts"). The original Mem0 paper describes a two-phase pipeline Chhikara+ 2025:

  1. Extraction: given the new message pair plus context (a conversation summary and the last \(m=10\) messages), the LLM proposes candidate facts.
  2. Update: for each candidate, retrieve the top \(s=10\) most similar existing memories, and ask the LLM (via a tool call) to choose one operation: ADD (new fact), UPDATE (augment or refine an existing one), DELETE (the new fact contradicts an old one), or NOOP (already known).
new turn ─► [LLM extract] ─► candidate facts f₁..fₙ │ for each fᵢ: top-s similar existing memories (vector search) │ [LLM decide via tool call] ┌──────────┬─────────┴──────────┬──────────┐ ADD UPDATE(id) DELETE(id) NOOP insert fᵢ rewrite memory id remove id skip

The paper reported, on LoCoMo, a 26% relative improvement on an LLM-as-judge metric over OpenAI's memory, 91% lower p95 latency than full-context (about 1.4 s vs 17 s), and over 90% fewer tokens; memories averaged about 7k tokens per conversation vs about 26k for the full conversation Chhikara+ 2025. A graph variant (Mem0g) adds entity and relation extraction and did better on temporal questions.

May be out of date

Mem0 changed its core algorithm in 2026. According to Mem0's own migration docs, the new algorithm does single-pass, ADD-only extraction (one LLM call, no UPDATE/DELETE), stores new facts alongside old ones with temporal metadata, and relies on retrieval to surface the current fact. Retrieval became hybrid: vector + BM25 + entity matching, fused into one score. Graph memory was removed from the open-source package and moved into the hosted platform Mem0 docs. Mem0 reports large benchmark gains from this change (LoCoMo roughly 71 → 92, LongMemEval roughly 68 → 93), but those are vendor-reported numbers. The design lesson stands either way: reconcile on write (ADD/UPDATE/DELETE) and append on write, resolve on read are two legitimate strategies, and the field has been moving toward the second.

Reconcile on write (ADD/UPDATE/DELETE)

+ Compact store, no stale facts to confuse the reader, cheap reads.
− Two LLM calls per write; destructive (a wrong DELETE loses data forever); no history ("where did the user live before?"); an extraction mistake overwrites a correct fact.

Append on write, resolve on read

+ One cheap write; non-destructive; full history; easy audit.
− Store grows; reader must rank by recency/validity; risk of stale and current facts both surfacing (the "knowledge update" category suffers first).

4.3 Knowledge-graph and temporal memory (Zep / Graphiti)

Facts are not independent: people work at companies, projects have owners, preferences change. A temporal knowledge graph stores entities as nodes and facts as edges (subject, predicate, object), and gives every edge a validity interval. Zep's engine, Graphiti, is the reference design Rasmussen+ 2025:

Zep reported 94.8% vs MemGPT's 93.4% on the older DMR benchmark (which the paper itself criticises as too easy), and on LongMemEval up to 18.5% accuracy improvement with about 90% lower latency, using about 1.6k context tokens instead of about 115k for full context Rasmussen+ 2025. Open-source Graphiti runs on Neo4j, FalkorDB or Amazon Neptune Graphiti repo.

Edget_validt_invalidt'_createdt'_expired
(Alice) –WORKS_AT→ (Acme)2023-012026-05-012025-02-102026-06-03
(Alice) –WORKS_AT→ (Globex)2026-05-01∅ (current)2026-06-03∅

In the example, on 3 June 2026 Alice says "I started at Globex at the start of May". The system learns it (transaction time 06-03) and back-dates validity to 05-01. The Acme edge is invalidated, not deleted. "Where did Alice work in April 2026?" still answers correctly.

Common mistake

Reaching for a graph database because "memory is a graph". Graph memory has the highest write cost (entity extraction, entity resolution/dedup, relation extraction, contradiction detection: several LLM calls per episode) and entity resolution errors ("Bob" vs "Robert" vs "my manager") compound. It pays off when questions are relational or temporal (multi-hop, "who on my team…", "what changed since…"), or for enterprise data with natural entities. For a consumer chatbot that mainly needs preferences, a fact table with timestamps is usually enough. The bi-temporal idea (valid-from/valid-to columns, invalidate rather than delete) is valuable even in plain Postgres.

4.4 OS-style hierarchical memory (MemGPT / Letta)

MemGPT treats the LLM like a CPU with a small, fast main memory (the context window) and large, slow external storage, and lets the LLM itself page data in and out with function calls ("virtual context management") Packer+ 2023.

Main context (the prompt, ~ RAM) System instructions read-only: persona rules, memory-function docs Working context / core memory blocks fixed-size, editable by the LLM: [persona] [human] ... e.g. "User: Priya, vegetarian, prefers concise answers" FIFO message queue [recursive summary of evicted msgs] + recent messages, tool results, system alerts ("memory pressure at 70%") Recall storage (~ disk) every message ever exchanged, searchable by text / date conversation_search() Archival storage (~ disk) arbitrary-length text objects, vector-indexed; docs, facts archival_memory_insert/search() evict LLM as processor: every action is a function call core_memory_append / core_memory_replace → edit working context (self-editing memory) archival_memory_insert / _search, conversation_search → page data in and out (paginated) send_message; request_heartbeat=true → chain another step without waiting for the user
MemGPT-style hierarchy. Solid arrows: writes and eviction; dashed arrows: retrieval back into the prompt. Function names follow the MemGPT/Letta convention; exact names vary by version.

Key mechanics from the paper Packer+ 2023:

Letta (the company built around MemGPT) generalised the working context into memory blocks: labelled, size-limited strings (default persona and human) that always sit in the context, each with a description that tells the agent what belongs there Letta docs. Blocks can be shared between agents (useful for multi-agent shared state). Letta later added "sleep-time" agents that rewrite memory in the background between interactions Lin+ 2025.

May be out of date

Letta's direction moved in 2026 toward git-backed memory files ("context repositories" / MemFS): an agent's memory is a folder of Markdown files with frontmatter descriptions, a system/ directory whose files are always loaded into the prompt, and git versioning so multiple sub-agents can write memory concurrently in separate worktrees and merge Letta 2026. Background "dreaming" sub-agents consolidate recent conversations into memory Letta docs. Secondary sources say Letta has de-emphasised its older server-side framework; check current docs before quoting tool names.

Intuition

The MemGPT contribution is less the tiers (any RAG system has "context + store") than agency over memory: the model decides what to remember and when to look things up, using the same tool-calling loop it uses for everything else. The cost is that memory quality now depends on the model's judgement and on spending tokens and steps on memory management. The trend toward files is the same idea with a better-trained interface.

4.5 Reflection-based memory (Generative Agents)

Park et al.'s simulated town of 25 agents introduced the memory stream: a time-ordered log of natural-language observations, each with a creation timestamp, last-access timestamp and an importance score Park+ 2023. Two mechanisms made it influential.

Retrieval scoring. For a query (the agent's current situation), every memory \(m\) gets

$$\text{score}(m) = \alpha_\text{rec}\,\widehat{\text{recency}}(m) + \alpha_\text{imp}\,\widehat{\text{importance}}(m) + \alpha_\text{rel}\,\widehat{\text{relevance}}(m, q)$$

where each component is min-max normalised to \([0,1]\) over the candidate set and all \(\alpha = 1\) in the paper's implementation. The components are Park+ 2023:

Back-of-envelope: with \(\gamma=0.995\), a memory untouched for 24 game-hours has recency \(0.995^{24} \approx 0.89\); after a week (168 h), \(\approx 0.43\); after a month (720 h), \(\approx 0.027\). So recency matters on a scale of days to weeks. In a real product you'd choose \(\gamma\) per hour of wall-clock time so the half-life matches how fast your users' facts go stale: half-life \(= \ln 2 / (-\ln\gamma)\), which is about 138 hours here.

Reflection. Raw observations are too low-level to drive behaviour ("Klaus is reading a book on gentrification"). When the sum of importance scores of recent observations exceeds a threshold (150 in the paper, which worked out to roughly 2–3 reflections per game day), the agent takes the 100 most recent memories, asks the LLM for the 3 most salient high-level questions, retrieves memories relevant to each, and asks for 5 high-level insights with citations to the supporting memories Park+ 2023. The insights ("Klaus is dedicated to his research on gentrification") go back into the memory stream as new memories, and can themselves be reflected on, forming a tree of abstractions.

observations ──► memory stream (time-ordered, importance-scored) │ │ │ Σ importance > 150 ? ──yes──► reflect: │ 100 recent → 3 questions │ retrieve per question │ → 5 insights (+ citations) │ │ └──────────── insights appended to stream ◄──┘ (can be reflected on again) retrieve(q) = top-k by reĉ + imp̂ + rel̂ → plan / act

This is the template for most "memory consolidation" jobs today: periodically distill many episodic memories into fewer semantic ones, with links back to the evidence.

4.6 Procedural memory (skills, instructions, learned behaviour)

Procedural memory stores how to do things. It is the least standardised and arguably the highest-leverage form, because it changes behaviour on every future task rather than answering one question.

Common mistake

Letting an agent rewrite its own system prompt or skills without evaluation. A procedural change is a code change that affects every future run: one bad "lesson" learned from a single weird interaction can degrade all users. Treat procedural updates like deploys: propose → evaluate on a regression set → human review for shared scopes → version and allow rollback. User-scoped procedural memory ("this user wants terse answers") is lower risk than agent- or org-scoped.

4.7 Comparison of approaches

ApproachWhat's storedWrite costRead costStrengthsFailure modes
Raw vector memoryChunks of past turns/sessions + embeddings + metadataEmbedding only (cheap, no LLM)ANN query; verbose injectionLossless, simple, cheap to build, good recall for "what did we discuss about X"No time or truth model; stale/contradictory facts; similarity ≠ relevance; poor on temporal and aggregation questions
Extracted facts (Mem0-style)Atomic facts with metadata; optionally ADD/UPDATE/DELETE reconciled1–2 LLM calls per turn (extraction + reconciliation)Hybrid search over short facts; compact injectionCompact, token-cheap reads; natural for preferences and profile factsExtraction misses or hallucinates; destructive updates lose history; append-only variant surfaces stale facts
Temporal KG (Zep/Graphiti)Episodes + entities + relation edges with bi-temporal validity; communitiesHighest: several LLM calls (entities, resolution, relations, contradictions)Hybrid search + graph traversal + rerank; moderateTemporal and relational questions; history preserved; structured enterprise dataEntity-resolution errors compound; complex ops; ingestion latency; over-engineering for simple use cases
OS-style tiers (MemGPT/Letta)Core blocks in context + recall (messages) + archival (vector) or filesAgent tool calls in the hot path (tokens, steps)Core: free (always present); archival: agent-initiated searchAgent controls salience; core blocks give reliable always-on personalisationDepends on model judgement; extra steps and latency; agent may forget to save or search; block bloat
Reflection (Generative Agents)Observation stream + importance + higher-level reflections with citationsImportance-scoring call per memory + periodic reflection jobsScore all candidates by rec+imp+relProduces abstractions and "beliefs"; good for simulation and long-lived personasReflections can be wrong and self-reinforcing; costly at scale; importance scores noisy
Procedural (skills, instructions)Code, prompts, instruction files, insightsLow frequency, but needs verificationIndex always in context, body on demandChanges behaviour across all future tasks; compounding improvementBad lessons generalise badly; prompt bloat; hard to evaluate; security risk if writable by untrusted input
Notes files (NOTES.md, progress logs)Agent-written free-form or JSON stateFile-tool calls in the hot pathFile read/grep; just-in-timeIn-distribution for coding models; inspectable; survives resetsUnstructured; can sprawl; no dedup or scoping unless the harness enforces it
Interview angle

"Which memory approach would you pick?" There's no single right answer; interviewers want a justified choice tied to the query distribution. Preference recall for a consumer assistant: extracted facts + a small always-in-context profile. "What did we decide in last month's meeting with Acme?": episodic store with time-aware retrieval. "Who owns the services Alice's team depends on and what changed since Q2?": temporal graph. Long autonomous coding: notes/progress files + compaction. Then say how you'd validate the choice with an eval built from real user questions.

Go deeper

5. Designing the write path

The write path decides what enters long-term memory. Most memory bugs in production are write-path bugs: the system stored the wrong thing, stored it twice, stored it without a date, or overwrote the right thing.

WRITE PATH (usually async, after the turn) Transcript / tool events Salience gate worth storing? PII? Extract atomic facts, entities Dedup + resolve vs similar memories Commit + ts, scope, source Memory store facts · episodes graph · profile consolidate · decay · delete READ PATH (hot path, per turn or per tool call) Trigger always / tool call Query formation rewrite, time, entities Hybrid retrieve vector+BM25+graph Rank + filter rel × rec × imp, valid Inject budgeted, delimited → LLM Scope keys (tenant, user, agent, session) are attached on write and enforced as hard filters on read, never left to the LLM.
Write and read pipelines around a memory store. The write path is usually asynchronous; the read path is on the latency-critical hot path.

5.1 What to remember: salience

Storing everything makes retrieval noisy and costs money. Storing too little misses what matters. Common salience signals:

5.2 Shape of a memory record

Whatever the backend, a good memory record looks roughly like this:

{ "id": "mem_8f2c...", "tenant_id": "acme", "user_id": "u_123", "agent_id": "support-bot", "scope": "user", "type": "preference", // fact | preference | event | instruction | insight "content": "Prefers email follow-ups over phone calls", "entities": ["user:u_123"], "valid_from": "2026-09-12", "valid_to": null, // world time "created_at": "2026-09-12T10:04Z", "superseded_at": null, // system time "source": {"kind": "user_message", "session_id": "s_77", "message_id": "m_912"}, "confidence": 0.86, "importance": 6, "last_accessed_at": "2026-09-30T08:11Z", "access_count": 4, "embedding": [...], "ttl": null, "sensitivity": "normal" }

Every field earns its place: scope keys for isolation; type for routing and schema; two time axes for temporal reasoning; provenance (which message produced this) for debugging, user trust ("why do you think that?"), poisoning forensics and cascading deletion; access stats for decay; sensitivity for policy.

5.3 Dedup and conflict resolution

New facts often duplicate or contradict old ones. The standard procedure: retrieve the top-\(s\) most similar existing memories for the same scope and entity, then classify the relationship.

RelationExample (old → new)Typical action
Duplicate"Likes hiking" → "Enjoys hiking"NOOP; bump access/confidence
Refinement"Has a dog" → "Has a beagle named Max"UPDATE (merge into richer fact), keep provenance of both
Contradiction (change over time)"Lives in London" → "Moved to Berlin last month"Invalidate old (set valid_to), add new. Don't hard-delete: "where did I live before?" is a valid question
Contradiction (error)"Allergic to peanuts" → user: "No, it's shellfish, not peanuts"Supersede and mark old as incorrect (not "past truth")
Context-dependent"Prefers Python" vs "Uses Go at work"Keep both, qualify with context
Unrelated-ADD

Precedence rules worth stating in an interview: explicit user statements beat inferred ones; newer beats older for changeable state; first-party statements beat third-party content; higher-confidence sources beat lower. When in doubt, keep both with timestamps and let the reader see the history, or ask the user.

Common mistake

Running reconciliation without a scope or entity filter. "Similar" memories from another user, or about a different person ("my sister is vegan" vs "I am vegan"), will get merged. Speaker attribution is a classic extraction error: facts about third parties get stored as facts about the user. Include "who is this fact about?" as an extracted field.

5.4 Hot path vs background writes

Hot path (in-turn, often agent-initiated)

The agent calls save_memory / edits a block / writes a notes file mid-turn. Memory is visible immediately, even later in the same session, and the user can see "Memory updated". Cost: added latency, tokens, and a distracted agent. Mostly suited to explicit "remember this" and task notes.

Background (async worker)

After the turn (or at session end, or on a schedule) a worker processes the transcript: extraction, dedup, graph updates, reflection. No user-facing latency; can use a cheaper model and batch. Cost: lag, plus another system to operate (queue, retries, idempotency). This is where consolidation ("sleep-time compute", "dreaming") lives Lin+ 2025.

A common hybrid: hot path for explicit memories and session notes; a background job for implicit extraction and consolidation. Make background writes idempotent (key them on message ID) so retries don't create duplicates.

Go deeper

6. Designing the read path

6.1 When to retrieve

StrategyHowProsCons
Always-in-contextProfile / core memory blocks / MEMORY.md index injected every turnNever "forgets to look"; zero retrieval latency; prompt-cacheable if stableToken cost every call; must be small (hundreds to a few thousand tokens); stale if not maintained
Retrieve every turnHarness runs a memory search on each user message before calling the LLMSimple, predictable, no agent judgement neededAdds latency to every turn (embedding + search, typically tens of ms; more with rerank); injects noise when nothing is relevant
On demand (tool call)Agent calls search_memory(query) when it decides it needs toOnly pays when needed; agent can reformulate and iterateAgent may not realise it lacks information; extra LLM round-trip
Event-triggeredRetrieve on session start, entity mention, or task typeTargetedNeeds routing logic

The usual production answer is a combination: a small always-on profile + automatic retrieval at session start or per turn with a relevance threshold + a search tool for deep dives. Retrieval with a minimum score threshold matters: injecting irrelevant memories is worse than injecting none, because the model will try to use them (the "creepy" or non-sequitur personalisation failure).

6.2 Query formation

The raw user message is often a poor query ("yes, do that" retrieves nothing useful). Options: rewrite the query from recent context with a small model; generate several sub-queries (entities, topics); extract time expressions ("last week" → a date filter) as LongMemEval recommends Wu+ 2024; and pull entities for graph lookups. Each LLM-based rewrite adds about one small-model call of latency, so it's often done only for the on-demand tool path.

6.3 Ranking

Generalise the Generative Agents score into a ranking function you can tune:

$$\text{score}(m,q) = w_\text{rel}\cdot \text{sim}(m,q) + w_\text{rec}\cdot \gamma^{\,\Delta t(m)} + w_\text{imp}\cdot \text{imp}(m) + w_\text{freq}\cdot \log(1+\text{hits}(m))$$

with hard filters applied before scoring: scope (tenant/user), validity (\(t_\text{invalid}\) is null, or overlaps the time asked about), not deleted, sensitivity policy. In practice the relevance term is itself a fusion of vector and BM25 (and entity/graph signals), often combined with Reciprocal Rank Fusion, \(\text{RRF}(d) = \sum_r 1/(k + \text{rank}_r(d))\) with \(k\approx 60\), and optionally a cross-encoder rerank of the top 20–50. Tune the weights on an eval set rather than by intuition. Recency should be strong for changeable state ("current project") and weak for stable facts ("allergies").

6.4 Injection format and placement

Interview angle

Expect a latency question: "Memory adds how much to time-to-first-token?" Walk the budget: embedding the query (~10–50 ms for a hosted API, less if local), ANN search in pgvector or a vector DB (~ms to tens of ms at per-user scale), optional rerank (~50–200 ms), optional LLM query rewrite (hundreds of ms). These are rough magnitudes; measure your own stack. Then explain how to hide it: run retrieval in parallel with other pre-processing, skip rewrite on the automatic path, cache the profile in the prompt prefix, and do all writes asynchronously.

Go deeper

7. Consolidation and forgetting

Memory that only grows gets worse over time: retrieval precision falls, stale facts accumulate, cost rises, and privacy risk grows. Mature systems actively maintain memory.

MechanismHow it worksNotes
DecayScore down memories by time since last access (\(\gamma^{\Delta t}\)); optionally reinforce on accessMemoryBank explicitly borrows the Ebbinghaus forgetting curve to forget and reinforce by elapsed time and significance Zhong+ 2023. Usually a ranking effect, not deletion.
Merging / dedupPeriodic job clusters near-duplicate memories and merges themKeep provenance of all merged sources so deletion still cascades
Summarisation hierarchiesEpisodes → session summaries → weekly/topic summaries → profileReflection trees (Generative Agents), Zep's community summaries. Read at the coarsest level that answers the question, drill down when needed
InvalidationMark facts no longer true (valid_to)Preserves history; excluded from "current" queries
TTLsHard expiry for inherently temporary data"I'm in Tokyo this week" → expire in 7 days. Mem0's API supports an expiration date per memory Mem0 docs. Anthropic's memory-tool docs suggest deleting long-unaccessed files Anthropic docs
Size capsPer-block character limits, per-user memory quotasLetta blocks have a size limit Letta docs; Claude Code caps the loaded MEMORY.md index and asks the model to rewrite it when it's over the limit Claude Code docs. Forces prioritisation
User-requested deletion"Forget that I…"; delete all memoryMust be hard deletion across all derived artifacts (see §9)

Consolidation is where "sleep-time compute" fits: use idle time to pre-process context so that query-time work is cheaper. Letta's paper reported that pre-computing over a known context cut the test-time compute needed for the same accuracy by roughly 5× on their stateful reasoning benchmarks Lin+ 2025. Product memory systems have adopted the same idea with background consolidation passes.

Common mistake

Confusing "forgetting for relevance" with "deleting for compliance". Decay and invalidation are soft and reversible; they make retrieval better. A user's deletion request or a retention policy needs hard, verifiable deletion, including copies inside summaries, reflections, graph edges, embeddings, caches, backups (on their own schedule) and any fine-tuning data. Design provenance links up front so you can find every derived artifact.

Go deeper

8. Memory in multi-agent systems

With several agents (orchestrator + workers, or peers), memory becomes a coordination problem. See the agent architectures page for the topologies; the memory-specific choices are:

PatternHowGood forRisks
Private memory per agentEach agent has its own context and store; communicates only by messagesIsolation, clean contexts, least privilegeDuplicated work; information lost in handoffs
Shared blackboardA shared store (document, DB table, shared memory block, files) all agents read and writeCoordination, shared plan and findingsWrite conflicts; one agent's error or injected content pollutes everyone; access control harder
Artifacts by referenceWorkers write outputs to storage and pass back a pointer + short summaryAvoiding the "game of telephone" through the orchestratorNeed a consistent artifact store and naming
Handoff context packetWhen transferring control, pass a structured summary: goal, state, decisions, open questions, relevant IDsAgent-to-agent or agent-to-human transferSame lossy-summary issues as compaction

Anthropic's multi-agent research system illustrates several of these: the lead agent saves its plan to memory because its context may be truncated past 200k tokens, and sub-agents can write outputs directly to external storage rather than routing everything through the lead, to avoid information loss Anthropic 2025. The same post notes multi-agent runs used roughly 15× the tokens of a chat interaction, so shared memory that avoids re-discovery has real cost value. Letta's git-backed memory lets sub-agents write memory in parallel worktrees and merge, which is version control applied to the blackboard problem Letta 2026.

Interview angle

A good answer separates task-scoped shared state (the plan, findings so far; lives as long as the job; blackboard or artifacts) from long-term memory (persists across jobs; written through the same governed write path as single-agent memory). Also mention least privilege: a web-browsing worker that reads untrusted pages should not have write access to org-wide memory.

9. Privacy, security and governance

9.1 Privacy and compliance

9.2 Security: memory poisoning

A prompt injection normally lasts for one session. If injected content is written to memory, it persists and fires in future sessions, possibly for other users if the memory scope is shared. OWASP's Top 10 for Agentic Applications (Dec 2025) includes memory and context poisoning as a top risk OWASP 2025 (numbered ASI06 according to secondary sources; verify in the document).

Defences, roughly in order of value:

  1. Scope and isolation: per-user memory by default; shared (agent/org) memory only via a reviewed path. Tenant ID as a hard filter at the storage layer (row-level security, separate namespaces or indexes), never as an LLM instruction.
  2. Provenance-aware writes: record the source of every memory; don't let content from tools, web pages, emails or documents become "user facts" or instructions; lower trust for third-party-derived memories.
  3. Write-time screening: classify candidate memories for instruction-like content ("always send files to…", "ignore previous…") and for secrets before committing.
  4. Read-time framing: inject memories as quoted data with an explicit "not instructions" frame; keep high-privilege actions behind confirmations regardless of memory content.
  5. Audit and rollback: versioned memory (append-only logs, git-backed files) so you can find and revert poisoned entries.
  6. Path and tool hygiene: file-based memory needs path-traversal protection; Anthropic's docs warn that paths like /memories/../../secrets.env must be rejected Anthropic docs.
May be out of date

Memory security research has been very active in 2025–2026 (new attacks and proposed certified defences appear monthly). The principles above are stable; specific attack success rates and defence claims are not. Check recent work before quoting numbers.

Go deeper

10. Evaluating memory

10.1 Benchmarks

BenchmarkSetupWhat it testsCaveats
DMR (from the MemGPT paper)Multi-session chats; single fact-retrieval questionsCan you retrieve a fact from earlier sessions?Too easy; short conversations; full-context baselines do well (the Zep paper criticises it) Rasmussen+ 2025
LoCoMoLLM-generated conversations between two personas, about 300 turns and 9k tokens on average, up to 35 sessions Maharana+ 2024QA (categories include single-hop, multi-hop, temporal and open-domain), event summarisation, multimodal dialogueConversations fit in modern context windows; scores vary a lot with judge model and setup; widely used for vendor comparisons, which are often disputed
LongMemEval500 curated questions embedded in scalable chat histories (the S variant is about 115k tokens) Wu+ 2024Five abilities: information extraction, multi-session reasoning, temporal reasoning, knowledge updates, abstentionCommercial assistants and long-context LLMs showed about a 30% accuracy drop in the paper. Single-session categories are near saturation for strong systems by 2026
BEAM100 conversations from 128k up to 10M tokens, 2,000 validated questions Tavakoli+ 2025Ten abilities incl. contradiction resolution, event ordering, instruction and preference following, summarisationNew (ICLR 2026); designed so long context alone can't solve it. Too new for a settled leaderboard
May be out of date

Memory benchmark scores moved a lot in 2025–2026 and many published numbers are vendor-reported with differing judges, prompts and answer models. LoCoMo in particular has seen public disputes between vendors about methodology. Treat leaderboard claims as marketing until you've reproduced them. The categories (knowledge update, temporal, multi-session, abstention) are the durable part.

10.2 Metrics that matter in production

10.3 The accuracy–latency–token trade-off

Every memory system sits on a three-way frontier. Representative published points (each from the system's own paper, so not directly comparable):

SetupContext tokens per queryLatencyAccuracy note
Full history (LoCoMo, per Mem0 paper)~26kp95 ~17 sStrong but slow and costly Chhikara+ 2025
Mem0 extracted facts (LoCoMo)~7k memory footprintp95 ~1.4 sClose to or above full context on their judge metric Chhikara+ 2025
Full history (LongMemEval-S, per Zep paper)~115k~29 s (gpt-4o)60.2% with gpt-4o Rasmussen+ 2025
Zep graph memory (LongMemEval-S)~1.6k~2.6 s (gpt-4o)71.2% with gpt-4o Rasmussen+ 2025

These numbers are from 2025 papers with 2024-era models; current long-context models and newer memory systems will score differently. The robust lesson: a good memory layer cuts injected context by one to two orders of magnitude and latency by about an order of magnitude while matching or beating full context on accuracy, because it removes distractors. What you pay instead is write-path cost (LLM calls per turn) and complexity.

Interview angle

"How would you evaluate the memory feature you just designed?" A strong answer: (1) build a golden set from real (consented, anonymised) multi-session histories with questions per category: recall, update, temporal, multi-session, abstention; (2) measure write quality separately from read quality, so you know which half failed; (3) track tokens and p95 latency alongside accuracy; (4) run ablations (no memory, full context, memory) to prove the feature earns its cost; (5) monitor online with correction rate and memory edits; (6) include adversarial cases (injection attempts, third-party facts, sensitive data).

Go deeper

11. Product examples

May be out of date

Consumer memory features change frequently. The descriptions below are based on vendor announcements and docs available as of late 2026 and describe behaviour only to the extent publicly documented. Internal architectures are generally not public; don't claim to know them in an interview.

ProductWhat's publicly describedDesign lessons
ChatGPT memoryLaunched as "saved memories" (explicit facts the user can view and delete) in 2024; in April 2025 expanded to also reference past chat history. Users can turn either off, and there is a temporary-chat mode intended to stay out of memory (verify current behaviour). In June 2026 OpenAI announced "Dreaming", described in coverage as a background process that consolidates and rewrites memory to address staleness, correctness and scale OpenAI 2026. (Details here are from secondary coverage; I couldn't load the original post, so verify specifics.)Explicit + implicit memory as two separately controllable layers; background consolidation as the scaling answer; an "off the record" mode
Claude memory (claude.ai apps)Announced Sept 2025 for Team/Enterprise, extended to Pro and Max in Oct 2025. Memory is project-scoped (a separate memory per project), users can view and edit a memory summary, and incognito chats aren't saved to memory Anthropic 2025Scope as a privacy boundary (work projects don't bleed into each other); an editable, human-readable summary for transparency
Claude memory tool (API)A client-side tool where Claude issues file operations (view, create, str_replace, insert, delete, rename) under /memories and your application executes them against storage you control; pairs with context editing and compaction Anthropic docsModel-driven, file-shaped memory with the developer owning storage, isolation and security
Claude CodeCLAUDE.md instruction files at org/user/project/local scope, plus agent-written auto memory: a MEMORY.md index loaded each session and topic files read on demand Claude Code docsProcedural memory as files; index-in-context, content-on-demand; size caps force curation
Memory frameworksMem0 (extracted facts, hybrid retrieval), Zep/Graphiti (temporal KG), Letta (tiered memory, now git-backed files), LangMem / LangGraph store (namespaced semantic/episodic/procedural)Use as references for patterns; evaluate on your own data before adopting

12. A reference architecture to draw in an interview

Here is a design you can sketch in five minutes for "add long-term memory to our assistant / support agent", and then adapt.

┌─────────────────────────── Agent runtime ───────────────────────────┐ user ──► API gateway ──►│ 1. load session state (Redis) 2. fetch profile + core memory │ (tenant, user auth) │ 3. auto-retrieve top-k memories (threshold) 4. build prompt │ │ 5. LLM loop (tools incl. search_memory / save_memory) │──► response │ 6. emit transcript + events ─────────────┐ │ └──────────────────────────────────────────┼──────────────────────────┘ ▼ queue (SQS / Kafka / Redis streams) │ ┌──────────────── Memory service (owns all reads/writes) ─────────────┐ │ write worker: salience → extract → dedup/resolve → commit │ │ consolidation cron: merge, summarise, reflect, decay, TTL expiry │ │ deletion API: cascade by provenance (facts, embeddings, summaries) │ │ policy: scope filters, PII classifier, injection screen, audit log │ └───────┬──────────────┬──────────────────┬──────────────────┬────────┘ ▼ ▼ ▼ ▼ Postgres + pgvector Graph DB Redis / KV Object store (S3) facts, episodes, (optional) session state, raw transcripts, profiles, RLS by entities + working memory, notes/skill files, tenant, HNSW index temporal edges rate limits audit archives
StoreHoldsWhy this choiceAlternatives
Postgres + pgvectorFacts, episodes, profiles, embeddings, provenance, validity intervalsOne transactional store for metadata + vectors; SQL filters for scope/time; row-level security for tenant isolation; HNSW or IVFFlat indexes and iterative scans for filtered search pgvectorDedicated vector DB (at very large scale or for advanced hybrid search); OpenSearch for BM25 + vectors
Graph DB (optional)Entities, relations, temporal edges, communitiesMulti-hop and temporal relational queriesNeo4j, FalkorDB, Neptune (Graphiti supports these); or adjacency tables in Postgres to start
Redis / KVSession buffers, running summary, scratchpad, locksLow-latency hot state with TTLsDynamoDB; or the agent framework's checkpointer
Object storeRaw transcripts (source of truth), notes and skill files, exportsCheap and durable; lets you re-run extraction when your extractor improvesGit repo for file-based memory (versioned, mergeable)
Queue + workersAsync write path and consolidationKeeps LLM extraction off the hot path; retries; idempotency by message IDDurable workflow engine (Temporal etc.) for multi-step consolidation

Back-of-envelope sizing (say it out loud, flag assumptions): 1M users × ~500 memories each = 5×108 facts. At 1,536-dim float32 embeddings that's about 6 KB per vector, so roughly 3 TB of raw vectors before index overhead. That pushes you toward smaller embeddings (e.g. 512–768 dims), halfvec or quantisation, partitioning by tenant, or a dedicated vector store. Per-query search, however, is per user (a few hundred rows), so with a good user_id filter or partition it can even be brute force. Write-path cost: if extraction is one small-model call of about 1–2k tokens per turn, at 10 turns per user-day that's 10–20k tokens per active user per day, which is often more than the read path, so batch per session rather than per turn when lag is acceptable.

Interview angle

The talking points that separate senior answers: (1) the memory service owns scope enforcement; the LLM never chooses tenant or user IDs; (2) raw transcripts are the source of truth and memories are a derived, rebuildable index, so you can re-extract when you improve the extractor and you can cascade deletions; (3) async writes, sync reads, with a latency budget for the read path; (4) bi-temporal fields even without a graph DB; (5) evaluation and observability: log which memories were injected into each prompt so you can debug "why did it say that?"; (6) a user-facing memory UI (view, edit, delete, incognito).

Go deeper
  • Graphiti repository: a production-grade open-source temporal graph memory you can read end to end.
  • pgvector: the default "start here" vector store for memory on Postgres.
  • Anthropic memory tool docs: a clean spec for model-driven file memory, including security guidance.

Interview question bank

1. LLMs are stateless. What exactly does "memory" mean for an LLM agent?

The model only conditions on the tokens in the current request, so "memory" is whatever the harness decides to put back into the prompt. Short-term or working memory is the content of the context window during a task or conversation, managed by buffering, truncation, summarisation and compaction. Long-term memory is information stored outside the model (databases, files, graphs) and selectively retrieved into future prompts. There's also parametric memory (knowledge in weights, changed by fine-tuning), but product memory is almost always external and textual because it must be inspectable, editable, scoped and deletable. So designing memory means designing a write path (what to store), storage, a read path (what to inject, when, where) and a forgetting policy.

2. Explain episodic, semantic and procedural memory with an agent example of each.

Episodic memory records specific experiences with time and context: "On 12 Sept the user and I debugged a failing migration; the cause was a missing index." Semantic memory holds distilled facts that aren't tied to one episode: "The user's production DB is Postgres 16 on RDS." Procedural memory holds how to act: a skill file for "how to run this repo's migrations safely", or an updated instruction like "always run the linter before committing". In CoALA terms these are long-term memories that the agent reads into working memory by retrieval and updates by learning. In practice, consolidation turns episodes into semantic facts and into procedural lessons, and good systems keep provenance links from facts back to episodes.

3. Why not just use a model with a 1M-token context and resend everything?

Cost: every token is prefilled on every call, so a long history multiplies cost and time-to-first-token; caching reduces price but not the window limit or all latency. Quality: models use the middle of long contexts poorly and degrade as input grows, especially with distractors, so a curated 2k-token memory can beat 100k tokens of raw history. Published examples: Zep reported about 1.6k tokens of context beating about 115k tokens of full history on LongMemEval with gpt-4o, at about a tenth of the latency. Scale: multi-year histories still don't fit. Semantics: raw history doesn't resolve contradictions, doesn't support deletion or scoping, and doesn't distil lessons. Long context is still useful: it lets you inject more retrieved memory and delay compaction.

4. Walk me through Mem0's original write path. What are its weaknesses?

For each new exchange, an LLM extracts candidate facts using a conversation summary plus the last ~10 messages as context. For each candidate, the system retrieves the top ~10 similar existing memories and asks the LLM, via a tool call, to choose ADD, UPDATE, DELETE or NOOP. That keeps the store compact and current. Weaknesses: two LLM calls per write (cost, latency); destructive operations, so an extraction or judgement error deletes a correct fact and you lose history; contradiction handling without time loses "what was true before"; and reconciliation quality depends on retrieving the right neighbours. Notably, Mem0 itself moved in 2026 to single-pass ADD-only extraction with temporal metadata and hybrid retrieval, resolving currency at read time instead, which trades a bigger store for non-destructive, cheaper writes.

5. What is bi-temporal modelling and why does it matter for agent memory?

Each fact carries two time axes: valid time (when it was true in the world: valid_from, valid_to) and transaction time (when the system recorded and retired it: created_at, expired_at). Zep's Graphiti stores both on graph edges. When a contradicting fact arrives, the old edge is invalidated (valid_to set to the new fact's valid_from) instead of deleted. This lets you answer "where does Alice work now?", "where did she work in March?", and "what did the system believe last Tuesday?" (useful for debugging and audits). It also handles late-arriving information: the user says today that they changed jobs a month ago, so transaction time is today but valid time is a month back. You can implement this with two pairs of timestamp columns in Postgres; you don't need a graph DB.

6. Describe MemGPT's architecture. What is "self-editing memory"?

MemGPT treats the context window like RAM and external stores like disk. Main context contains read-only system instructions, a fixed-size editable working context (later "core memory blocks" like persona and human), and a FIFO message queue whose evicted messages are folded into a recursive summary. External context has recall storage (all past messages, searchable) and archival storage (an arbitrary vector-indexed text store). The LLM manages memory itself through function calls: appending to or replacing text in core memory, inserting into and searching archival storage, searching conversation history. That's self-editing memory. A queue manager warns the model at about 70% context usage so it can save important things before a flush, and a heartbeat flag lets it chain several memory operations before replying. The trade-off: memory quality depends on the model's judgement and costs extra steps.

7. Write down the Generative Agents retrieval score and explain each term.

score = α_rec·recency + α_imp·importance + α_rel·relevance, with each component min-max normalised to [0,1] over candidates and all α = 1 in the paper. Recency is exponential decay γ^Δt with γ = 0.995 per game-hour since the memory was last accessed, so using a memory refreshes it. Importance is an LLM-assigned 1–10 poignancy score set at write time. Relevance is cosine similarity between memory and query embeddings. With γ = 0.995 the half-life is ln2/0.005 ≈ 138 hours, so recency falls to about 0.43 after a week. In a product you'd tune the weights and γ on an eval set, use wall-clock time, and add hard filters (scope, validity) before scoring.

8. What is reflection in Generative Agents and what's the modern equivalent?

When the summed importance of recent observations passes a threshold (150 in the paper, about 2–3 times per game day), the agent takes its 100 most recent memories, asks for the 3 most salient high-level questions, retrieves memories for each, and generates 5 insights with citations to the evidence. The insights are stored as new memories and can be reflected on again, forming a tree of abstractions. The modern equivalent is background consolidation: "sleep-time" or "dreaming" jobs that periodically turn episodic logs into semantic facts, profile updates and procedural lessons. Risks: reflections can be wrong and self-reinforcing, so keep citations to evidence, give them a lower confidence than direct statements, and let them be invalidated.

9. Your agent's compaction keeps losing important details. How do you fix it?

First find out what's lost by diffing pre- and post-compaction states on failing traces. Typical fixes: (1) a structured compaction prompt with required sections (goal and acceptance criteria, decisions and rationale, open tasks, exact identifiers such as paths, IDs, URLs and error strings, user constraints, dead ends); (2) keep the most recent turns verbatim and summarise only older ones; (3) move durable state out of the conversation into a notes/progress file or memory tool that the agent updates as it goes, so compaction doesn't have to carry it; (4) clear stale tool outputs first, which delays compaction; (5) keep the raw transcript retrievable so the agent can search it; (6) compact less often but at higher fidelity, since each round compounds loss and invalidates the prompt cache. Add an eval: continue tasks from compacted state and measure success vs uncompacted.

10. Hot-path vs background memory writes: when would you choose each?

Hot path, where the agent writes during the turn, is right for explicit requests ("remember that…"), for task notes needed later in the same session, and when the user benefits from seeing "memory updated". It costs latency, tokens and agent attention. Background writes, where an async worker processes transcripts after the turn or session, suit implicit extraction, dedup, graph building and consolidation: no user-facing latency, can batch and use cheaper models, but memory lags. Most production systems do both: explicit memories in the hot path, everything else in the background, with idempotency keyed on message ID and a short lag SLA (seconds to minutes) so the next session sees the update.

11. How do you handle the user saying "I moved to Berlin" when memory says "lives in London"?

Detect the conflict at write time: retrieve similar memories for the same user and entity, and have the extractor classify the relation as a change over time (not an error, not a duplicate). Then invalidate the London fact by setting valid_to to the move date (or now if unknown), and add the Berlin fact with valid_from, both with provenance to the message. On read, filter to currently valid facts for "where do I live?", but keep the history for "where did I live before?". If the system is append-only, make sure ranking strongly prefers newer facts for changeable attributes and that injected memories show dates so the model can reason. If it's ambiguous (a trip vs a move), store with lower confidence or a TTL, or ask.

12. Always-inject vs retrieve-on-demand: how do you decide what goes where?

Always-inject a small, high-value, slowly changing set: the user profile (name, role, key preferences, hard constraints like allergies or "never contact by phone"), agent persona, and an index of what else exists (like MEMORY.md or skill descriptions). It must be small (hundreds to low thousands of tokens) and stable so it stays in the prompt-cache prefix. Retrieve everything else: episodes, long-tail facts, documents, using automatic retrieval with a relevance threshold at session start or per turn, plus a search tool for deeper lookups. The rule of thumb: if forgetting it causes a serious failure and it's small, pin it; if it's large or only sometimes relevant, retrieve it.

13. Back-of-envelope: what does a memory layer cost per active user per day?

State the assumptions. Say 10 turns per day. Write path: one extraction call per turn with ~1.5k input tokens and ~150 output tokens on a small model gives ~15k input + 1.5k output tokens per day; reconciliation might double it, giving ~30k tokens/day, plus embeddings (negligible). Read path: ~1k injected memory tokens per turn gives ~10k extra input tokens per day on the main model, which often costs more per token than the small extraction model. Storage: a few hundred facts at ~6 KB per 1,536-dim float32 vector plus text is a few MB per user, which is cheap. Batching extraction per session instead of per turn can cut write cost several-fold. Compare against resending full history: by turn 50 of a conversation, a single request may carry tens of thousands of tokens.

14. How would you design tenant isolation for memory in a B2B agent platform?

Isolation must be enforced by infrastructure, not the prompt. Every memory record carries tenant_id (and user_id, agent_id, scope); the memory service derives these from the authenticated request, never from model output. In Postgres, use row-level security policies keyed on a session variable, or separate schemas or databases for high-sensitivity tenants; for vector stores, use per-tenant namespaces or indexes, or mandatory filters with tests that prove no cross-tenant results. Encryption at rest with per-tenant keys enables crypto-shredding on offboarding. Shared org-level memory goes through a reviewed write path. Add automated tests that attempt cross-tenant retrieval, log injected memory IDs per request for audit, and apply the same scoping to caches and background workers, which are a common leak point.

15. What is memory poisoning and how do you defend against it?

An attacker gets malicious content written into an agent's persistent memory, through injected instructions in a web page, email or document the agent reads, or just through crafted queries as in MINJA, so it influences future sessions, possibly other users' if memory is shared. AgentPoison showed tiny poison rates can achieve high attack success with little effect on normal behaviour. Defences: per-user scope by default and a reviewed path for shared memory; provenance on every memory, with content from tools or third parties never promoted to "user facts" or instructions; write-time classifiers for instruction-like content and secrets; read-time framing of memories as untrusted data; confirmations for high-impact actions regardless of memory; versioned, auditable memory with rollback; and red-team evals for persistence attacks.

16. A user asks you to "forget everything about my divorce". What has to happen technically?

Identify all affected memories: semantic search plus entity and topic matching over the user's memories, ideally with the user confirming the list. Then follow provenance links to delete derived artifacts: extracted facts, their embeddings, graph edges and entity summaries, reflections or session summaries that mention it (regenerate them without it, or delete them), cached profiles and prompt caches. Decide on the raw transcripts per policy and the user's request (delete or exclude the relevant sessions). Log the deletion event without the content. Backups expire on their own schedule, which should be documented. Add a "do not remember" rule so the topic isn't re-extracted from future chats. Verify by running retrieval probes afterwards. This is why provenance and raw-transcript-as-source-of-truth are designed in from day one.

17. Compare vector memory, extracted-fact memory and a temporal knowledge graph for a sales assistant.

Questions a sales assistant gets: "What did Acme's CTO say about pricing last call?" (episodic, temporal), "Who's the decision maker at Acme now?" (relational, changeable), "What are this rep's preferences?" (simple facts). Raw vector memory over call notes handles the first reasonably with time filters but fumbles "now" questions because old and new contacts both match. Extracted facts handle preferences cheaply but flatten relations and, if reconciled destructively, lose history. A temporal graph models people, companies, roles and deals with validity intervals and answers "who is the decision maker now" and "what changed since Q2" well, at the highest write cost and with entity-resolution risk. A pragmatic design: episodic store of call notes + entity/relationship tables with validity columns in Postgres, upgrading to a graph engine only if multi-hop queries dominate.

18. How do you evaluate a memory system before launch?

Build a golden dataset from realistic multi-session histories (synthetic plus consented real ones) with questions per category: single-session recall, multi-session aggregation, temporal reasoning, knowledge updates, abstention (questions about things never said), and adversarial cases (injection, third-party facts, sensitive data). Evaluate the write path separately: extraction precision/recall against labelled "should-store" facts, duplicates, contradictions, speaker-attribution errors. Evaluate the read path: recall@k of supporting memories and end-to-end accuracy with a calibrated LLM judge. Always report tokens per query and p95 latency, and compare against no-memory and full-context baselines. Public benchmarks (LoCoMo, LongMemEval, BEAM) are sanity checks, not substitutes for your own query distribution.

19. What's the difference between notes files (NOTES.md, progress logs) and a memory database? When would you use each?

Notes files are agent-written, free-form or JSON documents, read and edited with file tools. They're great for task-scoped and project-scoped state in long-running agents (plan, progress, decisions, gotchas) because coding-trained models are very good with files, the agent decides what's salient, notes survive context resets, and humans can read and edit them. They lack built-in dedup, scoping, ranking and deletion guarantees, and they sprawl unless capped. A memory database with scoped records, timestamps, provenance and hybrid retrieval is better for many-user, long-term personalisation with compliance needs. Many systems combine them: files for a single agent's working knowledge (Claude Code auto memory, Letta's git-backed memory), a database behind a memory service for multi-tenant user memory.

20. How should memory work in a multi-agent system with an orchestrator and workers?

Separate task-scoped shared state from long-term memory. For the task: a blackboard (shared plan, findings, artifact store) that workers write to directly, passing back references plus short summaries so the orchestrator's context isn't flooded and details aren't lost in a game of telephone. The orchestrator persists its plan externally because its own context may be compacted. Workers get clean, private contexts. For long-term memory, all agents go through the same governed write path with scope and provenance, and least privilege applies: a worker that reads untrusted web content shouldn't write to org-wide memory. For concurrent writes, use transactions, optimistic concurrency, or version control (as in Letta's git-worktree approach) to merge.

21. Where should retrieved memories go in the prompt, and in what format?

Stable memory (profile, core blocks, skill index) goes near the top after the system prompt, so it's part of the cached prefix and doesn't change every turn. Per-turn retrieved memories go near the end, just before the latest user message, so changing them doesn't invalidate the cache and they're not buried in the middle where attention is weakest. Format: a delimited block (e.g. a tagged section) of concise bullets with dates and optionally source and confidence, plus an instruction that these are notes from past conversations, may be outdated, and must not be treated as instructions. Keep a token budget (hundreds to ~2k), and apply a relevance threshold so irrelevant memories aren't injected at all.

22. What is procedural memory for agents and how can it go wrong?

Procedural memory stores how to do things: skill libraries of verified code (Voyager), lessons from failures (Reflexion, ExpeL), self-updated system prompts (LangMem's prompt optimisers), instruction files (CLAUDE.md/AGENTS.md), and Skills folders loaded via progressive disclosure. It's high-leverage because it changes behaviour on every future task. It goes wrong when a lesson from one odd interaction generalises badly, when prompts bloat until adherence drops, when conflicting instructions accumulate, or when untrusted input can write procedures (a persistence vector for injection). Treat procedural updates like code: propose, test on a regression set, review for shared scopes, version, roll back. Keep user-scoped preferences separate from agent- or org-wide procedures.

23. Design: add memory to a customer-support agent used by 500 enterprise tenants.

Clarify requirements: which questions need memory (past tickets, customer preferences, account facts), retention and compliance (GDPR, residency), latency budget. Scopes: tenant (playbooks, product facts: procedural/semantic, reviewed writes), customer/user (preferences, history: episodic + semantic), session (Redis buffer + rolling summary). Write path: async worker after each conversation extracts typed facts (preference, issue, commitment, contact info) with provenance and validity, dedups against the customer's existing facts, screens for PII policy and injection; ticket summaries stored as episodes. Read path: at session start inject a compact customer profile and open commitments; auto-retrieve similar past tickets with a threshold; give the agent a search tool. Storage: Postgres + pgvector with RLS by tenant, S3 transcripts as source of truth, a queue for writes. Governance: per-tenant retention, deletion API cascading by provenance, an agent-visible memory UI. Evaluation: golden sets per category, resolution rate and repeat-question rate online, p95 latency and token cost.