The previous chapter defined what memory is. This one shows how it is wired into a RAG agent, using the two designs that memory libraries such as LangMem (from the LangChain team) package up, and that you can build with any agent framework.
Both designs share one piece: a memory store (in development, often an in-memory store; in production, PostgreSQL, MongoDB or a vector database) holding records under a namespace per user and per memory type, for example ("user-001", "profile"), ("user-001", "semantic") and ("user-001", "episodic"). They differ in when and by whom memories are written.
Design 1: background extraction
The conversation graph stays lean and adds just one node at the start:
- load_memory: given the user ID and the question, search the store for the user's profile (one record), the most relevant semantic facts (say, top 5) and the most relevant episodes (say, top 3). Put them into the agent's state.
- retrieve: the usual document retrieval from the vector index.
- generate: answer using the question, the retrieved documents and the loaded memories.
Writing happens outside the graph. After each response, the application hands the latest exchange (question and answer) to a background executor, which runs a set of memory managers:
- a profile manager, with a schema such as name, role, organisation, expertise, which updates the single profile record;
- a semantic manager, with a schema such as subject, predicate, object, context ("user — works with — Pinecone — for RAG projects"), which adds or updates fact records;
- an episodic manager, with a schema such as observation, thoughts, action, result, which records experiences worth recalling.
Each manager is an LLM call that reads the conversation and the existing memories, and returns a set of create, update and delete operations, so it consolidates rather than blindly appending.
The executor is debounced: rather than extracting after every single message, it waits for a quiet period (10 seconds in a demo, minutes or the end of a session in production) and then processes everything since the last run. Rapid back-and-forth messages are processed once, with full context, and none of it slows down the user's turn.
In practice: a user says, "My name is Asha, I work at Acme Retail, and I want to understand the improvement types in this quality standard." The answer comes back at normal speed. A few seconds later, the background run fills the profile (name: Asha, organisation: Acme Retail), stores a semantic fact about the topic of interest, and records an episode. The next question benefits from all of it.
Design 2: hot-path memory tools
Here the agent itself manages memory, during the conversation, through tools:
- search tools (
search_profile,search_semantic,search_episodic) to look up what is known; - manage tools (
manage_profile,manage_semantic,manage_episodic) to create, update or delete records.
The graph is the standard agent loop of an agent node and a tools node: the model decides which tools to call (knowledge-base search, memory search, memory write), the tools run, and the model continues. The system prompt tells it when to use each one: "When the user shares personal details, update their profile. Before answering, search memory for relevant context."
In practice: a user says, "My name is Ravi, I work at Google, and I want to know what this standard says about improvements." The agent calls search_profile (empty), search_knowledge_base (the standard's chunks) and manage_profile (save name and employer), then answers. In a follow-up in the same session, it greets the user by name and, asked "what is my name?", answers from memory without calling the knowledge base.
There is a characteristic failure here too. The agent decides which tools to call, so sometimes it saves the profile but skips the semantic or episodic write that a background manager would always have attempted. Tool-driven memory is only as diligent as the model's tool choices, which is why prompts for hot-path agents usually spell out when each memory tool must be used, and why some systems pair hot-path writes with a background sweep.
Choosing
| Background | Hot path | |
|---|---|---|
| Latency added to the user's turn | None | One or more tool calls |
| Memories available | After the debounce, so in the next turn or session | Immediately, even within the same turn |
| Extraction quality | Sees whole exchanges and consolidates carefully | Sees the moment; may skip writes |
| Consistency | Every manager runs every time | Depends on the model's tool choices |
| Complexity | A separate worker or queue | All inside the agent loop |
A common production blend: read memory in a dedicated node at the start of every turn (cheap and deterministic), write in the background for most memory types, and expose a hot-path tool only for things the user explicitly asks to be remembered ("remember that I prefer metric units"), so that those take effect immediately.
Retrieval quality applies to memories too
Memory search is retrieval, and everything earlier in this course applies to it:
- Keep memory records small and single-topic, as with chunking.
- Store timestamps and types as metadata, and filter on them, for example recent episodes only, or never expired items.
- Evaluate it: build test conversations where a remembered fact should change the answer, and check that it does (and that stale facts do not).
- Trace memory reads and writes like any other step, so that "why did it call me by the wrong name?" can be answered.