Memory in Practice: Hot-Path and Background Pipelines

This chapter turns memory theory into architecture. It walks through the two standard designs, as implemented by memory frameworks such as LangMem: a background pipeline that extracts memories after the conversation using schema-driven memory managers, and a hot-path agent that saves and searches memories through tools. It then sets out how to choose between them and the ways each one fails.

Advanced RAG Masterclass

The previous chapter defined what memory is. This one shows how it is wired into a RAG agent, using the two designs that memory libraries such as LangMem (from the LangChain team) package up, and that you can build with any agent framework.

Both designs share one piece: a memory store (in development, often an in-memory store; in production, PostgreSQL, MongoDB or a vector database) holding records under a namespace per user and per memory type, for example ("user-001", "profile"), ("user-001", "semantic") and ("user-001", "episodic"). They differ in when and by whom memories are written.

Design 1: background extraction

The conversation graph stays lean and adds just one node at the start:

  1. load_memory: given the user ID and the question, search the store for the user's profile (one record), the most relevant semantic facts (say, top 5) and the most relevant episodes (say, top 3). Put them into the agent's state.
  2. retrieve: the usual document retrieval from the vector index.
  3. generate: answer using the question, the retrieved documents and the loaded memories.

Writing happens outside the graph. After each response, the application hands the latest exchange (question and answer) to a background executor, which runs a set of memory managers:

  • a profile manager, with a schema such as name, role, organisation, expertise, which updates the single profile record;
  • a semantic manager, with a schema such as subject, predicate, object, context ("user — works with — Pinecone — for RAG projects"), which adds or updates fact records;
  • an episodic manager, with a schema such as observation, thoughts, action, result, which records experiences worth recalling.

Each manager is an LLM call that reads the conversation and the existing memories, and returns a set of create, update and delete operations, so it consolidates rather than blindly appending.

The executor is debounced: rather than extracting after every single message, it waits for a quiet period (10 seconds in a demo, minutes or the end of a session in production) and then processes everything since the last run. Rapid back-and-forth messages are processed once, with full context, and none of it slows down the user's turn.

In practice: a user says, "My name is Asha, I work at Acme Retail, and I want to understand the improvement types in this quality standard." The answer comes back at normal speed. A few seconds later, the background run fills the profile (name: Asha, organisation: Acme Retail), stores a semantic fact about the topic of interest, and records an episode. The next question benefits from all of it.

Design 2: hot-path memory tools

Here the agent itself manages memory, during the conversation, through tools:

  • search tools (search_profile, search_semantic, search_episodic) to look up what is known;
  • manage tools (manage_profile, manage_semantic, manage_episodic) to create, update or delete records.

The graph is the standard agent loop of an agent node and a tools node: the model decides which tools to call (knowledge-base search, memory search, memory write), the tools run, and the model continues. The system prompt tells it when to use each one: "When the user shares personal details, update their profile. Before answering, search memory for relevant context."

In practice: a user says, "My name is Ravi, I work at Google, and I want to know what this standard says about improvements." The agent calls search_profile (empty), search_knowledge_base (the standard's chunks) and manage_profile (save name and employer), then answers. In a follow-up in the same session, it greets the user by name and, asked "what is my name?", answers from memory without calling the knowledge base.

There is a characteristic failure here too. The agent decides which tools to call, so sometimes it saves the profile but skips the semantic or episodic write that a background manager would always have attempted. Tool-driven memory is only as diligent as the model's tool choices, which is why prompts for hot-path agents usually spell out when each memory tool must be used, and why some systems pair hot-path writes with a background sweep.

Two architectures side by side. Background: graph load_memory → retrieve → generate, with a debounced background executor running profile, semantic and episodic memory managers after the response. Hot path: an agent node calling knowledge-base search, memory search and memory manage tools during the turn, all reading and writing the same per-user memory store
Two ways to write memory. Background: the graph only reads memory; a debounced executor extracts and consolidates after the response. Hot path: the agent reads and writes memory through tools, mid-conversation.

Choosing

BackgroundHot path
Latency added to the user's turnNoneOne or more tool calls
Memories availableAfter the debounce, so in the next turn or sessionImmediately, even within the same turn
Extraction qualitySees whole exchanges and consolidates carefullySees the moment; may skip writes
ConsistencyEvery manager runs every timeDepends on the model's tool choices
ComplexityA separate worker or queueAll inside the agent loop

A common production blend: read memory in a dedicated node at the start of every turn (cheap and deterministic), write in the background for most memory types, and expose a hot-path tool only for things the user explicitly asks to be remembered ("remember that I prefer metric units"), so that those take effect immediately.

Retrieval quality applies to memories too

Memory search is retrieval, and everything earlier in this course applies to it:

  • Keep memory records small and single-topic, as with chunking.
  • Store timestamps and types as metadata, and filter on them, for example recent episodes only, or never expired items.
  • Evaluate it: build test conversations where a remembered fact should change the answer, and check that it does (and that stale facts do not).
  • Trace memory reads and writes like any other step, so that "why did it call me by the wrong name?" can be answered.
MediumMemoryArchitecture

In a background memory design, why debounce the extraction instead of running it after every message?

MediumMemoryAgents

A hot-path memory agent saved the user's name but not their stated project goal. Why might this happen, and how would you fix it?

EasyMemorySecurity

Why should memory records be namespaced per user and per memory type?