• Compute
  • Customers
  • Pricing
Sign In
More Blog Posts
XDiscordLinkedInYouTube

Products

  • GPUs
  • Inference
  • Studio

Developers

  • Model library
  • Documentation
  • Glossary

Company

  • About Us
  • Blog
  • Events
  • Partnership
  • Scale
  • Career
  • Ambassador program
  • Mission & Vision

Popular models

    Stay in the loop

    By submitting, you acknowledge that we may collect and use the information you provide, which may include personal information.

    XDiscordLinkedInYouTube

    Copyright ©2026 All rights reserved.

    Privacy PolicyTerms of UseLegal Documentation
    More Blog Posts
    Engineering

    Agent Memory Architecture: Working, Session, and Long-Term Memory in Production

    A practical breakdown of working, session, and long-term memory in AI agents, including storage choices, retrieval ranking, write policies, consolidation costs, and how to keep memory off the hot path.

    October 04, 2026

    The intuitive fix for an agent that forgets is a longer context window, and it does not work. The BEAM benchmark exists partly to demonstrate this: it cannot be solved by expanding the context window, which is why it has become one of the more relevant evaluations for production deployments. The reason is that the memory problem is not storage capacity. It is deciding what is worth keeping, what has stopped being true, and what should surface in a given moment. A million-token context window holds more of the conversation and does nothing about the fact that a preference the user expressed in March was contradicted in July, or that the fifteenth most relevant thing in the history keeps outranking the thing that just happened. In 2026 memory is treated as a dedicated architectural component sitting beside the context window rather than as a longer prompt.

    • Three tiers with different storage, different latency, and different failure modes. Working memory is the context window itself. Session memory is per-thread state that survives a turn. Long-term memory is cross-session knowledge, and it is the only tier that is genuinely hard.

    • The dominant cost is not storage. Vector storage is cheap. The LLM calls for extraction and consolidation can exceed storage cost by 10 to 100 times, which makes batching consolidation at session boundaries rather than running it per message the single most effective cost decision.

    • GMI Cloud serves both halves of a memory architecture, since consolidation and embedding are inference workloads with a different profile from the agent’s hot path: batchable, latency-insensitive, and well suited to different capacity than the interactive endpoint.

    • Retrieval quality degrades in production for a counterintuitive reason. It is usually not that search is bad. It is that irrelevant past state keeps outranking fresh context, which is a ranking problem rather than an embedding problem.

    • The hot path should read from long-term memory and never write to it synchronously. Writes belong on a background path at session boundaries, because a user waiting on a memory extraction call is waiting for work that does not improve their current answer.

    • The hard problems named by practitioners in 2026 are cross-session identity, temporal abstraction at scale, and memory staleness. All three are write-policy problems rather than retrieval problems.

    The Three Tiers

    The tiers are distinguished by where the data lives, how long it survives, and what it costs to access.

    Tier

    Storage

    Persistence

    Retrieval latency

    Working

    Model context window

    Current turn or session

    Zero, already in context

    Session

    Redis, Postgres, checkpointer

    Thread lifetime

    Single-digit milliseconds

    Long-term

    Vector store, graph, or both

    Indefinite

    Tens to hundreds of milliseconds

    The framing that clarifies the relationship. The most influential idea in this space came from the MemGPT work, now folded into Letta: treat the context window as RAM and external storage as disk, with the agent responsible for paging between them through tool calls. The model decides when something should be evicted from context and written out, and when something stored should be pulled back in.

    That framing makes the tiers feel like one system with a hierarchy rather than three separate features, and it explains why the interesting engineering is in the paging policy rather than in any individual store.

    Working Memory Is the Context Window

    Working memory needs the least explanation and has one property worth stating explicitly: it is free to read and expensive to hold.

    Zero retrieval latency. Anything in the context window is already available to the model. There is no lookup, no network call, no ranking decision. This is why the instinct to put more in context is so persistent: for the thing you put there, access is perfect.

    The cost is the token budget and the attention budget. Every token in context is paid on every forward pass, and beyond a certain volume the model’s attention is diluted across content that is not relevant to the current step. This is the same effect that degrades tool selection when too many tool definitions are loaded, applied to conversational history.

    The management pattern. Production systems monitor token usage and, as limits approach, prompt the model to summarise and move details into external storage. The working context stays focused; the detail moves down a tier. Where this is done well, the summary preserves what matters for the current task and the full detail remains retrievable.

    The infrastructure consequence. Working memory on the serving side is the KV cache. A session whose requests route to the same serving instance reuses the cached prefix across turns; one that bounces between instances rebuilds it every turn, which costs both prefill latency and compute. Session affinity is not an optimisation for conversational agents, it is what makes working memory behave the way the architecture assumes.

    Session Memory: The Tier That Gets Skipped

    Many memory designs jump from the context window to a vector store and leave a gap where session memory should be.

    What belongs here. Per-thread state with a known shape: the current task, which steps have completed, intermediate results, the active tool context, the user’s stated goal for this interaction. This is structured state, not semantic knowledge, and semantic search is the wrong retrieval mechanism for it.

    Where it should live. Relational and key-value stores remain the right answer for working state, message buffers, and anything with a known schema. Redis is the most common backing store in production agent stacks, and a checkpointer in an orchestration framework frequently handles this tier without the team having to build it.

    Why skipping it causes problems. Task state pushed into a vector store becomes semantically searchable, which means it competes with genuine long-term knowledge in retrieval results. The step-four output of the current task should be available by key lookup, not surfaced because it happened to embed near the user’s question. This is one concrete source of the ranking degradation described later.

    The useful consequence of getting it right. When session state has its own store, long-term memory only has to handle cross-session knowledge, which is the part that is actually difficult. A framework checkpointer carrying thread state for free narrows the hard problem considerably.

    Long-Term Memory: Three Kinds, Several Stores

    Long-term memory is where the architecture becomes interesting, and the first useful distinction is that it holds three different kinds of thing.

    Episodic memory records what happened. A past conversation, an action taken, an outcome observed. It is inherently temporal and its value decays, which makes expiry policy relevant in a way it is not for the other two.

    Semantic memory records what is true. User preferences, entity facts, domain knowledge. It is the tier most likely to go stale and the one where contradiction matters, because a fact superseded in July should not surface alongside the March version as though both were current.

    Procedural memory records how to do something. Successful approaches to recurring tasks, learned workflows, patterns that worked. It is the least commonly implemented and the most valuable for agents that repeat similar work.

    Storage choices by shape of the problem.

    A vector store is the current production default and sufficient for a single agent with straightforward retrieval needs. Chroma, pgvector, and the managed options all work, with HNSW indexing scaling logarithmically as the corpus grows.

    A graph-native layer fits when entities must be connected, when memory is shared across agents, or when long-term user context at production scale involves relationships rather than isolated facts. Mem0 added graph memory in January 2026, storing memories as directed labelled graphs with entities as nodes and relationships as edges. Temporal knowledge graph approaches add the dimension that handles supersession explicitly.

    A hybrid is common at scale: vector retrieval for semantic similarity, graph traversal for relational and temporal queries, with the retrieval pipeline merging both.

    The agent-managed alternative. Rather than a memory layer that writes on the agent’s behalf, Letta agents manage their own memory through tool calls: writing to core memory that is always in context, searching archival memory for long-term storage, and recalling conversation memory for past interactions. This puts the paging decision in the model’s hands, which is more flexible and more expensive in tool calls.

    Write Policy: The Problem Most Stacks Miss

    The shift that separates a working memory system from a demo is moving from storing and retrieving context to governing what becomes true over time.

    The questions a write policy answers. What gets promoted from transient execution state into durable knowledge? What happens when a new fact contradicts a stored one? When does a stored memory expire? Who or what has authority to assert something as true?

    Why this is harder than retrieval. Retrieval is a well-understood problem with mature tooling. Write policy is a judgement about truth, and getting it wrong produces an agent that confidently recalls something the user corrected two months ago.

    The three hard problems practitioners name for 2026. Cross-session identity, meaning reliably knowing that this user is the same user. Temporal abstraction at scale, meaning representing how facts relate across time rather than as a flat set. And memory staleness, meaning detecting and handling facts that have stopped being true. All three are write-policy problems.

    Conflict resolution needs a defined strategy. When multiple sessions or devices write to the same memory store, something has to decide which write wins. Last-write-wins with vector clock ordering for causal consistency is one documented approach. The specific choice matters less than having made one deliberately, because the default is whichever write happened to land second.

    The practical minimum. Record a timestamp and a source on every memory. Supersede rather than accumulate when a new fact contradicts an old one in the same slot. Apply TTL-based expiry to episodic memory so the store does not grow without bound. These three do not solve the hard problems but they prevent the most common failure, which is a store that accumulates contradictions and surfaces them indiscriminately.

    The Consolidation Cost

    This is the number that most changes how a memory system should be built.

    Vector storage is cheap. Storing embeddings at production scale costs little, and managed vector databases with HNSW indexing scale logarithmically rather than linearly.

    The LLM calls dwarf it. Extraction and consolidation, meaning the model calls that read a conversation and decide what is worth storing, can exceed storage cost by 10 to 100 times. The memory system’s dominant line item is inference, not storage.

    The decision that follows. Batch consolidation at session boundaries rather than running it per message. A conversation of forty turns consolidated once at the end costs one extraction call. The same conversation consolidated per message costs forty, and the incremental information from running it forty times is small because most turns do not introduce durable facts.

    The second decision. Consolidation is a batch workload with no latency requirement, which means it does not need the same capacity as the agent’s interactive path. An endpoint reserved and tuned for low time-to-first-token is the wrong place to run a batch of extraction calls, and running them there competes with user-facing requests for the same GPU capacity.

    The embedding half. Writing to a vector store requires embedding every stored memory, and retrieving requires embedding the query. Embedding inference has a completely different resource profile from generation: small model, short input, fixed-size output, very high throughput per GPU. Sizing it like a generation workload over-provisions it substantially, and co-locating it on the generation endpoint consumes VRAM that the generation model needs for KV cache.

    Retrieval: Ranking Is the Problem, Not Search

    The failure mode that surprises teams in production is worth stating plainly.

    What it looks like. Memory retrieval returns results that are semantically related to the query and unhelpful. The agent’s answers get worse as the memory store grows, which feels like a search quality problem.

    What it usually is. Irrelevant past state outranking fresh context. The embedding found genuinely similar content; the problem is that similar is not the same as useful, and a three-month-old conversation about a related topic can score higher than what the user said two turns ago.

    The ranking order that works. The documented pattern ranks user memories above session context above raw history, with the retrieval pipeline merging the tiers rather than treating them as one undifferentiated pool. Recency and memory type both factor into the score alongside similarity.

    Reranking as a second pass. A cross-encoder reranker re-scores candidates against the query before anything enters the context window, using a dedicated reranking model or an LLM. This is where precision improves meaningfully, because the first-pass vector search optimises for recall and the second pass optimises for relevance.

    The simplest version of the fix. Score memories on similarity, recency, and importance together rather than similarity alone, and keep the tiers separate in retrieval so session state does not compete with durable knowledge for the same slots.

    The Hot Path Principle

    A clean separation that resolves several design questions at once.

    During the active interaction, the agent reads from long-term memory and writes only to the short-term session cache. The user is waiting, and every operation on this path should contribute to the answer they are waiting for.

    After the interaction, extraction and consolidation run on a background path. The agent does not wait on them, and in production this typically means a background queue rather than an inline call at the end of the turn.

    Why the distinction matters operationally. A memory write on the hot path adds its latency to the user’s wait for work that does not improve their current answer. An extraction call taking 800 milliseconds makes every turn 800 milliseconds slower in exchange for a memory the user will benefit from next week.

    The production footgun worth naming. Memory frameworks that default to synchronous writes put extraction on the hot path by default, which is why at least one framework moved asynchronous mode to the default in a major version specifically because synchronous writing was the most common production mistake.

    Progressive Adoption

    Memory architecture is frequently over-built at the start, and the published advice converges on a staged approach.

    Start with good system prompts and a short context. A surprising share of what teams build memory for is better handled by a well-constructed system prompt and the current conversation. Build memory when the absence of it is a demonstrated problem, not in anticipation.

    Add session state when tasks span multiple turns. A checkpointer or a Redis-backed store for thread state. This is cheap, structured, and solves the most common immediate complaint, which is an agent losing track of what it was doing.

    Add vector retrieval when the agent needs to reference past research or documentation. This is the first genuinely external memory tier and the first one with a meaningful cost.

    Add persistent facts when user preferences or long-running state actually matter. A key-value store or a small Postgres table, updated explicitly when the agent learns something trustworthy, is frequently sufficient and considerably more predictable than semantic retrieval over everything the user has ever said.

    Add graph structure when relationships between entities drive the answers. This is the last step and the one with the highest operational cost, justified when the queries are genuinely relational rather than similarity-based.

    Infrastructure for the Memory Layer

    Three workloads sit in an agent memory architecture, and they have different resource profiles.

    The interactive agent path needs consistent low latency and benefits from session affinity so that working memory, which is the KV cache, persists across turns rather than being rebuilt. GMI Prime Inference provides reserved capacity with pre-loaded weights for this path, and the cache continuity that session affinity enables is what keeps per-turn latency flat as a conversation grows rather than climbing with accumulated context.

    The embedding path is high-throughput and small-model. It serves both the write path, embedding memories as they are stored, and the read path, embedding each query. Sizing it alongside the generation model over-provisions it; co-locating it on the generation endpoint takes VRAM the generation model needs for KV cache. At volume it deserves its own allocation.

    The consolidation path is batch inference with no latency requirement, and it is the one that dominates cost. Because it is batchable and can tolerate interruption, it runs well on capacity that would be wasted on an interactive endpoint. Separating it from the interactive path both reduces its cost and stops it competing with user-facing requests.

    The broader cost mechanics across caching, batching, and routing that apply to these workloads are covered in LLM Inference Cost Optimization: Caching, Batching, and Routing. For teams deploying agents rather than assembling the infrastructure directly, GMI Agentbox provides the runtime, deployment, and per-session observability layer.

    Conclusion

    Longer context windows do not solve agent memory, which is why a benchmark explicitly designed to resist that solution has become a reference point for production work. The problem is deciding what to keep, detecting what has stopped being true, and surfacing the right thing at the right moment, and none of those are capacity problems.

    The architecture that works separates three tiers with different stores and different access patterns: the context window as working memory with zero retrieval cost and a real attention cost, a structured session store for thread state that should not compete with semantic retrieval, and long-term memory covering episodic, semantic, and procedural knowledge across sessions.

    Two decisions carry most of the economics. Batch consolidation at session boundaries rather than per message, because extraction calls can cost 10 to 100 times more than storage. And keep writes off the hot path, because a user waiting on a memory extraction is waiting for work that benefits a future conversation rather than the current one.

    When retrieval quality degrades as the store grows, the cause is usually ranking rather than search: stale and tangentially related material outranking fresh context. Scoring on recency and importance alongside similarity, keeping the tiers separate in retrieval, and adding a reranking pass addresses it more reliably than changing the embedding model.

    Build AI Without Limits

    GMI Cloud helps you architect, deploy, optimize, and scale your AI strategies

    FAQ

    Because the problem is selection and currency, not capacity. The BEAM benchmark was constructed so that it cannot be solved by expanding the context window, which is part of why it has become a relevant evaluation for production deployments. A larger window holds more conversation and does nothing about a preference expressed in March being contradicted in July, or about tangentially related history outranking what the user said two turns ago. It also carries a real cost: every token in context is paid on every forward pass, and beyond a certain volume attention is diluted across content irrelevant to the current step. In 2026 memory is treated as a dedicated architectural component beside the context window rather than as a longer prompt.

    Ready to build?

    Explore powerful AI models and launch your project in just a few clicks.

    Get Started