For agents needing both short-term and long-term memory, Mem0 is the go-to framework for personalized, persistent user data. If your goal is managing complex agent steps and persistent state, LangGraph is the best fit, often paired with databases like Redis or Pinecone for long-term storage. For multi-agent teams, CrewAI simplifies memory setup, while MemGPT functions like an operating system to manage infinite context through distinct storage tiers.
1LangChain LangGraphBest for controlling complex agent workflows that require built-in state persistence. It manages short-term history and connects to databases like Redis for long-term storage, offering high flexibility for stateful applications.64%
When building an AI agent that requires both volatile short-term context (recent conversation turns, immediate task variables) and persistent long-term memory (user preferences, historical facts, evolving knowledge graphs), there isn't a single "one-size-fits-all" framework. Instead, the ecosystem has specialized around a few top-tier tools depending on your architecture.
The leading frameworks for integrating short-term and long-term memory solutions include:
Letta (formerly MemGPT) : Best if you want an OS-inspired, memory-first agent runtime. It treats the LLM context window like physical RAM and external storage like disk memory, allowing the agent itself to manage its own memory edits, recalls, and archival storage autonomously.
Mem0 : Best as a drop-in, general-purpose memory layer. It provides seamless multi-level scopes (user, session, and agent) combining vector search, SQL, and entity extraction to automatically distill and persist personalization data across interactions.
Zep : Best for enterprise-grade, temporal knowledge graph memory. Zep excels at capturing facts that change over time, utilizing progressive summarization and temporal graphs (via Graphiti) to handle complex conversational histories and fast retrieval.
LangMem : Best if you are already heavily building within the LangChain / LangGraph ecosystem . It provides native primitives for optimizing, updating, and recalling agent reflections and long-term state across multi-turn workflows.
To help narrow down the best fit, let me know:
What LLM orchestration framework are you currently using (e.g., LangGraph, LlamaIndex, CrewAI, or raw API calls)?
Is this memory meant for a single-user personal assistant , or a multi-tenant enterprise application?
If you’re building an agent today, I’d separate “memory” into two layers rather than looking for one database/framework to do everything:
Short-term / working memory → the current thread, tool state, intermediate results, and recent messages.
Long-term memory → durable facts, preferences, past experiences, and learned procedures that survive across sessions.
The strongest general architecture I’d recommend is LangGraph + LangMem, with Zep/Graphiti or Mem0 added if you need a more specialized long-term memory backend.
LangGraph explicitly distinguishes thread-scoped short-term memory from cross-thread long-term memory, and its long-term store supports arbitrary namespaces.
LangMem adds the particularly useful piece: a memory manager that can extract, update, consolidate, and delete memories, including both synchronous “hot path” memory and background processing.
Why I wouldn't just use a vector database
A vector DB is useful, but “long-term memory = embeddings” is too simplistic.
I'd model at least these three categories:
Memory
Example
Retrieval
Semantic
“User prefers Python over JavaScript”
Direct/profile + semantic
Episodic
“Last month we debugged their Stripe integration”
Semantic + recency
Procedural
“When deploying this service, run tests X → Y → Z”
Structured retrieval
LangMem makes essentially this distinction between semantic, episodic, and procedural memory.
For an advanced agent, I'd also preserve the raw episode behind an extracted memory. That gives you:
Raw conversation/event
↓
Memory extraction
↓
Structured fact
↓
Provenance → original conversation/event
↓
Current validity / confidence
That becomes extremely valuable when a user says:
“Actually, I don't work at Acme anymore.”
You don't want two equally valid embeddings saying:
This is where Zep/Graphiti becomes particularly interesting. Graphiti represents facts as temporal relationships with validity windows and maintains provenance back to the episodes that produced them. It also combines semantic, keyword, and graph retrieval.
When I'd choose each framework
LangGraph + LangMem
Best starting point if you're already building an agent/workflow around LangGraph. You get a clean separation between checkpointed conversation state and persistent memory, plus memory-management primitives.
Zep / Graphiti
I'd look here when your agent needs to understand changing relationships and history:
“Who did I work with at Company X in 2024?”
“What was my previous preference before I changed it?”
“What led to this decision?”
Graphiti's temporal graph/provenance model is designed for precisely this kind of evolving knowledge.
Mem0
A good choice when you primarily want a memory layer you can bolt onto an existing agent, rather than adopting a larger agent runtime. Current comparisons consistently position it as a relatively framework-agnostic memory layer.
Letta
More appropriate when memory is fundamentally part of the agent runtime itself—the agent actively manages different memory tiers and decides what should stay in its working context versus archival memory.
What I'd build for a production agent
I'd use this pattern:
USER
│
▼
┌───────────────┐
│ Agent Runtime │
│ LangGraph │
└───────┬───────┘
│
┌───────────┴───────────┐
▼ ▼
Thread Checkpoint Memory Retrieval
│ │
current messages user profile
current task relevant facts
tool results past episodes
working state procedures
│ │
└───────────┬───────────┘
▼
LLM CALL
│
▼
RESPONSE
│
┌───────┴────────┐
│ │
immediate save background
important facts consolidation
│
▼
Long-term store
The important design choice is not to put the entire long-term memory into every prompt. Retrieve only what is relevant to the current task. LangGraph's documentation specifically notes that long conversation histories become expensive and can actually degrade model performance even when the context window technically accommodates them.
I'd also make memory writes asynchronous
Don't make every user request wait for an expensive “memory consolidation” operation.
LangMem explicitly supports this background-memory-manager approach.
For particularly important facts, however, I'd allow hot-path writes:
User: "Remember that my production database is in us-east-1."
Agent → explicit memory write → durable store
rather than relying exclusively on an asynchronous extractor.
The stack I'd pick
If I were starting from scratch:
Simple agent
LangGraph + LangMem + Postgres
Agent with sophisticated temporal/user memory
LangGraph + LangMem + Zep/Graphiti
Existing agent framework, want memory with minimal architectural change
Mem0
Highly autonomous, persistent agent where memory is part of the runtime
Letta
One important caveat: I wouldn't blindly trust vendor benchmark rankings. Recent comparisons themselves note that many memory benchmarks are vendor-produced and aren't necessarily independently reproduced.
The key architectural principle
The framework matters less than making these five operations explicit:
REMEMBER → What should be retained?
STORE → Where/how is it represented?
RETRIEVE → What matters for this task?
UPDATE → Has an old fact become obsolete?
FORGET → What should no longer be retained?
If you get those right, you can swap Mem0 ↔ Zep ↔ your own Postgres/graph implementation later without rewriting the agent.
If you tell me what kind of agent you're building (coding agent, personal assistant, customer-support agent, autonomous research agent, etc.) and your stack (Python/TypeScript, LangGraph/OpenAI/Anthropic/etc.), I can give you a concrete memory architecture and implementation, including the schema, retrieval flow, and when memories should be written.
If you're building an agent that needs to remember users and accumulate knowledge over weeks/months, I’d separate short-term working state from long-term memory rather than trying to make one database do both.
Short-term memory should answer "what is happening right now?" Long-term memory should answer "what have I learned about this user/world over time?"
LangGraph, for example, explicitly distinguishes persistence of interactions within and across threads, while Letta takes a more agent-centric approach with persistent memory blocks.
Frameworks I'd seriously consider
Framework
Architecture
Particularly useful when
LangGraph + LangMem
Agent state + semantic/episodic/procedural memory
You're already using LangChain/LangGraph
Mem0
Memory extraction/retrieval layer
You want the easiest drop-in long-term memory
Zep / Graphiti
Temporal knowledge graph
Facts change and when something was true matters
Letta
The important distinction is that these aren't merely competing vector databases. Letta is much closer to an agent runtime, whereas Mem0 is a memory layer and Graphiti is a temporal graph framework.
My default choice
For a new system, I'd build around:
LangGraph (or your existing agent runtime) + a dedicated long-term memory layer.
If you're already invested in LangGraph, LangMem is the most natural starting point because it is designed to provide semantic, episodic, and procedural long-term memory alongside LangGraph's state/persistence model.
If you aren't committed to LangGraph, I'd evaluate Mem0 vs. Zep/Graphiti first:
Mem0 → "Remember that Alice prefers concise answers and works in finance."
Zep/Graphiti → "Alice worked at Company A until March, then moved to Company B; the project ownership changed with her."
Letta → "The agent itself decides what information belongs in its persistent memory and manages what stays in its active context."
Graphiti is particularly interesting when your memory contains changing relationships and historical state: it explicitly models temporal validity and combines semantic, keyword, and graph retrieval. help.getzep.com Letta's memory blocks, meanwhile, are persistent sections of context that remain directly available to the agent, which makes them useful for things like persona, user preferences, and working state.
One architectural detail I'd strongly recommend
Don't automatically dump every conversation into "long-term memory."
Instead:
Every interaction
│
▼
Short-term state
│
├── relevant only to current task ──► discard
│
└── potentially durable information
│
▼
Memory extraction
│
┌───────────┼───────────┐
▼ ▼ ▼
Semantic Episodic Procedural
"facts" "events" "how to"
│ │ │
└───────────┼───────────┘
▼
Long-term store
│
▼
Retrieval at runtime
That separation is important because storing information isn't the same as remembering it. Production systems have to deal with stale facts, contradictions, deletion, scope, and deciding when a memory is actually worth retrieving.
And I would not make the memory framework responsible for the agent's entire architecture unless you specifically want that coupling.
One caveat: current vendor benchmark comparisons are difficult to interpret because results often use different models, datasets, and evaluation harnesses. Recent comparisons explicitly warn against treating published LongMemEval numbers as a universal ranking.
If you tell me what your agent actually does (coding agent, personal assistant, customer support, research agent, etc.) and your stack (Python/TypeScript + LangGraph/CrewAI/custom), I can give you a concrete production architecture—including the database choices, memory schema, write/retrieval pipeline, and which framework I'd use.
There is no single "magic bullet" framework, but the ecosystem has largely converged on specialized memory layers paired with orchestration frameworks . Short-term memory handles immediate conversational context (the active context window), while long-term memory manages persistent facts, preferences, and historical traversal across sessions.
The top frameworks for integrating both short-term and long-term memory depend on your exact architectural preference:
1. Dedicated Memory Layers (Best to plug into anything)
Mem0: Operates as a dedicated memory intelligence layer designed specifically for AI agents. It automatically extracts user preferences and salient facts from interactions, organizing them across user, session, and agent scopes using vector search and metadata filtering. It pairs exceptionally well with orchestrators like LangGraph.
Zep: A conversational memory store optimized for long-term persistence. It handles asynchronous fact and entity extraction, progressive summarization of older dialogue, and semantic/temporal search, reducing short-term window bloat.
2. OS-Inspired & Tiered Architectures (Best for autonomous agent control)
Letta (formerly MemGPT): Treats the LLM like an operating system with tiered memory management. It gives the agent explicit self-editing tools to manage its own memory hierarchy (moving things between core short-term blocks, archival long-term storage, and external recall storage) dynamically during execution.
3. Orchestrator-Native Memory (Best if building entirely within a single ecosystem)
LangGraph + LangMem: If you are building stateful, graph-based multi-agent flows, LangGraph manages short-term conversational checkpointing natively, while tools like LangMem or Mem0 integration handle background reflection and long-term semantic storage across threads.
LlamaIndex: Excellent if your long-term memory requires bridging conversational history with heavy document/knowledge base retrieval (RAG).
Recommendation on How to Choose
Choose Mem0 or Zep if you want an out-of-the-box API/database layer that quietly extracts and injects memories into your existing agent setup without forcing you into a rigid state machine.
Choose Letta if you want your agent to actively control and decide what it remembers, forgets, or updates via function-calling mimicking human memory consolidation.
To help narrow down the ideal choice, tell me:
What orchestration framework (if any) are you currently using (e.g., LangChain, CrewAI, raw API calls)?
Does your agent need to remember things across different users (multi-tenant) or just a single user over time?
LangGraph (paired with a persistence layer like Redis or a vector database) is currently the best and most robust framework for integrating both short-term and long-term memory for AI agents.
While libraries like LangChain offer basic memory wrappers, LangGraph provides native graph-based state management that cleanly separates ephemeral execution state from persistent storage.
Memory Architecture Breakdown
Short-Term Memory (Context/State): Handled via LangGraph's built-in checkpointer feature, which automatically saves the exact state, message history, and tool outputs of a running thread after every super-step.
Long-Term Memory (Semantic/Recall): Handled by plugging an external database—such as Redis Agent Memory Server or a vector store like Pinecone or Weaviate—into the graph nodes to perform semantic similarity searches across past sessions.
Recommended Hybrid Memory Stack
Memory Tier
Purpose
Recommended Tool / Tech
Short-Term
Current dialogue turns, active tool outputs, scratchpad variables.
LangGraph Checkpointer + Redis / In-Memory Cache
Long-Term (Semantic)
Past interactions, historical facts, and learned preferences.
Vector Database (Pinecone, Weaviate, or Qdrant)
Long-Term (Factual)
How to Implement It
Define State: Set up your agent state schema in LangGraph to track standard message arrays for immediate conversation.
Add Checkpointing: Attach a memory saver (like MemorySaver or a Redis checkpointer) to your compiled graph to handle short-term persistence across turns.
Create Memory Nodes: Add specific custom nodes in your graph workflow that trigger background extractions—saving important user facts or summaries to your vector/relational database when a session concludes or reaches a milestone.
If you'd like, let me know:
What LLM or orchestration framework you are currently using
Whether your agent is single-agent or multi-agent
I can provide a concrete code architecture or boilerplate for your setup.
For a production AI agent, the strongest pattern is usually not choosing one “memory framework”, but designing a memory architecture where short-term state, long-term memory, retrieval, and consolidation are separate layers. Modern agent architectures generally distinguish session memory from persistent memory because they have different lifecycles and storage needs.
A practical stack looks like this:
User request
|
v
+--------------------+
| Agent Runtime |
| (LangGraph/etc.) |
+--------------------+
|
+----------------+----------------+
| |
v v
Short-term memory Long-term memory
(working state) (persistent)
| |
Redis/Postgres Vector DB + Graph DB
Conversation state Facts, experiences,
Tool outputs preferences, skills
Scratchpad Past outcomes
|
v
Context builder
(select what enters prompt)
Recommended architecture
1. Short-term memory: session state
Use this for things that matter during the current task:
conversation history
current plan
tool outputs
intermediate reasoning artifacts
active goals
Good options:
LangGraph state/checkpointing
Redis
PostgreSQL
SQLite for local agents
The key idea: do not dump the whole chat history into every prompt. Summarize, compress, and maintain a working state. Microsoft’s agent architecture guidance similarly separates session-scoped memory from persistent memory.
2. Long-term memory: persistent knowledge
Split it into categories:
Memory type
Example
Storage
Semantic
"User prefers Python"
Vector DB + metadata
Episodic
"Last month we fixed deployment issue X"
Event store/vector DB
Procedural
"When debugging API failures, check logs first"
Rules/workflow store
Entity memory
"Customer A owns project B"
A common mistake is storing raw transcripts forever. Better systems extract useful memories and store compressed representations.
Framework choices
Option A — LangGraph + LangMem (good general choice)
Best when you are already building agent workflows.
Vector search alone often struggles with relationships and changing facts. Hybrid retrieval (vector + graph) is increasingly common.
My default recommendation in 2026
For most serious agents:
Agent Framework:
LangGraph
Short-term:
PostgreSQL checkpointing
+ Redis cache
Long-term:
Mem0 or LangMem
+
PostgreSQL for structured facts
+
pgvector / vector DB for semantic recall
Advanced:
Graph database for entities and relationships
Memory flow:
User talks to agent
Conversation goes into short-term state
At checkpoints:
extract important facts
remove duplicates
update existing memories
assign confidence and timestamps
Before future responses:
retrieve only relevant memories
inject into working context
Avoid these common mistakes
❌ “Store every conversation embedding”
You get a noisy memory dump.
❌ “Put all memories into the system prompt”
The agent becomes slower and less accurate.
❌ “Use only a vector database”
You lose relationships, updates, and structured facts.
❌ “Never forget anything”
Good memory requires decay, correction, and deletion policies.
If I were building a new long-lived agent today, I would start with LangGraph + PostgreSQL + pgvector + Mem0/LangMem, then add graph memory only when the domain requires relationship reasoning.
Choosing the "best" framework for integrating short-term and long-term memory depends heavily on your architecture and whether you want a memory-first runtime or a modular memory layer you can plug into an existing orchestrator.
The top dedicated frameworks managing both short-term (context windows, active session state) and long-term (vector/graph persistence, user profiles, self-editing) memory break down as follows:
Letta (formerly MemGPT) : Best if you want a memory-first agent runtime inspired by operating system virtual memory. It explicitly splits memory into core memory (short-term/active working memory that the agent can read and self-edit) and archival memory (long-term storage/recall). It is ideal for agents that need to manage extremely long running contexts or perform autonomous background processing.
Mem0 : Best if you need a lightweight, plug-and-play memory layer across multiple sessions. Mem0 acts as an intelligent memory SDK that extracts facts and user preferences from interactions, seamlessly updating both short-term context and a persistent long-term storage backend without forcing you into a specific agent orchestration framework.
Zep (with Graphiti) : Best for applications requiring deep temporal and relationship tracking . Zep excels at maintaining user history and long-term knowledge graphs, scoring very high on complex multi-session evaluation benchmarks by tracking how facts change over time (bi-temporality).
LangGraph (with LangMem) : Best if you are already building your agent workflows using LangChain . LangGraph handles robust short-term state management out-of-the-box, while LangMem provides native primitives for optimizing and updating long-term memory across threads.
Quick Recommendation
Choose Letta if your agent needs to autonomously manage and edit its own memory space.
Choose Mem0 if you want to drop a smart memory service into an existing stack with minimal friction.
Choose Zep if you need robust knowledge graphs and complex historical fact tracking.
To help narrow down the ideal choice, tell me:
What orchestration framework (if any) are you currently using (e.g., LangGraph, CrewAI, custom Python)?
Is the long-term memory meant for a single user (personalized assistant) or a shared multi-user environment?
If you’re building an agent that needs both conversation continuity and durable knowledge, I’d avoid treating “memory” as one database. The strongest architecture is a two-tier memory system with explicit promotion from short-term state into long-term memory.
Long-term memory = what is worth remembering later.
LangGraph explicitly separates these concepts: its checkpointer persists state within a conversation/thread, while its Store is intended for cross-thread long-term memory.
My default stack
For a new production agent, I'd start with:
LangGraph + LangMem + Postgres, and only add a specialized memory service if your requirements justify it.
LangGraph gives you durable execution and thread-level state. LangMem adds memory extraction, search, updating, deletion and consolidation, including both agent-driven “hot path” memory and background extraction.
This is particularly valuable because a future agent can reason from what happened, rather than merely retrieving isolated facts.
4. Entity/relationship memory
Useful when your agent operates over people, organizations, projects, products, etc.:
Alice
├── works_at → Acme
├── manages → Project X
└── prefers → weekly reports
Project X
├── owned_by → Acme
└── depends_on → Service Y
This is where a temporal knowledge graph such as Zep/Graphiti can become attractive. Zep's architecture specifically focuses on relationships and facts whose validity changes over time.
The critical part: memory retrieval
Don't inject the user's entire memory database into every prompt.
Instead:
current query
↓
memory retrieval
↓
filter by:
relevance
recency
confidence
scope
permissions
↓
top N memories
↓
agent context
I'd also use different retrieval mechanisms for different memory types:
Memory
Retrieval
Current conversation
direct state
User preferences
semantic/key lookup
Past episodes
semantic + recency
Entities
structured lookup
Relationships
graph traversal
Current facts
authoritative DB/API
That last distinction is important: your memory system should not become a shadow database.
If the agent needs today's account balance, query the banking system. Don't rely on something the agent remembered three weeks ago.
Memory writes should be selective
A useful rule is:
Read memory frequently; write memory conservatively.
For example, after a conversation:
User: I switched our production database to Postgres.
Memory manager:
→ existing DB = MySQL
→ conflict detected
→ update database preference/state
→ preserve historical value
Don't simply create:
"My database is MySQL"
"My database is Postgres"
because you'll eventually retrieve contradictory memories.
Temporal memory becomes especially useful in systems where facts evolve.
Where Mem0, Zep and Letta fit
I wouldn't think of these as simply competing databases.
Mem0 is attractive when you already have an agent/application and want to bolt on persistent memory without replacing your runtime.
Zep/Graphiti makes more sense when your agent needs rich relationships and time-aware facts—for example CRM, support, project-management or relationship-heavy agents.
Letta takes a different approach: memory is much more deeply integrated into the agent runtime, with the agent actively managing what is in working context versus archival memory. That's compelling for genuinely long-running autonomous agents.
So my decision tree would be:
Already using LangGraph?
│
yes
↓
LangMem + Store
│
├── Need sophisticated temporal relationships?
│ ↓
│ Consider Zep/Graphiti
│
└── Need a specialized managed memory layer?
↓
Mem0
Building an autonomous agent whose
memory management is part of its identity?
↓
Consider Letta
One thing I'd add regardless of framework
Create a Memory Service interface in your application:
In short: use the agent framework's state/checkpointing for short-term memory, a persistent Store for durable memory, and a dedicated memory-management layer for deciding what deserves to survive. LangMem is particularly clean if you're already in LangGraph because those pieces are designed to work together.
If you tell me what you're building the agent with—Python/TypeScript, LangGraph, OpenAI Agents SDK, CrewAI, custom loop, etc.—I can give you a concrete production architecture and schema for it.
Choosing the "best" memory framework depends heavily on your existing stack, architectural style, and whether you need simple user personalization or deep, autonomous state management.
The top dedicated memory frameworks stand out for blending short-term context with persistent long-term storage:
Mem0 — Best for Quick Integration & Personalization
How it works: Acts as a smart, drop-in memory layer that automatically extracts, updates, and retrieves user preferences and historical context across sessions.
Why choose it: It has massive community adoption, low friction to install, and works across various LLM providers without forcing you to refactor your core agent architecture.
Check it out: Explore the project via Mem0 on GitHub.
Zep (with Graphiti) — Best for Enterprise & Temporal Reasoning
How it works: Uses temporal knowledge graphs to track how facts, entities, and relationships evolve over time. It manages fast short-term conversational buffers alongside deep long-term storage.
Why choose it: Scores exceptionally high on complex memory benchmarks (like LongMemEval) and handles enterprise requirements like data governance, latency optimization, and multi-tenancy.
Check it out: Learn more at Zep AI.
Letta (formerly MemGPT) — Best for Autonomous, OS-style Memory Management
How it works: Treats the LLM like an operating system, giving the agent explicit "system calls" to manage its own memory hierarchy (moving things between core working memory and archival long-term storage during execution).
Why choose it: Ideal if your agent needs self-editing memory, long-running autonomy, or background "sleep-time" compute to consolidate memories.
Check it out: Read the docs via Letta Documentation.
LangMem — Best for LangGraph Ecosystems
How it works: Built natively by the LangChain team specifically to handle cross-session persistence and memory optimization inside LangGraph workflows.
Why choose it: If you are already orchestrating your multi-agent system using LangGraph, LangMem natively replaces older, deprecated memory patterns with clean background reflection and extraction.
To help narrow down the ideal choice for your project, tell me:
What orchestration framework or stack are you currently building your agent with (e.g., LangGraph, CrewAI, custom Python)?
Does your agent need to manage multiple users/tenants , or is it a single-user assistant?
Long-term memory: LangMem — extracts durable facts, preferences, summaries, and learned information from conversations.
Storage: Postgres (with vector search/pgvector if appropriate).
Retrieval: selectively inject only relevant long-term memories into the agent's context.
Background consolidation: periodically merge, update, or invalidate memories rather than blindly appending every conversation.
LangGraph explicitly separates durable execution/state from its long-term Store, while LangMem provides both agent-managed memory tools and background memory extraction/consolidation.
A useful mental model is:
┌─────────────────────┐
│ Agent │
└──────────┬──────────┘
│
┌──────────────┴──────────────┐
│ │
Short-term state Long-term memory
"What am I doing?" "What do I know?"
│ │
LangGraph Checkpointer LangMem
│ │
Postgres Postgres + vector
│ │
┌───────┴────────┐ ┌──────┴─────────┐
│ thread/session │ │ user/agent/org │
│ state │ │ memories │
└────────────────┘ └────────────────┘
The important distinction is that short-term memory is state, whereas long-term memory is knowledge that has been promoted from previous interactions. LangMem's documentation makes the same distinction: long-term memory requires deciding what is worth remembering, when to create it, and where/how it should be retrieved.
When I'd choose something else
Your situation
I'd use
Already building with LangGraph
LangGraph + LangMem
Want the easiest memory layer to bolt onto an existing agent
Mem0
Agent needs to track facts changing over time
Zep/Graphiti
Building a genuinely autonomous, persistent agent
Mem0 is particularly attractive when you don't want to change your agent architecture: it's essentially a persistent memory layer you call from your existing application. Letta, by contrast, is much closer to an entire stateful-agent runtime, with the agent actively managing different levels of memory.
Zep/Graphiti becomes compelling when temporal relationships matter: e.g. "Alice was VP of Sales in 2024, but became CTO in 2025" rather than simply storing "Alice is VP of Sales."
The architecture I'd actually build
I'd use four memory layers, not just "chat history + vector DB":
Working memory
Current conversation
Current task/plan
Tool outputs
Ephemeral state
Episodic memory
Important past interactions
Decisions
Completed tasks
Significant events
Semantic memory
User preferences
Stable facts
Relationships
Learned domain knowledge
Procedural memory
"When doing X, use Y"
Learned workflows
Agent-specific strategies/instructions
Then give the agent two operations:
search_memory(query)
save_memory(memory)
But don't let every piece of conversation become a memory. A background process should determine whether something is:
This is where newer memory frameworks are substantially better than simply dumping conversation embeddings into a vector database. LangMem, for example, supports both explicit "hot path" memory management and background extraction/consolidation.
One important warning
Don't use long-term memory for information that should actually live in a source-of-truth database.
For example:
"The customer's subscription expires September 30."
If that's billing data, query Stripe/your DB. Don't rely on an LLM memory saying the date is September 30.
Memory should answer questions like:
"The customer prefers email over phone."
or
"Last time we discussed this project, the customer decided to postpone the migration."
That distinction dramatically improves reliability.
So, if you're starting from scratch, my pick is: LangGraph + LangMem + Postgres, with a deliberately designed memory schema and retrieval policy. It's more controllable than handing the whole agent to a memory runtime, while avoiding the work of inventing your own extraction/consolidation system.
If you tell me what your agent is built with (OpenAI Agents SDK, LangGraph, CrewAI, Claude, custom Python/TypeScript, etc.), I can give you a concrete production architecture and code for the short-term + long-term memory layer.