Google AI ModeSep 29, 2026
Data as of Oct 5, 2026Based on 54 AI responses from ChatGPT Search and Google AI Mode
Reviewed by Dimitry Apollonsky ·
PostgreSQL is named in 9 of the latest 10 answers. Mem0 follows, in 6.
Recommended in · Sep 13 – Sep 29, 2026
Google AI ModeSep 29, 2026
ChatGPT SearchSep 29, 2026
ChatGPT SearchSep 25, 2026
Google AI ModeSep 25, 2026
ChatGPT SearchSep 21, 2026
Google AI ModeSep 21, 2026
ChatGPT SearchSep 17, 2026
ChatGPT SearchSep 13, 2026
48% of citations to these sources link to brands' own websites.
redis.io
getzep.com
machinelearningmastery.com
mem0.ai
redis.io
getzep.com
Giving an AI agent persistent long-term memory across sessions requires shifting from a purely stateless request-response model to a tiered memory architecture . Because LLMs have a fixed context window and forget everything once a session ends, you must offload persistence to an external storage layer and give the agent mechanisms to read, write, and compress its own history.
If you'd like to dive deeper, let me know:
I can provide a tailored code architecture or setup guide for your specific scenario.
The key idea is: don’t try to make the model itself remember. Make memory an external, persistent system that your agent reads from and writes to.
A solid architecture looks like this:
┌─────────────────────┐
User ────────────►│ Agent / LLM │
└──────────┬──────────┘
│
retrieve relevant memories
▼
┌─────────────────────┐
│ Memory Layer │
├─────────────────────┤
│ Semantic memories │
│ Episodic memories │
│ User preferences │
│ Facts / entities │
└──────────┬──────────┘
│
persistent DBShort-term memory is the current conversation/session: recent messages, tool calls, intermediate state, etc.
Long-term memory is information that should survive the session:
OpenAI's current APIs can persist conversation/session state, but that's not quite the same thing as a carefully designed long-term memory system. The Conversations API, for example, can persist conversation items across sessions and devices.
A simple production schema might be:
memories
---------
id
user_id
type
content
importance
created_at
updated_at
expires_at
source
embeddingWhere type might be:
semantic → facts/preferences
episodic → things that happened
procedural → how the user/system likes something doneYou don't necessarily need a dedicated vector database. PostgreSQL + pgvector is often enough to start.
This is one of the biggest mistakes.
Instead, after a conversation/task, have a memory-extraction step:
Conversation
│
▼
Memory extractor
│
├── Is this actually useful later?
├── Is it a fact or preference?
├── Is it temporary?
├── Does it contradict an existing memory?
└── How important is it?
│
▼
Persistent storeFor example:
User: "I'm using TypeScript for this project from now on."
Store:
{
"type": "semantic",
"content": "User prefers TypeScript for this project.",
"importance": 0.8
}But don't necessarily store:
"User said hello."
At the beginning of a new session:
query = current_user_message
memories = memory_store.search(
user_id=user_id,
query=query,
limit=8
)
context = build_memory_context(memories)
response = llm(
system_prompt + context,
user_message
)The important part is selective retrieval.
Don't dump 50,000 memories into the context window. Search for memories relevant to the current task and provide the model with only those.
This "retrieve relevant memory → reason → write new memory" pattern is a common architecture for persistent agents.
For a genuinely long-lived agent, I'd use two layers:
┌──────────────┐
│ Raw episodes │
└──────┬───────┘
│
consolidation
│
▼
┌──────────────────┐
│ Semantic memory │
└──────────────────┘Suppose over three months the agent observes:
User chose PostgreSQL.
User rejected MongoDB.
User repeatedly uses SQL migrations.
User's production environment is AWS RDS.Instead of keeping all four observations forever, consolidation can produce:
User's primary database technology is PostgreSQL,
deployed on AWS RDS.You retain the original episodes for provenance, while the semantic memory becomes the thing you normally retrieve.
This is extremely important if multiple users exist.
A memory should generally have a scope such as:
global
organization
project
user
conversationFor example:
user:123
project:website-redesign
"Uses Tailwind CSS"should not accidentally become:
global
"Everyone uses Tailwind CSS"Memory systems should also have explicit lifecycle, retention, provenance, and isolation rules; this is increasingly treated as an architectural concern rather than simply "put embeddings in a vector DB."
There are actually three different levels you can use.
Conversation continuity: use the Conversations API / Responses API when you mainly want the agent to continue an existing conversation. OpenAI describes Conversations as durable objects that can be reused across sessions, devices, or jobs.
Agent sessions: the Agents SDK provides sessions that maintain conversation state over time.
True long-term memory: add your own persistent memory layer—or use the newer sandbox memory capabilities where appropriate. OpenAI's sandbox-agent memory, for example, distinguishes session history from distilled memories that can be reused in later runs.
A good production architecture is therefore:
┌────────────────────┐
│ User │
└─────────┬──────────┘
▼
┌────────────────────┐
│ Agent │
└─────────┬──────────┘
│
┌─────────────┴─────────────┐
▼ ▼
Conversation/session Long-term memory
state service
│ │
▼ ┌──────┴──────┐
recent context │ │
▼ ▼
semantic episodic
memories memories
│ │
└──────┬──────┘
▼
PostgreSQL/vector DBDon't over-engineer it:
memories tableThat gives you a real persistent-memory system without requiring a complicated agent framework.
If you tell me what you're building it with—Python, TypeScript/Node, LangGraph, OpenAI Agents SDK, Claude, CrewAI, etc.—I can show you a concrete implementation, including the database schema and the memory read/write loop.
Yes. The key is to separate session context from long-term memory.
A good architecture looks like this:
User
↓
New session
↓
Retrieve relevant long-term memories
↓
Agent + current conversation + retrieved memories
↓
Response
↓
Extract durable facts/preferences/decisions
↓
Persist them in a databaseInstead, extract things that are likely to remain useful:
pnpm test.”This lets you retrieve a few relevant memories rather than dumping months of conversation into every prompt.
For example:
user_id = "usr_12345"Store durable memories against that ID:
{
"user_id": "usr_12345",
"memory": "Prefers TypeScript over JavaScript",
"type": "preference",
"confidence": 0.95
}When the user starts a new session, use the same user_id to retrieve relevant memories. This is the fundamental mechanism behind cross-session persistence.
Use something like:
user_id = usr_12345 ← permanent identity
session_id = ses_987 ← current conversationSession memory contains things such as:
“We're currently debugging the authentication bug.”
Long-term memory contains:
“User's application uses OAuth.”
Don't automatically promote everything from a session into permanent memory. Define what deserves retention and what should expire.
Your memory layer essentially needs:
memory.add(user_id, information)
memory.search(user_id, query)
memory.update(memory_id, new_information)
memory.delete(memory_id)Then your agent loop becomes:
def run_agent(user_id, session_id, message):
memories = memory.search(
user_id=user_id,
query=message
)
context = {
"current_message": message,
"relevant_memories": memories
}
response = agent.run(context)
new_memories = extract_durable_facts(
message,
response
)
for m in new_memories:
memory.add(user_id, m)
return responseThe important part is that memory retrieval happens before generation and memory extraction happens after the interaction.
You have several approaches:
| Approach | Good for |
|---|---|
| PostgreSQL | Structured facts, preferences, relationships |
| Vector database | Semantic retrieval |
| Graph database | Relationships between people/projects/entities |
| Files/object storage | Agent-generated summaries and workspace knowledge |
| Dedicated memory layer | Getting all of the above without building it yourself |
A common production architecture combines structured metadata + semantic/vector retrieval.
For example:
PostgreSQL
├── users
├── sessions
└── memory_metadata
Vector store
└── embeddings of durable memoriesIf you're using OpenAI's current agent infrastructure, there are now several levels of persistence.
The Agents API can maintain session state across turns, so you can continue work without reconstructing the conversation yourself.
For sandbox agents, OpenAI also has a distinct memory mechanism that distills useful lessons from previous runs rather than merely replaying conversation history.
For provider-independent long-term user memory, systems such as Mem0 provide persistent memory with user-scoped retrieval and operations for adding, searching, updating, and deleting memories.
For a typical personal AI agent, I'd use:
┌─────────────────┐
│ User / App │
└────────┬────────┘
│
stable user_id
│
┌────────▼────────┐
│ Memory Layer │
│ │
│ semantic facts │
│ preferences │
│ past decisions │
│ useful episodes │
└────────┬────────┘
│
relevant memories
│
┌──────────────▼──────────────┐
│ Agent │
│ │
│ system instructions │
│ + memory │
│ + current conversation │
│ + tools │
└──────────────┬──────────────┘
│
response
│
┌────────▼────────┐
│ Memory extractor │
└────────┬────────┘
│
durable facts
│
Memory LayerThe biggest design principle: don't think of long-term memory as "saving the conversation." Think of it as maintaining a continually updated knowledge base about the user, their projects, and what the agent has learned.
If you're building this with Python/TypeScript + OpenAI, I can also give you a concrete production-ready implementation (database schema + retrieval + memory extraction + code) using either Postgres/pgvector or a dedicated memory layer.
You give an AI agent persistent long-term memory by decoupling storage from the LLM's context window , using an external database or memory layer (such as Mem0 or Zep ) to extract, store, and selectively retrieve relevant facts at runtime.
Appending unlimited past chat logs directly into prompts causes high latency, token bloat, and context degradation. Instead, implement a multi-layered architecture that manages state across sessions efficiently.
To see how automated archiving and context injection work in practice, watch this overview on setting up an agent memory bank:
If you can share which framework or LLM orchestration tool you are using (e.g., LangChain, LlamaIndex, custom Python), I can provide a specific code snippet to hook up your long-term memory layer.
The key idea is: don’t try to make the model itself remember. Give the agent an external memory layer, then retrieve the right memories into each new session. LLM calls are effectively stateless unless your application supplies prior state.
┌──────────────────────┐
│ AI Agent │
│ LLM + tools + logic │
└──────────┬───────────┘
│
┌───────────▼───────────┐
│ Memory Controller │
│ │
│ retrieve → inject │
│ extract → store │
│ update → consolidate │
└───────┬───────┬───────┘
│ │
┌──────────▼─┐ ┌─▼────────────┐
│ PostgreSQL │ │ Vector index │
│ source of │ │ semantic │
│ truth │ │ retrieval │
└─────────────┘ └──────────────┘A good production design usually separates working, episodic, semantic, and sometimes procedural memory rather than dumping the entire conversation into a vector database.
For example:
{
"user_id": "user_123",
"type": "semantic",
"content": "User prefers concise technical explanations.",
"source": "conversation_8472",
"created_at": "2026-09-21T08:00:00Z",
"confidence": 0.94
}You might maintain:
The distinction matters because “what happened?” and “what is generally true?” are different retrieval problems.
When a new session starts:
query = current_user_message
memories = memory.search(
user_id=user_id,
query=query,
limit=10
)
context = format_memories(memories)
response = agent.run(
user_message,
memory_context=context
)Don't retrieve everything. Retrieve memories based on semantic relevance + recency + importance, and impose a token budget. This prevents your "memory" from eventually becoming an enormous prompt that hurts rather than helps the agent.
After the agent finishes:
Conversation
↓
Memory extractor
↓
"Is anything here worth remembering?"
↓
┌───────────────┬────────────────┐
│ New fact │ Update existing│
│ │ fact │
└───────────────┴────────────────┘
↓
persistent storeFor example, if the conversation contains:
“I switched our backend from MongoDB to PostgreSQL.” the memory system could create/update:
project.backend_database = PostgreSQLrather than storing the entire conversation forever.
This extract → deduplicate → update/consolidate pipeline is important. Simply appending every conversation to a vector store eventually produces duplicates, stale information, and contradictory memories.
You don't necessarily need a specialized "AI memory database."
A very reasonable starting stack is:
PostgreSQL
├── users
├── sessions
├── messages
├── memories
└── memory_versions
+ pgvector
└── embeddings for semantic retrievalFor a memory record:
CREATE TABLE memories (
id UUID PRIMARY KEY,
user_id UUID NOT NULL,
type TEXT NOT NULL,
content TEXT NOT NULL,
embedding VECTOR(1536),
importance FLOAT,
confidence FLOAT,
created_at TIMESTAMPTZ,
last_accessed_at TIMESTAMPTZ,
expires_at TIMESTAMPTZ,
source_session UUID
);Then retrieve something like:
SELECT *
FROM memories
WHERE user_id = $1
ORDER BY embedding <=> $2
LIMIT 10;You can later add keyword search, metadata filtering, recency weighting, and reranking.
A vector database is useful, but vector similarity should not be your only mechanism. Structured facts are often better represented as ordinary database records, while episodic memories benefit from semantic retrieval.
Suppose the agent remembers:
User's database = MongoDBSix months later:
User's database = PostgreSQLYou don't want two equally authoritative memories.
Instead, maintain versions:
MongoDB
↓ superseded
PostgreSQL
↓ currentOr update the existing semantic fact while retaining the historical event.
For high-value systems, temporal/versioned memory is especially useful because you can answer both:
“What does the user use now?” and:
“What did they use last year?” Recent research on agent memory is increasingly treating revision, forgetting, provenance, and temporal state as first-class problems rather than merely storing embeddings.
If you're building your first version, I'd start with just this:
EACH SESSION
User message
│
▼
Retrieve top 5–10 relevant memories
│
▼
Add memories to agent context
│
▼
Run agent
│
▼
Extract candidate memories
│
▼
Deduplicate / update
│
▼
Persist to PostgreSQL + vector indexAnd give every memory:
user_id
memory_type
content
embedding
importance
confidence
created_at
updated_at
sourceThat gets you surprisingly far.
The important conceptual shift is that long-term memory isn't "a longer context window." It's a separate system that decides what is worth remembering, how it should be represented, when it should be recalled, and when old information should be revised or forgotten.
If you're using LangGraph, LangChain, OpenAI Agents SDK, CrewAI, AutoGen, or a custom Python/TypeScript agent, the exact implementation differs considerably; tell me which one you're using and I can show the concrete architecture and code.
Giving your AI agent persistent long-term memory across sessions requires moving beyond the standard short-term context window . This is typically achieved by coupling an external database/vector store with an extraction and retrieval architecture so the agent can remember, update, and recall past interactions.
user_id_123 -> favorite_color: blue, dietary_restrictions: none).{"learned_fact": "User is building a React Native app"}).text-embedding-3-small ), and write them to your vector store with metadata tags (user_id, timestamp).[Long-Term Memory Context: User previously mentioned they prefer TypeScript over JavaScript]).If you want to dive deeper, tell me:
I can provide a custom code snippet or architecture tailored to your stack.
Giving an AI agent persistent long-term memory across sessions requires moving beyond simple active chat history and implementing a three-tier memory architecture (Working memory → Session state → Long-term memory store).
Instead of writing custom pipelines from scratch for extraction, embedding, and deduplication, utilize proven production-ready memory layers:
You will need a combination of two databases:
When a session ends (or incrementally during a conversation), do not just dump the raw transcript into long-term storage—it bloats context windows and introduces noise. Run an extraction routine:
user_id metadata.At the start of a new session or during incoming turns, dynamically query the long-term memory store:
user_id.To help tailor this to your stack, tell me:
The key idea is: don’t try to make the model itself remember. Give the agent an external memory system that survives process/session restarts, and retrieve only the relevant memories into each new context.
┌─────────────────────┐
New session ─────►│ Memory Retriever │
└──────────┬──────────┘
│ relevant memories
▼
┌────────────────┐
│ LLM / Agent │
└───────┬────────┘
│
conversation
│
▼
┌─────────────────────┐
│ Memory Extractor │
│ + dedup + conflict │
└──────────┬──────────┘
│
▼
┌─────────────────────┐
│ Persistent Database │
└─────────────────────┘The important part is that memory has both a write path and a read path. A common production pattern is to save interaction episodes, extract durable facts/preferences, and periodically consolidate or prune them.
I recommend at least these three:
This distinction is useful because you don't want to retrieve a thousand old conversation snippets when a single structured fact will do.
For example:
{
"type": "semantic",
"user_id": "user_123",
"key": "preferred_language",
"value": "Python",
"confidence": 0.97,
"updated_at": "2026-09-17T08:00:00Z"
}versus:
{
"type": "episodic",
"user_id": "user_123",
"summary": "User and agent decided to use PostgreSQL for the project.",
"timestamp": "2026-09-16T21:14:00Z",
"embedding": "..."
}For a straightforward implementation, PostgreSQL + pgvector is a very good starting point:
PostgreSQL
├── users
├── memories
│ ├── id
│ ├── user_id
│ ├── type
│ ├── content
│ ├── importance
│ ├── created_at
│ ├── updated_at
│ └── embedding
└── memory_versionsThe relational database gives you durable structured data and metadata, while vector search lets you find memories by meaning rather than exact keywords. Hybrid architectures using structured storage plus vector retrieval are increasingly common.
You can also use a dedicated vector database such as Qdrant, Weaviate, or Chroma.
Suppose the user says:
"Can you continue the project we discussed last week?" Don't dump the entire history into the prompt.
Instead:
query = "project discussed last week"
memories = memory.search(
user_id=user_id,
query=query,
limit=8
)
context = build_memory_context(memories)
response = llm(
system_prompt=SYSTEM_PROMPT,
memory=context,
current_messages=messages
)Your retrieval function should generally combine:
semantic similarity
+ recency
+ importance
+ user/agent scope
+ metadata filtersrather than relying on vector similarity alone. Recent work specifically highlights that blindly injecting semantically similar memories can produce inappropriate or even unsafe context.
After a session, run a memory-extraction step:
Conversation
↓
"What should survive this session?"
↓
┌─────────────────────────────┐
│ User preference │
│ New fact │
│ Important decision │
│ Recurring procedure │
│ Temporary information │
│ Nothing worth remembering │
└─────────────────────────────┘
↓
Persistent memoryFor example:
User: I always use tabs, not spaces.
→ semantic memory:
coding_indentation = tabsBut:
User: What's the weather today?
→ don't create permanent memoryThis filtering is crucial. Otherwise your database becomes an enormous transcript rather than useful memory.
This is one of the hardest parts.
Suppose the agent remembers:
User works at Acme.Six months later:
User works at Globex.You don't want both facts injected indiscriminately.
Give memories versions/status:
Acme employment
valid_until: 2026-04-01
status: superseded
Globex employment
valid_from: 2026-04-01
status: activeFor more sophisticated systems, bitemporal/versioned memory models can preserve both when a fact was true and when the system learned it.
Over time:
1000 episodic memories
↓
consolidation
↓
50 useful semantic memoriesFor example:
Monday:
"I prefer dark mode."
Wednesday:
"Dark mode is easier on my eyes."
Friday:
"Please keep everything dark."
↓
Semantic memory:
"User strongly prefers dark-mode interfaces."You can run this asynchronously once a day or after a certain number of interactions. Consolidation, deduplication, and forgetting prevent long-running agents from accumulating an unmanageable memory store.
If you're building your first agent, you don't need a complicated cognitive architecture.
I'd start with:
PostgreSQL
│
┌─────────┴─────────┐
│ │
structured facts pgvector
│ │
└─────────┬─────────┘
│
Memory API
/ \
retrieve write
│ │
▼ ▼
Agent ←────── conversationAnd expose only two operations to your agent:
memories = memory.retrieve(
user_id,
query=current_task,
limit=10
)
memory.write(
user_id,
memories_to_keep
)Then progressively add importance scoring, deduplication, conflict resolution, expiration, and consolidation.
A useful rule of thumb is:
The context window is working memory; your database is long-term memory; retrieval is recall; consolidation is learning. That architecture lets the same agent come back tomorrow, next month, or after a process restart and still have access to information from previous sessions without putting the entire historical conversation into every prompt.
If you're building this with Python, LangGraph, OpenAI API, Claude, or a custom agent loop, I can also show you a concrete implementation with the database schema and read/write code.
The core pattern is external persistent storage + selective retrieval + memory updates. Don’t try to make the model itself remember.
user_id, agent, project, etc., and attach timestamps/provenance so facts can be corrected later.A simple architecture is:
User
↓
Agent
├── retrieve(query, user_id)
│ ↓
│ memory DB / vector index
│ ↓
│ relevant memories
│
├── LLM + current context + memories
│
└── memorize(session)
↓
extract → deduplicate → update
↓
persistent storeA Postgres + pgvector setup is often a good starting point: relational fields handle structured facts and metadata, while embeddings handle semantic retrieval. You can add a graph later if relationships become important.
The crucial design principle is “retrieve → act → reflect → memorize,” not “dump the entire history into the prompt.”
Giving an AI agent persistent, long-term memory across sessions requires moving away from treating the LLM context window as your sole storage . You need an external, durable storage layer combined with a systematic ingestion, indexing, and retrieval workflow.
The standard architectural blueprint to achieve this involves four core components:
user_id or recency) to fetch only what is relevant.Instead of building this infrastructure from scratch, you can leverage dedicated memory orchestration frameworks:
To help tailor this architecture, tell me: