Every morning the same ritual plays out. A customer opens a support chat, and the assistant asks for an account number, a plan tier, and a reason for contacting them. It has done this exact dance with the same person 40 times before. It has no idea, because nobody gave it a way to remember.
That’s the default design of retrieval-augmented generation, or RAG, and for a long time it was fine. AI memory layers are starting to change that. A RAG system is an archivist. It indexes your documents, fetches the chunks a question needs, and hands them to the model with no recollection of yesterday. Every session starts from a clean slate.
That clean slate is now under pressure from three directions. Meta’s researchers published work on trainable memory layers that sit inside the model itself. Startups raised real money to turn memory into infrastructure. Letta announced a $10 million seed round led by Felicis Ventures in April 2025, and Mem0 announced $24 million in October 2025. The model vendors shipped memory as a product too. At OpenAI’s DevDay in October 2025, Sam Altman announced 800 million weekly ChatGPT users. With less fanfare, he also announced a rebuilt memory system rolling out first to paid plans. Anthropic added memory to Claude the same year.
The most interesting story in AI right now is not a bigger context window or a higher benchmark score. It’s AI memory layers. Memory spent most of the RAG era as an afterthought, and now it’s a part of the stack that teams design for, buy, and operate. Retrieval no longer owns the knowledge problem by itself.
What follows is practical. First, the research lineage, from MemGPT to Meta’s memory layers. Then a way to sort a confusing vendor market into three categories, plus an honest look at where memory beats retrieval and where it loses. Last, five shifts worth planning for if you run a RAG system in production. Expect concrete names, published numbers where they exist, and skepticism about benchmark claims. There has been plenty to be skeptical about.
How AI Memory Layers Went From a Research Paper to a Product Line
The current wave started with an operating systems metaphor. In October 2023, a team at UC Berkeley published MemGPT, a paper that treats the context window like paged memory in an OS. The main context is RAM. Long-term storage is the disk. The model itself decides what to page in and out. It sounds like an academic exercise. In practice, it became the mental model behind Letta, the startup two of the paper’s authors founded. It also shaped the memory designs in most agent frameworks today.
Meta’s researchers then built AI memory layers into the weights themselves. Memory Layers at Scale, a paper they published in December 2024, describes a sparse, trainable key-value store. It sits inside the network alongside attention and feed-forward layers. The values update through gradient descent like any other parameter. A routing scheme keeps only a small fraction of them active per token, so latency stays low. On fact-based question answering, the memory-equipped models beat compute-matched mixture-of-experts baselines. Every RAG team eventually asks how much of the index the model should just know. This is the most direct answer anyone has published.
2025 turned the research into product. OpenAI and Anthropic shipped consumer memory. Letta and Mem0 raised funding. A July 2025 paper called MemOS went further. It proposed a full operating system that manages memory as a scheduled resource, moving facts between hot and cold storage the way an OS moves pages.
Strip the lineage down and one question ties it together. How does a system carry what it learned yesterday into today? Context memory, weight memory, and external stores are the three kinds of AI memory layers answering it. Here’s how to tell them apart.
Three Kinds of AI Memory Layers, One Confusing Market
Every company selling AI memory layers uses the same word for three different things. Telling them apart keeps you from buying the wrong layer twice.
Weight memory: layers inside the model
This is Meta’s approach. Knowledge lives in the parameters, learned during training or fine-tuning and retrieved at full speed during inference. It’s also the least controllable option on the list. You can’t query it, you can’t delete a single fact, and you can’t see why the model said what it said. For enterprise RAG teams, it’s a research thread today and a signal about where the model vendors are heading.
Context memory: the model runs its own workspace
This is MemGPT’s approach, commercialized by Letta. The agent gets a working memory block it can read, write, and summarize. The model decides what to keep and what to archive. Letta has also been experimenting with sleep-time compute, letting agents reorganize their memory while the user is away. The pattern suits long-running agents: research assistants, coding agents, support agents that need continuity across sessions. The trade-off is that the model is now the librarian, and models are sloppy librarians.
External memory: a store next to your index
Mem0, Zep, and the memory features OpenAI and Anthropic ship all live here. Facts get extracted from conversations, stored in a database, and retrieved when relevant. Zep’s open-source Graphiti builds a temporal knowledge graph, so every fact carries a validity interval and history survives updates. If you run a serious RAG deployment, this is the slice of the AI memory layers market to watch. It composes with retrieval instead of trying to replace it.
One warning before you evaluate anything in this space: the public benchmark claims are a mess. Mem0’s paper claims the strongest published results on LOCOMO, a long-term conversation benchmark from 2024 by Maharana and colleagues. In April 2025, Zep published a post called “Lies, Damn Lies, and Statistics” that pulled those claims apart. Letta published its own critique of the comparison setups. All three companies sell memory. Build an eval set from your own conversations before trusting any vendor number, including the ones in this post.
Where Memory Wins, and Where Retrieval Wins
The decision between retrieval and AI memory layers is less mysterious than vendor blogs make it look. It comes down to what kind of knowledge you’re storing.
Give memory the small, stable facts
The facts that make an assistant feel human are short, stable, and personal. Think of a user’s name, a preference for metric units, or the fact that this customer already rebooted the router twice. Retrieving that from a document index technically works, but it costs a pipeline hop and usually misses. Nobody writes “prefers tabs” into a wiki. Memory stores exist for exactly this class of knowledge. OpenAI’s rebuilt memory is designed to surface it from past chats without the user restating anything.
Latency matters here too. A memory lookup is one indexed query against a small store. A full RAG hop means embedding, search, reranking, and prompt assembly. On high-volume assistant traffic, pulling the hundred most recurring facts out of the retrieval path saves real milliseconds. It also cuts retrieval calls per session.
Keep documents in the index
Everything long, fresh, or legally sensitive stays in retrieval. Three reasons.
Citation. Enterprise buyers want sources, and the EU AI Act’s transparency rules make provenance a compliance question, not a nice-to-have. A retrieved chunk has a document ID and a paragraph number. A memory has, at best, a fuzzy trail back to some conversation from months ago.
Change. Contracts get renegotiated, prices get revised, policies get rewritten. An index handles that with a re-crawl. Most memory stores handle it badly, because consolidation overwrites instead of versioning. Temporal edges are the exception, which is why Graphiti is worth studying even if you buy something else. We’ve covered index drift on this blog before. Memory drifts faster, because consolidation runs automatically while nobody is watching.
Isolation. In a multi-tenant system, if user A’s facts can surface in user B’s answers, you have a privacy incident, not a bug. Vector indexes solved tenant filtering years ago. Most memory products are still working on it.
There’s a security angle too, and it’s ugly. Say an attacker gets text into your memory store, through a prompt injection or a poisoned document. They’ve now changed every future answer the system gives. Injection becomes persistent. Any memory you add needs write controls and provenance. Read controls alone won’t cut it.
5 Shifts to Plan For
For a team running production RAG, AI memory layers change five things.
-
Memory becomes a budget line, not a feature flag. Consolidation jobs, storage, lookups, and the eval work that keeps it honest all cost engineering time. Platform memory is bundled into consumer subscriptions. A self-hosted memory layer for an enterprise assistant is a real line item. Price it before it goes on a roadmap.
-
Split your knowledge into facts and documents. Use a blunt test. If a piece of knowledge fits in one sentence and rarely changes, it belongs in memory. If it runs longer than a paragraph and changes with any frequency, it belongs in the index. Account state, preferences, and glossary entries go to one side. Contracts, policies, and codebases go to the other.
-
Consolidation becomes a scheduled job with rules you wrote. What happens when a customer downgrades from Pro to Free, but old memories still say Pro? Someone has to decide the conflict policy, the retention window, and what counts as a contradiction. Temporal knowledge graphs make this tractable by storing validity intervals on every fact. Flat key-value stores make it invisible until a user complains.
-
Evaluation grows a memory leg. Your test suite now needs cases for stale facts, contradictions between memory and fresh documents, and cross-user leakage. It also needs to test whether the system forgets what it should. The LOCOMO fight is your warning not to outsource this. Build a small eval from your own chat logs and run it weekly.
-
Every memory needs a delete button and an audit trail. Under GDPR, inferred facts about a person count as personal data. Erasure requests reach your memory store, not just your user table. Store the source conversation ID with each extracted fact so you can explain where a statement came from. Make per-user deletion an API call you’ve tested yourself.
Where to Start
Pick one workload with high volume and recurring facts. Internal help desks and customer support agents are the usual suspects. Run a memory store beside your existing retrieval for 30 days. Let memory hold the stable facts and the index handle everything else. Measure three things: task completion without follow-up questions, the rate of stale or contradictory answers, and how many retrieval calls you avoided. If the numbers hold up, expand. If they don’t, you learned cheaply.
Don’t treat this as a rip-and-replace. AI memory layers don’t retire your index. They take the small, stable, personal knowledge out of it, so retrieval can stick to the documents it’s good at.
The assistant asking for that account number for the 41st time is still the default state of most RAG deployments. It won’t stay that way. Memory went from a Berkeley paper about paged context to a funded category to a checkbox in ChatGPT and Claude. All of that took about two years. Meta’s results say a slice of it will eventually live in the weights. The practical move today is external AI memory layers: a memory store next to your index. It holds the short, stable facts your retrieval path wastes time relearning every session.
If you’re sorting out what this means for your stack, we go deep on the adjacent problems too. Think RAG evaluation, index drift, and multi-tenant access control. Subscribe for weekly breakdowns like this one. And if your team is shipping a memory or retrieval feature that needs documentation people can follow, that’s the work we do at Rag About It.



