The capability boundary of a large model is defined by the context window, but 'able to accommodate' does not mean 'able to remember'. As soon as the conversation ends, the model no longer holds any state; Even if the window is long enough to cram all the history into prompt words, it means high computational costs and increasingly slow responses. More importantly, the truly useful memory is not "replay of the original text", but knowledge that has been refined, retrievable, and updated with use. This is the watershed between AI agents and chatbots: agents need to accumulate experience across tasks and conversations in order to evolve from "starting from scratch every time" to "understanding you more and more as you use them".
1、 The Three Forms of Memory
Drawing on cognitive science, researchers typically classify agent memory into three categories. Episodic memory records "what happened" - such as the user's last request or the execution process of a task; Semantic memory precipitates "what you know" - preferences, facts, and knowledge extracted from conversations; Procedural memory stores "how to do better" - workflow, strategy, and tool usage experience. The three types of memory correspond to different storage and retrieval methods, collectively forming the "long-term memory" of the agent.
2、 From MemGPT to Memory Framework
In 2023, researchers at the University of California, Berkeley proposed MemGPT, titled "MemGPT: Towards LLMs as Operating Systems," which introduces the "layered storage" idea of operating systems into the big model: treating the big context as "memory" and external storage as "hard disk," allowing the model to decide when to archive and retrieve information on its own. This idea later evolved into the open-source Letta project, becoming a representative practice in the field of Agent memory. At the same time, RAG (Retrieval Enhanced Generative) technology provides the infrastructure for memory: vector databases segment, embed, and index text, and agents retrieve the most relevant fragments when needed, rather than reading the entire history. The combination of memory and retrieval allows agents to have almost unlimited "background knowledge" at a limited cost.
3、 Memory Practice on the Product Side
Memory ability is moving from research to products. In 2024, OpenAI will add a cross session memory function to ChatGPT, which can remember user preferences and call them in subsequent conversations; Google's Gemini has also introduced similar capabilities. In China, various AI assistants and intelligent agent platforms also regard "long-term memory" as a differentiated selling point. However, memory also brings new problems: privacy and forgetting. Users should be able to view and delete the content remembered by the agent, and both product design and regulation need to leave space for the 'right to be forgotten'.
Conclusion
If the contextual window determines how far a large model can see at once, then the memory system determines how long it can accompany you. From the proposal of MemGPT to the follow-up of major products, memory is becoming the next main battlefield for Agent capabilities. For developers, understanding the classification and retrieval architecture of memory is the first step in building truly "useful and reliable" agents.
【 Reference source 】 MemGPT paper (University of California, Berkeley, 2023); OpenAI official blog (ChatGPT memory function introduction); Google I/O 2024 Gemini Memory Function Public Information