Quickstart · Architecture · Review guide
An interactive playground to explore how different memory strategies work in LLM applications. Built for developers who want to understand the tradeoffs before implementing memory in production.
When building chatbots or AI agents, you need to decide how to manage conversation history. This project implements 8 common approaches, with full transparency into what gets sent to the LLM at each turn.
Think of it as a debugging tool that shows you exactly what's happening inside the "memory" component.
| Area | Implementation signal |
|---|---|
| Product problem | Help developers choose the right memory strategy before building an agent or chatbot. |
| Interactive surface | Streamlit playground with side-by-side memory-state inspection. |
| Agent systems depth | Implements short-term, summary, vector, graph, hierarchical, and OS-inspired memory patterns. |
| Debuggability | Exposes the exact context sent to the model, plus retrieval scores, summaries, and extracted facts. |
Basic
- Full Memory - Keep everything (until you hit token limits)
- Sliding Window - Only keep last N messages
Smart Filtering
- Relevance Filter - Use embeddings to find similar past messages
- Summary Memory - Compress old messages into summaries
External Storage
- Vector Memory (RAG) - Store in vector DB, retrieve by similarity
- Knowledge Graph - Extract facts as triples, query the graph
Hybrid
- Hierarchical - Combine vector search + recent messages
- OS-Inspired - Simulate RAM/disk with explicit paging
# clone and setup
git clone https://github.com/yueyang0120/llm-agent-memory-patterns.git
cd llm-agent-memory-patterns
python3.10 -m venv .venv
source .venv/bin/activate
pip install -r requirements.txt
# add your openai key
cp env.example .env
# edit .env with your OPENAI_API_KEY
# run
streamlit run app/playground.py- Pick a strategy from the sidebar
- Chat with the bot
- Click "Internal Monologue & Memory State" to see:
- What context was sent to the LLM
- Which messages were kept/dropped
- Similarity scores, summaries, graph facts, etc.
| Strategy | Context selection | Main tradeoff |
|---|---|---|
| Full memory | All recorded messages | Prompt grows with the conversation |
| Sliding window | Last N messages | Older context is excluded |
| Relevance filter | Messages above similarity threshold | Can lose chronology |
| Summary | Generated summary + recent buffer | Lossy; summary size is not formally bounded |
| Vector | Top-k retrieved messages | Retrieval may miss useful context |
| Knowledge graph | Selected entities and relationships | Extraction can introduce incorrect facts |
| Hierarchical | Retrieved messages + recent window | Combines two retrieval policies |
| OS-inspired | Bounded active list + simulated archive | Simple recall rules; no durable paging |
These describe model context, not total storage complexity. Message length also affects token usage.
- Python 3.10+
- Streamlit (UI)
- LangChain (LLM wrapper)
- OpenAI GPT-4o-mini
- ChromaDB (vector store)
- NetworkX (graph)
flowchart TD
UI["Streamlit playground"] --> M["Selected BaseMemory implementation"]
M --> B["Full history / sliding window"]
M --> A["Relevance / summary"]
M --> E["Vector / graph"]
M --> H["Hierarchical / OS-inspired"]
A -.-> L["OpenAI · embeddings / summarization"]
E -.-> S["Chroma collection / NetworkX graph"]
H -.-> S
B --> C["get_context(query)"]
A --> C
E --> C
H --> C
C --> P["final_prompt → chat model"]
C --> D["debug_info → inspection panel"]
P --> U["add_message · update selected memory"]
U --> M
classDef out fill:#fce7f3,stroke:#db2777,color:#831843;
classDef storage fill:#e0f2fe,stroke:#0284c7,color:#0c4a6e;
class P,D out;
class L,S storage;
All strategies inherit from BaseMemory:
def add_message(role: str, content: str)
def get_context(query: str) -> dict # returns {final_prompt, debug_info}
def clear()Uses GPT-4o-mini with structured output to extract triples:
"I work at Google" → (user, work_at, google)Stores each message as an embedding in ChromaDB, retrieves top-k by similarity.
Combines two layers:
- Long-term: Vector search (top 2)
- Short-term: Sliding window (last 3)
Use Full Memory if:
- Conversations are short (<10 turns)
- You need perfect recall
Use Sliding Window if:
- Each message is independent
- You only care about recent context
Use Vector Memory if:
- Users ask about things mentioned long ago
- You have a knowledge base to search
Use Hierarchical if:
- Building a production chatbot
- Need balance between recency and relevance
The Knowledge Graph uses Pydantic structured output for reliable extraction (no regex, no eval):
class Triple(BaseModel):
subject: str
relation: str
object: strEntity extraction for graph queries also uses structured output to avoid hardcoded patterns.
If you are scanning this repository, the most relevant implementation areas are:
core/base.py: shared memory interface and debug contract.core/basic_memory.py: full-memory and sliding-window baselines.core/advanced_memory.py: relevance filtering and summary memory.core/external_memory.py: vector memory and knowledge-graph memory.core/hybrid_memory.py: hierarchical and OS-inspired strategies.app/playground.py: interactive UI that makes memory tradeoffs inspectable.
PRs welcome. Ideas:
- Add more strategies (e.g., attention-based, learned compression)
- Better visualization
- Benchmark suite
- Multi-turn evaluation metrics
MIT
Most tutorials show you how to add memory to an LLM app, but don't explain which strategy to use. This project lets you see the tradeoffs firsthand by inspecting the actual context sent to the model.
Import a strategy from core/ and use the same interface as the playground:
from core.basic_memory import SlidingWindowMemory
memory = SlidingWindowMemory(window_size=3)
memory.add_message("user", "Use EUR for prices.")
context = memory.get_context("Which currency should the answer use?")
print(context["final_prompt"])
print(context["debug_info"])Pass final_prompt to your chat model and record subsequent messages with add_message. To add a strategy, implement BaseMemory and register it in the playground's strategy selection.
These are educational implementations. Bounded prompt selection does not imply bounded total storage: implementations retain message history, and vector/graph stores can grow. The OS-inspired strategy uses Python lists and simple recall rules; it does not implement durable disk paging or autonomous memory management. Summaries and retrieved facts can omit or distort information.