Skip to content

About

Interactive Streamlit playground comparing eight LLM memory strategies with transparent context, retrieval, summaries, and graph state.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

 

History

8 Commits

Folders and files

Repository files navigation

LLM Agent Memory Patterns

LLM Agent Memory Patterns

python ui patterns

Quickstart · Architecture · Review guide

An interactive playground to explore how different memory strategies work in LLM applications. Built for developers who want to understand the tradeoffs before implementing memory in production.

What is this?

When building chatbots or AI agents, you need to decide how to manage conversation history. This project implements 8 common approaches, with full transparency into what gets sent to the LLM at each turn.

Think of it as a debugging tool that shows you exactly what's happening inside the "memory" component.

At a Glance

Area Implementation signal
Product problem Help developers choose the right memory strategy before building an agent or chatbot.
Interactive surface Streamlit playground with side-by-side memory-state inspection.
Agent systems depth Implements short-term, summary, vector, graph, hierarchical, and OS-inspired memory patterns.
Debuggability Exposes the exact context sent to the model, plus retrieval scores, summaries, and extracted facts.

The 8 Strategies

Basic

  • Full Memory - Keep everything (until you hit token limits)
  • Sliding Window - Only keep last N messages

Smart Filtering

  • Relevance Filter - Use embeddings to find similar past messages
  • Summary Memory - Compress old messages into summaries

External Storage

  • Vector Memory (RAG) - Store in vector DB, retrieve by similarity
  • Knowledge Graph - Extract facts as triples, query the graph

Hybrid

  • Hierarchical - Combine vector search + recent messages
  • OS-Inspired - Simulate RAM/disk with explicit paging

Quick Start

# clone and setup
git clone https://github.com/yueyang0120/llm-agent-memory-patterns.git
cd llm-agent-memory-patterns
python3.10 -m venv .venv
source .venv/bin/activate
pip install -r requirements.txt

# add your openai key
cp env.example .env
# edit .env with your OPENAI_API_KEY

# run
streamlit run app/playground.py

How to Use

  1. Pick a strategy from the sidebar
  2. Chat with the bot
  3. Click "Internal Monologue & Memory State" to see:
    • What context was sent to the LLM
    • Which messages were kept/dropped
    • Similarity scores, summaries, graph facts, etc.

Strategy Comparison

Strategy Context selection Main tradeoff
Full memory All recorded messages Prompt grows with the conversation
Sliding window Last N messages Older context is excluded
Relevance filter Messages above similarity threshold Can lose chronology
Summary Generated summary + recent buffer Lossy; summary size is not formally bounded
Vector Top-k retrieved messages Retrieval may miss useful context
Knowledge graph Selected entities and relationships Extraction can introduce incorrect facts
Hierarchical Retrieved messages + recent window Combines two retrieval policies
OS-inspired Bounded active list + simulated archive Simple recall rules; no durable paging

These describe model context, not total storage complexity. Message length also affects token usage.

Tech Stack

  • Python 3.10+
  • Streamlit (UI)
  • LangChain (LLM wrapper)
  • OpenAI GPT-4o-mini
  • ChromaDB (vector store)
  • NetworkX (graph)

Architecture

flowchart TD
  UI["Streamlit playground"] --> M["Selected BaseMemory implementation"]
  M --> B["Full history / sliding window"]
  M --> A["Relevance / summary"]
  M --> E["Vector / graph"]
  M --> H["Hierarchical / OS-inspired"]
  A -.-> L["OpenAI · embeddings / summarization"]
  E -.-> S["Chroma collection / NetworkX graph"]
  H -.-> S
  B --> C["get_context(query)"]
  A --> C
  E --> C
  H --> C
  C --> P["final_prompt → chat model"]
  C --> D["debug_info → inspection panel"]
  P --> U["add_message · update selected memory"]
  U --> M
  classDef out fill:#fce7f3,stroke:#db2777,color:#831843;
  classDef storage fill:#e0f2fe,stroke:#0284c7,color:#0c4a6e;
  class P,D out;
  class L,S storage;
Loading

Architecture Notes

Base Class

All strategies inherit from BaseMemory:

def add_message(role: str, content: str)
def get_context(query: str) -> dict  # returns {final_prompt, debug_info}
def clear()

Knowledge Graph

Uses GPT-4o-mini with structured output to extract triples:

"I work at Google" → (user, work_at, google)

Vector Memory

Stores each message as an embedding in ChromaDB, retrieves top-k by similarity.

Hierarchical

Combines two layers:

  • Long-term: Vector search (top 2)
  • Short-term: Sliding window (last 3)

When to Use What

Use Full Memory if:

  • Conversations are short (<10 turns)
  • You need perfect recall

Use Sliding Window if:

  • Each message is independent
  • You only care about recent context

Use Vector Memory if:

  • Users ask about things mentioned long ago
  • You have a knowledge base to search

Use Hierarchical if:

  • Building a production chatbot
  • Need balance between recency and relevance

Implementation Details

The Knowledge Graph uses Pydantic structured output for reliable extraction (no regex, no eval):

class Triple(BaseModel):
    subject: str
    relation: str
    object: str

Entity extraction for graph queries also uses structured output to avoid hardcoded patterns.

Review Guide

If you are scanning this repository, the most relevant implementation areas are:

  • core/base.py: shared memory interface and debug contract.
  • core/basic_memory.py: full-memory and sliding-window baselines.
  • core/advanced_memory.py: relevance filtering and summary memory.
  • core/external_memory.py: vector memory and knowledge-graph memory.
  • core/hybrid_memory.py: hierarchical and OS-inspired strategies.
  • app/playground.py: interactive UI that makes memory tradeoffs inspectable.

Contributing

PRs welcome. Ideas:

  • Add more strategies (e.g., attention-based, learned compression)
  • Better visualization
  • Benchmark suite
  • Multi-turn evaluation metrics

License

MIT

Why I Built This

Most tutorials show you how to add memory to an LLM app, but don't explain which strategy to use. This project lets you see the tradeoffs firsthand by inspecting the actual context sent to the model.

Integration and limits

Import a strategy from core/ and use the same interface as the playground:

from core.basic_memory import SlidingWindowMemory

memory = SlidingWindowMemory(window_size=3)
memory.add_message("user", "Use EUR for prices.")
context = memory.get_context("Which currency should the answer use?")
print(context["final_prompt"])
print(context["debug_info"])

Pass final_prompt to your chat model and record subsequent messages with add_message. To add a strategy, implement BaseMemory and register it in the playground's strategy selection.

These are educational implementations. Bounded prompt selection does not imply bounded total storage: implementations retain message history, and vector/graph stores can grow. The OS-inspired strategy uses Python lists and simple recall rules; it does not implement durable disk paging or autonomous memory management. Summaries and retrieved facts can omit or distort information.

About

Interactive Streamlit playground comparing eight LLM memory strategies with transparent context, retrieval, summaries, and graph state.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages