A production-grade, Graph-Vector Hybrid memory microservice for autonomous AI agents. Stop feeding your agents massive conversation logs. Teach them dense, actionable rules.
Frameworks like CrewAI, LangGraph, and AutoGen are powerful, but they have amnesia. If an agent fails a task on Tuesday, it will make the exact same mistake on Wednesday.
Current memory solutions (like Mem0) solve this by dumping past conversation logs into a Vector DB.
The Problem: Vector DBs store what happened, not what was learned. The agent retrieves a massive wall of text, burning tokens and context window without actually improving behavior.
Instead of saving logs, this microservice uses a Reflexion Loop:
- Evaluate: A strict reviewer checks the agent's output for failures.
- Distill: A lightweight LLM extracts the failure into a single, dense Rule (e.g., "Always include a 10s timeout").
- Categorize: The LLM extracts a Concept Category (e.g.,
HTTP_REQUEST_BEST_PRACTICES), with an optional parent concept for hierarchical inheritance. - Store: The rule is stored in ChromaDB (Vector) and Neo4j (Graph), linking related rules together conceptually.
- Retrieve: Before any new task, the agent retrieves the top relevant rules — including inherited rules from parent concepts — saving the bulk of context tokens.
- Graph-Vector Hybrid: ChromaDB for semantic search, Neo4j for conceptual hierarchies.
- Multi-Tenancy: Pass an
agent_idto give every distinct agent its own isolated memory namespace. - Hierarchical Concept Inheritance [NOVEL-1]: Rules stored under a concept automatically surface to queries on related parent/child concepts.
- Cross-Agent Confidence Reinforcement [NOVEL-2]: When one agent learns a rule, semantically matching rules in other agents' namespaces get their confidence reinforced — without duplicating storage.
- Temporal Rule Decay [NOVEL-3]: Rules have a confidence score. It increases on success, decreases on failure or staleness. Rules hitting 0 are automatically deleted.
- Secured Microservice: Built with FastAPI, Dockerized, and secured with API Key authentication and
agent_idinput validation.
- Python 3.11+
- Docker & Docker Compose
- A free Groq API key
git clone https://github.com/YallaNuthan/agent-reflexion-memory.git
cd agent-reflexion-memory
cp .env.example .env
Open .env and fill in:
GROQ_API_KEY=your_groq_key_here
API_ACCESS_KEY=choose_a_secret_key
NEO4J_URI=bolt://neo4j:7687
NEO4J_USER=neo4j
NEO4J_PASSWORD=password123
docker-compose up --build
python -m venv myenv
myenv\Scripts\activate # Windows
source myenv/bin/activate # macOS/Linux
pip install -r requirements.txt
# Start Neo4j separately (or via docker-compose up -d neo4j)
uvicorn api:app --reload
Open the interactive API docs:
http://localhost:8000/docs
Or check health:
curl http://localhost:8000/health
python -m pytest tests/ -v
All 24 tests should pass.
All endpoints (except /health) require an X-API-Key header matching your API_ACCESS_KEY.
curl -X POST http://localhost:8000/v1/reflect \
-H "X-API-Key: your_key" \
-H "Content-Type: application/json" \
-d '{
"agent_id": "agent_alpha",
"task_description": "Fetch user data from REST API",
"failure_reason": "Request timed out after 30s"
}'
curl -X POST http://localhost:8000/v1/rules \
-H "X-API-Key: your_key" \
-H "Content-Type: application/json" \
-d '{
"agent_id": "agent_alpha",
"task_description": "Call a third-party weather API"
}'
curl -X POST http://localhost:8000/v1/reinforce \
-H "X-API-Key: your_key" \
-H "Content-Type: application/json" \
-d '{
"agent_id": "agent_alpha",
"rule_ids": ["rule_agent_alpha_123"],
"success": true
}'
curl -X POST http://localhost:8000/v1/decay \
-H "X-API-Key: your_key" \
-H "Content-Type: application/json" \
-d '{"agent_id": "agent_alpha"}'
curl -X POST http://localhost:8000/v1/concepts/link \
-H "X-API-Key: your_key" \
-H "Content-Type: application/json" \
-d '{
"agent_id": "agent_alpha",
"child_concept": "ASYNC_HTTP_TIMEOUT",
"parent_concept": "HTTP_REQUEST_BEST_PRACTICES"
}'
curl "http://localhost:8000/v1/concepts/hierarchy?agent_id=agent_alpha" \
-H "X-API-Key: your_key"
curl http://localhost:8000/health
- FastAPI — REST API layer
- ChromaDB — vector embeddings + semantic search
- Neo4j — concept graph + hierarchy traversal
- Groq (Llama) — LLM-powered rule distillation and concept categorization
- APScheduler — autonomous temporal decay job
- pytest — 24-test suite covering all endpoints, auth, error sanitization, hierarchy cycle protection, and atomic confidence updates
- GitHub Actions — CI pipeline running on every push/PR
- Docker / Docker Compose — containerized deployment
Agent Reflexion Memory achieves 37–54% token savings vs raw-log memory (Mem0-style), with all 3 novel mechanisms verified live against real Neo4j + ChromaDB.
See benchmarks/RESULTS.md for the full breakdown including NOVEL-1 hierarchy retrieval, NOVEL-2 cross-agent reinforcement, and NOVEL-3 temporal decay verification.
Contributions are welcome! See CONTRIBUTING.md for setup details, coding guidelines, and how the three novel mechanisms (NOVEL-1, NOVEL-2, NOVEL-3) are protected from breaking changes.
Built by YallaNuthan.