Pain-weighted hybrid memory for AI agents — yours by design.
The only agent memory that’s genuinely yours. SQLite on your disk, provider your choice, zero vendor lock-in.
NOX-Supermem packages the nox-mem engine for standalone, self-hosted use — CLI · MCP server · HTTP API.
🏗️ How it works · 🚀 Install · 👤 Humans · 🤖 Agents · 📊 Numbers
Long-term memory engine that any agent (OpenClaw, Hermes, Claude Code, custom) can use to remember decisions, search past context, and never ask "where were we?" again. The engine lives in nox-mem/ and ships with no data — your memory starts empty.
Real terminal — stats · hybrid search · reflect (RAG with cited sources) · HTTP API for agents.
Five layers, one SQLite file:
- Ingest — router auto-detects entity files (
compiled/frontmatter/timelinesections), plain markdown, or graphify input. A privacy filter applies redaction patterns before anything is stored. - Store — chunks land in SQLite with an FTS5 index plus a 3072-d Gemini vector via
sqlite-vec. Retention is typed:feedback/personnever decay,lesson180d,decision/project365d, default 90d. - Retrieve — the query runs in parallel through FTS5 BM25 and Gemini semantic; RRF fusion (k=60) merges them, with language-aware weights.
- Rank — salience (
recency × pain × importance) composes additively with section and temporal boosts. Shadow discipline: ranking changes ship in shadow mode for 7 days before they ever touch a live query. - Answer — CLI, MCP, and HTTP surfaces with citation footers and an anti-hallucination guard.
Copy the SQLite file, you copy the memory. Switch the embedding provider, the store doesn't care.
Pick your interface — same engine, one SQLite file behind all three:
| Interface | Entry point | Best for |
|---|---|---|
| CLI | nox-mem <cmd> |
humans, scripts, cron |
| MCP server | nox-mem-mcp |
agents (OpenClaw, Hermes, Claude Code) — 21 tools |
| HTTP API | nox-mem-api |
services, dashboards, remote agents |
Prerequisites (Linux / macOS): Node 20+. build-essential and python3 only matter if better-sqlite3 cannot download a prebuilt binary for your platform and has to compile (inotify-tools is not needed).
node --version # must be >= 20
# Debian/Ubuntu, only if the install below asks for a compiler:
sudo apt-get update && sudo apt-get install -y build-essential python3npm install -g nox-mem # published on npm — that's the whole install
export GEMINI_API_KEY=AIza... # https://aistudio.google.com/apikey
nox-mem stats # first run creates ~/.nox-mem/nox.db
nox-mem ingest notes/*.md # one or many files
nox-mem search "hello"1. Install — from npm (easiest), or from source if you want the agent profiles/templates too:
# From npm (recommended)
npm install -g nox-mem
nox-mem --help
# …or from source (also gets perfis/ + templates/)
git clone https://github.com/totobusnello/nox-mem.git
cd nox-mem/nox-mem && npm ci && npm run build && npm install -g .2. Configure — create a .env (template in nox-mem/.env.example):
# Required
GEMINI_API_KEY=AIza... # Google AI Studio key
# Optional
NOX_DB_PATH=$HOME/.nox-mem/nox.db # SQLite database; this is already the default (missing folder is created)
# NOX_MEM_DIR is NOT your notes folder: it only sets where pre-op snapshots go
# ($NOX_MEM_DIR/.nox-snapshots) and widens the op-audit allowlist. Leave it unset.
# HTTP API (optional) — default port 18802
NOX_API_PORT=18802
NOX_API_HOST=127.0.0.1
# NOX_API_TOKEN=change-me # if set, API requires Authorization: Bearer <token>Load it before running the CLI in any shell, cron, or service:
set -a; source "$HOME/.nox-mem/.env"; set +a
⚠️ Without sourcing the env,vectorize/kg-*fail silently ("Done: 0 embedded").
3. Initialize & verify
nox-mem stats # first run auto-creates the schema — no migrations to run
nox-mem doctor # diagnostic: SQLite, FTS5, vector extension, config4. Ingest, embed, search
nox-mem ingest notes/*.md # plain markdown is fine; one or many files
nox-mem vectorize # embeds new chunks (needs GEMINI_API_KEY)
nox-mem search "what did we decide about pricing"
nox-mem primer # ~500-token context-recovery summary
# keep indexing a folder as you save into it (absolute paths, comma-separated)
NOX_WATCH_DIRS="$HOME/notes" nox-mem watchA directory passed to ingest is an error naming it (the other files still run, exit 1). Past 10,000 chunks ingest and watch ask you to confirm: --allow-prod or NOX_ALLOW_PROD_INGEST=1. That guard stops a test script from writing into your real database by accident.
Agents connect over MCP (preferred) or the HTTP API. The bootstrap is idempotent — each step verifies before continuing.
1. Deterministic bootstrap (run in order; stop on first failure)
# preconditions
node --version | grep -qE 'v(2[0-9]|[3-9][0-9])' || { echo "need Node >=20"; exit 1; }
# install from npm
npm install -g nox-mem
# config
export GEMINI_API_KEY="<key>" NOX_DB_PATH="/data/nox/nox.db" # the folder is created if missing
# verify schema
nox-mem stats | grep -q "Chunks:" || { echo "schema init failed"; exit 1; }
# the MCP server is a bin on PATH
command -v nox-mem-mcp || { echo "nox-mem-mcp not on PATH"; exit 1; }2. Wire it as an MCP server (recommended) — 21 tools (nox_mem_search, nox_mem_answer, nox_mem_ingest, nox_mem_primer, nox_mem_reflect, nox_mem_kg_query, nox_mem_decision_*, nox_mem_cross_search, …). Add to your agent's MCP config (Claude Code .mcp.json, OpenClaw/Hermes equivalent):
{
"mcpServers": {
"nox-mem": {
"command": "nox-mem-mcp",
"env": {
"GEMINI_API_KEY": "AIza...",
"NOX_DB_PATH": "/data/nox/nox.db"
}
}
}
}Or in one line for Claude Code: claude mcp add nox-mem -e GEMINI_API_KEY="$GEMINI_API_KEY" -e NOX_DB_PATH=/data/nox/nox.db -- nox-mem-mcp.
Temporal filters work on every surface: nox_mem_search and nox_mem_answer take as_of / changed_since (on the CLI, --as-of / --changed-since; over HTTP, ?as_of= / ?changed_since= on /api/search and the same keys in the POST /api/answer body). A bad or blank date gives exit 2 / HTTP 400 / MCP isError.
The agent calls nox_mem_search to recall and nox_mem_ingest to store. Run nox_mem_primer at session start for context recovery. Reusable agent profiles (assistente-pessoal, financeiro, pesquisador) live in perfis/; generic SOUL/HEARTBEAT/IDENTITY templates in templates/.
3. Or wire it as an HTTP API
set -a; source /data/nox/.env; set +a
nox-mem-api # listens on 127.0.0.1:18802 (NOX_API_PORT overrides)| Endpoint | Purpose |
|---|---|
GET /api/health |
status + vectorCoverage ({embedded, total, orphans, indexOnly}) |
GET /api/search?q=... |
hybrid search |
GET /api/brief |
salience-ranked session priming |
POST /api/answer |
RAG answer over memory (body accepts as_of / changed_since) |
GET /api/kg, /api/kg/path |
knowledge graph |
GET /api/reflect?q=... |
synthesis over memory + KG (q is required; 400 without it) |
If NOX_API_TOKEN is set, send Authorization: Bearer <token>.
The engine is the same core benchmarked in memoria-nox. All results 5-batch + 95% CI verified.
| Benchmark | nox-mem | Best competitor | Δ |
|---|---|---|---|
| EverMemBench Overall (Gemini-3-flash) | 63.28% | MemOS 42.55% | +20.73pp |
| EverMemBench MA composite | 88.42% | MemOS 55.68% | +32.74pp |
| LoCoMo retrieval@10 strict | 74.52% | Mem0 SOTA F1 66.88% | above |
| MuSiQue F1 (n=2,417, single-shot) | 58.62% | IRCoT 35.80% / EX(SA) 49.70% | +22.82pp / +8.92pp |
| HotPotQA ans_F1 (n=7,405 distractor) | 73.37% | DPR+FiD reader 65–72% | above band |
| Dimension | nox-mem | Comparison |
|---|---|---|
| KG path latency | 2.5ms p50 | none sub-10ms published |
| KG path cost/query | $0.00 | Mem0 Cloud $0.001 → 769× cheaper |
| Self-hosted footprint | 399MB single-process | Zep/Mem0/MemOS run 4+ services |
| Backbone portability | −10.54pp on backbone swap | MemOS −16.72pp → 1.6× more portable |
| Monthly OPEX (embed + KG + VPS) | < $11/mo all-in | — |
• LoCoMo is retrieval@10 strict (a retrieval metric) shown next to Mem0's reported F1 — different metrics, for scale, not a like-for-like claim.
• LongMemEval 1.0 is the oracle retrieval ceiling (gold answers in-corpus ⇒ nDCG@10 = 1.0), not an end-to-end inference score; standalone accuracy is ~68%.
• KG-path $0 / 769× cheaper compares nox-mem's pure-SQL graph path (no LLM call) to Mem0 Cloud's per-query price (which includes inference) — true for that path only, apples-to-oranges by design.
• < $11/mo assumes the cheapest Hostinger VPS + Google AI Studio free tier.
• +78.8% nDCG@10 is vs an internal local-embedding baseline.
Step-by-step on what's reproducible from this package vs the research harness: REPRODUCE.md. Methodology, paper, and full competitive analysis:
memoria-nox. MemOS arXiv:2602.01313 · MuSiQue (Trivedi 2022) · HotPotQA (Yang 2018).
Default is Gemini via Google AI Studio for both LLM and embeddings (GEMINI_API_KEY). The RAG answer/reflect layer and embeddings are provider-pluggable at runtime — no rebuild:
- LLM (
NOX_LLM_PROVIDER) —gemini(default) oropenai, whereopenaidrives any OpenAI-compatible endpoint: OpenAI, DeepSeek, OpenRouter, Together, or a local Ollama/vLLM.anthropicis interface-ready but not yet implemented. - Embedding (
NOX_EMBEDDING_PROVIDER) —gemini(default, 3072-d) oropenai(any OpenAI-compatible embeddings endpoint).voyageis interface-ready but not yet implemented.
# DeepSeek LLM + OpenAI embeddings (both OpenAI-compatible)
NOX_LLM_PROVIDER=openai
NOX_LLM_BASE_URL=https://api.deepseek.com/v1
NOX_LLM_MODEL=deepseek-chat
NOX_LLM_API_KEY=sk-...
NOX_EMBEDDING_PROVIDER=openai
NOX_EMBEDDING_BASE_URL=https://api.openai.com/v1
NOX_EMBEDDING_MODEL=text-embedding-3-large
NOX_EMBEDDING_DIM=3072 # MUST equal the vec0 table dim
NOX_EMBEDDING_API_KEY=sk-...ℹ️ Scope of provider routing today: the RAG answer layer (
reflect//api/answer) and embeddings honor the env vars above. Some internal LLM operations — knowledge-graph extraction, consolidation, digest, and query expansion — still call Gemini directly and requireGEMINI_API_KEYeven when another provider is set. Routing every path through the provider layer is on the roadmap.
⚠️ Dimension lock: the sqlite-vec table is created with a fixed dimension. Switching embedding provider or model requires re-embedding the entire corpus with a single model at a single dimension matching thevec0table.text-embedding-3-largesupportsdimensions=3072(same as the default Gemini table). Vectors from different models are not comparable — mixing silently corrupts semantic search. Full env reference:nox-mem/README.md.
nox-mem-api &
curl -s "http://127.0.0.1:${NOX_API_PORT:-18802}/api/health" | jq '.vectorCoverage | .embedded/.total'
# close to 1.0 = all chunks embedded; below 0.99 → run `nox-mem vectorize`| Symptom | Fix |
|---|---|
vectorize says "0 embedded" |
env not sourced — set -a; source .env; set +a |
vec0 ... cannot open shared object |
platform binary missing — reinstall nox-mem (npm i -g nox-mem) so npm fetches sqlite-vec-<os>-<arch>; installing sqlite-vec by itself does not help |
better-sqlite3 build error |
install build-essential + python3, then npm ci again |
| API port in use | set NOX_API_PORT (default is 18802) |
ingest aborts with "Large-DB ingest guard" |
the database has more than 10,000 chunks: pass --allow-prod or set NOX_ALLOW_PROD_INGEST=1 (or point NOX_DB_PATH at another file if it is the wrong database) |
nox-mem watch says "Watching 0 directories" |
set NOX_WATCH_DIRS to your notes folder (absolute path) |
reindex refuses or finds "0 files" |
reindex rebuilds from $OPENCLAW_WORKSPACE/memory and /shared, not from your notes; on a standalone install use nox-mem ingest |
| path rejected by op-audit guard | set NOX_OP_AUDIT_ALLOWED_PREFIXES, or keep DB under NOX_DB_PATH/NOX_MEM_DIR (auto-allowed) |
Full env-var reference and per-command notes: nox-mem/README.md.
MIT © 2026 Luiz Antonio Busnello (Toto). Use it, fork it, ship it.
