Skip to content

Sakushi-Dev/nexus_memory

Folders and files

NameName
Last commit message
Last commit date

Latest commit

 

History

47 Commits
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Nexus Memory

Nexus Memory

A local-first, dependency-light agent-memory library for Python. It gives an LLM application a persistent, self-managing long-term memory backed by SQLite + sqlite-vec — one .db file, no server, no network, and no model download on the default path. Everything flows through a single entry point, NexusMemory.process(), which validates every request with pydantic and never raises to the caller.

Features

  • Offline & deterministic — the default HashingEmbedder (768-dim, blake2b feature hashing, L2-normalized) needs no downloads and produces stable, reproducible vectors. Nothing leaves the machine.
  • Provider-agnostic — for semantic recall, drop in the local FastEmbedEmbedder (fastembed/ONNX, no torch — downloads a small model once, then runs offline, no vendor), or the lazy SentenceTransformerEmbedder / OpenAIEmbedder adapters; bring your own fact extractor, summarizer, or directive detector. The diary and procedural directive extraction run on any model through a handoff outbox — the library itself never calls an LLM.
  • Five-layer cognitive memory — working, episodic, semantic, and procedural layers fan out from a single ingest; an optional hierarchical diary (Layer V) is off by default.
  • One entry point — every request is a dict (or JSON string) routed through NexusMemory.process(); all errors come back as {"status": "error", "error": ...}.
  • Transparent & sovereigninspect, forget, pin, and update your own memories.
  • Privacy by design — an opt-in regex PII filter masks emails/phones/names before embedding (off by default on the local path; turn it on for external embedding APIs). An optional SQLCipher encryption hook stays off the critical path.
  • User-centric — by default only the user's statements become semantic facts; assistant prose goes to the episodic diary, not the vector store.

Install

Clone the repo, create a virtual environment, and install it editable with pip. Requires Python ≥ 3.11.

# 1. Get the code.
git clone https://github.com/Sakushi-Dev/nexus_memory.git
cd nexus_memory

# 2. Create and activate a virtual environment (recommended).
python -m venv .venv
.venv\Scripts\Activate.ps1        # Windows (PowerShell)  ·  cmd.exe: .venv\Scripts\activate.bat
# source .venv/bin/activate       # macOS / Linux

# 3. Install the package editable, into the active venv (source edits take effect immediately).
pip install -e .

This reads pyproject.toml, pulls in the dependencies (sqlite-vec, pydantic, numpy), and registers the src/nexus_memory/ package so import nexus_memory works from anywhere with that interpreter:

import nexus_memory  # importable after `pip install -e .`

Multiple Pythons? pip installs into the interpreter it belongs to. To target a specific one (e.g. the one your editor runs), call pip through it: path/to/python.exe -m pip install -e .

The default install is fully offline (the lexical HashingEmbedder). For semantic recall and other backends, install the optional extras into the same activated venv — only what you need:

pip install -e ".[local-embeddings]"         # local semantic embedder (fastembed + bge-base, offline after one download)
pip install -e ".[sentence-transformers]"    # SentenceTransformer embedder (heavier; needs torch)
pip install -e ".[openai]"                   # OpenAI embedder (external API)

See Embedders for details.

Updating

Because the install is editable, pulling the latest code is usually all you need:

git pull

Re-run pip install -e . only if the dependencies changed (e.g. after a new release bumps requirements in pyproject.toml):

pip install -e .   # picks up new/updated dependencies

Driving it from another language? Every request and response is a plain dict (or JSON string) through one process() entry point, so the module can sit behind a thin Python bridge (subprocess, stdin/stdout JSON, or a small socket/IPC shim) and be called from any language — no Python-specific objects cross the boundary.

Quickstart

from nexus_memory import NexusMemory

memory = NexusMemory(db_path="my_agent.db")

memory.process({
    "action": "ingest",
    "interaction": {
        "query": "where do I keep my keys?",
        "response": "You always keep your house keys in the blue ceramic bowl.",
    },
})
memory.wait()  # ingest is async; wait() makes a script deterministic

result = memory.process({"action": "assemble", "query": "where are my keys?", "top_k": 3})
print(result["context_xml"])  # prompt-ready <memory_context> XML

memory.close()

A runnable version lives in examples/basic_usage.py; the diary handoff loop is in examples/diary_outbox.py.

Live demo

A runnable chat app — a terminal UI that puts Nexus to work end-to-end against a real LLM (via OpenRouter) — lives on the demo branch. Clone just that branch into its own folder and install it into a virtualenv:

# 1. Clone only the demo branch.
git clone --branch demo https://github.com/Sakushi-Dev/nexus_memory.git nexus-demo
cd nexus-demo

# 2. Create a virtualenv, then install the bundled module + the app's deps into it.
python -m venv .venv
.\.venv\Scripts\python.exe -m pip install -e .\nexus_memory_pkg
.\.venv\Scripts\python.exe -m pip install -r requirements.txt

# 3. Add your OpenRouter key (get one at https://openrouter.ai/keys).
copy .env.example .env        # then edit .env and paste your key

# 4. Run — full-screen TUI (add --classic for the line UI, or --selftest for an offline check, no key needed).
.\.venv\Scripts\python.exe chat.py

macOS/Linux: use ./.venv/bin/python instead of .\.venv\Scripts\python.exe, and cp .env.example .env.

For semantic recall in the demo, first install the local embedder, then run with --embedder fastembed (it downloads a small model once, then runs offline):

.\.venv\Scripts\python.exe -m pip install "fastembed[cpu]"   # or: -e ".\nexus_memory_pkg[local-embeddings]"
.\.venv\Scripts\python.exe chat.py --embedder fastembed

A store is bound to the embedder it was created with — to switch, start a fresh DB or re-embed with python -m nexus_memory.reindex --db data\chat_memory.db --backend fastembed. See the demo's own README.md for model settings + the slash-command reference.

The layer model

A single ingest consolidates across layers, and assemble returns one unified, layer-aware <memory_context>:

  • I. Working (working.py) — a volatile RAM ring buffer of the last N turns for fast recency context.
  • II. Episodic (episodic/) — persistent raw dialogue plus deterministic narrative day-summaries.
  • III. Semantic (semantic/, core/db.py) — decontextualized fact vectors retrieved by cosine KNN, then re-ranked by similarity × importance × exp(-λ · days).
  • IV. Procedural (procedural.py) — standing behavioral directives (e.g. "Keep answers concise.") mined from the conversation by the aux LLM (Mem0-style add/update/delete, with an offline regex fallback) and injected into the assembled context.
  • V. Diary (optional, off by default) (diary/) — a bounded session-pyramid of model-written summaries (one rolling summary per session, folded into a single growing persistent summary) that carries the conversation across session boundaries; enable with NexusMemory(diary=True), then drain pending_summaries() and return text via submit_summary().

Actions

Every payload carries an action, passed to memory.process(...):

action key fields returns
assemble query, top_k=5, min_score=0.6 {status, context_xml, raw_facts, directives, recent_dialogue, meta, latency_ms}
ingest interaction:{query, response}, metadata?, priority? (1-10 importance floor) {status:"processing", task_id, estimated_completion_ms}
forget exactly one of fact_id / query (a query below forget_min_similarity returns not_found) {status, deleted_id}
pin content, importance=10.0 {status, id, content, importance}
update target_id, new_content {status, updated_id, content}
inspect type:"health"|"episodic"|"semantic"|"working"|"procedural", filter? {status, data}
optimize {before_bytes, after_bytes, facts}
diary day?, time_range?, store? {status, period, summary, turn_count}
rule op:"add"|"list"|"deactivate", directive?, priority?, rule_id? add: {status, rule} · list: {status, rules} · deactivate: {status, rule_id, deactivated}
distill {status, promoted:[rule,...]}
pending_summaries limit? (Layer V only) {status, jobs:[{job_id, kind, session, prompt, prior_summary, input}, ...]}
submit_summary job_id, summary (Layer V only) {status, applied?:"session"|"summary"}

Convenience wrappers: inspect(...), forget(...), pin(...), update(...), wait(...), close(), remember_rule(...), list_rules(), diary(...), working_snapshot(), reconstruct(...), history(...), distill(), the unified background-job seam drain_aux(...) / pending_aux_jobs(...) / submit_aux_job(...), plus pending_summaries(...) / submit_summary(...) when the diary is enabled.

Tune everything (scoring, dedup threshold, cache, privacy, per-layer settings) via NexusConfig.

Documentation

Full documentation lives under docs/ — start at docs/index.md.

  • Architecture — the five-layer design, retrieval & scoring, persistence, extension points.
  • Usage & API — getting started, every action, embedders, transparency.
  • ConfigurationNexusConfig, DiaryConfig, tuning.
  • Use cases — agent memory, multiple isolated agents, behavioral rules, the diary, privacy.

Run the tests

pip install -e ".[dev]"   # once, to get pytest
python -m pytest -q

License

MIT — see LICENSE.

About

A local-first, dependency-light agent-memory library for Python. It gives an LLM application a persistent, self-managing long-term memory backed by SQLite + sqlite-vec — no server, no network, no model download required for the default path.

Topics

Resources

Stars

Watchers

Forks

Releases

Packages

Contributors

Languages