A local-first reference architecture for building practical AI agents.
π§ Intelligence β’ π οΈ Tools β’ π‘οΈ Verification
πΉ Watch or Download Full 1080p HD Demo Video with Audio (MP4)
Most agent frameworks today start with a fundamental compromise: they send your private code, files, and prompts to closed-source cloud APIs, while giving the LLM unconstrained or poorly-bounded tool access to the host machine.
When building autonomous software, generative text alone is rarely enough. Practical agents must:
- Decompose complex objectives into discrete, trackable milestones.
- Route model requests locally across hardware-accelerated runners without cloud lock-in.
- Execute tools inside strict sandboxes where path traversal and destructive commands are physically intercepted.
- Empirically verify results before reporting completion.
I built Nexus Agent as a clean, minimal reference architecture to explore this exact pipeline: Goal βββΊ Plan βββΊ Route βββΊ Tool βββΊ Verify.
- Cloud Lock-in & Privacy: Decouples agent reasoning from proprietary remote APIs by running natively against local models (Ollama, LM Studio, vLLM, or offline mock engines).
- Uncontrolled Tool Execution: Enforces strict workspace containment and command sanitization to prevent accidental system corruption or path escapes.
- State & Memory Loss: Maintains a deterministic two-tier memory model combining active conversational context with indexed episodic memory for factual recall.
User Goal βββΊ Agent βββΊ Planner βββΊ Model Router βββΊ Tools βββΊ Execution βββΊ Verification βββΊ Memory
| Layer | Implementation File | Purpose & Guarantees |
|---|---|---|
| Model Router | core/router.py |
Dispatches prompts across Ollama, local OpenAI-compatible endpoints (vLLM / LM Studio), and deterministic offline Mock engines. |
| Task Planner | core/planner.py |
Decomposes high-level user objectives into ordered milestone steps (PENDING βββΊ IN_PROGRESS βββΊ COMPLETED). |
| Memory Subsystem | core/memory.py |
Two-tier architecture: sliding context window for active turns + indexed episodic key-value store for tool observations. |
| Sandboxed Tools | tools/workspace.py |
JSON Schema-validated tools strictly confined within target workspace_root (parent directory traversal ../ is blocked). |
| Safety Policy | safety/policy.py |
Multi-tier security engine (STRICT, BALANCED, PERMISSIVE) with regex pattern filters for destructive shell commands. |
# Clone repository
git clone https://github.com/dax0056/nexus-agent.git
cd nexus-agent
# Install dependencies in editable mode
pip install -e .[dev]No external model server, API keys, or internet connection required:
from pathlib import Path
from nexus_agent import NexusAgent, InferenceBackend, SecurityLevel
# Initialize agent inside a sandboxed workspace
agent = NexusAgent(
workspace_dir=Path("./workspace"),
backend=InferenceBackend.MOCK,
security_level=SecurityLevel.BALANCED
)
# Run autonomous task
result = agent.run("Analyze workspace structure and save executive summary")
print(f"Execution Status: {result.status}")
print(f"Steps executed : {len(result.executed_steps)}")If you have Ollama installed locally:
# Pull your preferred local model
ollama run llama3:8bfrom pathlib import Path
from nexus_agent import NexusAgent, InferenceBackend, SecurityLevel
agent = NexusAgent(
workspace_dir=Path("./workspace"),
backend=InferenceBackend.OLLAMA,
model_name="llama3:8b",
security_level=SecurityLevel.BALANCED
)
result = agent.run("Inspect workspace and generate structured summary")
print(f"Status: {result.status}")Here is the actual execution trace produced by the agent run harness:
[NexusAgent] Initialized with backend: MOCK | Workspace: ./workspace
[Planner] Decomposed goal into 3 steps:
1. [inspect_workspace] List directory structure
2. [extract_data] Read key documentation files
3. [write_summary] Generate analysis report
[Router] Generating step execution for step 1 via MOCK backend...
[Tools] Executing 'list_files' inside workspace -> 8 files found.
[Safety] Action permitted by BALANCED security policy.
[Memory] Stored 8 items in episodic memory buffer.
[Router] Generating step execution for step 3 via MOCK backend...
[Tools] Executing 'write_file' -> 'workspace_summary.txt' (248 bytes written).
[Agent] Task completed successfully with status: SUCCESS
Nexus Agent is covered by 7 automated unit tests validating all core subsystems:
============================= test session starts =============================
tests/test_agent_flow.py::test_nexus_agent_execution_flow PASSED [ 14%]
tests/test_router_memory.py::test_model_router_mock PASSED [ 28%]
tests/test_router_memory.py::test_agent_memory_buffer PASSED [ 42%]
tests/test_router_memory.py::test_agent_memory_retrieval PASSED [ 57%]
tests/test_tools_safety.py::test_safety_policy_blocking PASSED [ 71%]
tests/test_tools_safety.py::test_sandboxed_workspace_tools PASSED [ 85%]
tests/test_tools_safety.py::test_browser_mock_tool PASSED [100%]
============================== 7 passed in 0.06s ==============================
- Passed:
7 / 7(100%) - Test Command:
pytest -v
- Strict Root Boundary: Path resolution enforces relative workspace containment. Any attempt to access
/etc,C:\Windows, or parent directories (../) raises an immediate security violation. - Command Sanitization: Destructive system patterns (
format,del,rmdir,shutdown,powershell -enc) are blocked prior to tool invocation. - For vulnerability reporting, see SECURITY.md.
- Milestone development plan: ROADMAP.md
- Release notes and version history: CHANGELOG.md
- Contribution guidelines: CONTRIBUTING.md
- Code of conduct: CODE_OF_CONDUCT.md
Nexus Agent serves as the Intelligence flagship of the series:
- π§ nexus-agent β Intelligence & Local Agent Core
- π» micro-coding-agent β Deterministic Code AST & Patch Engine
- π₯οΈ desktop-action-agent β Safe Sandboxed Desktop Automation
Distributed under the MIT License. See LICENSE for details.
