Sub-Microsecond Agent Memory β’ openCypher Knowledge Graphs β’ Vector Search β’ Relational SQL β’ Documents
Single Encrypted .tapir File β’ 100% Safe Rust β’ < 4 MB Idle RAM β’ Zero Cloud Daemons
"Build private AI memory without operating a data stack."
TapirusDB is the high-performance, embedded cognitive memory engine for local AI agents, robotics, and sovereign edge hardware. It collapses vector similarity, knowledge graphs, relational metadata, and JSON documents into a single encrypted.tapirfile with sub-microsecond in-process retrieval ($0.51\ \mu\text{s}$ ) and zero memory corruption risk.
π₯οΈ Need a Visual Database Manager (like phpMyAdmin or Supabase Studio)?
Use Tapirus Studio β our free visual GUI companion for TapirusDB!
β’ π Run Instant In-Browser: tapirusdb.com/studio (Zero installation required)
β’ β¬οΈ Official Download Landing Page: tapirusdb.com/download.html (Windows, macOS, Linux, CLI)
β’ π¦ Releases & Binary Downloads: github.com/tapiruslab/TapirusDB/releases
Modern AI and edge developers are forced into Fragmented Polyglot Persistenceβgluing together multiple complex, heavy databases across network boundaries:
- π¦ 100% Pure Safe Rust (
#![forbid(unsafe_code)]): Guaranteed memory safety at compile-time. Zero buffer overflows, zero dangling pointers, zero use-after-free vulnerabilities, and zero C/C++ memory corruption CVEs. - π¦ True In-Process Architecture (Zero-IPC): Compiles and links directly into your binary (Rust, Python, TypeScript, C/C++, Go). No background database servers (
mysqld,postgres,mongod), zero network serialization overhead, and sub-microsecond in-memory query traversal. - β‘ Quad-Model Data Consolidation: Seamlessly unifies Relational SQL-92, HNSW & IVF Vector Search, openCypher Property Graphs, and MongoDB-style JSON Documents inside a single B+Tree slotted-page file.
- π Advanced SQL Window Functions & Graph Algorithms: Built-in ANSI SQL window operations (
ROW_NUMBER(),RANK(),DENSE_RANK(),NTILE(),LAG(),LEAD()) and native graph topology algorithms (GRAPH ALGORITHM louvain,betweenness,connected_components,pagerank). - π§ Production GraphRAG & AI Memory: Built-in seed-and-traverse GraphRAG with Tri-Modal Reciprocal Rank Fusion (RRF), episodic memory with exponential temporal decay, and isolated agent namespaces.
- β‘ Tap Sub-Millisecond Cognitive Instinct Engine: In-database non-autoregressive System-1 decision core (
Tap::classify,Tap::score,Tap::verify,Tap::route). Run 1,000 deterministic agent decisions per second directly inside SQL queries (TAP_CLASSIFY(),TAP_VERIFY()) in pure Safe Rust (< 2ms) with zero cloud tokens and zero API costs. - π― Dynamic 16-Lane SIMD & RaBitQ 32x Quantization: Parallel AVX-512 / AVX2 / NEON vector kernels combined with Fast Walsh-Hadamard 1-bit/2-bit random rotation quantization, reducing 1536-D embeddings from 6,144 bytes to 196 bytes with single-cycle
POPCNTdistance evaluation. - πΈοΈ Compressed Sparse Row (CSR) Topology & openCypher: Contiguous adjacency arrays on disk and in memory for zero-allocation slice neighbor sweeps, paired with standard declarative openCypher syntax (
MATCH ... WHERE ... RETURN ...). - βοΈ Atomic Four-Model Transactions: Single ACID transaction committing or rolling back across SQL rows, JSON documents, openCypher graph edges, and Vector embeddings simultaneously with zero torn states.
- π Multi-Tenant & Role-Scoped GraphRAG: Native tenant isolation (
tenant_id) and role-based ACL filtering (allowed_roles) across knowledge graph traversal and vector scoring. - π‘οΈ Physical Integrity Audit & Safe Hot Backups: Zero-downtime atomic backup snapshots, KCV key validation, and slotted-page CRC32 consistency verification (
tapirus backup,tapirus restore,tapirus verify). - π Streaming Data Importer & Protected REST Daemon: High-throughput streaming ingest for CSV (auto-inferred schema), JSONL, and Markdown straight into tables and AI memory; secure embedded REST server (
tapirus serve) with constant-time Bearer/API-key verification. - π S3/R2 Remote Range Streaming: On-demand 4KB page streaming directly from cloud object stores via HTTP Range requests with zero local disk footprint.
- π€ Native Model Context Protocol (MCP): Out-of-the-box stdio JSON-RPC 2.0 server (
tapirus mcp) for Claude Desktop, Cursor, and Gemini autonomous agents.
- Why TapirusDB? (Kill the "Frankenstack")
- Highlights & Technical Advantages
- Tap Decision Core: Sub-Millisecond In-Database Instinct Engine
- Quickstart & 30-Second Code
- Beyond AI: Classic Applications (SQLite & Mongo Alternative)
- High-Impact Domains (Research, Analytics, IoT)
- Architectural Comparison vs Polyglot Frankenstack
- Core Technical Pillars (SIMD, CSR, RaBitQ, ChaCha20)
- Industrial Edge & Autonomous Systems
- Developer Tooling & MCP Server
- Verified Benchmarks & Latency Comparison
- When (and When NOT) to Use TapirusDB
- Formal Safety Verification (TLA+)
- Documentation & Architectural Specs
Traditionally, when autonomous AI agents make structured decisions (categorizing a support ticket, checking a policy claim, scoring urgency, or choosing a graph execution branch), developers have been forced to pay an exorbitant "LLM Latency & Cost Tax": sending database payloads over HTTP to an autoregressive model like GPT-4o-mini, waiting 400msβ800ms, risking JSON formatting hallucinations, and paying per-token API bills.
TapirusDB changes this paradigm with Tap.
Tap is TapirusDB's native System-1 cognitive instinct subsystem. Built directly in Safe Rust with zero external services and zero background Python runtimes, Tap executes deterministic single-pass decision projections directly over database records in under 2 milliseconds.
[ Traditional LLM API (e.g. GPT-4o-mini) ] ββββββββββββββββββββββββββββββββββ 450ms - 800ms
[ Remote Decision Microservice (HTTP) ] ββββββββββββββββββ 85ms - 150ms
[ Python Decision Server (PyTorch) ] ββββββββ 35ms - 65ms
| Primitive | Purpose | Rust API | SQL Syntax |
|---|---|---|---|
classify |
Categorical selection with calibrated softmax distribution | conn.tap().classify(text, candidates) |
SELECT TAP_CLASSIFY(body, '["fraud", "legit"]') |
score |
Continuous rubric evaluation in |
conn.tap().score(text, criteria) |
SELECT TAP_SCORE(incident, 'emergency_severity') |
verify |
Calibrated boolean truth validation with strict margin | conn.tap().verify(premise, hypothesis) |
SELECT id FROM claims WHERE TAP_VERIFY(claim, 'active_policy') = 1 |
route |
Autonomous graph & workflow branch selection | conn.tap().route(state, routes) |
SELECT TAP_ROUTE(task_state, 'retry, escalate, resolve') |
Tap functions can be executed directly inside standard SQL SELECT projections and WHERE filtering clauses:
-- Categorize and score incoming tickets in a single database pass (< 2ms)
SELECT
id,
customer,
TAP_CLASSIFY(message, '["billing", "technical", "sales"]') AS category,
TAP_SCORE(message, 'critical system outage emergency') AS urgency_score
FROM support_inbox;
-- Filter fraud or compliance violations directly in SQL WHERE clause
SELECT id, transaction_amount, merchant
FROM transaction_audit
WHERE TAP_VERIFY(notes, 'unauthorized account takeover attempt') = 1;use tapirus::{Connection, Result};
fn main() -> Result<()> {
let conn = Connection::open_in_memory()?;
// 1. Categorical Decision (< 1.5ms)
let decision = conn.tap().classify(
"Refund requested because package arrived damaged",
&["refund", "billing", "sales_inquiry"]
)?;
println!("Action: {} (Confidence: {:.2}%)", decision.top_choice, decision.confidence * 100.0);
// 2. Truth Verification (< 1.2ms)
let verified = conn.tap().verify(
"User confirmed receipt of digital product license",
"digital license successfully received"
)?;
if verified.is_verified {
println!("Claim verified with margin: {:.3}", verified.margin);
}
// 3. Autonomous Graph Branch Routing (< 1.8ms)
let step = conn.tap().route(
"Payment gateway returned code 504 gateway timeout",
&["retry_transaction", "fallback_processor", "cancel_order"]
)?;
println!("Next Workflow Step: {}", step.selected_route);
Ok(())
}import tapirus
# Query with embedded Tap decision functions
conn = tapirus.connect(":memory:")
conn.execute("CREATE TABLE claims (id INTEGER PRIMARY KEY, details TEXT);")
conn.execute("INSERT INTO claims VALUES (1, 'Claim filed for broken windshield from hailstorm');")
rows = conn.query("SELECT id, TAP_VERIFY(details, 'weather damage claim') AS valid_weather FROM claims;")
print(rows) # [{'id': 1, 'valid_weather': 1}]
# Or call direct decision helpers
verdict, conf = tapirus.tap_classify("Server disk full emergency", ["infrastructure", "billing", "general"])
print(f"Top Category: {verdict} ({conf*100:.1f}%)")- β¬οΈ Official Download Landing Page: tapirusdb.com/download.html (Windows, macOS, Linux, CLI)
- π€ Edge AI, Robotics & Smart IoT Portal: tapirusdb.com/edge.html (Robotics SLAM, Wear-Leveling, Mobile SLMs)
- π₯οΈ Tapirus Studio (Free GUI): Run Live in Browser (Zero installation required)
- π¦ Pre-Compiled Binary Assets: GitHub Releases
TapirusDB is 100% self-contained with zero cloud or daemon dependencies. Pre-compiled binaries run out-of-the-box across:
| Platform Family | Architecture | Supported Operating Systems & Distros |
|---|---|---|
| Linux (Universal glibc) | x86_64, aarch64 |
Debian, Ubuntu, Fedora, RHEL, CentOS, Rocky Linux, AlmaLinux, Arch Linux, openSUSE, Amazon Linux 2/2023 |
| Linux (musl & Containers) | x86_64, aarch64 |
Alpine Linux, Docker / OCI (ghcr.io/tapiruslab/tapirusdb), Embedded Linux / IoT |
| macOS | Apple Silicon & Intel | macOS 12+ (Monterey, Ventura, Sonoma, Sequoia) |
| Windows | x86_64 |
Windows 10, Windows 11, Windows Server 2019/2022/2025 |
| WebAssembly (WASM) | wasm32 |
All modern web browsers (Chrome, Edge, Safari, Firefox) via Tapirus Studio |
# macOS & Linux (Homebrew)
brew install https://raw.githubusercontent.com/tapiruslab/TapirusDB/main/Formula/tapirus.rb
# Windows (Windows Package Manager)
winget install --manifest https://raw.githubusercontent.com/tapiruslab/TapirusDB/main/winget/tapirus.yaml
# (Or 'winget install tapirus' once indexed in Microsoft community repo)
# Linux / macOS Automated Script
curl -fsSL https://raw.githubusercontent.com/tapiruslab/TapirusDB/main/install.sh | bash
# Windows PowerShell Automated Script
irm https://raw.githubusercontent.com/tapiruslab/TapirusDB/main/install.ps1 | iex# Rust Engine
cargo add tapirus
# Python SDK (Python 3.9+)
pip install tapirus
# Node.js & TypeScript SDK
npm install tapirus
# Bun Runtime
bun add tapirus
# Go SDK
go get github.com/tapiruslab/TapirusDB/sdks/go
# PHP Composer
composer require tapiruslab/tapirusdb
# OCI Container (Docker & Podman)
docker pull ghcr.io/tapiruslab/tapirusdb:latestπ‘ Looking for FFI integration, C-ABI bindings, or native shared libraries? See the Multi-Language SDK & FFI Guide.
use tapirus::{Connection, DistanceMetric, Result};
fn main() -> Result<()> {
// Open in-memory or single-file database: "production.tapir"
let db = Connection::open_in_memory()?;
// 1. Create table with structured columns and dense vector embedding
db.execute("
CREATE TABLE documents (
id INTEGER PRIMARY KEY,
title TEXT NOT NULL,
category TEXT NOT NULL,
embedding VECTOR(4)
);
")?;
db.execute("
INSERT INTO documents VALUES
(1, 'Safe Systems in Rust', 'tech', [0.95, 0.05, 0.0, 0.0]),
(2, 'Neural Vector Databases', 'ai', [0.10, 0.90, 0.15, 0.0]);
")?;
// 2. Hybrid Vector Search with Single-Pass SQL Pre-Filtering (Exact k Recall)
let rows = db.query("
SELECT id, title
FROM documents
VECTOR NEAR embedding = [0.92, 0.08, 0.0, 0.0] TOP 1
WHERE category = 'tech';
")?;
for row in rows {
println!("Match: {}", row.get::<String>("title")?);
}
Ok(())
}use tapirus::{Connection, Result};
fn main() -> Result<()> {
let db = Connection::open_in_memory()?;
// 1. Ingest entities and relationships
db.execute("GRAPH INSERT NODE 1 LABEL 'Person' PROPERTIES '{\"name\": \"Alice\"}';")?;
db.execute("GRAPH INSERT NODE 2 LABEL 'Person' PROPERTIES '{\"name\": \"Bob\"}';")?;
db.execute("GRAPH INSERT NODE 3 LABEL 'Company' PROPERTIES '{\"name\": \"TapirusTech\"}';")?;
db.execute("GRAPH INSERT EDGE 1 -> 2 LABEL 'KNOWS' WEIGHT 0.9;")?;
db.execute("GRAPH INSERT EDGE 2 -> 3 LABEL 'WORKS_AT' WEIGHT 1.0;")?;
// 2. Query graph patterns using industry-standard openCypher
let rows = db.query("
MATCH (a:Person)-[r:KNOWS]->(b:Person)
WHERE b.name = 'Bob'
RETURN a.name, b.name, r.weight;
")?;
for row in rows {
println!("{} knows {} (weight: {})",
row.get::<String>("a.name")?,
row.get::<String>("b.name")?,
row.get::<f64>("r.weight")?
);
}
Ok(())
}use tapirus::{Connection, MemoryRecallFilter, Result};
use tapirus::vector::DistanceMetric;
fn main() -> Result<()> {
let conn = Connection::open_in_memory()?;
// 1. Graph-to-Vector Chaining (Sub-microsecond 0.55 Β΅s retrieval)
// Constrains vector distance calculations strictly to local graph neighborhood O(M Β· D)
let candidates = conn
.chain(1) // Seed Patient Node
.out(Some("TREATS")) // Traverse outgoing relationships
.filter_label("Medicine") // Target node label
.vector_near(&[0.90, 0.10, 0.0, 0.0], 5, DistanceMetric::Cosine)?;
// 2. Vector-to-Graph Chaining (Seed-and-Traverse)
// Seeds from query vector, then traverses adjacent knowledge subgraph
let discovered = conn
.chain_from_vector(&[0.85, 0.15, 0.0, 0.0], 1)?
.out(Some("AUTHORED_BY"))
.collect_nodes();
// 3. Autonomous AI Agent Long-Term Memory (LTM)
// Multi-modal recall: Dense Vector + BM25 Lexical + Recency Decay (e^-Ξ»Ξt)
let memory_id = conn.memory_remember(
"User prefers sovereign on-device processing and strict privacy",
Some(&[0.92, 0.08, 0.0, 0.0]),
0.95, // Importance priority score
&["preferences", "privacy"],
)?;
let filter = MemoryRecallFilter::default(); // Balanced Vector + BM25 + Recency
let recalled = conn.memory_recall(Some("sovereign privacy"), None, 3, &filter);
println!("Recalled Agent Memory: {}", recalled[0].entry.content);
Ok(())
}import tapirus
# Connect directly to local encrypted vault or in-memory
conn = tapirus.connect("app.tapir")
# 1. Relational SQL & Window Functions
conn.execute("CREATE TABLE telemetry (id INTEGER PRIMARY KEY, sensor TEXT, value REAL);")
conn.execute("INSERT INTO telemetry VALUES (1, 'temp', 23.8), (2, 'temp', 24.1), (3, 'temp', 22.9);")
records = conn.query("""
SELECT id, sensor, value,
ROW_NUMBER() OVER (ORDER BY value DESC) as rank
FROM telemetry;
""")
print(records) # [{'id': 2, 'sensor': 'temp', 'value': 24.1, 'rank': 1}, ...]
# 2. Native Vector Search
conn.execute("CREATE TABLE docs (id INTEGER PRIMARY KEY, vec VECTOR(3));")
conn.execute("INSERT INTO docs VALUES (1, [0.9, 0.1, 0.0]), (2, [0.1, 0.9, 0.0]);")
top_docs = conn.vector_search("docs", "vec", [0.85, 0.15, 0.0], top_k=1)
# 3. Native Graph Algorithms
community_map = conn.graph_algorithm("louvain")
conn.checkpoint()import { TapirusClient, open } from "tapirusdb";
// Connect to single-file database
const db = new TapirusClient({ dbPath: "production.tapir" });
// 1. Relational SQL with Window Functions
await db.execute("CREATE TABLE users (id INT PRIMARY KEY, name TEXT, score REAL);");
await db.execute("INSERT INTO users VALUES (1, 'Alice', 95.5), (2, 'Bob', 88.0);");
const ranked = await db.query(`
SELECT name, score,
RANK() OVER (ORDER BY score DESC) as leaderboard_rank
FROM users;
`);
console.log(ranked);
// 2. Built-in SIMD Vector Search & Graph Clustering
const neighbors = await db.vectorSearch("docs", "vec", [0.9, 0.1, 0.0], 5);
const communities = await db.graphAlgorithm("louvain");While TapirusDB is the premier memory engine for sovereign AI and robotics, you do not need AI to benefit from TapirusDB. It is also a first-class, zero-configuration embedded database for general applications, edge systems, and analytics:
Need reliable relational tables, transactions, and foreign keys without AI? TapirusDB provides standard SQL with pure Safe Rust reliability:
// Standard Relational SQL with ACID transactions
db.execute("CREATE TABLE accounts (id INTEGER PRIMARY KEY, email TEXT, balance REAL);")?;
db.execute("INSERT INTO accounts VALUES (1, 'alice@example.com', 1250.50);")?;
// Complex queries with Subqueries & CTEs
let rows = db.query("
WITH active_accounts AS (
SELECT id, email, balance FROM accounts WHERE balance > 1000.0
)
SELECT * FROM active_accounts;
")?;- Advanced Query Engine: Built-in subqueries, CTEs (
WITH ... AS),INNER/LEFT JOIN, and Cost-Based Optimizer (CBO). - Transparent Encryption Included: SIMD-accelerated ChaCha20-Poly1305 AEAD encryption at rest (RFC 8439) without paying for proprietary SQLite commercial extensions.
Need to store dynamic payloads, user settings, or sensor telemetry with flexible schemas?
let collection = db.collection("telemetry")?;
let doc_id = collection.insert_one(&serde_json::json!({
"sensor_id": "temp_probe_09",
"reading_celsius": 24.3,
"calibration": { "offset": 0.05, "certified": true },
"tags": ["factory_floor", "zone_b"]
}))?;- SIMD Aggregations: Vectorized
SUM,AVG,COUNTprocessing multi-megabyte datasets in microseconds. - Transparent Compression: Built-in pure Safe Rust LZ4 page compression reduces disk footprint by 50%β70%.
- Developer CLI: Fast code search tool
tapirus tgbuilt right into the binary.
TapirusDB's zero-dependency single-file architecture is purpose-built for environments where spinning up complex database server clusters is impossible, expensive, or counterproductive:
- 100% Reproducible Research Bundles: Peer reviewers and researchers no longer need to configure Docker containers, PostgreSQL, Neo4j, and Milvus just to run a paper's code. Package an entire multimodal datasetβmolecular/protein graphs, high-dimensional vector embeddings, and assay measurement SQL tablesβinto a single verifiable
experiment.tapirfile. - Zero-Setup Python & Jupyter Workflows: Install in seconds (
pip install tapirus) and query directly inside Jupyter notebooks without starting any background daemons. - Guaranteed Memory Determinism: 100% Pure Safe Rust (
#![forbid(unsafe_code)]) guarantees zero memory leaks, buffer overruns, or segfault crashes during 72-hour batch computation runs.
# Python/Jupyter Research Workflow
import tapirus
# Open single research dataset container
db = tapirus.open("paper_dataset.tapir")
# Query molecular knowledge graph combined with chemical vector distance
results = db.query("""
MATCH (c:Compound)-[:BINDS_TO]->(p:Protein {id: 'EGFR'})
WHERE c.smiles_vector <-> $query_vec < 0.15
RETURN c.id, c.affinity_score;
""", query_vec=target_embedding)- Zero-IPC Columnar Aggregations: Vectorized
SUM,AVG, andCOUNTaccumulators run directly across local memory pages with sub-microsecond execution times, eliminating network hop overhead completely. - Transparent LZ4 Disk Compression: Built-in page compression slashes disk space by 50%β70%, allowing edge gateways and industrial PCs to retain months of historical sensor telemetry locally.
- Zero Cloud Egress Costs: Query, aggregate, and analyze high-frequency telemetry at the edge without paying exorbitant bandwidth and ingress bills to cloud data warehouses.
// In-Process Telemetry Aggregation with Common Table Expressions (CTEs)
let summary = db.query("
WITH sensor_rollup AS (
SELECT sensor_id, AVG(reading) AS avg_reading, COUNT(*) AS samples
FROM telemetry_logs
WHERE timestamp >= NOW() - 3600
GROUP BY sensor_id
)
SELECT * FROM sensor_rollup WHERE avg_reading > 85.0;
")?;- 100% Sovereign & Local-First: Run entirely offline on a Raspberry Pi 4/5 or Intel NUC with < 4 MB idle RAM. Your private camera triggers, sensor logs, and home conversations never leak to external cloud servers.
- Mesh Network Topology (openCypher Graph): Model Zigbee, Matter, and Thread device hierarchies natively (
MATCH (s:Switch)-[:CONTROLS]->(l:Light)). - Offline Voice Intent Matching (Vector Engine): Store speech and intent embeddings locally for sub-millisecond local voice assistant recognition (Whisper / Home Assistant Voice).
- Blackout Resilience (ACID WAL): If your home experiences an abrupt power outage, TapirusDB's Write-Ahead Log guarantees zero database corruption upon reboot.
// Local Voice Intent Resolution + Zigbee Mesh Pathfinding
let intent_vector = local_whisper.embed("turn off kitchen lights");
// 1. Semantic voice intent match (Vector)
let matched_action = db.vector_search("voice_intents", &intent_vector, 1)?;
// 2. Resolve Zigbee device relay path (openCypher Graph)
let route = db.graph_query("
MATCH path = (hub:Gateway)-[:ROUTES_THROUGH*1..3]->(d:Device {name: 'kitchen_main_light'})
RETURN path LIMIT 1;
")?;| Capability | TapirusDB v1.0.0 | Traditional Relational (SQLite / DuckDB) | Dedicated Vector DBs | Graph Databases (Neo4j) | Document Stores (MongoDB) |
|---|---|---|---|---|---|
| Runtime Architecture | In-Process Single File | In-Process Single File | Server Daemon / Cloud | Server Daemon (JVM) | Server Daemon (mongod) |
| Memory Safety Model | 100% Safe Rust (forbid) |
C / C++ (Manual memory) | Rust / Go / C++ | Java / JVM | C++ |
| Data Models Supported | Quad-Model (SQL+Vec+Graph+Doc) | Relational SQL only | Vector embeddings only | Graph only | JSON Document only |
| AI Vector Search | Native HNSW, IVF & RaBitQ | None (or slow extension) | Native ANN | Basic / Extension | Add-on Atlas Vector |
| Vector Quantization | RaBitQ 32x (1-Bit/2-Bit) + SQ8 | None | PQ / SQ | None | None |
| Graph Query Engine | openCypher + CSR + GraphRAG | Recursive CTE only | None | Native Cypher | $graphLookup |
| Encrypted At-Rest | ChaCha20-Poly1305 (Zero-Cost) | Commercial Add-on ($$$) | Cloud KMS only | Enterprise Tier ($$$) | Enterprise KMS |
| Cold Start / Idle RAM | < 4 MB RAM | ~4 MB (SQLite) / ~35 MB | > 500 MB | > 1,200 MB | > 350 MB |
| Binary Size | ~3.8 MB | ~1.5 MB β 42 MB | > 150 MB | > 300 MB | > 200 MB |
| Multi-Service Sync Drift | Zero (Single Container) | High (manual ETL) | High (CDC pipelines) | High (sync lag) | High (glue code) |
ββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β YOUR APPLICATION HOST PROCESS β
β (Rust β’ Python β’ TypeScript β’ Bun β’ Go β’ PHP β’ WebAssembly β’ C/C++) β
β β
β ββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ β
β β TapirusDB Core Engine (In-Process) β β
β β 100% Safe Rust β’ Idle RAM < 4 MB β β
β βββββββββ¬βββββββββββββββββββββ¬ββββββββββββββββββββββββ¬ββββββββββββββββββββ¬ββββββββ β
β β β β β β
β βββββββββΌβββββββββ βββββββββΌβββββββββ βββββββββΌβββββββββ βββββββββΌβββββββββ β
β β 1. Relational β β 2. Schema-less β β 3. AI Vector β β 4. Knowledge β β
β β SQL Tables β β JSON Document β β HNSW + IVF β β Graph Engine β β
β β Slotted B+Tree β β Collection β β (SIMD/RaBitQ) β β (openCypher) β β
β βββββββββ¬βββββββββ βββββββββ¬βββββββββ βββββββββ¬βββββββββ βββββββββ¬βββββββββ β
β ββββββββββββββββββββββ΄ββββββββββββββββββββββββ΄ββββββββββββββββββββ β
β β Direct In-Memory Traversal β
β βΌ (Sub-Microsecond Zero-IPC Chaining) β
β ββββββββββββββββββββββββββββββββββββββββββββββββ β
β β Anthropic Model Context Protocol (MCP) Tools β β
β β tapirus_remember β’ tapirus_recall β’ SQL β β
β ββββββββββββββββββββββββ¬ββββββββββββββββββββββββ β
β β Direct File I/O (WAL + 4KB Slotted Pages) β
β βΌ β
β ββββββββββββββββββββββββββββββββββββββββββββββββ β
β β Single Encrypted Database File Container β β
β β β’ app.tapir (Authenticated Ciphertext)β β
β β β’ app.tapir-wal (ACID Append-Only Log) β β
β ββββββββββββββββββββββββββββββββββββββββββββββββ β
ββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
For multi-million vector scale, TapirusDB pairs IvfIndex) with RaBitQ (Random Rotation Quantization):
-
Fast Walsh-Hadamard Transform ($O(N \log N)$): Orthogonal sign-flip rotation equalizes coordinate variance across high dimensions without the
$O(N^2)$ memory overhead of dense projection matrices. -
Extreme Compression Ratio: 1-bit binary sign packing shrinks vectors into
u64bitmasks:$$\text{1536 Dimensions (FP32)} = 6,144 \text{ bytes} \longrightarrow \mathbf{196 \text{ bytes}} \quad (\mathbf{>31\times \text{ Reduction}})$$ -
Single-Cycle Hardware POPCNT: Vector distance evaluation executes in single-cycle CPU instructions using hardware
POPCNT(count_ones()). -
Multi-Probe Search: Multi-probe clustering inspects only
$n_{\text{probe}} \ll K$ Voronoi cells, pruning ~95% of the vector search space before scoring.
Traditional graph systems suffer from pointer indirection and binary join memory explosion. TapirusDB implements:
-
Contiguous Slice Adjacency: Adjacency lists are stored in contiguous flat memory arrays (
outgoing_offsets,outgoing_targets,outgoing_weights). Callingcsr.outgoing_neighbors(node_id)returns a contiguous slice&[u64]with zero heap allocations and instant hardware prefetching. -
Worst-Case Optimal Join (WCOJ) Primitives: Rapid edge existence checks in
$O(\log d)$ via binary search on sorted neighbor slices, with two-pointer intersection sweeps for triangle counting (triangle_count()). -
Declarative openCypher: Full support for standard pattern matching:
MATCH (u:User)-[:FOLLOWS*1..3]->(v:User) WHERE u.id = 1 AND v.active = true RETURN v.name, count(*)
Rather than executing expensive unconstrained global vector scans across gigabytes of embeddings, TapirusDB executes Seed-and-Traverse GraphRAG:
User Query βββΊ [IVF/PQ Asymmetric Seeding] βββΊ Top 2-3 Seed Entities
β β
(Sub-millisecond βΌ
Centroid Pruning) [Micro-Hop CSR Traversal]
(BFS 1-2 Hops, Contiguous Memory:
Extract Factual Knowledge Subgraph)
β
βΌ
[Tri-Modal RRF Fusion]
(Vector + BM25 Lexical + Graph Proximity)
β
βΌ
[Prompt Context Synthesizer]
(Compact, Hallucination-Free Markdown)
- Enterprise Multi-Tenant & Scoped ACLs: Filter graph traversals and vector candidate ranking on-the-fly using
tenant_idandallowed_roles(conn.graph_rag_query_scoped()), ensuring sensitive contextual subgraphs never leak across tenants or privilege tiers.
TapirusDB features an automated cost-based query optimizer (src/sql/planner.rs):
- Computes disk I/O page fetch costs and CPU tuple comparison costs.
- Automatically selects between Sequential Scan, B+Tree Secondary Index Scan, and Primary Key Point Lookup.
EXPLAIN QUERY PLANoutputs estimated execution cost and expected row cardinality.
Traditional distributed stacks decouple graph databases and vector stores, causing high network serialization latency and memory-prohibitive global vector scans. TapirusDB executes native bidirectional in-memory chaining at 1,606,037 ops/sec (0.55 Β΅s):
-
Graph-to-Vector (Targeted Scored Neighborhoods): Traverses structured entity relationships first (
$A \to B$ ), then restricts vector distance scoring strictly to candidate neighborhood nodes ($M \ll N$ ). Yields 100% exact Recall with zero approximation loss ($O(M \cdot D)$ instead of $O(N \log N)$). - Vector-to-Graph (Seed-and-Traverse GraphRAG): Uses ANN centroids to locate seed nodes, then instantly expands 1-hop and 2-hop CSR slices to extract factual context, eliminating LLM hallucinations.
-
Autonomous Agent Long-Term Memory (LTM): Automatically balances semantic vector similarity (
$S_v$ ), BM25 lexical precision ($S_l$ ), and exponential temporal recency decay:$$\text{RecallScore}(m) = w_v \cdot S_v + w_l \cdot S_l + w_r \cdot e^{-\lambda \Delta t} + w_i \cdot \text{Importance}$$
Unlike fragmented multi-database architectures where cross-model consistency is impossible without complex distributed consensus (2PC/Sagas), TapirusDB provides true single-transaction atomicity across all four data models:
- Single Atomic Commit: A transaction can update a Relational SQL state row, insert an unstructured JSON document, link openCypher knowledge graph nodes and edges, and index a vector embedding within a single
conn.begin_transaction(). - Zero Torn States: If any operation fails or the host process loses power, Write-Ahead Log (WAL) crash recovery rolls back all four models simultaneously to their exact pre-transaction state, eliminating cross-model state corruption forever.
TapirusDB's quad-model engine (Relational SQL + Vector Search + openCypher Graph + JSON Documents) inside a single encrypted .tapir container solves mission-critical industrial challenges without multi-database operational overhead:
| Industrial Domain | How Quad-Model Solves It Without Server Clusters |
|---|---|
| Financial Fraud Detection & AML |
Graph traverses money-mule rings and cyclic transactions ( |
| Cybersecurity Threat Hunting & SIEM | Graph traces Active Directory lateral movement attack vectors; Vector detects polymorphic binary and syscall sequence anomalies; SQL queries firewall events and access control lists in microsecond windows. |
| Supply Chain & Bill-of-Materials (BOM) | Graph manages multi-tiered supplier dependency trees and failure propagation; Vector clusters sensor telemetry patterns; Document ingests unstructured customs and logistics manifests. |
| Healthcare, Genomics & Life Sciences |
Graph traverses Disease |
| Scientific Research & Academic Labs |
Single-file .tapir dataset container guarantees 100% reproducible paper workflows; Graph + Vector + SQL models molecular pathways and tabular metrics in Python/Jupyter with zero Docker dependencies. |
| In-Process Telemetry & Edge BI |
SIMD vectorized accumulators compute AVG/SUM/COUNT across millions of sensor readings in microseconds; Transparent LZ4 cuts disk usage by 70% with zero cloud egress cost. |
| Privacy-First Smart Home & Home Assistant | Graph maps Zigbee/Matter/Thread device meshes; Vector performs local voice intent matching offline; WAL guarantees crash durability across home power outages on Raspberry Pi (<4MB RAM). |
| Air-Gapped Sovereign Hardware & Edge IoT | Operates on Raspberry Pi, avionics, drones, and naval vessels with zero server daemons, < 4 MB idle RAM, and SIMD-accelerated ChaCha20-Poly1305 AEAD encryption at rest. |
TapirusDB ships as a single zero-dependency standalone binary (tapirus):
High-throughput in-process developer code search combining regex matching, BM25 token overlap, and local semantic vector similarity:
# Search codebase with semantic vector ranking enabled
tapirus tg --vector "transaction rollback wal" src/
# Case-insensitive search filtered by file extensions
tapirus grep -i --ext rs,toml "quantization" .tapirus production.tapirtapirus> CREATE TABLE users (id INTEGER PRIMARY KEY, name TEXT);
Query OK, 1 row(s) affected
tapirus> INSERT INTO users VALUES (1, 'Alex Chen');
Query OK, 1 row(s) affected
tapirus> SELECT * FROM users;
+----+------------+
| id | name |
+----+------------+
| 1 | Alex Chen |
+----+------------+
(1 row(s))
Launch an embedded database as a high-throughput, secure REST API with zero external dependencies:
# Launch with token authentication and database encryption
tapirus serve --port 3005 --api-key "your_secret_api_key" --passphrase "vault_secret" production.tapir- Constant-Time Verification: Prevents timing side-channel attacks via
subtle::ConstantTimeEq. - Flexible Authentication: Provide the key via
Authorization: Bearer <KEY>,X-API-Key: <KEY>, or query?api_key=<KEY>. - Healthcheck Probe:
GET /healthorGET /api/healthreturns operational status without authentication. - Query Execution:
POST /api/sqlorPOST /sqlaccepts SQL queries, graph traversals, and document queries.
Connect Claude Desktop, Cursor, or Gemini to TapirusDB over stdio:
{
"mcpServers": {
"tapirus": {
"command": "tapirus",
"args": ["mcp", "agent_memory.tapir"]
}
}
}Stream massive datasets directly into TapirusDB with automatic schema inference and transactional batching:
# 1. Ingest CSV with automated type inference (INTEGER, REAL, TEXT) into SQL tables
tapirus import csv data/telemetry.csv --table sensors --batch 1000 --db production.tapir
# 2. Ingest streaming JSON Lines (JSONL) into Document collections
tapirus import jsonl data/products.jsonl --collection catalog --batch 500 --db production.tapir
# 3. Semantic Markdown Ingestion into AI Episodic Memory & Knowledge Graph
tapirus import md docs/spec.md --namespace robotics --tags "hardware,specs" --session-id "session_01" --db production.tapirEnterprise-grade durability, snapshotting, and auditing tools for mission-critical edge deployments:
# Atomic online backup snapshot (validates encryption KCV prior to copying)
tapirus backup production.tapir backups/prod_2026_snapshot.tapir
# Safe restore (validates page headers and file geometry)
tapirus restore backups/prod_2026_snapshot.tapir restored_production.tapir
# Deep physical integrity audit (scans slotted pages, verifies CRC32 checksums, checks encryption keys)
tapirus verify production.tapirBenchmarks executed on native NVMe SSD hardware (cargo bench --bench tapirus_bench):
| Operation | Throughput | Mean Latency | Median (p50) | Tail (p99) |
|---|---|---|---|---|
| Relational Primary Key Point Lookup | 362,733 ops/sec | 2.61 Β΅s | 2.37 Β΅s | 4.68 Β΅s |
| CSR Graph Adjacency Sweep | 3,493,852 ops/sec | 0.23 Β΅s | 0.21 Β΅s | 0.37 Β΅s |
| Graph-to-Vector Bidirectional Chaining | 1,606,037 ops/sec | 0.55 Β΅s | 0.51 Β΅s | 1.01 Β΅s |
| Relational B+Tree Inserts | 149,176 ops/sec | 6.19 Β΅s | 4.66 Β΅s | 61.13 Β΅s |
| JSON Document Path Lookups | 355,004 docs/sec | 2.70 Β΅s | 2.58 Β΅s | 5.45 Β΅s |
| HNSW Vector Search (32D, k=5) | 51,060 QPS | 19.53 Β΅s | 17.06 Β΅s | 50.36 Β΅s |
| RaBitQ Asymmetric POPCNT Distance | > 12,000,000 ops/sec | 0.08 Β΅s | 0.08 Β΅s | 0.12 Β΅s |
| WAL Durable Disk Writes | 107,875 writes/sec | 9.15 Β΅s | 6.71 Β΅s | 62.21 Β΅s |
| AI Memory Ingest (BM25 Indexing) | 416,529 ops/sec | 2.30 Β΅s | 1.77 Β΅s | 4.38 Β΅s |
Cloud Vector DB (gRPC Roundtrip) [ββββββββββββββββββββββββββββββββββββββββ] 25,000 Β΅s (25.0 ms - WAN Network Hop)
Dedicated Graph DB (HTTP/JVM) [ββββββββββββββββββββββββ] 15,000 Β΅s (15.0 ms - TCP / JVM GC)
Relational SQL Server (TCP IPC) [ββββββββ] 5,000 Β΅s (5.0 ms - Unix Socket / IPC)
TapirusDB Combined Graph-Vector [β] 0.51 Β΅s (In-Process CPU Memory Bus)
Tested on native NVMe SSD hardware with true Write-Ahead Log (WAL) durability:
| Dimension | TapirusDB (In-Process) | Traditional Network DBs | Concrete Operational Value |
|---|---|---|---|
| Graph-Vector Retrieval | 0.51 Β΅s (median) | ~25,000 Β΅s (25 ms) | Sub-microsecond local reasoning vs. WAN gRPC network serialization hop. |
| Durable Disk Writes | 107,875 writes/sec | ~2,000β8,000 ops/sec | Real ACID WAL disk commits, not volatile in-memory caching. |
| Idle Memory Footprint | < 4 MB RAM | > 1.2 GB (multi-daemon) | Fits comfortably in Raspberry Pi, edge robotics, and local desktop apps. |
| Initial File Footprint | 4,096 Bytes | Server cluster required | Single encrypted .tapir container; zero cloud daemons to configure. |
π¬ Transparent & Peer-Reviewed Methodology:
We publish complete hardware specifications, statistical variance ($\sigma$ ), and cache-miss analysis in our Systems Architecture Paper (PAPER_TAPIRUSDB.md).Verify & run the benchmark suite yourself on your machine (1 command):
git clone https://github.com/tapiruslab/TapirusDB.git cd TapirusDB cargo bench --bench tapirus_benchFull tail percentiles (p50, p95, p99, Min, Max) will be automatically exported to
target/tapirus_bench_results.jsonfor independent peer review.
Engineering honesty is paramount. Choosing the right storage engine requires understanding boundary trade-offs:
| Workload & Scenario | Recommended Engine | Architectural Rationale |
|---|---|---|
| Local AI Agents & LLM RAG Memory | β TapirusDB | Microsecond episodic retrieval, combined vector + openCypher graph in one atomic .tapir file. |
| Embedded Edge, Robotics & IoT Hardware | β TapirusDB | < 4 MB idle RAM, 100% Safe Rust core, zero background daemon processes or JVM runtimes. |
| Desktop Apps, CLI Tools & Local-First Web | β TapirusDB | Single-file portability, zero server configuration, pure client-side SQLite/Mongo alternative. |
| Petabyte Distributed Big Data Warehousing | β ClickHouse / Snowflake | TapirusDB is optimized for operational single-node/in-process workloads, not massive multi-rack OLAP scans. |
| Multi-Region Active-Active Distributed Writes | β CockroachDB / Spanner | For global multi-master write replication, use dedicated distributed consensus databases. |
| Complex Analytical BI Cubes over Billions of Rows | β DuckDB / ClickHouse | DuckDB is superior for vectorized columnar OLAP; TapirusDB excels at transactional, graph, vector, and episodic AI memory. |
- TLA+ Specifications: Write-Ahead Logging (WAL) state transitions and crash recovery are formally modeled under TLA+ in
docs/formal_verification/. - Memory Safety Contract: Strict
#![forbid(unsafe_code)]enforced across all core modules insrc/lib.rs.
- π Production API Cookbook & Code Recipes (Simple Website/Desktop RAG, 1M+ Low-Spec Research Analytics, Hybrid RRF Search, GraphRAG, Bulk Ingestion, Time-Travel, Window Functions)
- π§ Tutorial: Autonomous AI Agent Memory in 30 Minutes
- π Multi-Language SDK & C-ABI Integration Guide
- π Systems Architecture Deep Dives:
- βοΈ GraphRAG & Cloud S3/R2 Remote Storage
- π Scientific Systems Architecture Paper
- βοΈ C ABI & Native Foreign Function Interface
- π‘οΈ Formal Verification Suite (TLA+)
- βοΈ Software License (BUSL 1.1)
Developed & Maintained by TapirusDB Contributors β’ Tapirus Tech Lab (tapirusdb.com)
Contact: contact@tapirusdb.com