Skip to content

Latest commit

Β 

History

69 Commits

Folders and files

NameName
Last commit message
Last commit date
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 

Repository files navigation

TapirusDB

The Embedded Cognitive Memory & Multi-Model Engine for Sovereign AI & Edge Systems

Sub-Microsecond Agent Memory β€’ openCypher Knowledge Graphs β€’ Vector Search β€’ Relational SQL β€’ Documents
Single Encrypted .tapir File β€’ 100% Safe Rust β€’ < 4 MB Idle RAM β€’ Zero Cloud Daemons


Crates.io Memory Safety License: BUSL-1.1 Documentation Downloads Hub Tapirus Studio Edge AI & Robotics GitHub Releases


"Build private AI memory without operating a data stack."
TapirusDB is the high-performance, embedded cognitive memory engine for local AI agents, robotics, and sovereign edge hardware. It collapses vector similarity, knowledge graphs, relational metadata, and JSON documents into a single encrypted .tapir file with sub-microsecond in-process retrieval ($0.51\ \mu\text{s}$) and zero memory corruption risk.

πŸ–₯️ Need a Visual Database Manager (like phpMyAdmin or Supabase Studio)?
Use Tapirus Studio β€” our free visual GUI companion for TapirusDB!
β€’ 🌐 Run Instant In-Browser: tapirusdb.com/studio (Zero installation required)
β€’ ⬇️ Official Download Landing Page: tapirusdb.com/download.html (Windows, macOS, Linux, CLI)
β€’ πŸ“¦ Releases & Binary Downloads: github.com/tapiruslab/TapirusDB/releases


Why TapirusDB? (Kill the "Frankenstack")

Modern AI and edge developers are forced into Fragmented Polyglot Persistenceβ€”gluing together multiple complex, heavy databases across network boundaries:

TapirusDB Architecture vs The Fragile Frankenstack


Highlights

  • πŸ¦€ 100% Pure Safe Rust (#![forbid(unsafe_code)]): Guaranteed memory safety at compile-time. Zero buffer overflows, zero dangling pointers, zero use-after-free vulnerabilities, and zero C/C++ memory corruption CVEs.
  • πŸ“¦ True In-Process Architecture (Zero-IPC): Compiles and links directly into your binary (Rust, Python, TypeScript, C/C++, Go). No background database servers (mysqld, postgres, mongod), zero network serialization overhead, and sub-microsecond in-memory query traversal.
  • ⚑ Quad-Model Data Consolidation: Seamlessly unifies Relational SQL-92, HNSW & IVF Vector Search, openCypher Property Graphs, and MongoDB-style JSON Documents inside a single B+Tree slotted-page file.
  • πŸ“ˆ Advanced SQL Window Functions & Graph Algorithms: Built-in ANSI SQL window operations (ROW_NUMBER(), RANK(), DENSE_RANK(), NTILE(), LAG(), LEAD()) and native graph topology algorithms (GRAPH ALGORITHM louvain, betweenness, connected_components, pagerank).
  • 🧠 Production GraphRAG & AI Memory: Built-in seed-and-traverse GraphRAG with Tri-Modal Reciprocal Rank Fusion (RRF), episodic memory with exponential temporal decay, and isolated agent namespaces.
  • ⚑ Tap Sub-Millisecond Cognitive Instinct Engine: In-database non-autoregressive System-1 decision core (Tap::classify, Tap::score, Tap::verify, Tap::route). Run 1,000 deterministic agent decisions per second directly inside SQL queries (TAP_CLASSIFY(), TAP_VERIFY()) in pure Safe Rust (< 2ms) with zero cloud tokens and zero API costs.
  • 🎯 Dynamic 16-Lane SIMD & RaBitQ 32x Quantization: Parallel AVX-512 / AVX2 / NEON vector kernels combined with Fast Walsh-Hadamard 1-bit/2-bit random rotation quantization, reducing 1536-D embeddings from 6,144 bytes to 196 bytes with single-cycle POPCNT distance evaluation.
  • πŸ•ΈοΈ Compressed Sparse Row (CSR) Topology & openCypher: Contiguous adjacency arrays on disk and in memory for zero-allocation slice neighbor sweeps, paired with standard declarative openCypher syntax (MATCH ... WHERE ... RETURN ...).
  • βš›οΈ Atomic Four-Model Transactions: Single ACID transaction committing or rolling back across SQL rows, JSON documents, openCypher graph edges, and Vector embeddings simultaneously with zero torn states.
  • πŸ” Multi-Tenant & Role-Scoped GraphRAG: Native tenant isolation (tenant_id) and role-based ACL filtering (allowed_roles) across knowledge graph traversal and vector scoring.
  • πŸ›‘οΈ Physical Integrity Audit & Safe Hot Backups: Zero-downtime atomic backup snapshots, KCV key validation, and slotted-page CRC32 consistency verification (tapirus backup, tapirus restore, tapirus verify).
  • πŸš€ Streaming Data Importer & Protected REST Daemon: High-throughput streaming ingest for CSV (auto-inferred schema), JSONL, and Markdown straight into tables and AI memory; secure embedded REST server (tapirus serve) with constant-time Bearer/API-key verification.
  • 🌐 S3/R2 Remote Range Streaming: On-demand 4KB page streaming directly from cloud object stores via HTTP Range requests with zero local disk footprint.
  • πŸ€– Native Model Context Protocol (MCP): Out-of-the-box stdio JSON-RPC 2.0 server (tapirus mcp) for Claude Desktop, Cursor, and Gemini autonomous agents.

Table of Contents


⚑ Tap Decision Core: Sub-Millisecond In-Database Instinct Engine

Traditionally, when autonomous AI agents make structured decisions (categorizing a support ticket, checking a policy claim, scoring urgency, or choosing a graph execution branch), developers have been forced to pay an exorbitant "LLM Latency & Cost Tax": sending database payloads over HTTP to an autoregressive model like GPT-4o-mini, waiting 400ms–800ms, risking JSON formatting hallucinations, and paying per-token API bills.

TapirusDB changes this paradigm with Tap.

Tap is TapirusDB's native System-1 cognitive instinct subsystem. Built directly in Safe Rust with zero external services and zero background Python runtimes, Tap executes deterministic single-pass decision projections directly over database records in under 2 milliseconds.

[ Traditional LLM API (e.g. GPT-4o-mini) ]  ══════════════════════════════════ 450ms - 800ms
[ Remote Decision Microservice (HTTP)   ]  ══════════════════ 85ms - 150ms
[ Python Decision Server (PyTorch)      ]  ════════ 35ms - 65ms

TapirusDB Tap Decision Core vs External LLM Stack

The 4 Native Decision Primitives

Primitive Purpose Rust API SQL Syntax
classify Categorical selection with calibrated softmax distribution conn.tap().classify(text, candidates) SELECT TAP_CLASSIFY(body, '["fraud", "legit"]')
score Continuous rubric evaluation in $[0.0, 1.0]$ conn.tap().score(text, criteria) SELECT TAP_SCORE(incident, 'emergency_severity')
verify Calibrated boolean truth validation with strict margin conn.tap().verify(premise, hypothesis) SELECT id FROM claims WHERE TAP_VERIFY(claim, 'active_policy') = 1
route Autonomous graph & workflow branch selection conn.tap().route(state, routes) SELECT TAP_ROUTE(task_state, 'retry, escalate, resolve')

1. In-Database SQL Integration

Tap functions can be executed directly inside standard SQL SELECT projections and WHERE filtering clauses:

-- Categorize and score incoming tickets in a single database pass (< 2ms)
SELECT 
    id, 
    customer, 
    TAP_CLASSIFY(message, '["billing", "technical", "sales"]') AS category,
    TAP_SCORE(message, 'critical system outage emergency') AS urgency_score
FROM support_inbox;

-- Filter fraud or compliance violations directly in SQL WHERE clause
SELECT id, transaction_amount, merchant
FROM transaction_audit
WHERE TAP_VERIFY(notes, 'unauthorized account takeover attempt') = 1;

2. Rust Bare-Metal Native Instincts

use tapirus::{Connection, Result};

fn main() -> Result<()> {
    let conn = Connection::open_in_memory()?;

    // 1. Categorical Decision (< 1.5ms)
    let decision = conn.tap().classify(
        "Refund requested because package arrived damaged", 
        &["refund", "billing", "sales_inquiry"]
    )?;
    println!("Action: {} (Confidence: {:.2}%)", decision.top_choice, decision.confidence * 100.0);

    // 2. Truth Verification (< 1.2ms)
    let verified = conn.tap().verify(
        "User confirmed receipt of digital product license", 
        "digital license successfully received"
    )?;
    if verified.is_verified {
        println!("Claim verified with margin: {:.3}", verified.margin);
    }

    // 3. Autonomous Graph Branch Routing (< 1.8ms)
    let step = conn.tap().route(
        "Payment gateway returned code 504 gateway timeout", 
        &["retry_transaction", "fallback_processor", "cancel_order"]
    )?;
    println!("Next Workflow Step: {}", step.selected_route);

    Ok(())
}

3. Python SDK Native Integration

import tapirus

# Query with embedded Tap decision functions
conn = tapirus.connect(":memory:")
conn.execute("CREATE TABLE claims (id INTEGER PRIMARY KEY, details TEXT);")
conn.execute("INSERT INTO claims VALUES (1, 'Claim filed for broken windshield from hailstorm');")

rows = conn.query("SELECT id, TAP_VERIFY(details, 'weather damage claim') AS valid_weather FROM claims;")
print(rows)  # [{'id': 1, 'valid_weather': 1}]

# Or call direct decision helpers
verdict, conf = tapirus.tap_classify("Server disk full emergency", ["infrastructure", "billing", "general"])
print(f"Top Category: {verdict} ({conf*100:.1f}%)")

Quickstart

1. Installation

Official Downloads Hub & Visual Studio

Supported Platforms & Distributions

TapirusDB is 100% self-contained with zero cloud or daemon dependencies. Pre-compiled binaries run out-of-the-box across:

Platform Family Architecture Supported Operating Systems & Distros
Linux (Universal glibc) x86_64, aarch64 Debian, Ubuntu, Fedora, RHEL, CentOS, Rocky Linux, AlmaLinux, Arch Linux, openSUSE, Amazon Linux 2/2023
Linux (musl & Containers) x86_64, aarch64 Alpine Linux, Docker / OCI (ghcr.io/tapiruslab/tapirusdb), Embedded Linux / IoT
macOS Apple Silicon & Intel macOS 12+ (Monterey, Ventura, Sonoma, Sequoia)
Windows x86_64 Windows 10, Windows 11, Windows Server 2019/2022/2025
WebAssembly (WASM) wasm32 All modern web browsers (Chrome, Edge, Safari, Firefox) via Tapirus Studio

Package Managers (Terminal & CLI)

# macOS & Linux (Homebrew)
brew install https://raw.githubusercontent.com/tapiruslab/TapirusDB/main/Formula/tapirus.rb

# Windows (Windows Package Manager)
winget install --manifest https://raw.githubusercontent.com/tapiruslab/TapirusDB/main/winget/tapirus.yaml
# (Or 'winget install tapirus' once indexed in Microsoft community repo)

# Linux / macOS Automated Script
curl -fsSL https://raw.githubusercontent.com/tapiruslab/TapirusDB/main/install.sh | bash

# Windows PowerShell Automated Script
irm https://raw.githubusercontent.com/tapiruslab/TapirusDB/main/install.ps1 | iex

Language SDKs & Client Libraries

# Rust Engine
cargo add tapirus

# Python SDK (Python 3.9+)
pip install tapirus

# Node.js & TypeScript SDK
npm install tapirus

# Bun Runtime
bun add tapirus

# Go SDK
go get github.com/tapiruslab/TapirusDB/sdks/go

# PHP Composer
composer require tapiruslab/tapirusdb

# OCI Container (Docker & Podman)
docker pull ghcr.io/tapiruslab/tapirusdb:latest

Official Ecosystem & Registry Matrix

πŸ’‘ Looking for FFI integration, C-ABI bindings, or native shared libraries? See the Multi-Language SDK & FFI Guide.

Ecosystem Registry / Package Installation Command License
Rust Rust Crates.io cargo add tapirus BUSL-1.1
Python Python PyPI pip install tapirus MIT
Node.js Node.js / TS npm npm install tapirus MIT
Go Go Go Reference go get github.com/tapiruslab/TapirusDB/sdks/go MIT
PHP PHP Packagist composer require tapiruslab/tapirusdb MIT
Docker Docker Docker docker pull ghcr.io/tapiruslab/tapirusdb:latest BUSL-1.1
Linux Linux Linux curl -fsSL .../install.sh | bash BUSL-1.1
Windows Windows Winget winget install --manifest ... BUSL-1.1
Homebrew macOS (Brew) Homebrew brew install .../tapirus.rb BUSL-1.1

2. Code in 30 Seconds

Rust: Relational SQL & AI Vector Search

use tapirus::{Connection, DistanceMetric, Result};

fn main() -> Result<()> {
    // Open in-memory or single-file database: "production.tapir"
    let db = Connection::open_in_memory()?;

    // 1. Create table with structured columns and dense vector embedding
    db.execute("
        CREATE TABLE documents (
            id INTEGER PRIMARY KEY,
            title TEXT NOT NULL,
            category TEXT NOT NULL,
            embedding VECTOR(4)
        );
    ")?;

    db.execute("
        INSERT INTO documents VALUES 
        (1, 'Safe Systems in Rust', 'tech', [0.95, 0.05, 0.0, 0.0]),
        (2, 'Neural Vector Databases', 'ai', [0.10, 0.90, 0.15, 0.0]);
    ")?;

    // 2. Hybrid Vector Search with Single-Pass SQL Pre-Filtering (Exact k Recall)
    let rows = db.query("
        SELECT id, title 
        FROM documents 
        VECTOR NEAR embedding = [0.92, 0.08, 0.0, 0.0] TOP 1
        WHERE category = 'tech';
    ")?;

    for row in rows {
        println!("Match: {}", row.get::<String>("title")?);
    }

    Ok(())
}

Rust: Declarative openCypher Graph Pattern Matching

use tapirus::{Connection, Result};

fn main() -> Result<()> {
    let db = Connection::open_in_memory()?;

    // 1. Ingest entities and relationships
    db.execute("GRAPH INSERT NODE 1 LABEL 'Person' PROPERTIES '{\"name\": \"Alice\"}';")?;
    db.execute("GRAPH INSERT NODE 2 LABEL 'Person' PROPERTIES '{\"name\": \"Bob\"}';")?;
    db.execute("GRAPH INSERT NODE 3 LABEL 'Company' PROPERTIES '{\"name\": \"TapirusTech\"}';")?;

    db.execute("GRAPH INSERT EDGE 1 -> 2 LABEL 'KNOWS' WEIGHT 0.9;")?;
    db.execute("GRAPH INSERT EDGE 2 -> 3 LABEL 'WORKS_AT' WEIGHT 1.0;")?;

    // 2. Query graph patterns using industry-standard openCypher
    let rows = db.query("
        MATCH (a:Person)-[r:KNOWS]->(b:Person) 
        WHERE b.name = 'Bob' 
        RETURN a.name, b.name, r.weight;
    ")?;

    for row in rows {
        println!("{} knows {} (weight: {})", 
            row.get::<String>("a.name")?, 
            row.get::<String>("b.name")?, 
            row.get::<f64>("r.weight")?
        );
    }

    Ok(())
}

Rust: Bidirectional Graph-Vector Chaining & AI Agent Memory

use tapirus::{Connection, MemoryRecallFilter, Result};
use tapirus::vector::DistanceMetric;

fn main() -> Result<()> {
    let conn = Connection::open_in_memory()?;

    // 1. Graph-to-Vector Chaining (Sub-microsecond 0.55 Β΅s retrieval)
    // Constrains vector distance calculations strictly to local graph neighborhood O(M Β· D)
    let candidates = conn
        .chain(1)                              // Seed Patient Node
        .out(Some("TREATS"))                  // Traverse outgoing relationships
        .filter_label("Medicine")             // Target node label
        .vector_near(&[0.90, 0.10, 0.0, 0.0], 5, DistanceMetric::Cosine)?;

    // 2. Vector-to-Graph Chaining (Seed-and-Traverse)
    // Seeds from query vector, then traverses adjacent knowledge subgraph
    let discovered = conn
        .chain_from_vector(&[0.85, 0.15, 0.0, 0.0], 1)?
        .out(Some("AUTHORED_BY"))
        .collect_nodes();

    // 3. Autonomous AI Agent Long-Term Memory (LTM)
    // Multi-modal recall: Dense Vector + BM25 Lexical + Recency Decay (e^-λΔt)
    let memory_id = conn.memory_remember(
        "User prefers sovereign on-device processing and strict privacy",
        Some(&[0.92, 0.08, 0.0, 0.0]),
        0.95, // Importance priority score
        &["preferences", "privacy"],
    )?;

    let filter = MemoryRecallFilter::default(); // Balanced Vector + BM25 + Recency
    let recalled = conn.memory_recall(Some("sovereign privacy"), None, 3, &filter);
    println!("Recalled Agent Memory: {}", recalled[0].entry.content);

    Ok(())
}

Python: Clean Native Integration

import tapirus

# Connect directly to local encrypted vault or in-memory
conn = tapirus.connect("app.tapir")

# 1. Relational SQL & Window Functions
conn.execute("CREATE TABLE telemetry (id INTEGER PRIMARY KEY, sensor TEXT, value REAL);")
conn.execute("INSERT INTO telemetry VALUES (1, 'temp', 23.8), (2, 'temp', 24.1), (3, 'temp', 22.9);")
records = conn.query("""
    SELECT id, sensor, value, 
           ROW_NUMBER() OVER (ORDER BY value DESC) as rank 
    FROM telemetry;
""")
print(records)  # [{'id': 2, 'sensor': 'temp', 'value': 24.1, 'rank': 1}, ...]

# 2. Native Vector Search
conn.execute("CREATE TABLE docs (id INTEGER PRIMARY KEY, vec VECTOR(3));")
conn.execute("INSERT INTO docs VALUES (1, [0.9, 0.1, 0.0]), (2, [0.1, 0.9, 0.0]);")
top_docs = conn.vector_search("docs", "vec", [0.85, 0.15, 0.0], top_k=1)

# 3. Native Graph Algorithms
community_map = conn.graph_algorithm("louvain")
conn.checkpoint()

Node.js & TypeScript: Zero-Daemon Embedded Database

import { TapirusClient, open } from "tapirusdb";

// Connect to single-file database
const db = new TapirusClient({ dbPath: "production.tapir" });

// 1. Relational SQL with Window Functions
await db.execute("CREATE TABLE users (id INT PRIMARY KEY, name TEXT, score REAL);");
await db.execute("INSERT INTO users VALUES (1, 'Alice', 95.5), (2, 'Bob', 88.0);");
const ranked = await db.query(`
  SELECT name, score, 
         RANK() OVER (ORDER BY score DESC) as leaderboard_rank 
  FROM users;
`);
console.log(ranked);

// 2. Built-in SIMD Vector Search & Graph Clustering
const neighbors = await db.vectorSearch("docs", "vec", [0.9, 0.1, 0.0], 5);
const communities = await db.graphAlgorithm("louvain");

Beyond AI: An Ultra-Fast Embedded Database for Classic Applications

While TapirusDB is the premier memory engine for sovereign AI and robotics, you do not need AI to benefit from TapirusDB. It is also a first-class, zero-configuration embedded database for general applications, edge systems, and analytics:

1. Modern Drop-In Replacement for SQLite (Full SQL-92 + ACID)

Need reliable relational tables, transactions, and foreign keys without AI? TapirusDB provides standard SQL with pure Safe Rust reliability:

// Standard Relational SQL with ACID transactions
db.execute("CREATE TABLE accounts (id INTEGER PRIMARY KEY, email TEXT, balance REAL);")?;
db.execute("INSERT INTO accounts VALUES (1, 'alice@example.com', 1250.50);")?;

// Complex queries with Subqueries & CTEs
let rows = db.query("
    WITH active_accounts AS (
        SELECT id, email, balance FROM accounts WHERE balance > 1000.0
    )
    SELECT * FROM active_accounts;
")?;
  • Advanced Query Engine: Built-in subqueries, CTEs (WITH ... AS), INNER/LEFT JOIN, and Cost-Based Optimizer (CBO).
  • Transparent Encryption Included: SIMD-accelerated ChaCha20-Poly1305 AEAD encryption at rest (RFC 8439) without paying for proprietary SQLite commercial extensions.

2. Embedded MongoDB Alternative (Schema-less JSON Documents)

Need to store dynamic payloads, user settings, or sensor telemetry with flexible schemas?

let collection = db.collection("telemetry")?;
let doc_id = collection.insert_one(&serde_json::json!({
    "sensor_id": "temp_probe_09",
    "reading_celsius": 24.3,
    "calibration": { "offset": 0.05, "certified": true },
    "tags": ["factory_floor", "zone_b"]
}))?;

3. In-Process Analytics & SIMD Aggregations

  • SIMD Aggregations: Vectorized SUM, AVG, COUNT processing multi-megabyte datasets in microseconds.
  • Transparent Compression: Built-in pure Safe Rust LZ4 page compression reduces disk footprint by 50%–70%.
  • Developer CLI: Fast code search tool tapirus tg built right into the binary.

🎯 High-Impact Real-World Domains: Research, Analytics & Smart Home

TapirusDB's zero-dependency single-file architecture is purpose-built for environments where spinning up complex database server clusters is impossible, expensive, or counterproductive:

1. πŸ”¬ Scientific Research & Academic Laboratories

  • 100% Reproducible Research Bundles: Peer reviewers and researchers no longer need to configure Docker containers, PostgreSQL, Neo4j, and Milvus just to run a paper's code. Package an entire multimodal datasetβ€”molecular/protein graphs, high-dimensional vector embeddings, and assay measurement SQL tablesβ€”into a single verifiable experiment.tapir file.
  • Zero-Setup Python & Jupyter Workflows: Install in seconds (pip install tapirus) and query directly inside Jupyter notebooks without starting any background daemons.
  • Guaranteed Memory Determinism: 100% Pure Safe Rust (#![forbid(unsafe_code)]) guarantees zero memory leaks, buffer overruns, or segfault crashes during 72-hour batch computation runs.
# Python/Jupyter Research Workflow
import tapirus

# Open single research dataset container
db = tapirus.open("paper_dataset.tapir")

# Query molecular knowledge graph combined with chemical vector distance
results = db.query("""
    MATCH (c:Compound)-[:BINDS_TO]->(p:Protein {id: 'EGFR'})
    WHERE c.smiles_vector <-> $query_vec < 0.15
    RETURN c.id, c.affinity_score;
""", query_vec=target_embedding)

2. πŸ“Š High-Performance In-Process Analytics & Edge BI

  • Zero-IPC Columnar Aggregations: Vectorized SUM, AVG, and COUNT accumulators run directly across local memory pages with sub-microsecond execution times, eliminating network hop overhead completely.
  • Transparent LZ4 Disk Compression: Built-in page compression slashes disk space by 50%–70%, allowing edge gateways and industrial PCs to retain months of historical sensor telemetry locally.
  • Zero Cloud Egress Costs: Query, aggregate, and analyze high-frequency telemetry at the edge without paying exorbitant bandwidth and ingress bills to cloud data warehouses.
// In-Process Telemetry Aggregation with Common Table Expressions (CTEs)
let summary = db.query("
    WITH sensor_rollup AS (
        SELECT sensor_id, AVG(reading) AS avg_reading, COUNT(*) AS samples
        FROM telemetry_logs
        WHERE timestamp >= NOW() - 3600
        GROUP BY sensor_id
    )
    SELECT * FROM sensor_rollup WHERE avg_reading > 85.0;
")?;

3. 🏠 Privacy-First Smart Home & Local Automation (Home Assistant / IoT)

  • 100% Sovereign & Local-First: Run entirely offline on a Raspberry Pi 4/5 or Intel NUC with < 4 MB idle RAM. Your private camera triggers, sensor logs, and home conversations never leak to external cloud servers.
  • Mesh Network Topology (openCypher Graph): Model Zigbee, Matter, and Thread device hierarchies natively (MATCH (s:Switch)-[:CONTROLS]->(l:Light)).
  • Offline Voice Intent Matching (Vector Engine): Store speech and intent embeddings locally for sub-millisecond local voice assistant recognition (Whisper / Home Assistant Voice).
  • Blackout Resilience (ACID WAL): If your home experiences an abrupt power outage, TapirusDB's Write-Ahead Log guarantees zero database corruption upon reboot.
// Local Voice Intent Resolution + Zigbee Mesh Pathfinding
let intent_vector = local_whisper.embed("turn off kitchen lights");

// 1. Semantic voice intent match (Vector)
let matched_action = db.vector_search("voice_intents", &intent_vector, 1)?;

// 2. Resolve Zigbee device relay path (openCypher Graph)
let route = db.graph_query("
    MATCH path = (hub:Gateway)-[:ROUTES_THROUGH*1..3]->(d:Device {name: 'kitchen_main_light'})
    RETURN path LIMIT 1;
")?;

Architectural Comparison

Capability TapirusDB v1.0.0 Traditional Relational (SQLite / DuckDB) Dedicated Vector DBs Graph Databases (Neo4j) Document Stores (MongoDB)
Runtime Architecture In-Process Single File In-Process Single File Server Daemon / Cloud Server Daemon (JVM) Server Daemon (mongod)
Memory Safety Model 100% Safe Rust (forbid) C / C++ (Manual memory) Rust / Go / C++ Java / JVM C++
Data Models Supported Quad-Model (SQL+Vec+Graph+Doc) Relational SQL only Vector embeddings only Graph only JSON Document only
AI Vector Search Native HNSW, IVF & RaBitQ None (or slow extension) Native ANN Basic / Extension Add-on Atlas Vector
Vector Quantization RaBitQ 32x (1-Bit/2-Bit) + SQ8 None PQ / SQ None None
Graph Query Engine openCypher + CSR + GraphRAG Recursive CTE only None Native Cypher $graphLookup
Encrypted At-Rest ChaCha20-Poly1305 (Zero-Cost) Commercial Add-on ($$$) Cloud KMS only Enterprise Tier ($$$) Enterprise KMS
Cold Start / Idle RAM < 4 MB RAM ~4 MB (SQLite) / ~35 MB > 500 MB > 1,200 MB > 350 MB
Binary Size ~3.8 MB ~1.5 MB – 42 MB > 150 MB > 300 MB > 200 MB
Multi-Service Sync Drift Zero (Single Container) High (manual ETL) High (CDC pipelines) High (sync lag) High (glue code)

Core Technical Pillars

β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚                               YOUR APPLICATION HOST PROCESS                            β”‚
β”‚           (Rust β€’ Python β€’ TypeScript β€’ Bun β€’ Go β€’ PHP β€’ WebAssembly β€’ C/C++)          β”‚
β”‚                                                                                        β”‚
β”‚   β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”   β”‚
β”‚   β”‚                         TapirusDB Core Engine (In-Process)                     β”‚   β”‚
β”‚   β”‚                        100% Safe Rust β€’ Idle RAM < 4 MB                        β”‚   β”‚
β”‚   β””β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”˜   β”‚
β”‚           β”‚                    β”‚                       β”‚                   β”‚           β”‚
β”‚   β”Œβ”€β”€β”€β”€β”€β”€β”€β–Όβ”€β”€β”€β”€β”€β”€β”€β”€β”   β”Œβ”€β”€β”€β”€β”€β”€β”€β–Όβ”€β”€β”€β”€β”€β”€β”€β”€β”      β”Œβ”€β”€β”€β”€β”€β”€β”€β–Όβ”€β”€β”€β”€β”€β”€β”€β”€β”  β”Œβ”€β”€β”€β”€β”€β”€β”€β–Όβ”€β”€β”€β”€β”€β”€β”€β”€β”  β”‚
β”‚   β”‚ 1. Relational  β”‚   β”‚ 2. Schema-less β”‚      β”‚  3. AI Vector  β”‚  β”‚  4. Knowledge  β”‚  β”‚
β”‚   β”‚   SQL Tables   β”‚   β”‚  JSON Document β”‚      β”‚   HNSW + IVF   β”‚  β”‚  Graph Engine  β”‚  β”‚
β”‚   β”‚ Slotted B+Tree β”‚   β”‚   Collection   β”‚      β”‚  (SIMD/RaBitQ) β”‚  β”‚  (openCypher)  β”‚  β”‚
β”‚   β””β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”˜   β””β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”˜      β””β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”˜  β””β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”˜  β”‚
β”‚           β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜           β”‚
β”‚                                           β”‚ Direct In-Memory Traversal                 β”‚
β”‚                                           β–Ό (Sub-Microsecond Zero-IPC Chaining)        β”‚
β”‚                    β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”                    β”‚
β”‚                    β”‚ Anthropic Model Context Protocol (MCP) Tools β”‚                    β”‚
β”‚                    β”‚ tapirus_remember β€’ tapirus_recall β€’ SQL      β”‚                    β”‚
β”‚                    β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜                    β”‚
β”‚                                           β”‚ Direct File I/O (WAL + 4KB Slotted Pages)  β”‚
β”‚                                           β–Ό                                            β”‚
β”‚                    β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”                    β”‚
β”‚                    β”‚ Single Encrypted Database File Container     β”‚                    β”‚
β”‚                    β”‚   β€’ app.tapir      (Authenticated Ciphertext)β”‚                    β”‚
β”‚                    β”‚   β€’ app.tapir-wal  (ACID Append-Only Log)    β”‚                    β”‚
β”‚                    β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜                    β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜

Pillar 1: Inverted File (IVF) Clustering & RaBitQ 32x Quantization

For multi-million vector scale, TapirusDB pairs $k$-means Voronoi partitioning (IvfIndex) with RaBitQ (Random Rotation Quantization):

  1. Fast Walsh-Hadamard Transform ($O(N \log N)$): Orthogonal sign-flip rotation equalizes coordinate variance across high dimensions without the $O(N^2)$ memory overhead of dense projection matrices.
  2. Extreme Compression Ratio: 1-bit binary sign packing shrinks vectors into u64 bitmasks: $$\text{1536 Dimensions (FP32)} = 6,144 \text{ bytes} \longrightarrow \mathbf{196 \text{ bytes}} \quad (\mathbf{&gt;31\times \text{ Reduction}})$$
  3. Single-Cycle Hardware POPCNT: Vector distance evaluation executes in single-cycle CPU instructions using hardware POPCNT (count_ones()).
  4. Multi-Probe Search: Multi-probe clustering inspects only $n_{\text{probe}} \ll K$ Voronoi cells, pruning ~95% of the vector search space before scoring.

Pillar 2: Compressed Sparse Row (CSR) & openCypher

Traditional graph systems suffer from pointer indirection and binary join memory explosion. TapirusDB implements:

  • Contiguous Slice Adjacency: Adjacency lists are stored in contiguous flat memory arrays (outgoing_offsets, outgoing_targets, outgoing_weights). Calling csr.outgoing_neighbors(node_id) returns a contiguous slice &[u64] with zero heap allocations and instant hardware prefetching.
  • Worst-Case Optimal Join (WCOJ) Primitives: Rapid edge existence checks in $O(\log d)$ via binary search on sorted neighbor slices, with two-pointer intersection sweeps for triangle counting (triangle_count()).
  • Declarative openCypher: Full support for standard pattern matching:
    MATCH (u:User)-[:FOLLOWS*1..3]->(v:User) 
    WHERE u.id = 1 AND v.active = true 
    RETURN v.name, count(*)

Pillar 3: Seed-and-Traverse GraphRAG Engine

Rather than executing expensive unconstrained global vector scans across gigabytes of embeddings, TapirusDB executes Seed-and-Traverse GraphRAG:

User Query ──► [IVF/PQ Asymmetric Seeding] ──► Top 2-3 Seed Entities
                         β”‚                                β”‚
                  (Sub-millisecond                        β–Ό
                   Centroid Pruning)             [Micro-Hop CSR Traversal]
                                                 (BFS 1-2 Hops, Contiguous Memory:
                                                  Extract Factual Knowledge Subgraph)
                                                          β”‚
                                                          β–Ό
                                                 [Tri-Modal RRF Fusion]
                                                 (Vector + BM25 Lexical + Graph Proximity)
                                                          β”‚
                                                          β–Ό
                                            [Prompt Context Synthesizer]
                                            (Compact, Hallucination-Free Markdown)

$$\text{RRF}(e) = \sum_{m \in {\text{vec}, \text{lex}, \text{graph}}} \frac{w_m}{k_{\text{rrf}} + \text{rank}_m(e)}$$

  • Enterprise Multi-Tenant & Scoped ACLs: Filter graph traversals and vector candidate ranking on-the-fly using tenant_id and allowed_roles (conn.graph_rag_query_scoped()), ensuring sensitive contextual subgraphs never leak across tenants or privilege tiers.

Pillar 4: Cost-Based Query Optimizer (CBO) & Statistics

TapirusDB features an automated cost-based query optimizer (src/sql/planner.rs):

  • Computes disk I/O page fetch costs and CPU tuple comparison costs.
  • Automatically selects between Sequential Scan, B+Tree Secondary Index Scan, and Primary Key Point Lookup.
  • EXPLAIN QUERY PLAN outputs estimated execution cost and expected row cardinality.

Pillar 5: Bidirectional Graph-Vector Chaining & Agent Long-Term Memory

Traditional distributed stacks decouple graph databases and vector stores, causing high network serialization latency and memory-prohibitive global vector scans. TapirusDB executes native bidirectional in-memory chaining at 1,606,037 ops/sec (0.55 Β΅s):

  • Graph-to-Vector (Targeted Scored Neighborhoods): Traverses structured entity relationships first ($A \to B$), then restricts vector distance scoring strictly to candidate neighborhood nodes ($M \ll N$). Yields 100% exact Recall with zero approximation loss ($O(M \cdot D)$ instead of $O(N \log N)$).
  • Vector-to-Graph (Seed-and-Traverse GraphRAG): Uses ANN centroids to locate seed nodes, then instantly expands 1-hop and 2-hop CSR slices to extract factual context, eliminating LLM hallucinations.
  • Autonomous Agent Long-Term Memory (LTM): Automatically balances semantic vector similarity ($S_v$), BM25 lexical precision ($S_l$), and exponential temporal recency decay: $$\text{RecallScore}(m) = w_v \cdot S_v + w_l \cdot S_l + w_r \cdot e^{-\lambda \Delta t} + w_i \cdot \text{Importance}$$

Pillar 6: Unified Quad-Model Atomic Transactions (ACID)

Unlike fragmented multi-database architectures where cross-model consistency is impossible without complex distributed consensus (2PC/Sagas), TapirusDB provides true single-transaction atomicity across all four data models:

  • Single Atomic Commit: A transaction can update a Relational SQL state row, insert an unstructured JSON document, link openCypher knowledge graph nodes and edges, and index a vector embedding within a single conn.begin_transaction().
  • Zero Torn States: If any operation fails or the host process loses power, Write-Ahead Log (WAL) crash recovery rolls back all four models simultaneously to their exact pre-transaction state, eliminating cross-model state corruption forever.

🌐 Industrial Applications: AI & Beyond

TapirusDB's quad-model engine (Relational SQL + Vector Search + openCypher Graph + JSON Documents) inside a single encrypted .tapir container solves mission-critical industrial challenges without multi-database operational overhead:

Industrial Domain How Quad-Model Solves It Without Server Clusters
Financial Fraud Detection & AML Graph traverses money-mule rings and cyclic transactions ($A \to B \to C \to A$); Vector identifies anomalous spending behavior signatures; SQL enforces immutable balance reconciliation and strict ACID transactions.
Cybersecurity Threat Hunting & SIEM Graph traces Active Directory lateral movement attack vectors; Vector detects polymorphic binary and syscall sequence anomalies; SQL queries firewall events and access control lists in microsecond windows.
Supply Chain & Bill-of-Materials (BOM) Graph manages multi-tiered supplier dependency trees and failure propagation; Vector clusters sensor telemetry patterns; Document ingests unstructured customs and logistics manifests.
Healthcare, Genomics & Life Sciences Graph traverses Disease $\to$ Gene $\to$ Symptom $\to$ Drug pathways; Vector performs chemical fingerprint similarity (SMILES) for drug repurposing; SQL guarantees HIPAA/clinical record integrity.
Scientific Research & Academic Labs Single-file .tapir dataset container guarantees 100% reproducible paper workflows; Graph + Vector + SQL models molecular pathways and tabular metrics in Python/Jupyter with zero Docker dependencies.
In-Process Telemetry & Edge BI SIMD vectorized accumulators compute AVG/SUM/COUNT across millions of sensor readings in microseconds; Transparent LZ4 cuts disk usage by 70% with zero cloud egress cost.
Privacy-First Smart Home & Home Assistant Graph maps Zigbee/Matter/Thread device meshes; Vector performs local voice intent matching offline; WAL guarantees crash durability across home power outages on Raspberry Pi (<4MB RAM).
Air-Gapped Sovereign Hardware & Edge IoT Operates on Raspberry Pi, avionics, drones, and naval vessels with zero server daemons, < 4 MB idle RAM, and SIMD-accelerated ChaCha20-Poly1305 AEAD encryption at rest.

Developer Tooling & CLI

TapirusDB ships as a single zero-dependency standalone binary (tapirus):

1. Accelerated Workspace Search (tapirus grep / tapirus tg)

High-throughput in-process developer code search combining regex matching, BM25 token overlap, and local semantic vector similarity:

# Search codebase with semantic vector ranking enabled
tapirus tg --vector "transaction rollback wal" src/

# Case-insensitive search filtered by file extensions
tapirus grep -i --ext rs,toml "quantization" .

2. Interactive Terminal Shell

tapirus production.tapir
tapirus> CREATE TABLE users (id INTEGER PRIMARY KEY, name TEXT);
Query OK, 1 row(s) affected

tapirus> INSERT INTO users VALUES (1, 'Alex Chen');
Query OK, 1 row(s) affected

tapirus> SELECT * FROM users;
+----+------------+
| id | name       |
+----+------------+
| 1  | Alex Chen  |
+----+------------+
(1 row(s))

3. Built-in Protected HTTP REST Server (tapirus serve)

Launch an embedded database as a high-throughput, secure REST API with zero external dependencies:

# Launch with token authentication and database encryption
tapirus serve --port 3005 --api-key "your_secret_api_key" --passphrase "vault_secret" production.tapir
  • Constant-Time Verification: Prevents timing side-channel attacks via subtle::ConstantTimeEq.
  • Flexible Authentication: Provide the key via Authorization: Bearer <KEY>, X-API-Key: <KEY>, or query ?api_key=<KEY>.
  • Healthcheck Probe: GET /health or GET /api/health returns operational status without authentication.
  • Query Execution: POST /api/sql or POST /sql accepts SQL queries, graph traversals, and document queries.

4. Autonomous AI Agent MCP Server (tapirus mcp)

Connect Claude Desktop, Cursor, or Gemini to TapirusDB over stdio:

{
  "mcpServers": {
    "tapirus": {
      "command": "tapirus",
      "args": ["mcp", "agent_memory.tapir"]
    }
  }
}

5. High-Throughput Streaming Data Importer (tapirus import)

Stream massive datasets directly into TapirusDB with automatic schema inference and transactional batching:

# 1. Ingest CSV with automated type inference (INTEGER, REAL, TEXT) into SQL tables
tapirus import csv data/telemetry.csv --table sensors --batch 1000 --db production.tapir

# 2. Ingest streaming JSON Lines (JSONL) into Document collections
tapirus import jsonl data/products.jsonl --collection catalog --batch 500 --db production.tapir

# 3. Semantic Markdown Ingestion into AI Episodic Memory & Knowledge Graph
tapirus import md docs/spec.md --namespace robotics --tags "hardware,specs" --session-id "session_01" --db production.tapir

6. Hot Backup, Safe Restore & Physical Integrity Audit (tapirus backup, restore, verify)

Enterprise-grade durability, snapshotting, and auditing tools for mission-critical edge deployments:

# Atomic online backup snapshot (validates encryption KCV prior to copying)
tapirus backup production.tapir backups/prod_2026_snapshot.tapir

# Safe restore (validates page headers and file geometry)
tapirus restore backups/prod_2026_snapshot.tapir restored_production.tapir

# Deep physical integrity audit (scans slotted pages, verifies CRC32 checksums, checks encryption keys)
tapirus verify production.tapir

Verified Benchmarks

Benchmarks executed on native NVMe SSD hardware (cargo bench --bench tapirus_bench):

Operation Throughput Mean Latency Median (p50) Tail (p99)
Relational Primary Key Point Lookup 362,733 ops/sec 2.61 Β΅s 2.37 Β΅s 4.68 Β΅s
CSR Graph Adjacency Sweep 3,493,852 ops/sec 0.23 Β΅s 0.21 Β΅s 0.37 Β΅s
Graph-to-Vector Bidirectional Chaining 1,606,037 ops/sec 0.55 Β΅s 0.51 Β΅s 1.01 Β΅s
Relational B+Tree Inserts 149,176 ops/sec 6.19 Β΅s 4.66 Β΅s 61.13 Β΅s
JSON Document Path Lookups 355,004 docs/sec 2.70 Β΅s 2.58 Β΅s 5.45 Β΅s
HNSW Vector Search (32D, k=5) 51,060 QPS 19.53 Β΅s 17.06 Β΅s 50.36 Β΅s
RaBitQ Asymmetric POPCNT Distance > 12,000,000 ops/sec 0.08 Β΅s 0.08 Β΅s 0.12 Β΅s
WAL Durable Disk Writes 107,875 writes/sec 9.15 Β΅s 6.71 Β΅s 62.21 Β΅s
AI Memory Ingest (BM25 Indexing) 416,529 ops/sec 2.30 Β΅s 1.77 Β΅s 4.38 Β΅s

Architectural Latency Breakdown: Network/IPC Middleware vs. In-Process Memory Traversal

Cloud Vector DB (gRPC Roundtrip)  [β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ] 25,000 Β΅s (25.0 ms - WAN Network Hop)
Dedicated Graph DB (HTTP/JVM)     [β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ]                 15,000 Β΅s (15.0 ms - TCP / JVM GC)
Relational SQL Server (TCP IPC)   [β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ]                                  5,000 Β΅s (5.0 ms - Unix Socket / IPC)
TapirusDB Combined Graph-Vector   [β–Œ]                                          0.51 Β΅s (In-Process CPU Memory Bus)

⚑ Verified Tail-Latency & Physical Resource Footprint (Anti-Placebo)

Tested on native NVMe SSD hardware with true Write-Ahead Log (WAL) durability:

Dimension TapirusDB (In-Process) Traditional Network DBs Concrete Operational Value
Graph-Vector Retrieval 0.51 Β΅s (median) ~25,000 Β΅s (25 ms) Sub-microsecond local reasoning vs. WAN gRPC network serialization hop.
Durable Disk Writes 107,875 writes/sec ~2,000–8,000 ops/sec Real ACID WAL disk commits, not volatile in-memory caching.
Idle Memory Footprint < 4 MB RAM > 1.2 GB (multi-daemon) Fits comfortably in Raspberry Pi, edge robotics, and local desktop apps.
Initial File Footprint 4,096 Bytes Server cluster required Single encrypted .tapir container; zero cloud daemons to configure.

πŸ”¬ Transparent & Peer-Reviewed Methodology:
We publish complete hardware specifications, statistical variance ($\sigma$), and cache-miss analysis in our Systems Architecture Paper (PAPER_TAPIRUSDB.md).

Verify & run the benchmark suite yourself on your machine (1 command):

git clone https://github.com/tapiruslab/TapirusDB.git
cd TapirusDB
cargo bench --bench tapirus_bench

Full tail percentiles (p50, p95, p99, Min, Max) will be automatically exported to target/tapirus_bench_results.json for independent peer review.


When (and When NOT) to Use TapirusDB

Engineering honesty is paramount. Choosing the right storage engine requires understanding boundary trade-offs:

Workload & Scenario Recommended Engine Architectural Rationale
Local AI Agents & LLM RAG Memory βœ… TapirusDB Microsecond episodic retrieval, combined vector + openCypher graph in one atomic .tapir file.
Embedded Edge, Robotics & IoT Hardware βœ… TapirusDB < 4 MB idle RAM, 100% Safe Rust core, zero background daemon processes or JVM runtimes.
Desktop Apps, CLI Tools & Local-First Web βœ… TapirusDB Single-file portability, zero server configuration, pure client-side SQLite/Mongo alternative.
Petabyte Distributed Big Data Warehousing ❌ ClickHouse / Snowflake TapirusDB is optimized for operational single-node/in-process workloads, not massive multi-rack OLAP scans.
Multi-Region Active-Active Distributed Writes ❌ CockroachDB / Spanner For global multi-master write replication, use dedicated distributed consensus databases.
Complex Analytical BI Cubes over Billions of Rows ❌ DuckDB / ClickHouse DuckDB is superior for vectorized columnar OLAP; TapirusDB excels at transactional, graph, vector, and episodic AI memory.

Formal Safety Verification

  • TLA+ Specifications: Write-Ahead Logging (WAL) state transitions and crash recovery are formally modeled under TLA+ in docs/formal_verification/.
  • Memory Safety Contract: Strict #![forbid(unsafe_code)] enforced across all core modules in src/lib.rs.

Documentation & Architecture


TapirusDB β€” Engineered in Safe Rust for Sovereign AI & Edge Systems.
Developed & Maintained by TapirusDB Contributors β€’ Tapirus Tech Lab (tapirusdb.com)
Contact: contact@tapirusdb.com

About

The 100% Safe-Rust Embedded Quad-Model AI Database & Cognitive Memory Engine (SQL, Vectors, GraphRAG, Documents)

Topics

Resources

Contributing

Security policy

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages