Skip to content
shubhamshettyyPublic

About

Building a relational database engine from scratch in C++, implementing core storage and indexing internals that underpin production systems like PostgreSQL. Components include a buffer pool manager, B+ tree index, and query execution engine.

Resources

Stars

2 stars

Watchers

0 watching

Forks

Latest commit

 

History

30 Commits

Folders and files

Repository files navigation

KodaDB

KodaDB is a relational database engine I am building from scratch in C++17. The goal is not to ship a production-ready server, but to understand how storage, indexing, query execution, and a small SQL shell fit together in one coherent system and to have something I can demo, test, and extend over time.

The engine is single-process and single-user today. Data persists across restarts through heap files, index files, and a text catalog on disk. I can drive it interactively from a shell or run a scripted demo end to end.

Demo

Real captures from the interactive shell and the CRUD demo script.

Shell CRUD demo script
KodaDB interactive shell Demo script output
Query result table After restart
SELECT output Persistence check

To reproduce the demo locally:

cmake -S . -B build
cmake --build build
.\scripts\run_demo.ps1

On macOS or Linux, use scripts/run_demo.sh instead. The script creates a fresh demo_db/ directory, runs scripts/demo_crud.sql, then opens the database again to show that rows survived shutdown.

What works today

  • DDL: CREATE TABLE with INT, BOOL, and TEXT / TEXT(n); CREATE INDEX on one column per index
  • DML: INSERT, SELECT (with optional WHERE), UPDATE, and DELETE — all with a required single equality predicate on update/delete
  • Indexes: B+ tree equality lookups; the planner uses an index scan when exactly one WHERE clause matches an indexed column
  • Joins: block nested-loop join via SELECT … FROM left JOIN right ON … (no WHERE on join queries yet)
  • Catalog: table and index metadata saved in catalog.meta under the database directory
  • Shell: interactive REPL, styled terminal output, meta commands (.tables, .schema, .indexes), and batch mode via -f script.sql

Architecture (high level)

SQL / meta commands (shell)
        ↓
   parseCommand → Statement
        ↓
   Planner::execute → Operator tree
        ↓
   BufferManager ↔ DiskManager
        ↓
   heap files (*.heap) + index files (*_*.idx) + catalog.meta

Each layer is documented in more depth under docs/.

Documentation

Topic File
Heap pages, records, tombstones docs/storage.md
Buffer pool and disk I/O docs/buffer-and-disk.md
B+ tree indexes docs/indexing.md
Catalog and Database facade docs/catalog.md
Volcano operators and CRUD docs/execution.md
Planner, parser, shell docs/planner-and-shell.md
Tests and how to run them docs/testing.md
Design decision notes docs/decisions/

Build and test

Requirements: CMake 3.16+, a C++17 compiler.

cmake -S . -B build
cmake --build build
ctest --test-dir build --output-on-failure

Run the shell (default database directory kodadb_data/):

.\build\kodadb_app.exe
.\build\kodadb_app.exe my_db -f scripts\demo_crud.sql

Set NO_COLOR=1 to disable ANSI styling (useful for logs and CI).

Repository layout

include/kodadb/   public headers
src/              implementation (storage, buffer, index, execution, planner, shell, database)
tests/            unit and integration tests
scripts/          demo SQL and runner scripts
docs/             subsystem documentation
dataset/          sample CSV for loader tests

License

MIT — see LICENSE.

About

Building a relational database engine from scratch in C++, implementing core storage and indexing internals that underpin production systems like PostgreSQL. Components include a buffer pool manager, B+ tree index, and query execution engine.

Resources

Stars

2 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages