Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
47 changes: 46 additions & 1 deletion README.md
Original file line number Diff line number Diff line change
Expand Up @@ -66,6 +66,51 @@ What it is worth, summing one integer column of a hundred thousand rows on an M-

A row at a time is a boundary crossing a cell, and a hundred crossings cost about what one borrowed buffer costs. Both surfaces are there because both are the right answer to a different question, but a loop over a million rows should be reading a column.

## Getting rows in

Two ways, and which one you want follows from whether the database exists yet.

A loader builds one out of whole columns. It is the fastest way values get in and, while the engine has no DDL, it is the only way a table comes into being at all:

```java
try (Loader loader = Loader.create(Path.of("social.zu1"))) {
loader.table("Person", "Follows", 3);
loader.column("id", 1L, 2L, 3L);
loader.column("name", "ada", "grace", "alan");
loader.edges(new int[] {0, 1}, new int[] {1, 2});
loader.finish();
}
```

Columns go in as arrays or as `java.nio` buffers, and which you pass is the difference between a copy and no copy. A direct buffer is read where it lies, so the engine sees the memory your program already filled and nothing crosses the boundary but a pointer. An array is memory nothing outside the JVM can address, so it is copied off-heap first. `Linker.Option.critical(true)` would let a Java array through without either, at the price of blocking the collector for the length of the copy, and on a column this size that is not a trade worth making.

An appender adds rows to a table that already exists, a value at a time, with no statement anywhere near it:

```java
try (Appender rows = conn.appender("Person")) {
rows.append(4L).append("hedy").endRow();
rows.append(5L).append("katherine").endRow();
rows.finish();
}
```

Values are written in the order the table declares its columns, which `columnName(int)` will tell you, and a row is a row once `endRow()` has ended it. A value the column will not take ends its row there and rolls back the values already written into it, so a refused append never leaves half a row behind. Closing an appender that was never finished writes what it has anyway, because a loop that threw halfway should keep the rows it managed; `discard()` is there for when it should not.

What each is worth on an M-series laptop, JDK 25:

| How | Per row |
|---|---|
| `loader.column(name, direct LongBuffer)` | 0.44 ns |
| `loader.column(name, long[])` | 1.1 ns |
| `loader.column(name, List<String>)` | 108 ns |
| a whole two-column load, write included | 630 ns |
| `appender.append(...).endRow()` | 87 ns |
| the same row as an `INSERT` statement | 4.0 ms |

The first two lines are the copy: 0.68 ns a row over a hundred thousand rows is 68 microseconds to move 800 KB, which is about what a memcpy costs and about what a direct buffer saves. It is a small share of a load that also writes a file, and it is the whole difference at the boundary itself.

The last line is the one to read twice. A statement per row parses, plans, runs and commits per row, and none of that work says anything the row before it did not already say. That is what an appender is for.

## How it binds

The Foreign Function and Memory API is the primary path. The downcall handles are written by hand against `zu.h` rather than generated with `jextract`, because the C ABI here is around seventy functions with a stable shape, and a hand-written layer is where the interesting decisions live: which calls are `Linker.Option.critical` because they are short pure accessors, where the out-parameter scratch space comes from so that a query does not allocate, and how a `zu_error` becomes a typed Java exception exactly once. There is no native code in this repository beyond `libzu` itself.
Expand Down Expand Up @@ -99,7 +144,7 @@ catch (ZuSyntaxException e) {

## What works today

The engine has no DDL yet, so there is no `CREATE NODE TABLE` and nothing in this client writes a schema. What runs against a fresh database is the expression and projection surface: `RETURN`, `UNWIND`, parameters, lists, records, and the temporal types. The example at the top of this file describes the intended shape and needs a graph that some other tool built.
The engine has no DDL yet, so there is no `CREATE NODE TABLE` and no statement in this client writes a schema. A table comes into being through `Loader`, which is why the loader example above builds the graph the example at the top of this file reads. What runs against a fresh database with nothing in it is the expression and projection surface: `RETURN`, `UNWIND`, parameters, lists, records, and the temporal types.

## Building

Expand Down
109 changes: 109 additions & 0 deletions zudb-bench/src/main/java/dev/zudb/bench/AppendBench.java
Original file line number Diff line number Diff line change
@@ -0,0 +1,109 @@
package dev.zudb.bench;

import dev.zudb.Appender;
import dev.zudb.Connection;
import dev.zudb.Database;
import dev.zudb.Loader;
import java.io.IOException;
import java.nio.file.Files;
import java.nio.file.Path;
import java.util.concurrent.TimeUnit;
import org.openjdk.jmh.annotations.Benchmark;
import org.openjdk.jmh.annotations.BenchmarkMode;
import org.openjdk.jmh.annotations.Fork;
import org.openjdk.jmh.annotations.Level;
import org.openjdk.jmh.annotations.Measurement;
import org.openjdk.jmh.annotations.Mode;
import org.openjdk.jmh.annotations.OutputTimeUnit;
import org.openjdk.jmh.annotations.Scope;
import org.openjdk.jmh.annotations.Setup;
import org.openjdk.jmh.annotations.State;
import org.openjdk.jmh.annotations.TearDown;
import org.openjdk.jmh.annotations.Warmup;

/**
* What adding a row to a table that already exists costs.
*
* <p>One benchmark invocation is one row, so the score is the row, which is
* the unit a caller writes. Every iteration starts from a copy of a database
* with one row in it, so a long run measures appending rather than a table
* growing under it.
*/
@State(Scope.Benchmark)
@BenchmarkMode(Mode.AverageTime)
@OutputTimeUnit(TimeUnit.NANOSECONDS)
@Warmup(iterations = 3, time = 2)
@Measurement(iterations = 5, time = 2)
@Fork(value = 1, jvmArgs = {"--enable-native-access=ALL-UNNAMED"})
public class AppendBench {

private Path dir;
private Path template;
private long counter;

private Database db;
private Connection conn;
private Appender appender;

@Setup
public void build() throws IOException {
dir = Files.createTempDirectory("zu-append-bench");
// The table an appender appends to has to exist, and a bulk load is the
// only thing that makes one.
template = dir.resolve("template.zu");
try (Loader loader = Loader.create(template)) {
loader.table("Person", "Knows", 1);
loader.column("id", -1L);
loader.column("name", "seed");
loader.finish();
}
}

@TearDown
public void clean() throws IOException {
Temp.deleteTree(dir);
}

@Setup(Level.Iteration)
public void open() throws IOException {
Path path = dir.resolve("append-" + counter++ + ".zu");
Files.copy(template, path);
db = Database.open(path);
conn = db.connect();
appender = conn.appender("Person");
}

@TearDown(Level.Iteration)
public void close() {
if (appender != null) {
appender.close();
appender = null;
}
if (conn != null) {
conn.close();
conn = null;
}
if (db != null) {
db.close();
db = null;
}
}

/** One row of two columns, written the way a loop that knows its schema writes it. */
@Benchmark
public void row() {
appender.append(counter++).append("n").endRow();
}

/** The same row through the dynamic path, which costs a type test a value. */
@Benchmark
public void rowOfObjects() {
appender.row(counter++, "n");
}

/** What the appender is worth, against the statement it replaces. */
@Benchmark
public void rowThroughAStatement() {
conn.execute("INSERT (:Person {id: " + counter++ + ", name: 'n'})");
}
}
135 changes: 135 additions & 0 deletions zudb-bench/src/main/java/dev/zudb/bench/LoadBench.java
Original file line number Diff line number Diff line change
@@ -0,0 +1,135 @@
package dev.zudb.bench;

import dev.zudb.Loader;
import java.io.IOException;
import java.nio.ByteBuffer;
import java.nio.ByteOrder;
import java.nio.LongBuffer;
import java.nio.file.Files;
import java.nio.file.Path;
import java.util.ArrayList;
import java.util.Comparator;
import java.util.List;
import java.util.concurrent.TimeUnit;
import org.openjdk.jmh.annotations.Benchmark;
import org.openjdk.jmh.annotations.BenchmarkMode;
import org.openjdk.jmh.annotations.Fork;
import org.openjdk.jmh.annotations.Level;
import org.openjdk.jmh.annotations.Measurement;
import org.openjdk.jmh.annotations.Mode;
import org.openjdk.jmh.annotations.OperationsPerInvocation;
import org.openjdk.jmh.annotations.OutputTimeUnit;
import org.openjdk.jmh.annotations.Scope;
import org.openjdk.jmh.annotations.Setup;
import org.openjdk.jmh.annotations.State;
import org.openjdk.jmh.annotations.TearDown;
import org.openjdk.jmh.annotations.Warmup;

/**
* What building a database out of columns costs, per row of a hundred
* thousand.
*
* <p>The loader itself is made in an invocation fixture rather than in the
* measured region, because a whole load writes a file and the file would
* drown out everything else. What is measured is handing a column over, which
* is where the difference between an array and a direct buffer lives, and
* {@code wholeLoad} is there so the rest can be read against what a load
* really costs.
*/
@State(Scope.Benchmark)
@BenchmarkMode(Mode.AverageTime)
@OutputTimeUnit(TimeUnit.NANOSECONDS)
@Warmup(iterations = 3, time = 2)
@Measurement(iterations = 5, time = 2)
@Fork(value = 1, jvmArgs = {"--enable-native-access=ALL-UNNAMED"})
public class LoadBench {

/**
* How many rows one load carries. A constant rather than a parameter
* because the per-row score is scaled by it, and JMH wants that scale as a
* literal in an annotation.
*/
private static final int ROWS = 100_000;

private Path dir;
private long counter;

private long[] ids;
private LongBuffer direct;
private List<String> names;

private Loader loader;
private Path path;

@Setup
public void fill() throws IOException {
dir = Files.createTempDirectory("zu-load-bench");
ids = new long[ROWS];
for (int i = 0; i < ROWS; i++) {
ids[i] = i;
}
direct =
ByteBuffer.allocateDirect(ROWS * Long.BYTES).order(ByteOrder.nativeOrder()).asLongBuffer();
direct.put(ids);
direct.flip();
names = new ArrayList<>(ROWS);
for (int i = 0; i < ROWS; i++) {
names.add("n" + i);
}
}

@TearDown
public void clean() throws IOException {
Temp.deleteTree(dir);
}

/** A loader with its table named and no column in it yet. */
@Setup(Level.Invocation)
public void open() {
path = dir.resolve("load-" + counter++ + ".zu");
loader = Loader.create(path);
loader.table("Person", "Knows", ROWS);
}

@TearDown(Level.Invocation)
public void close() throws IOException {
if (loader != null) {
loader.close();
loader = null;
}
if (path != null) {
Files.deleteIfExists(path);
path = null;
}
}

/** One integer column as a Java array, which has to be copied off-heap. */
@Benchmark
@OperationsPerInvocation(ROWS)
public void columnFromArray() {
loader.column("id", ids);
}

/** The same column as a direct buffer, which is read where it lies. */
@Benchmark
@OperationsPerInvocation(ROWS)
public void columnFromDirectBuffer() {
loader.column("id", direct.duplicate());
}

/** A string column, which has no zero-copy shape and is checked for UTF-8 besides. */
@Benchmark
@OperationsPerInvocation(ROWS)
public void columnOfStrings() {
loader.column("name", names);
}

/** Two columns and the write, which is what a load costs a caller. */
@Benchmark
@OperationsPerInvocation(ROWS)
public void wholeLoad() {
loader.column("id", ids);
loader.column("name", names);
loader.finish();
}
}
30 changes: 30 additions & 0 deletions zudb-bench/src/main/java/dev/zudb/bench/Temp.java
Original file line number Diff line number Diff line change
@@ -0,0 +1,30 @@
package dev.zudb.bench;

import java.io.IOException;
import java.nio.file.Files;
import java.nio.file.Path;
import java.util.Comparator;

/** The files a benchmark that writes to disk leaves behind. */
final class Temp {

private Temp() {}

/** Removes a directory and everything under it, sidecars included. */
static void deleteTree(Path root) throws IOException {
if (root == null) {
return;
}
try (var walk = Files.walk(root)) {
walk.sorted(Comparator.reverseOrder()).forEach(Temp::delete);
}
}

private static void delete(Path path) {
try {
Files.deleteIfExists(path);
} catch (IOException e) {
throw new IllegalStateException(e);
}
}
}
Loading
Loading