Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion .github/workflows/ci.yml
Original file line number Diff line number Diff line change
Expand Up @@ -59,7 +59,7 @@ jobs:
# griffe reads the stub and inspects the installed extension, so
# the check runs against the wheel rather than against the
# checkout it was built from.
- run: pip install pytest griffe
- run: pip install pytest griffe ipython
- run: pytest

# The shared corpus, which is the same 945 cases the engine runs
Expand Down
22 changes: 21 additions & 1 deletion README.md
Original file line number Diff line number Diff line change
Expand Up @@ -126,6 +126,26 @@ result.record_batches() # a reader, for a result larger than memory

`Result` implements `__arrow_c_stream__`, so anything that reads the protocol reads a result directly and none of the four methods above is needed: `pyarrow.table(result)` and `polars.DataFrame(result)` both work. Batches are 65,536 rows. A column holds one type, which the values decide, and integers beside floats are the one mixture that widens rather than being refused. Nodes, rels and paths go across as structs. The copy runs with the GIL released, and on this machine 300,000 rows across three columns take 44 ms as Arrow against 67 ms as Python objects, and a single integer column takes 13.8 ms against 44.5 ms.

## In a notebook

A result in a cell draws itself as a table, because Jupyter asks an object for `_repr_html_` before it falls back to `repr` and a line saying how many rows there are is a strictly worse answer than the rows. Nodes, rels and paths draw themselves too, a path as the walk it is: `(person #0) -[knows]-> (person #1)`.

```python
%load_ext zudb.magic
%gql social.zu1
```

```python
%%gql
MATCH (p:person) WHERE p.score > 40 RETURN p.name AS name, p.score AS score
```

The cell is one statement, it runs on the current connection, and the result is the value of the cell, so `_` is a `zudb.Result` and everything a result can do is still there. `%gql` is about which connection and `%%gql` is about the statement, and neither guesses the other's job. A notebook that already called `zudb.connect` needs no `%gql` at all: if exactly one connection is lying about the namespace `%%gql` uses it, and if there is more than one it names them and asks which. `%%gql --conn other --params args --out rows` says which connection, where the parameters are, and where to put the result instead of showing it, each of them naming a variable because a notebook has the values already.

The markup is a table, a stylesheet and no script, so it survives `nbconvert`, an exported HTML file and a notebook diff, and there is nothing to install for any of it. Colours come from the notebook through `currentColor` and opacity, because a light theme and a dark one are both in the room. Values are escaped, since a string column holding `<script>` is a string column and a client that pasted one into the page would run a caller's data as code in their notebook. The first hundred rows are drawn and the note underneath says how many there were, and a value longer than two hundred characters is cut with a mark where it was cut, because a million rows of markup is a notebook file that will not open again.

IPython is not a dependency. It is what a notebook already has, and nothing here imports it until `%load_ext` does.

## Stopping a statement

A statement that is running can be stopped two ways, and neither of them closes the connection: the session, its plans and its warm readers are all there afterwards, which is the whole difference between stopping a statement and starting again.
Expand Down Expand Up @@ -183,7 +203,7 @@ The stub is checked against the module it describes in CI: griffe reads the stub

## What works today

The list above is what this client is for. What it does so far is the core of it: `connect`, `execute` and `sql` with named parameters, results that iterate and fetch, values as Python objects both ways including dates, times, datetimes and durations, `Node`, `Rel` and `Path` as classes, `load` for building a graph with edges in it, an appender for growing one, transactions as a context manager that commits at the end of a block and rolls back when it raises, every condition as an exception class carrying its code, its position and its documentation link, results as Arrow columns and as pandas and polars frames, `register` for putting a frame under a name a statement can match on and reading it where it lies, stubs inside the wheel with a gate that keeps them true, the GIL released around every statement, every load and every copy out, `Ctrl-C` and `interrupt()` stopping a statement without touching the connection under it, and `zudb.aio` for the same calls awaited on an event loop. A DB-API 2.0 wrapper is next, and each one lands with the tests that say it works.
The list above is what this client is for. What it does so far is the core of it: `connect`, `execute` and `sql` with named parameters, results that iterate and fetch, values as Python objects both ways including dates, times, datetimes and durations, `Node`, `Rel` and `Path` as classes, `load` for building a graph with edges in it, an appender for growing one, transactions as a context manager that commits at the end of a block and rolls back when it raises, every condition as an exception class carrying its code, its position and its documentation link, results as Arrow columns and as pandas and polars frames, `register` for putting a frame under a name a statement can match on and reading it where it lies, stubs inside the wheel with a gate that keeps them true, the GIL released around every statement, every load and every copy out, `Ctrl-C` and `interrupt()` stopping a statement without touching the connection under it, `zudb.aio` for the same calls awaited on an event loop, and results, nodes, rels and paths that draw themselves in a notebook with `%gql` and `%%gql` to run statements in one. A DB-API 2.0 wrapper is next, and each one lands with the tests that say it works.

## Wheels

Expand Down
5 changes: 5 additions & 0 deletions pyproject.toml
Original file line number Diff line number Diff line change
Expand Up @@ -47,6 +47,11 @@ dev = [
# Reads the stub as text and the built extension by inspection, which
# is what checks one against the other.
"griffe>=2",
# For `%gql` and `%%gql`, which are tested against a real shell
# rather than against a stub of one. Not a dependency of the wheel:
# anyone with a notebook has IPython already, and nobody else needs
# it.
"ipython>=8",
"pyarrow>=14",
"pandas>=2.0",
"polars>=1.3",
Expand Down
12 changes: 12 additions & 0 deletions python/zudb/_zudb.pyi
Original file line number Diff line number Diff line change
Expand Up @@ -197,6 +197,9 @@ class Result:
def __arrow_c_stream__(self, requested_schema: object | None = None) -> Any:
"""The rows as an Arrow stream, for anything that speaks Arrow."""

def _repr_html_(self) -> str:
"""The rows as an HTML table, which is what a notebook shows."""

def __len__(self) -> int: ...
def __iter__(self) -> Iterator[tuple[Value, ...]]: ...
def __repr__(self) -> str: ...
Expand All @@ -213,6 +216,9 @@ class Node:
def offset(self) -> int:
"""The row this node sits at in that table, counting from zero."""

def _repr_html_(self) -> str:
"""The node as a notebook draws it, which is the pair that names it."""

def __hash__(self) -> int: ...
def __repr__(self) -> str: ...

Expand All @@ -237,6 +243,9 @@ class Rel:
"""Where the edge's properties sit, which is its place in
the order the table was loaded in."""

def _repr_html_(self) -> str:
"""The edge as a notebook draws it, with the rows it joins on either side."""

def __hash__(self) -> int: ...
def __repr__(self) -> str: ...

Expand All @@ -256,6 +265,9 @@ class Path:
def rels(self) -> list[Rel]:
"""The edges of the walk, in the order it crosses them."""

def _repr_html_(self) -> str:
"""The walk as a notebook draws it, nodes and arrows alternating."""

def __len__(self) -> int: ...
def __repr__(self) -> str: ...

Expand Down
187 changes: 187 additions & 0 deletions python/zudb/magic.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,187 @@
"""`%%gql` in a notebook.

%load_ext zudb.magic
%gql social.zu1

%%gql
MATCH (p:person) WHERE p.score > 40 RETURN p.name AS name, p.score AS score

The cell is one statement, it runs on the current connection, and what
comes back is a `zudb.Result`, which the notebook draws as a table
because a result knows how to draw itself. It is the value of the cell
as well, so `_` is the result and everything a result can do is still
there.

Two magics and no more. `%gql` is about which connection, `%%gql` is
about the statement, and neither of them tries to guess the other's
job: a line magic that took a path or a statement depending on what it
looked like would guess wrong on the day somebody named a file
`MATCH`.

`%gql` with a path opens a connection and makes it the current one.
`%gql` with nothing says which connection that is. `%gql --close` shuts
it. A notebook that made its own connection does not need any of that:
if exactly one `zudb.Connection` is lying about the namespace, `%%gql`
uses it, and if there is more than one it says so and asks which,
because picking one of them would be picking somebody's read-only
connection half the time.

The cell magic takes three options, all of them naming variables rather
than holding values, since a notebook has the values already:

%%gql --conn other --params args --out rows
MATCH (p:person) WHERE p.uid = $uid RETURN p.name AS name

`--conn` is the connection to run on, `--params` is a dict of
parameters, and `--out` is where to put the result instead of showing
it.
"""

from __future__ import annotations

import shlex
from pathlib import Path
from typing import Any

from IPython.core.error import UsageError
from IPython.core.magic import Magics, cell_magic, line_magic, magics_class

from ._zudb import Connection, connect

__all__ = ["ZuMagics", "load_ipython_extension"]


@magics_class
class ZuMagics(Magics):
"""The two magics, and the connection `%gql` opened if it opened one."""

def __init__(self, shell: Any) -> None:
super().__init__(shell)
self.connection: Connection | None = None

@line_magic("gql")
def gql(self, line: str) -> Connection | None:
"""Opens a connection and makes it the one `%%gql` runs on.

%gql social.zu1
%gql --read-only social.zu1
%gql
%gql --close

The connection comes back so that a cell can keep it, and it is
closed at the end of the session like any other.
"""
words = shlex.split(line)
read_only = "--read-only" in words
closing = "--close" in words
paths = [word for word in words if not word.startswith("-")]
unknown = [word for word in words if word.startswith("-")]
for word in unknown:
if word not in ("--read-only", "--close"):
raise UsageError(f"%gql does not take {word}")
if closing:
if self.connection is not None:
self.connection.close()
self.connection = None
return None
if not paths:
if self.connection is None:
raise UsageError("no connection is open here: %gql <path> opens one")
return self.connection
if len(paths) > 1:
raise UsageError("%gql opens one database, and this named more than one")
# The old one is closed rather than left holding a file, since
# a notebook that opens a second database has finished with the
# first and nothing else here has a name for it.
if self.connection is not None:
self.connection.close()
self.connection = connect(Path(paths[0]).expanduser(), read_only=read_only)
return self.connection

@cell_magic("gql")
def gql_cell(self, line: str, cell: str) -> Any:
"""Runs the cell as one statement on the current connection."""
options = self.options(line)
conn = self.current(options.get("conn"))
statement = cell.strip()
# A person types the semicolon out of habit and a statement
# does not take one, which is a syntax error about a character
# nobody meant to type.
while statement.endswith(";"):
statement = statement[:-1].rstrip()
if not statement:
raise UsageError("this cell holds no statement")
result = conn.execute(statement, self.params(options.get("params")))
out = options.get("out")
if out is None:
return result
self.namespace()[out] = result
return None

def options(self, line: str) -> dict[str, str]:
"""`--conn`, `--params` and `--out`, each naming a variable."""
words = shlex.split(line)
taken: dict[str, str] = {}
while words:
word = words.pop(0)
name = word.removeprefix("--")
if name == word or name not in ("conn", "params", "out"):
raise UsageError(f"%%gql takes --conn, --params and --out, and not {word}")
if not words:
raise UsageError(f"--{name} names a variable, and this named none")
taken[name] = words.pop(0)
return taken

def namespace(self) -> dict[str, Any]:
"""Where the notebook's own names live."""
if self.shell is None:
raise UsageError("there is no notebook here to read names out of")
return self.shell.user_ns

def current(self, named: str | None) -> Connection:
"""The connection to run on: the one named, the one `%gql`
opened, or the one the notebook made if it made exactly one.
"""
if named is not None:
found = self.namespace().get(named)
if not isinstance(found, Connection):
raise UsageError(f"'{named}' is not a zudb.Connection in this notebook")
return found
if self.connection is not None:
return self.connection
theirs = sorted(
name
for name, value in self.namespace().items()
if isinstance(value, Connection) and not name.startswith("_")
)
if len(theirs) == 1:
return self.namespace()[theirs[0]]
if not theirs:
raise UsageError(
"no connection is open here: %gql <path> opens one, or make one with "
"zudb.connect and %%gql will find it"
)
raise UsageError(
"this notebook holds more than one connection ("
+ ", ".join(theirs)
+ "), so %%gql --conn <name> has to say which"
)

def params(self, named: str | None) -> dict[str, Any] | None:
"""The parameters, out of the variable `--params` named."""
if named is None:
return None
found = self.namespace().get(named)
if not isinstance(found, dict):
raise UsageError(f"'{named}' is not a dict of parameters in this notebook")
return found


def load_ipython_extension(ipython: Any) -> None:
"""Registers `%gql` and `%%gql`, which is what `%load_ext` calls.

Registered by hand rather than on import, because a library that
reaches into the interpreter it was imported into is a library that
surprises somebody, and `%load_ext zudb.magic` is one line.
"""
ipython.register_magics(ZuMagics)
41 changes: 41 additions & 0 deletions src/conn.rs
Original file line number Diff line number Diff line change
Expand Up @@ -20,6 +20,7 @@ use zudb::{Config, Database, Interrupt};
use crate::appender::Appender;
use crate::columns;
use crate::error::{closed, programming, to_py_err};
use crate::html;
use crate::interrupt;
use crate::register;
use crate::txn::Transaction;
Expand Down Expand Up @@ -571,6 +572,46 @@ impl Result {
self.result.rows.len()
)
}

/// The rows as an HTML table, which is what a notebook shows.
///
/// Jupyter asks for this before it falls back to `repr`, so a
/// result in a cell is the rows and not a line about how many
/// there are. Only the first hundred are drawn and the note
/// underneath says how many there were, because a million rows of
/// markup is a notebook file nobody can open.
///
/// Reading it does not move the cursor `fetchone` uses, for the
/// same reason iterating does not: a person who looked at a result
/// has not taken any of its rows.
fn _repr_html_(&self, py: Python<'_>) -> PyResult<String> {
let rows = &self.result.rows;
let columns = &self.result.columns;
if columns.is_empty() {
return Ok(html::wrap(
"<span class=\"zu-note\">no columns, which is what a statement that \
writes gives back</span>",
));
}
let mut out = String::from("<table class=\"zu-table\"><thead><tr>");
for name in columns {
out.push_str(&format!("<th>{}</th>", html::escape(name)));
}
out.push_str("</tr></thead><tbody>");
for row in rows.iter().take(html::SHOWN) {
out.push_str("<tr>");
for value in row {
out.push_str(&html::cell(&to_py(py, value, &self.names)?)?);
}
out.push_str("</tr>");
}
out.push_str("</tbody></table>");
out.push_str(&format!(
"<div class=\"zu-note\">{}</div>",
html::note(rows.len(), columns.len())
));
Ok(html::wrap(&out))
}
}

impl Result {
Expand Down
Loading
Loading