From c4e9e1f26cbf5575a545e2e60ee8ce83ef2a2554 Mon Sep 17 00:00:00 2001 From: Tam Nguyen Duc <1218621+tamnd@users.noreply.github.com> Date: Wed, 19 Aug 2026 03:14:04 +0700 Subject: [PATCH] a result a notebook can read Jupyter asks an object for _repr_html_ before it falls back to repr, so a result printed in a cell is a line saying how many rows it has. The rows are already in memory and a table of them is a strictly better answer, so this draws one, and draws nodes, edges and walks the same way whether they turn up on their own or inside a cell. The markup is a table and a stylesheet and no script, which means it survives nbconvert, an exported HTML file and a notebook diff, and there is nothing to install for it to work. Colours come from the notebook through currentColor and opacity, because a light theme and a dark one are both in the room and a guess about which would be wrong half the time. A hundred rows is the cut, since a million rows of markup is a notebook file nobody can open again, and the note underneath says how many rows there really were. Two hundred characters is the cut on a cell, so a column holding a document does not become a page holding one. Values are escaped, because a string column holding a tag is a string column and pasting one into the page unescaped would run a caller's data as code in their notebook. zudb.magic goes with it: %gql opens a connection and %%gql runs a statement on it, with --conn to say which of several, --params to name a dict, and --out to put the result in a variable instead of showing it. A notebook that already called zudb.connect needs no %gql at all, since a cell with one connection in the namespace finds it, and a cell with two says which names it saw. The magic is a way in and not a layer: the statement goes straight to execute and the engine's exception arrives as the engine's exception. IPython is a dev dependency and not a wheel dependency, and the tests skip without it the way the pyarrow and pandas ones do. They run against a real InteractiveShell with %load_ext rather than a mock, so what is checked is what a notebook does. --- .github/workflows/ci.yml | 2 +- README.md | 22 +++- pyproject.toml | 5 + python/zudb/_zudb.pyi | 12 +++ python/zudb/magic.py | 187 ++++++++++++++++++++++++++++++++++ src/conn.rs | 41 ++++++++ src/html.rs | 215 +++++++++++++++++++++++++++++++++++++++ src/lib.rs | 1 + src/value.rs | 21 ++++ tests/test_html.py | 155 ++++++++++++++++++++++++++++ tests/test_magic.py | 198 +++++++++++++++++++++++++++++++++++ 11 files changed, 857 insertions(+), 2 deletions(-) create mode 100644 python/zudb/magic.py create mode 100644 src/html.rs create mode 100644 tests/test_html.py create mode 100644 tests/test_magic.py diff --git a/.github/workflows/ci.yml b/.github/workflows/ci.yml index 973357a..3e96859 100644 --- a/.github/workflows/ci.yml +++ b/.github/workflows/ci.yml @@ -59,7 +59,7 @@ jobs: # griffe reads the stub and inspects the installed extension, so # the check runs against the wheel rather than against the # checkout it was built from. - - run: pip install pytest griffe + - run: pip install pytest griffe ipython - run: pytest # The shared corpus, which is the same 945 cases the engine runs diff --git a/README.md b/README.md index eab13c2..8a38d54 100644 --- a/README.md +++ b/README.md @@ -126,6 +126,26 @@ result.record_batches() # a reader, for a result larger than memory `Result` implements `__arrow_c_stream__`, so anything that reads the protocol reads a result directly and none of the four methods above is needed: `pyarrow.table(result)` and `polars.DataFrame(result)` both work. Batches are 65,536 rows. A column holds one type, which the values decide, and integers beside floats are the one mixture that widens rather than being refused. Nodes, rels and paths go across as structs. The copy runs with the GIL released, and on this machine 300,000 rows across three columns take 44 ms as Arrow against 67 ms as Python objects, and a single integer column takes 13.8 ms against 44.5 ms. +## In a notebook + +A result in a cell draws itself as a table, because Jupyter asks an object for `_repr_html_` before it falls back to `repr` and a line saying how many rows there are is a strictly worse answer than the rows. Nodes, rels and paths draw themselves too, a path as the walk it is: `(person #0) -[knows]-> (person #1)`. + +```python +%load_ext zudb.magic +%gql social.zu1 +``` + +```python +%%gql +MATCH (p:person) WHERE p.score > 40 RETURN p.name AS name, p.score AS score +``` + +The cell is one statement, it runs on the current connection, and the result is the value of the cell, so `_` is a `zudb.Result` and everything a result can do is still there. `%gql` is about which connection and `%%gql` is about the statement, and neither guesses the other's job. A notebook that already called `zudb.connect` needs no `%gql` at all: if exactly one connection is lying about the namespace `%%gql` uses it, and if there is more than one it names them and asks which. `%%gql --conn other --params args --out rows` says which connection, where the parameters are, and where to put the result instead of showing it, each of them naming a variable because a notebook has the values already. + +The markup is a table, a stylesheet and no script, so it survives `nbconvert`, an exported HTML file and a notebook diff, and there is nothing to install for any of it. Colours come from the notebook through `currentColor` and opacity, because a light theme and a dark one are both in the room. Values are escaped, since a string column holding `'})") + html = empty.execute("MATCH (p:person) RETURN p.name AS name")._repr_html_() + assert "