Skip to content

a frame under a name a statement can match on - #14

Merged
tamnd merged 1 commit into
mainfrom
register-frames
Aug 18, 2026
Merged

tamnd merged 1 commit into
mainfrom
register-frames

Conversation

@tamnd

@tamnd tamnd commented Aug 18, 2026

Copy link
Copy Markdown
Owner

conn.register("people", frame) puts a DataFrame into the database under a name, and MATCH (p:people) reads it. Anything that speaks Arrow goes in, which is pandas, polars, pyarrow and a reader over one, and a dictionary of lists is there for a caller with none of them installed. unregister takes the rows back out and conn.registered says what is registered here.

The frame arrives over the same C Data Interface a result leaves by, so a column costs a memcpy and no Python object per cell: a million rows of one integer column read in 5 ms, and a million rows of an integer, a float and a string in 138 ms, which is the string column being the only one that allocates. The write behind it is the engine's own, and those million rows take 9.3 seconds in all, which is the same 0.1 to 0.2 million rows a second load and the appender manage.

It is a copy and not a scan, which is the one thing about it a caller has to know, and the module says why. DuckDB's register is zero copy because its executor can call back out to the Python object holding the data, and this engine has no such callback yet, so a registered frame is a snapshot: changing the DataFrame afterwards changes nothing until it is registered again. The call is the one it would be either way, and the day the engine grows a scan it becomes the cheap thing under the same name.

Two more places the engine shows through, both said rather than worked around. A name that has been used keeps the columns it was used with, because a table's columns are declared by its first row and no statement alters them, and unregister empties the table rather than removing it, because no statement drops one. The first row goes in as a statement with its values written out, since a table nothing declares is declared by the row written into it and a parameter is worked out rather than written; every row after it goes through the appender, in the order the table declares its columns.

Registering inside a transaction is refused for the reason an appender is, since its batches are commits of their own that no rollback reaches. A null anywhere is refused by column and row, a zoned timestamp with what to do about it, a column of bytes because no statement reads one back, and a name a statement could not carry before anything is written at all.

Twenty-eight tests over the four ways in, every column kind a row can hold, a stream of several batches, the snapshot, the replacements and every refusal, plus a budget on the read. README section and stub entries included. Local is green: pytest, ruff, cargo fmt, clippy and both other ABI features.

`conn.register("people", frame)` puts a DataFrame into the database
under a name, and `MATCH (p:people)` reads it. Anything that speaks
Arrow goes in, which is pandas, polars, pyarrow and a reader over one,
and a dictionary of lists is there for a caller with none of them
installed. `unregister` takes the rows back out and `conn.registered`
says what is registered here.

The frame arrives over the same C Data Interface a result leaves by, so
a column costs a memcpy and no Python object per cell: a million rows of
one integer column read in 5 ms, and a million rows of an integer, a
float and a string in 138 ms, which is the string column being the only
one that allocates. The write behind it is the engine's own, and those
million rows take 9.3 seconds in all.

It is a copy and not a scan, which is the one thing about it a caller
has to know, and the module says why: DuckDB's register is zero copy
because its executor can call back out to the Python object holding the
data, and this engine has no such callback yet. So a registered frame is
a snapshot, and the day the engine grows a scan it becomes the cheap
thing under the same name.

Two more places the engine shows through, both said rather than worked
around. A name that has been used keeps the columns it was used with,
because a table's columns are declared by its first row and no statement
alters them, and `unregister` empties the table rather than removing it,
because no statement drops one. The first row goes in as a statement
with its values written out, since a table nothing declares is declared
by the row written into it and a parameter is worked out rather than
written; every row after it goes through the appender.

Registering inside a transaction is refused for the reason an appender
is, since its batches are commits of their own. A null anywhere is
refused by column and row, a zoned timestamp with what to do about it, a
column of bytes because no statement reads one back, and a name a
statement could not carry before anything is written.

Twenty-eight tests over the four ways in, every column kind a row can
hold, a stream of several batches, the snapshot, the replacements and
every refusal, plus a budget on the read.
@tamnd
tamnd merged commit d7cdfc0 into main Aug 18, 2026
6 of 8 checks passed
@tamnd
tamnd deleted the register-frames branch August 18, 2026 15:30
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant