a frame under a name a statement can match on - #14
Merged
Merged
Conversation
`conn.register("people", frame)` puts a DataFrame into the database
under a name, and `MATCH (p:people)` reads it. Anything that speaks
Arrow goes in, which is pandas, polars, pyarrow and a reader over one,
and a dictionary of lists is there for a caller with none of them
installed. `unregister` takes the rows back out and `conn.registered`
says what is registered here.
The frame arrives over the same C Data Interface a result leaves by, so
a column costs a memcpy and no Python object per cell: a million rows of
one integer column read in 5 ms, and a million rows of an integer, a
float and a string in 138 ms, which is the string column being the only
one that allocates. The write behind it is the engine's own, and those
million rows take 9.3 seconds in all.
It is a copy and not a scan, which is the one thing about it a caller
has to know, and the module says why: DuckDB's register is zero copy
because its executor can call back out to the Python object holding the
data, and this engine has no such callback yet. So a registered frame is
a snapshot, and the day the engine grows a scan it becomes the cheap
thing under the same name.
Two more places the engine shows through, both said rather than worked
around. A name that has been used keeps the columns it was used with,
because a table's columns are declared by its first row and no statement
alters them, and `unregister` empties the table rather than removing it,
because no statement drops one. The first row goes in as a statement
with its values written out, since a table nothing declares is declared
by the row written into it and a parameter is worked out rather than
written; every row after it goes through the appender.
Registering inside a transaction is refused for the reason an appender
is, since its batches are commits of their own. A null anywhere is
refused by column and row, a zoned timestamp with what to do about it, a
column of bytes because no statement reads one back, and a name a
statement could not carry before anything is written.
Twenty-eight tests over the four ways in, every column kind a row can
hold, a stream of several batches, the snapshot, the replacements and
every refusal, plus a budget on the read.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
conn.register("people", frame)puts a DataFrame into the database under a name, andMATCH (p:people)reads it. Anything that speaks Arrow goes in, which is pandas, polars, pyarrow and a reader over one, and a dictionary of lists is there for a caller with none of them installed.unregistertakes the rows back out andconn.registeredsays what is registered here.The frame arrives over the same C Data Interface a result leaves by, so a column costs a memcpy and no Python object per cell: a million rows of one integer column read in 5 ms, and a million rows of an integer, a float and a string in 138 ms, which is the string column being the only one that allocates. The write behind it is the engine's own, and those million rows take 9.3 seconds in all, which is the same 0.1 to 0.2 million rows a second
loadand the appender manage.It is a copy and not a scan, which is the one thing about it a caller has to know, and the module says why. DuckDB's
registeris zero copy because its executor can call back out to the Python object holding the data, and this engine has no such callback yet, so a registered frame is a snapshot: changing the DataFrame afterwards changes nothing until it is registered again. The call is the one it would be either way, and the day the engine grows a scan it becomes the cheap thing under the same name.Two more places the engine shows through, both said rather than worked around. A name that has been used keeps the columns it was used with, because a table's columns are declared by its first row and no statement alters them, and
unregisterempties the table rather than removing it, because no statement drops one. The first row goes in as a statement with its values written out, since a table nothing declares is declared by the row written into it and a parameter is worked out rather than written; every row after it goes through the appender, in the order the table declares its columns.Registering inside a transaction is refused for the reason an appender is, since its batches are commits of their own that no rollback reaches. A null anywhere is refused by column and row, a zoned timestamp with what to do about it, a column of bytes because no statement reads one back, and a name a statement could not carry before anything is written at all.
Twenty-eight tests over the four ways in, every column kind a row can hold, a stream of several batches, the snapshot, the replacements and every refusal, plus a budget on the read. README section and stub entries included. Local is green: pytest, ruff, cargo fmt, clippy and both other ABI features.