Skip to content

fix(streaming): stop writing a response cache nothing can read (#150) - #234

Merged
ptimizeroracle merged 1 commit into
mainfrom
fix/150-no-orphan-cache
Aug 7, 2026
Merged

ptimizeroracle merged 1 commit into
mainfrom
fix/150-no-orphan-cache

Conversation

@ptimizeroracle

@ptimizeroracle ptimizeroracle commented Aug 7, 2026 •

Copy link
Copy Markdown
Owner

Closes #150.

The question the issue asked

Does streaming actually support resume from checkpoint, or only execute() / execute_async()?

It does not, and it never could. None of execute_stream(), execute_stream_async() or execute_stream_pipelined() accept resume_from, so there is no way to ask for one.

Measured, not reasoned

The durable cache is keyed (session_id, row_index). Streaming builds a fresh Pipeline per chunk, and each new ExecutionContext takes a new uuid4 session — so every row a chunk writes is unreadable by the next run, which generates new ids again.

A 6-row stream over 3 chunks, run twice against a counting client:

RUN1  llm_calls=6  chunks=3  db_rows=6   sessions=3
RUN2  llm_calls=6  chunks=3  db_rows=12  sessions=6

Second run re-called every row. responses.db grew to 12 rows across 6 dead sessions, none of which anything can read.

That is outcome (2) in the issue: the shared cache is accidental, not load-bearing.

Why it mattered beyond waste

Up to max_pending_chunks chunk pipelines wrote that single SQLite file concurrently. That is the database is locked contention #147 patched with busy_timeout. This removes the writes rather than the symptom — the mitigation stays, but it now has nothing to mitigate on this path.

The change

Sub-pipelines — streaming chunks and the auto-retry pass — no longer attach a durable cache.

After:

STREAM run1: calls=6  db: no responses.db written
STREAM run2: calls=6  db: no responses.db written
NON-STREAM:  calls=6  db: 6 rows / 1 session

The non-streaming path is untouched, and resume still skips completed rows — verified by killing a run after 3 of 6 rows and resuming: 3 new calls, not 6.

Documented, since the absence is the surprising part

docs/guides/checkpointing.md now states plainly that streaming does not resume, why it is structural rather than an oversight, and the workaround: shard the data and run one resumable pipeline per shard. A missing parameter is not documentation.

Tests

Both halves are pinned, deliberately:

  • streaming leaves no responses.db
  • non-streaming still writes exactly one session
  • resume still reuses cached rows

Asserting only the first would be satisfied by disabling the cache everywhere — which would silently make every resume re-pay for work already done. The first test fails when the sub-pipeline skip is removed.

Full suite: 1164 passing, mypy clean.

Not fixed here

Streaming still cannot resume. Making it resumable is a feature — chunks would need to share one session, or the top level would need to own a cache keyed across them — not something to smuggle into a cleanup. This makes the current behaviour honest instead of half-implemented.

ai assistance: i directed this work with help from claude code.

Summary by CodeRabbit

  • Bug Fixes

    • Streaming executions no longer create response-cache records.
    • Streaming completion handles zero-duration runs correctly.
    • Standard executions continue to support response caching and checkpoint resume.
    • Automatic retries avoid creating unnecessary durable cache entries.
  • Documentation

    • Clarified checkpoint resume support, streaming limitations, restart behavior, and potential duplicate processing costs.
    • Added guidance for handling large datasets with resumable execution options.

The durable cache is keyed (session_id, row_index) and is the source of truth
for resume. Streaming builds a fresh Pipeline per chunk, and each new
ExecutionContext takes a new uuid4 session — so every row a chunk wrote was
unreadable by the next run, which generated new ids again.

Measured on a 6-row stream over 3 chunks, run twice: 6 LLM calls each time,
and responses.db grew to 12 rows across 6 dead sessions. Nothing was ever read
back.

Those writes were not merely wasted. Up to max_pending_chunks chunks wrote
that one SQLite file concurrently, which is the contention #147 papered over
with busy_timeout. This removes the writes rather than the symptom.

Sub-pipelines — streaming chunks and the auto-retry pass — no longer attach a
durable cache. The non-streaming path is untouched: it still writes one
session per run, and resume still skips rows already answered.

This answers the question #150 asked. Streaming has no resume: none of
execute_stream(), execute_stream_async() or execute_stream_pipelined() accept
resume_from, so there is no way to ask for one, and the shared cache was
accidental rather than load-bearing. The checkpointing guide now says so
plainly, including the workaround — shard the data and run one resumable
pipeline per shard — rather than leaving readers to infer it from a missing
parameter.

Tests pin both halves: streaming writes nothing, and resume still reuses
cached rows. Only asserting the first would be satisfied by disabling the
cache everywhere, which would make every resume re-pay for completed work.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@coderabbitai

coderabbitai Bot commented Aug 7, 2026 •

Copy link
Copy Markdown
Contributor

Review Change Stack

📝 Walkthrough

Walkthrough

The pipeline now skips durable response caches for streaming chunks and retry sub-pipelines. Standard executions retain SQLite caching and resume support. Tests verify these behaviors, and the checkpointing guide documents streaming limitations.

Changes

Streaming response-cache ownership

Layer / File(s) Summary
Sub-pipeline cache ownership
ondine/api/pipeline.py
Pipeline tracks sub-pipelines. Streaming chunks and retry pipelines skip durable cache creation. Top-level runs retain SQLite response caching.
Cache and resume validation
tests/unit/test_streaming_no_orphan_cache.py
Tests verify that streaming creates no cache rows, standard execution creates one cache session, and resumed execution reuses cached rows.
Checkpointing behavior documentation
docs/guides/checkpointing.md
The guide documents that streaming methods do not support resume_from and restart failed runs from the beginning.

Estimated code review effort: 3 (Moderate) | ~20 minutes

Sequence Diagram(s)

sequenceDiagram
  participant Caller
  participant TopLevelPipeline
  participant ChunkPipeline
  participant SqliteResponseCache
  Caller->>TopLevelPipeline: execute_stream_pipelined
  TopLevelPipeline->>ChunkPipeline: create streaming chunk sub-pipeline
  ChunkPipeline->>SqliteResponseCache: skip durable cache creation
  Caller->>TopLevelPipeline: execute or execute_async with resume_from
  TopLevelPipeline->>SqliteResponseCache: read and reuse cached rows
Loading

Possibly related PRs

🚥 Pre-merge checks | ✅ 5
✅ Passed checks (5 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title clearly identifies the streaming response-cache fix and matches the primary change.
Linked Issues check ✅ Passed The PR confirms streaming resume is unsupported, documents the behavior, removes unusable sub-pipeline caching, and adds tests for issue [#150].
Out of Scope Changes check ✅ Passed The documentation, pipeline changes, and tests directly support the linked issue and stated PR objectives.
Docstring Coverage ✅ Passed No functions found in the changed files to evaluate docstring coverage. Skipping docstring coverage check.
✨ Finishing Touches
📝 Generate docstrings
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch fix/150-no-orphan-cache

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 2

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@docs/guides/checkpointing.md`:
- Around line 60-62: Update the checkpointing documentation’s duplicate-cost
statement to clarify that only rows processed before the failure are charged
again on restart; replace “every row is paid for again” with wording such as
“previously processed rows are paid for again” or “the restarted run reprocesses
all rows.”
- Around line 72-74: Update the streaming checkpointing guidance in the section
around line 284 to remove the claim that checkpoints support chunk-level resume.
State that streaming does not support resuming, or link readers to the existing
large-dataset guidance near “execute()” and checkpointing, while preserving the
rest of the documentation.
🪄 Autofix

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Pro Plus

Run ID: 79d3cf00-bdc8-4d0c-80db-d6b08a0eaf65

📥 Commits

Reviewing files that changed from the base of the PR and between 6ebf2a7 and bf730e4.

📒 Files selected for processing (3)
  • docs/guides/checkpointing.md
  • ondine/api/pipeline.py
  • tests/unit/test_streaming_no_orphan_cache.py

Comment on lines +60 to +62
None of them accept `resume_from`, so there is no way to ask for one. A
streamed run that dies must be restarted from the beginning, and every row is
paid for again.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🎯 Functional Correctness | 🟡 Minor | ⚡ Quick win

Qualify the duplicate-cost statement.

After a failure at row N, only previously processed rows were charged in the failed run. Restarting repeats those rows. Later rows incur their first charge. Replace “every row is paid for again” with “previously processed rows are paid for again” or “the restarted run reprocesses all rows.”

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@docs/guides/checkpointing.md` around lines 60 - 62, Update the checkpointing
documentation’s duplicate-cost statement to clarify that only rows processed
before the failure are charged again on restart; replace “every row is paid for
again” with wording such as “previously processed rows are paid for again” or
“the restarted run reprocesses all rows.”

Comment on lines +72 to +74
**If you need resume on a large dataset**, prefer `execute()` with
checkpointing over streaming, or split the data yourself and run one pipeline
per shard so each shard has a session you can resume.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🎯 Functional Correctness | 🟡 Minor | ⚡ Quick win

Remove the conflicting chunk-level checkpointing guidance.

Line 284 still says that streaming checkpointing belongs at the chunk level. This conflicts with the new statement that streaming has no resume support. Update that section to describe the no-resume limitation or link to this section.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@docs/guides/checkpointing.md` around lines 72 - 74, Update the streaming
checkpointing guidance in the section around line 284 to remove the claim that
checkpoints support chunk-level resume. State that streaming does not support
resuming, or link readers to the existing large-dataset guidance near
“execute()” and checkpointing, while preserving the rest of the documentation.

@ptimizeroracle
ptimizeroracle merged commit a7cc6d7 into main Aug 7, 2026
40 checks passed
@ptimizeroracle
ptimizeroracle deleted the fix/150-no-orphan-cache branch August 7, 2026 09:46
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Verify streaming resume semantics + document or fix

1 participant