Skip to content

fix(result): blank lost rows to None instead of three different markers (#262) - #263

Merged
ptimizeroracle merged 2 commits into
mainfrom
feat/unify-loss-markers
Aug 13, 2026
Merged

ptimizeroracle merged 2 commits into
mainfrom
feat/unify-loss-markers

Conversation

@ptimizeroracle

@ptimizeroracle ptimizeroracle commented Aug 12, 2026 •

Copy link
Copy Markdown
Owner

Decision

Resolves the marker half of #254, tracked as #262: one representation (None) on all three loss paths, materialized in a single pass keyed off errors — not by matching cell strings.

A lost row used to carry [SKIPPED] (skip policy), "null" (failed batch), or "" (empty response) depending on which path lost it. None are NaN, so df.isna() reported holes as complete and a caller had to know all three magic strings.

Mechanism (why keyed off errors)

One pass blanks the output cells of every row recorded in errors. Keying off the error list — not cell strings — buys two properties a blanket string-replace can't:

  • No collision. A model may legitimately answer "null" or ""; only rows Ondine recorded as lost are touched.
  • Recovery-aware. After auto-retry a formerly-failed row may hold a real value; a cell is blanked only if it still carries a marker.

Authority on which rows were lost stays errors / is_complete (from #260); the None in the cell is its visible shadow. New: result.lost_row_indices for one-step survivor/loss slicing.

Bonus correctness fix

A run where every row failed used to return a frame of "null" strings that the quality check counted as valid → success=True. Those cells are now None → zero valid output → the whole-run guard fires (same guard that already caught ""/[SKIPPED]). Two tests unknowingly relying on that masking are corrected.

⚠️ Behavioral change — your versioning call

Lost cells are now None/NaN, not "[SKIPPED]"/"null"/"". Anyone detecting loss by those strings switches to is_complete / errors / df.isna(). Behaviorally breaking — I did not add a BREAKING CHANGE: footer (that would force release-please to a major bump); say the word if you want a major, otherwise it ships as a fix with a prominent note.

The ambiguous empty-but-successful response (not recorded as an error) is deliberately left as-is — data, not a tagged loss.

Verification

  • 1501 unit + e2e tests pass; both marker paths (skip + batch-null) verified to blank to None, df.isna() truthful, survivors intact
  • New unit tests pin collision-safety and recovery-awareness
  • mypy + ruff clean; docs-name checker 292 snippets

ai assistance: i directed this work with help from claude code.

Summary by CodeRabbit

  • Bug Fixes

    • Skipped or failed rows now appear as None instead of the [SKIPPED] placeholder.
    • Recovered retry results, legitimate null responses, and empty responses are preserved correctly.
    • Added clearer error handling when no valid outputs are produced.
  • New Features

    • Execution results now expose the sorted indices of rows that were lost during processing.
  • Documentation

    • Updated error-handling guidance to explain missing values and how to identify failed rows.

…rs (#262)

A lost row used to carry one of three strings in its output cell — "[SKIPPED]"
from the skip policy, "null" from a failed batch response, "" from an empty
one — depending on which path lost it. None of them are NaN, so df.isna()
reported a frame full of holes as complete, and a caller had to know all three
magic strings to find the losses by value.

Converge them on None, in one place. A single pass over the result blanks the
output cells of every row recorded in `errors`, keyed off the error list rather
than by matching cell strings. That keying is deliberate:

* No collision. A model may legitimately answer "null" or ""; only rows Ondine
  actually recorded as lost are touched, so a real answer that looks like a
  marker is never blanked.
* Recovery-aware. After auto-retry a formerly-failed row may hold a real value
  again; a cell is blanked only if it still carries a loss marker.

The authority on which rows were lost stays `errors` / `is_complete`; the None
in the cell is the visible shadow of that. Add `result.lost_row_indices` so
slicing survivors from losses is one step.

Consequence worth noting: a run where *every* row failed used to hand back a
frame of "null" strings that the quality check counted as valid, so it returned
success. Those cells are now None, correctly counted as zero valid output, so
such a run trips the whole-run guard and fails loudly.

Behavioral change: lost cells are now None/NaN, not "[SKIPPED]"/"null"/"".
Callers detecting loss by those strings should use is_complete / errors /
df.isna() instead. The ambiguous empty-but-successful response (a row that was
not recorded as an error) is left untouched — it is data, not a tagged loss.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@coderabbitai

coderabbitai Bot commented Aug 12, 2026 •

Copy link
Copy Markdown
Contributor

Review Change Stack

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Pro Plus

Run ID: e8b6195a-75ac-4aaf-a310-5bd5fa7835ff

📥 Commits

Reviewing files that changed from the base of the PR and between 93cce0b and fbdc15a.

📒 Files selected for processing (1)
  • tests/integration/test_auto_retry_with_batching.py

📝 Walkthrough

Walkthrough

Skipped output markers are normalized to None. ExecutionResult exposes lost row indices. Tests and documentation verify missing-value detection, row alignment, retry recovery, and zero-output failures.

Changes

Lost-row output handling

Layer / File(s) Summary
Lost-row output contract
ondine/api/pipeline.py, ondine/core/models.py, docs/guides/error-handling.md
Defines recognized lost-output markers, exposes lost_row_indices, and documents None outputs for skipped rows.
Pipeline output normalization
ondine/api/pipeline.py, tests/unit/test_silent_failure.py
Normalizes lost-row cells after retries while preserving recovered values and unaffected marker-like values.
Conformance and streaming validation
tests/e2e/*, tests/integration/*, tests/unit/test_pipelined_streaming.py
Updates output, alignment, error, retry, and streaming tests for the new behavior.

Estimated code review effort: 3 (Moderate) | ~20 minutes

Mergeability Score: 🟡 Moderate · up to fbdc1

This PR standardizes lost cells to None and adds lost-row indexing, but recovered retry results can still be marked as lost, causing valid rows to be discarded and completeness to be reported incorrectly; merge should wait for that correctness issue to be fixed or explicitly accepted.

Possibly related issues

  • ptimizeroracle/ondine#262: Covers replacing lost-row markers with None and detecting them through missing-value and error metadata.

Possibly related PRs

🚥 Pre-merge checks | ✅ 4 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 44.44% which is insufficient. The required threshold is 80.00%. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (4 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title clearly and concisely describes the main behavioral change: representing lost-row outputs as None instead of multiple markers.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
✨ Finishing Touches 💡 1
📝 Generate docstrings 💡
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch feat/unify-loss-markers

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 2

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@ondine/api/pipeline.py`:
- Around line 1731-1743: The retry flow around _auto_retry_failed_rows() must
reconcile recovered rows with result.errors and row metrics before returning.
Remove recovered indices from lost_row_indices, update is_complete and
skipped/failed metrics consistently, and add a regression assertion verifying
recovered rows are absent from lost_row_indices.

In `@tests/unit/test_pipelined_streaming.py`:
- Line 133: Strengthen the assertions in the test around
pipeline.execute_stream_pipelined so each yielded chunk is verified to contain
exactly chunk_size rows and all expected input rows, rather than only checking
the number of result objects. Preserve the existing chunk-count assertion while
validating completeness and row counts for every result.
🪄 Autofix

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Pro Plus

Run ID: d480c8ee-e714-4697-92d3-dfffd2c10280

📥 Commits

Reviewing files that changed from the base of the PR and between b7116db and 93cce0b.

📒 Files selected for processing (8)
  • docs/guides/error-handling.md
  • ondine/api/pipeline.py
  • ondine/core/models.py
  • tests/e2e/test_pipeline_conformance.py
  • tests/e2e/test_router_conformance.py
  • tests/integration/test_end_to_end.py
  • tests/unit/test_pipelined_streaming.py
  • tests/unit/test_silent_failure.py

Comment thread ondine/api/pipeline.py
Comment on lines +1731 to +1743
This converges them on ``None``, in one place, keyed off ``errors``
rather than by matching cell strings. Two reasons for that:

* **No collision.** A model may legitimately answer ``"null"`` or
``""``; only rows Ondine actually recorded as lost are touched, so a
real answer that happens to look like a marker is never blanked.
* **Recovery-aware.** After auto-retry a formerly-failed row may hold a
real value again; a cell is only blanked if it still carries a loss
marker, so recovered rows keep their answer.

The authority on *which* rows were lost stays ``errors`` /
``is_complete``; the ``None`` in the cell is the visible shadow of that,
not a second source of truth.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🗄️ Data Integrity & Integration | 🟠 Major | 🏗️ Heavy lift

Reconcile recovered rows with loss metadata.

_auto_retry_failed_rows() replaces recovered cells but leaves the original result.errors and skipped or failed metrics unchanged. This method then preserves the recovered value. A recovered row can therefore contain a valid answer while lost_row_indices still includes it and is_complete remains False.

Remove or reclassify errors for recovered rows. Update the row metrics before returning the retried result. Add a regression assertion that a recovered row is absent from lost_row_indices.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@ondine/api/pipeline.py` around lines 1731 - 1743, The retry flow around
_auto_retry_failed_rows() must reconcile recovered rows with result.errors and
row metrics before returning. Remove recovered indices from lost_row_indices,
update is_complete and skipped/failed metrics consistently, and add a regression
assertion verifying recovered rows are absent from lost_row_indices.


with patch("litellm.acompletion", side_effect=client.ainvoke):
results = list(pipeline.execute_stream_pipelined(chunk_size=chunk_size))
results = list(pipeline.execute_stream_pipelined(chunk_size=chunk_size))

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🎯 Functional Correctness | 🟡 Minor | ⚡ Quick win

Assert that each chunk contains all input rows.

Line 133 only proves that two chunk result objects were yielded. A regression that drops rows inside a chunk still passes.

Assert that each result has chunk_size rows and that it is complete.

Proposed test assertions
         for i, chunk_result in enumerate(results):
             assert chunk_result.success, f"Chunk {i} failed"
             assert hasattr(chunk_result, "data"), f"Chunk {i} missing data"
+            assert len(chunk_result.to_list()) == chunk_size
+            assert chunk_result.is_complete
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@tests/unit/test_pipelined_streaming.py` at line 133, Strengthen the
assertions in the test around pipeline.execute_stream_pipelined so each yielded
chunk is verified to contain exactly chunk_size rows and all expected input
rows, rather than only checking the number of result objects. Preserve the
existing chunk-count assertion while validating completeness and row counts for
every result.

The single-row retry branch returned raw text ("Processed_5"), which the retry
sub-pipeline's JSON batch parser cannot read, so row 5 never actually
recovered — the test only passed because the old failure marker was a non-null
string that satisfied notna(). With lost cells now None (#262) the masking is
gone; return the array shape the retry parses so the row genuinely recovers.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@ptimizeroracle
ptimizeroracle merged commit 1f2ddc9 into main Aug 13, 2026
41 checks passed
@ptimizeroracle
ptimizeroracle deleted the feat/unify-loss-markers branch August 13, 2026 08:09
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant