Skip to content

feat(benchmark): net out retried attempts by trial id in task-level routing stats - #120

Merged
ayushag-nv merged 1 commit into
mainfrom
rlempka/routing-log-retry-netting
Jul 23, 2026
Merged

ayushag-nv merged 1 commit into
mainfrom
rlempka/routing-log-retry-netting

Conversation

@ryan-lempka

@ryan-lempka ryan-lempka commented Jul 23, 2026 •

Copy link
Copy Markdown
Collaborator

Extends the task-level routing stats (#116) so retried attempts can be netted out of cost, plus a token schema that matches the global stats file.

What changed

  1. Retry netting by trial. The proxy sidecar stamps x-switchyard-trial-id (Harbor's trial name, read from the container env) alongside the existing task and session headers. At finalize, requests are grouped by trial, and within a trial the latest attempt is the one Harbor kept (it retries sequentially and discards earlier attempts). So tokens from retried-away attempts land in a retries section and can be subtracted. No timestamps, no result.json parsing.
  2. Six-field token schema. Records and the rollup now carry cached_tokens, cache_creation_tokens, and reasoning_tokens alongside prompt/completion/total, matching routing_stats_final.json so the existing cost estimator prices either file. model is the served model id, so buckets reconcile with the global file.

Schema, before → after (per task in routing_stats_by_task.json):

// before
"build-system-task-ordering": {
  "requests": 24,
  "sessions": ["8c54..."],
  "models": [
    {"model": "gpt-5.4-mini", "tier": "strong", "calls": 12,
     "prompt_tokens": 349072, "completion_tokens": 22239, "total_tokens": 371311}
  ]
}

// after
"build-system-task-ordering": {
  "requests": 29,
  "n_retries": 1,
  "final": {
    "calls": 24,
    "totals": {"prompt_tokens": 349072, "cached_tokens": 300112, "cache_creation_tokens": 0,
               "completion_tokens": 22239, "reasoning_tokens": 18110, "total_tokens": 371311},
    "models": [
      {"model": "gpt-5.4-mini", "tier": "strong", "calls": 12,
       "prompt_tokens": 349072, "cached_tokens": 300112, "cache_creation_tokens": 0,
       "completion_tokens": 22239, "reasoning_tokens": 18110, "total_tokens": 371311}
    ]
  },
  "retries": { "calls": 5, "totals": { /* same six fields */ }, "models": [ /* ... */ ] }
}

final and retries share the same shape, so the cost estimator runs on either: net cost = price(final), retry overhead = price(retries).

Tested end to end: Ran a real Harbor benchmark, forced a task to time out and retry, and confirmed the stamped trial id matches Harbor's trial name and that the retried attempt's tokens land in retries while the kept attempt's stay in final.

@ryan-lempka
ryan-lempka requested a review from a team as a code owner July 23, 2026 03:57
@coderabbitai

coderabbitai Bot commented Jul 23, 2026 •

Copy link
Copy Markdown
Contributor

Review Change Stack

Walkthrough

Changes

Routing accounting

Layer / File(s) Summary
Detailed routing records
switchyard/lib/processors/routing_log_response_processor.py, tests/test_routing_log_response_processor.py
Routing records now prefer the backend-served model and capture prompt, cache, completion, reasoning, and total token fields across supported provider usage shapes.
Final and retry manifest rollup
benchmark/run_manifest.py, tests/test_run_manifest.py, benchmark/README.md
Manifest finalization classifies requests using Harbor trial windows, aggregates final/retry statistics, updates tests, and documents the resulting artifacts.

Estimated code review effort: 4 (Complex) | ~45 minutes

Poem

A rabbit logs tokens, both fluffy and bright,
With models served true and retries in sight.
Final hops stay final, old hops gently retire,
Cache crumbs and thoughts join the routing wire.
Hop, hop—clean stats climb higher!

🚥 Pre-merge checks | ✅ 5
✅ Passed checks (5 passed)
Check name Status Explanation
Docstring Coverage ✅ Passed No functions found in the changed files to evaluate docstring coverage. Skipping docstring coverage check.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title accurately describes the main change: task-level routing stats now separate and net out retries by trial window.

Comment @coderabbitai help to get the list of available commands.

@ryan-lempka
ryan-lempka marked this pull request as draft July 23, 2026 04:21
…outing stats

Signed-off-by: Ryan Lempka <rlempka@nvidia.com>
@ryan-lempka ryan-lempka self-assigned this Jul 23, 2026
@ryan-lempka
ryan-lempka force-pushed the rlempka/routing-log-retry-netting branch from 9d216e5 to 78c6ab8 Compare July 23, 2026 04:40
@ryan-lempka ryan-lempka changed the title feat(benchmark): net out retried attempts in task-level routing stats feat(benchmark): net out retried attempts by trial id in task-level routing stats Jul 23, 2026
@ryan-lempka
ryan-lempka marked this pull request as ready for review July 23, 2026 04:41

@ayushag-nv ayushag-nv left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Looks good. Aligned with what I was trying to do with logging per request routing decisions

@ayushag-nv
ayushag-nv merged commit ee5e96d into main Jul 23, 2026
18 checks passed
@ayushag-nv
ayushag-nv deleted the rlempka/routing-log-retry-netting branch July 23, 2026 04:47
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants