Repository navigation
fix(translation): report Anthropic context-window stops as a token limit - #918
bharadwaj-pendyala wants to merge 1 commit into
Conversation
Signed-off-by: Bharadwaj Pendyala <bharadwajpendyala@gmail.com>
|
No actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info⚙️ Run configuration
📒 Files selected for processing (4)
Included review availability: This review used your included allowance. Your plan provides up to 12 included reviews per hour; 11 remain after this review. WalkthroughChangesThe translation maps Anthropic’s Anthropic stop mapping
Priority: ➖ Normal Estimated code review effort: 2 (Simple) | ~8 minutes Merge Risk: ⚪ Minimal · up to The supplied review identifies no issue that needs resolution before merging. 🚥 Pre-merge checks | ✅ 4 | ❌ 1❌ Failed checks (1 warning)
✅ Passed checks (4 passed)
A rabbit checks the stop reason at dawn Comment |
What
Treat Anthropic's
model_context_window_exceededstop reason as a token limit, likemax_tokens. This updatesmap_anthropic_stop_reasonfor buffered responses and the Anthropic stream decoder'smessage_deltahandling.Why
Sonnet 4.5 and newer return
model_context_window_exceededwithout a beta header when the answer runs into the context window. Anthropic's docs say to treat it as truncated. Switchyard doesn't know the spelling, so a cut-off answer reaches other clients like this onb9e7ccce:finish_reason: "stop"status: "completed"finish_reason: "model_context_window_exceeded"response.completedIn the streamed Chat row, the raw Anthropic string isn't one of the values the OpenAI spec allows for
finish_reason. The other three rows report a truncated answer as complete.max_tokenson the same inputs giveslengthandincompleteeverywhere, as set up in #300.The stream decoder change also covers streams buffered by the server.
ResponseAccumulatorpasses the chunk reason tostop_reason_from_strinprotocol/src/stream.rs, which returnsUnknownfor this spelling today. Emittingmax_tokensinstead producesMaxTokenswithout changingprotocol. The Responses decoder already mapsresponse.incompletetomax_tokens.Notes for reviewers
max_tokenswhere it used to sayend_turn.pause_turn, which also passes through raw to Chat stream clients. It means "send the turn back to continue", and there's no Chat or Responses equivalent, so it needs its own decision.stop_reason_from_strfor Bedrock. This PR doesn't touch that file, so the two don't conflict.Validation
Two regression tests, one buffered and one streamed, fail on
b9e7ccceand pass on this branch. The streamed test also folds the decoded chunk throughResponseAccumulatorand checks forStopReason::MaxTokens.cargo fmt --all --check: cleancargo clippy --workspace --all-targets --exclude switchyard-py -- -D warnings: cleancargo test --workspace --exclude switchyard-py: 913 passed, 0 failedI didn't build
switchyard-pylocally. Nothing in it changes here, and CI runs it.Summary by CodeRabbit