Skip to content

gateway: a client that stalls mid-upload times out its own request, no failover (D260) - #123

Merged
jaredLunde merged 3 commits into
mainfrom
stall-upload-no-failover
Oct 4, 2026
Merged

jaredLunde merged 3 commits into
mainfrom
stall-upload-no-failover

Conversation

@jaredLunde

Copy link
Copy Markdown
Contributor

A client that stalled partway through a Content-Length upload on a catalog walk was failed over: when the upstream read timeout fired, error_while_proxy saw Upstream ReadTimedout with the body not done and took the walk's failover branch, so the next candidate got a connection and the abandoned candidate's breaker recorded a failure. Reproduced over HTTP/1.1 (when read_timeout_secs < pingora's 60 s body timeout) and h2c (at the default 600 s; pingora's h2 server has no body-read timeout).

Fix (error path only, stateless): client_stalled_upload — a ReadTimedout while client_still_uploading holds (some body fed, body not done) means nothing moved in either direction with the provider waiting on the client (body chunks restart the read timeout; a blocked body write is WriteTimedout). It is retagged downstream, so no retry/failover, the breaker permit is released, no estimate is billed (client_cancelled), and the client gets 408 "request body timed out".

Already correct before: no key cooling (transport errors never cool), no billed estimate (body_delivered requires the body done), provider routes kept the breaker clean.

Tests (cancellation.rs): a_client_stalled_mid_upload_is_not_failed_over (HTTP/1.1 + h2c: 408, no estimate, breaker threshold 1 still reaches the primary, zero failovers/key failures) and a_client_stalled_mid_chunked_upload_times_out_as_its_own. All hand mutations of the new code caught. D260 in verify/defects.toml; ARCHITECTURE.md updated.

Separate, not fixed here: an h2c client stalling during the gateway's up-front body read (peek_body_model) has no timeout at all (pingora's h2 server sets none), so it can hold a tenant slot indefinitely.

🤖 Generated with Claude Code

https://claude.ai/code/session_01JimHGjsfk2Ktm5GxyZJKKk

jaredLunde and others added 3 commits October 4, 2026 09:35
…260)

A managed client that stopped sending its body with the connection open
was failed over to the next catalog candidate when the upstream read
timeout fired, and the walk charged the candidate it left a breaker
failure from upstream_peer. The provider was only waiting on the client.

error_while_proxy now retags a ReadTimedout while the client still owes
body bytes (client_stalled_upload: some fed, is_body_done false) as a
downstream error: no failover, the breaker permit is released, nothing is
billed (body_delivered is false), and fail_to_proxy answers 408 (request
body timed out), as pingora's own HTTP/1.1 body read timeout does.
Pingora re-arms the upstream read timeout on each body chunk and an
upstream that stops reading is a WriteTimedout, so the timeout can only
mean the client went silent. h2c clients reach this at the default
600s (pingora's HTTP/2 server has no body read timeout); HTTP/1.1 ones
when read_timeout_secs is under pingora's 60s. Error path only.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JimHGjsfk2Ktm5GxyZJKKk
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JimHGjsfk2Ktm5GxyZJKKk
@jaredLunde
jaredLunde enabled auto-merge (squash) October 4, 2026 16:36
@jaredLunde
jaredLunde merged commit f470b2b into main Oct 4, 2026
20 checks passed
@jaredLunde
jaredLunde deleted the stall-upload-no-failover branch October 4, 2026 16:47
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant