Skip to content

gateway tests: make the load-only failures deterministic (port theft, GOAWAY mock, fixed sleeps) - #124

Merged
jaredLunde merged 2 commits into
mainfrom
load-flakes
Oct 4, 2026
Merged

jaredLunde merged 2 commits into
mainfrom
load-flakes

Conversation

@jaredLunde

Copy link
Copy Markdown
Contributor

Five integration-test failures that appeared only under heavy host load (during the full mutation run, where they also produced false "caught" verdicts). All five are harness/mock assumptions, not gateway defects.

  • Port theft (most of them). free_port() bound :0, released it, and handed it to a subprocess; under load another process's bind(:0) took it in the gap. Signatures: "did not come up on port", ai_allowance_ready never 1 (the shared nats-server lost its port and the bare TCP readiness probe connected to someone else), a "dead" upstream that answered. Now: ports drawn at random from the 16k just below ip_local_port_range (never auto-assigned), NATS readiness checks the server_name in its INFO greeting and respawns on a fresh port, a gateway that logs "is in use" is respawned (up to 5×), closed_port() holds a bound non-listening socket for the test's life. Repro: a neighbour churning 6,000 bind(:0) listeners failed 16/81 tests before; 190/190 pass after with 17,000.
  • upstream_goaway_is_handled. The mock closed its socket the instant graceful shutdown ended, so the kernel sent RST and the in-flight GOAWAY was lost (the gateway correctly doesn't resend a POST it can't prove unprocessed). Real servers linger (Go waits 1 s; nginx lingering close); the mock now drains for up to 1 s. 0/384 after (was ~1%). Explanation added to D248's note.
  • Fixed sleeps. the_tenant_cap_is_checked_before_a_large_body_is_read and sigterm_drains_an_in_flight_request_and_bills_it used fixed head starts; they now hold the upstream reply and wait on an observable event.
  • Startup wait raised to 60 s (measured startup: 0.1 s idle, ~1.3 s at 200 busy-loops); failures now print the child's log tail.

Full beyond-ai suite 1222/1222 unloaded; stress loop (160 busy-loops, -j 48, load ~175, 3 runs over large_bodies, model_routing, billing_*, claims_reliability, reliability_breaker) 3/3 green, previously failing every run.

🤖 Generated with Claude Code

https://claude.ai/code/session_01JimHGjsfk2Ktm5GxyZJKKk

jaredLunde and others added 2 commits October 4, 2026 10:02
…c under host load

Five families of gateway tests failed only under heavy host load (a long
local mutation run, load 60-230 on 16 cores) and passed alone. None is a
gateway defect; each was a harness or mock assumption.

- Ports stolen inside the free_port window. free_port drew from the
  kernel's ephemeral range via bind(:0), and Linux readily re-hands a
  just-released port to the next bind(:0) on the host, so concurrent runs'
  mocks and servers took gateway, NATS and "dead upstream" ports before the
  subprocess bound them. Pingora then retried the bind for 30 s ("did not
  come up"), the shared nats-server exited while a bare TCP connect
  succeeded against the thief (ai_allowance_ready never 1), and a "dead"
  candidate answered. With a neighbour churning 6,000 bind(:0) listeners,
  16 of 81 tests failed on unchanged code; after, 190 of 190 passed with
  17,000. free_port now draws at random below ip_local_port_range, where
  the kernel never auto-assigns. A gateway whose port is taken anyway is
  respawned on fresh ports (Pingora's "is in use" line); nats-server
  readiness matches its server_name in the INFO greeting and respawns a
  server that exited; closed_port/dead_authority hold the port bound and
  never listening for the test's life.
- Startup bound. 20 s became STARTUP_BUDGET (60 s): the gateway does no
  network or disk work before it binds, and a batch of starts stalled past
  20 s while freshly linked binaries were exec'd under IO pressure. A child
  that exits or a start that times out now fails with its own log, and a
  metric wait that times out prints the log tail.
- the_tenant_cap_is_checked_before_a_large_body_is_read and
  sigterm_drains_an_in_flight_request_and_bills_it slept a fixed 300/400 ms
  and failed every run at load ~175. Both now hold the upstream reply
  (Reply::Held) and synchronize on the upstream hit; the drain test
  releases it only after the gateway logs "SIGTERM received".
- upstream_goaway_is_handled: the mock closed its socket the instant
  hyper's graceful shutdown ended, with client bytes unread, so the kernel
  sent RST and destroyed the GOAWAY in flight; the gateway saw a broken
  pipe and rightly did not resend the POST (502). Real servers linger after
  GOAWAY; the mock now does too. 3/368 contended runs failed before, 0/384
  after. D248's open note records the explanation.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JimHGjsfk2Ktm5GxyZJKKk
@jaredLunde
jaredLunde enabled auto-merge (squash) October 4, 2026 17:24
@jaredLunde
jaredLunde merged commit 149b324 into main Oct 4, 2026
16 of 19 checks passed
@jaredLunde
jaredLunde deleted the load-flakes branch October 4, 2026 17:24
jaredLunde added a commit that referenced this pull request Oct 4, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant