Skip to content

perf(relay): reconcile early SSE header flush with pre-output failover #434

Description

@LIghtJUNction

Upstream source

Upstream #7509 adds CommitEventStreamHeaders: after an upstream response has been accepted as a stream, it writes and flushes the downstream SSE headers before scanning the first data frame. The goal is to avoid making clients wait for the first model token before receiving HTTP 200 / text/event-stream. It also treats header-flush failures as writer errors and updates stream handlers that bypass the shared scanner.

Current LMM behavior

Current main (997a1b673dcf718d75cbd957c437828a86c1a319) has the same basic transport gap:

  • apps/api-go/relay/helper/common.go has SetEventStreamHeaders, but it only sets headers and does not flush them.
  • apps/api-go/relay/helper/stream_scanner.go calls copyCodexSSEHeaders then SetEventStreamHeaders before reading the stream, so clients may not observe the HTTP/SSE response until the first body write.

However, LMM deliberately diverged from upstream with merged PR #312 (480cfc329c881a7670148b5aa335fa60f0d0babf): OpenAI-compatible streams can use OPENAI_FIRST_OUTPUT_TIMEOUT to fail over a channel that returns upstream headers but never produces visible output.

Current OaiStreamHandler:

  1. starts FirstResponseTimeout after upstream headers;
  2. only marks the first response after visible content/reasoning/tool/function output is successfully handled;
  3. on timeout before visible output, marks ContextKeyUpstreamChannelFailure and returns 504 upstream_timeout, allowing the existing channel retry/failover path;
  4. avoids retry after visible output.

StreamScannerHandler also intentionally prevents locally generated heartbeat writes before the first business response.

Why not copy #7509 directly

Flushing the downstream writer normally commits the HTTP response. If LMM commits 200 text/event-stream immediately after upstream headers, the later first-visible-output timeout can no longer reliably return a real 504 before output, and the existing pre-output failover boundary may be weakened or become dependent on already-committed downstream state.

That is a product/transport-policy conflict, not a mechanical backport. The upstream latency improvement is useful, but LMM should not trade away its stalled-stream failover silently.

Suggested LMM adaptation

Treat downstream header commitment as part of the retry boundary rather than applying CommitEventStreamHeaders unconditionally:

  1. Add a small explicit helper for committing/flushing SSE headers, but call it only when the current attempt is no longer eligible for pre-output failover.
  2. For OpenAI-compatible streams with FirstResponseTimeout > 0, keep the response uncommitted while the attempt is still in the pre-visible-output retry window. Commit immediately before the first downstream-visible event, then permanently disable retry for that attempt.
  3. When the first-output timeout is disabled, allow early header flush after the upstream response has been validated as a stream, matching the latency benefit of upstream #7509.
  4. Audit dedicated streaming handlers that do not use StreamScannerHandler; only adopt early commitment where there is no equivalent retry/failover contract to preserve.
  5. If committing/flushing fails, record a bounded writer/client transport failure and stop the stream. Do not classify a downstream writer failure as an upstream channel failure and do not retry it against another provider.
  6. Preserve existing copyCodexSSEHeaders, Cache-Control: no-cache, no-transform, X-Accel-Buffering: no, write deadlines, bounded queue/backpressure and heartbeat ordering.
  7. Keep this transport-only. No pricing, group, model-price-lock, /fast, OAuth, routing-weight, quota, or administrator-assistant behavior may change.

Acceptance criteria

  • With OPENAI_FIRST_OUTPUT_TIMEOUT=0, a validated SSE upstream can flush HTTP 200 + SSE headers before the first model frame, and a regression test observes the flush before any data callback.
  • With first-output timeout enabled, no downstream response is committed while the attempt is still eligible for failover; headers-only / role-only / usage-only upstream activity can still end in the existing retryable 504 upstream_timeout behavior.
  • The first successfully forwarded visible content/reasoning/tool/function event commits the stream and retires the pre-output retry boundary.
  • No request is retried after downstream headers/body have been committed in a way that could duplicate user-visible output.
  • A downstream flush/write error is terminal for that client request and is not counted as an upstream provider/channel failure.
  • Dedicated stream handlers are covered by an explicit audit: each either uses the safe shared policy or documents why immediate header commit is safe.
  • Tests cover: timeout disabled early flush, timeout enabled headers-then-stall, role-only-then-stall, visible output before timeout, flush failure, and existing heartbeat ordering.
  • Relevant Go relay/helper/channel tests and repository CI/release qualification pass.

Related LMM work

This issue is intentionally separate from those: it tracks the new upstream TTFT/header-commit optimization and the exact boundary needed to adopt it without regressing LMM's failover semantics.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions