Upstream source
Upstream #7509 adds CommitEventStreamHeaders: after an upstream response has been accepted as a stream, it writes and flushes the downstream SSE headers before scanning the first data frame. The goal is to avoid making clients wait for the first model token before receiving HTTP 200 / text/event-stream. It also treats header-flush failures as writer errors and updates stream handlers that bypass the shared scanner.
Current LMM behavior
Current main (997a1b673dcf718d75cbd957c437828a86c1a319) has the same basic transport gap:
apps/api-go/relay/helper/common.go has SetEventStreamHeaders, but it only sets headers and does not flush them.
apps/api-go/relay/helper/stream_scanner.go calls copyCodexSSEHeaders then SetEventStreamHeaders before reading the stream, so clients may not observe the HTTP/SSE response until the first body write.
However, LMM deliberately diverged from upstream with merged PR #312 (480cfc329c881a7670148b5aa335fa60f0d0babf): OpenAI-compatible streams can use OPENAI_FIRST_OUTPUT_TIMEOUT to fail over a channel that returns upstream headers but never produces visible output.
Current OaiStreamHandler:
- starts
FirstResponseTimeout after upstream headers;
- only marks the first response after visible content/reasoning/tool/function output is successfully handled;
- on timeout before visible output, marks
ContextKeyUpstreamChannelFailure and returns 504 upstream_timeout, allowing the existing channel retry/failover path;
- avoids retry after visible output.
StreamScannerHandler also intentionally prevents locally generated heartbeat writes before the first business response.
Why not copy #7509 directly
Flushing the downstream writer normally commits the HTTP response. If LMM commits 200 text/event-stream immediately after upstream headers, the later first-visible-output timeout can no longer reliably return a real 504 before output, and the existing pre-output failover boundary may be weakened or become dependent on already-committed downstream state.
That is a product/transport-policy conflict, not a mechanical backport. The upstream latency improvement is useful, but LMM should not trade away its stalled-stream failover silently.
Suggested LMM adaptation
Treat downstream header commitment as part of the retry boundary rather than applying CommitEventStreamHeaders unconditionally:
- Add a small explicit helper for committing/flushing SSE headers, but call it only when the current attempt is no longer eligible for pre-output failover.
- For OpenAI-compatible streams with
FirstResponseTimeout > 0, keep the response uncommitted while the attempt is still in the pre-visible-output retry window. Commit immediately before the first downstream-visible event, then permanently disable retry for that attempt.
- When the first-output timeout is disabled, allow early header flush after the upstream response has been validated as a stream, matching the latency benefit of upstream #7509.
- Audit dedicated streaming handlers that do not use
StreamScannerHandler; only adopt early commitment where there is no equivalent retry/failover contract to preserve.
- If committing/flushing fails, record a bounded writer/client transport failure and stop the stream. Do not classify a downstream writer failure as an upstream channel failure and do not retry it against another provider.
- Preserve existing
copyCodexSSEHeaders, Cache-Control: no-cache, no-transform, X-Accel-Buffering: no, write deadlines, bounded queue/backpressure and heartbeat ordering.
- Keep this transport-only. No pricing, group, model-price-lock,
/fast, OAuth, routing-weight, quota, or administrator-assistant behavior may change.
Acceptance criteria
- With
OPENAI_FIRST_OUTPUT_TIMEOUT=0, a validated SSE upstream can flush HTTP 200 + SSE headers before the first model frame, and a regression test observes the flush before any data callback.
- With first-output timeout enabled, no downstream response is committed while the attempt is still eligible for failover; headers-only / role-only / usage-only upstream activity can still end in the existing retryable
504 upstream_timeout behavior.
- The first successfully forwarded visible content/reasoning/tool/function event commits the stream and retires the pre-output retry boundary.
- No request is retried after downstream headers/body have been committed in a way that could duplicate user-visible output.
- A downstream flush/write error is terminal for that client request and is not counted as an upstream provider/channel failure.
- Dedicated stream handlers are covered by an explicit audit: each either uses the safe shared policy or documents why immediate header commit is safe.
- Tests cover: timeout disabled early flush, timeout enabled headers-then-stall, role-only-then-stall, visible output before timeout, flush failure, and existing heartbeat ordering.
- Relevant Go relay/helper/channel tests and repository CI/release qualification pass.
Related LMM work
This issue is intentionally separate from those: it tracks the new upstream TTFT/header-commit optimization and the exact boundary needed to adopt it without regressing LMM's failover semantics.
Upstream source
c58cfc1b4577a9fc129b0fc6e5dffda82038cf8e9c293e8c02371bda844af79e3500ff2d516d1ddaUpstream #7509 adds
CommitEventStreamHeaders: after an upstream response has been accepted as a stream, it writes and flushes the downstream SSE headers before scanning the first data frame. The goal is to avoid making clients wait for the first model token before receiving HTTP 200 /text/event-stream. It also treats header-flush failures as writer errors and updates stream handlers that bypass the shared scanner.Current LMM behavior
Current
main(997a1b673dcf718d75cbd957c437828a86c1a319) has the same basic transport gap:apps/api-go/relay/helper/common.gohasSetEventStreamHeaders, but it only sets headers and does not flush them.apps/api-go/relay/helper/stream_scanner.gocallscopyCodexSSEHeadersthenSetEventStreamHeadersbefore reading the stream, so clients may not observe the HTTP/SSE response until the first body write.However, LMM deliberately diverged from upstream with merged PR #312 (
480cfc329c881a7670148b5aa335fa60f0d0babf): OpenAI-compatible streams can useOPENAI_FIRST_OUTPUT_TIMEOUTto fail over a channel that returns upstream headers but never produces visible output.Current
OaiStreamHandler:FirstResponseTimeoutafter upstream headers;ContextKeyUpstreamChannelFailureand returns504 upstream_timeout, allowing the existing channel retry/failover path;StreamScannerHandleralso intentionally prevents locally generated heartbeat writes before the first business response.Why not copy #7509 directly
Flushing the downstream writer normally commits the HTTP response. If LMM commits
200 text/event-streamimmediately after upstream headers, the later first-visible-output timeout can no longer reliably return a real 504 before output, and the existing pre-output failover boundary may be weakened or become dependent on already-committed downstream state.That is a product/transport-policy conflict, not a mechanical backport. The upstream latency improvement is useful, but LMM should not trade away its stalled-stream failover silently.
Suggested LMM adaptation
Treat downstream header commitment as part of the retry boundary rather than applying
CommitEventStreamHeadersunconditionally:FirstResponseTimeout > 0, keep the response uncommitted while the attempt is still in the pre-visible-output retry window. Commit immediately before the first downstream-visible event, then permanently disable retry for that attempt.StreamScannerHandler; only adopt early commitment where there is no equivalent retry/failover contract to preserve.copyCodexSSEHeaders,Cache-Control: no-cache, no-transform,X-Accel-Buffering: no, write deadlines, bounded queue/backpressure and heartbeat ordering./fast, OAuth, routing-weight, quota, or administrator-assistant behavior may change.Acceptance criteria
OPENAI_FIRST_OUTPUT_TIMEOUT=0, a validated SSE upstream can flush HTTP 200 + SSE headers before the first model frame, and a regression test observes the flush before any data callback.504 upstream_timeoutbehavior.Related LMM work
This issue is intentionally separate from those: it tracks the new upstream TTFT/header-commit optimization and the exact boundary needed to adopt it without regressing LMM's failover semantics.