Skip to content

feat: expose Gemini 429 retry delay and failed-message retry state #87

Description

@vitorhugo-dotnet

Context

The Ask ApplyWell assistant currently maps provider failures to a generic SSE error (Assistant provider is unavailable), which hides useful Gemini rate-limit information and makes it impossible for the frontend to identify which user message failed.

Grafana/Loki logs show Gemini 429 RESOURCE_EXHAUSTED responses containing a provider-provided retry delay, for example:

Quota exceeded for metric: generativelanguage.googleapis.com/generate_content_free_tier_requests
Please retry in 35.878391973s.

Goal

Expose structured provider error information through SSE so the frontend can show the correct failed message state and allow the user to retry it.

Backend requirements

  • Detect Gemini 429 RESOURCE_EXHAUSTED / quota errors explicitly.
  • If Gemini provides a retry delay, extract it and propagate it to the SSE error event.
  • Do not invent a retry delay when the provider does not send one.
  • Preserve the generic provider unavailable behavior for non-rate-limit failures.
  • Keep provider-specific parsing isolated from the controller when possible.

Suggested SSE payload

{
  "code": "RATE_LIMITED",
  "message": "Gemini rate limit exceeded",
  "retryAfterSeconds": 36
}

For generic provider failures:

{
  "code": "PROVIDER_UNAVAILABLE",
  "message": "Assistant provider is unavailable"
}

Frontend / chat UX requirements

The frontend must be able to detect exactly which user message originated the failed request.

When a request fails:

  • Mark the corresponding user message as failed.
  • Show a visible error/status icon next to that message.
  • Show a retry icon/action next to the failed message.
  • Retrying must resend only that failed message.
  • Do not duplicate the user message in chat when retrying.
  • When the retry succeeds, clear the failed state.

Rate-limit UX

When the error contains retryAfterSeconds:

  • Show a message such as Rate limit reached. Try again in 36s.
  • Show a countdown using the provider-provided value.
  • Disable the retry action while the countdown is active.
  • Enable retry when the countdown reaches zero.
  • Do not automatically resend the message when the countdown ends.

When no retry delay is provided:

  • Keep the manual retry action available.
  • Do not fabricate a timeout.

Tests

Add coverage for:

  • Gemini 429 with retry delay.
  • Gemini 429 without retry delay.
  • Generic provider failure.
  • Structured SSE error payload.
  • Failed user-message state in the frontend.
  • Retry resending the original message only once.
  • Retry not duplicating the user message.
  • Countdown disabling retry until the provider-provided delay expires.
  • Successful retry clearing the failed state.

Notes

The backend should expose structured error metadata instead of requiring the frontend to parse human-readable provider messages.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Labels

No labels
No labels

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions