Context
The Ask ApplyWell assistant currently maps provider failures to a generic SSE error (Assistant provider is unavailable), which hides useful Gemini rate-limit information and makes it impossible for the frontend to identify which user message failed.
Grafana/Loki logs show Gemini 429 RESOURCE_EXHAUSTED responses containing a provider-provided retry delay, for example:
Quota exceeded for metric: generativelanguage.googleapis.com/generate_content_free_tier_requests
Please retry in 35.878391973s.
Goal
Expose structured provider error information through SSE so the frontend can show the correct failed message state and allow the user to retry it.
Backend requirements
- Detect Gemini
429 RESOURCE_EXHAUSTED / quota errors explicitly.
- If Gemini provides a retry delay, extract it and propagate it to the SSE error event.
- Do not invent a retry delay when the provider does not send one.
- Preserve the generic provider unavailable behavior for non-rate-limit failures.
- Keep provider-specific parsing isolated from the controller when possible.
Suggested SSE payload
{
"code": "RATE_LIMITED",
"message": "Gemini rate limit exceeded",
"retryAfterSeconds": 36
}
For generic provider failures:
{
"code": "PROVIDER_UNAVAILABLE",
"message": "Assistant provider is unavailable"
}
Frontend / chat UX requirements
The frontend must be able to detect exactly which user message originated the failed request.
When a request fails:
- Mark the corresponding user message as failed.
- Show a visible error/status icon next to that message.
- Show a retry icon/action next to the failed message.
- Retrying must resend only that failed message.
- Do not duplicate the user message in chat when retrying.
- When the retry succeeds, clear the failed state.
Rate-limit UX
When the error contains retryAfterSeconds:
- Show a message such as
Rate limit reached. Try again in 36s.
- Show a countdown using the provider-provided value.
- Disable the retry action while the countdown is active.
- Enable retry when the countdown reaches zero.
- Do not automatically resend the message when the countdown ends.
When no retry delay is provided:
- Keep the manual retry action available.
- Do not fabricate a timeout.
Tests
Add coverage for:
- Gemini
429 with retry delay.
- Gemini
429 without retry delay.
- Generic provider failure.
- Structured SSE error payload.
- Failed user-message state in the frontend.
- Retry resending the original message only once.
- Retry not duplicating the user message.
- Countdown disabling retry until the provider-provided delay expires.
- Successful retry clearing the failed state.
Notes
The backend should expose structured error metadata instead of requiring the frontend to parse human-readable provider messages.
Context
The Ask ApplyWell assistant currently maps provider failures to a generic SSE error (
Assistant provider is unavailable), which hides useful Gemini rate-limit information and makes it impossible for the frontend to identify which user message failed.Grafana/Loki logs show Gemini
429 RESOURCE_EXHAUSTEDresponses containing a provider-provided retry delay, for example:Goal
Expose structured provider error information through SSE so the frontend can show the correct failed message state and allow the user to retry it.
Backend requirements
429 RESOURCE_EXHAUSTED/ quota errors explicitly.Suggested SSE payload
{ "code": "RATE_LIMITED", "message": "Gemini rate limit exceeded", "retryAfterSeconds": 36 }For generic provider failures:
{ "code": "PROVIDER_UNAVAILABLE", "message": "Assistant provider is unavailable" }Frontend / chat UX requirements
The frontend must be able to detect exactly which user message originated the failed request.
When a request fails:
Rate-limit UX
When the error contains
retryAfterSeconds:Rate limit reached. Try again in 36s.When no retry delay is provided:
Tests
Add coverage for:
429with retry delay.429without retry delay.Notes
The backend should expose structured error metadata instead of requiring the frontend to parse human-readable provider messages.