Rewritten after looking at production (the original plan here, shortening the 30-day identity cache, would make this worse).
What production shows (cloudflare-api-mcp, last hour, Workers Observability)
|
per hour |
POST /mcp requests |
~1.24M |
api_token_identity_probe log lines |
~23.6k (0 from the OAuth callback probe) |
…of which 429 retries (fetchWithRetry: 429) |
~18.5k |
| probes that failed after all retries |
~5.3k |
MCP requests answered 429 temporarily_unavailable |
~4.1k |
All-time on the Issues tab: 4.79M fetchWithRetry failed … /accounts … 429 and 1.59M … /user … 429. Both the account-token path (per_page=5, ~6.4k/h) and the user/legacy path (/user + per_page=31, ~8.9k/h each) are being rate-limited.
Why it amplifies itself
- Retrying 429s. Every identity probe goes through
fetchWithRetry (up to 4 attempts with about 1 s backoff), and user tokens probe two endpoints in parallel. One rate-limited request becomes up to 8 more requests against the same quota, spending the user's own Cloudflare API budget, which their real tool calls then also run out of.
- Failures aren't cached. Only successful identities are cached (30 days), so every request from a rate-limited token probes again.
- Concurrent misses aren't coalesced. A new token's first burst of requests all miss the cache together.
Fix
- Don't retry 429s in identity probes. Surface
temporarily_unavailable with Cloudflare's Retry-After straight away.
- Cache a 429 briefly (for the
Retry-After window) so repeat requests from that token back off instead of re-probing.
- Keep the 30-day success cache. Revoked-token staleness is a real trade-off, but reducing it by probing more often is the wrong lever while probes are the top error. Revisit once probe volume is under control.
Related: about 24k/h of the probes are for tokens in workers-oauth-provider's own format (expired MCP access tokens), which the library hands to resolveExternalToken. That's being fixed in the library (cloudflare/workers-oauth-provider, linked from here).
Rewritten after looking at production (the original plan here, shortening the 30-day identity cache, would make this worse).
What production shows (
cloudflare-api-mcp, last hour, Workers Observability)POST /mcprequestsapi_token_identity_probelog linesfetchWithRetry: 429)429 temporarily_unavailableAll-time on the Issues tab: 4.79M
fetchWithRetry failed … /accounts … 429and 1.59M… /user … 429. Both the account-token path (per_page=5, ~6.4k/h) and the user/legacy path (/user+per_page=31, ~8.9k/h each) are being rate-limited.Why it amplifies itself
fetchWithRetry(up to 4 attempts with about 1 s backoff), and user tokens probe two endpoints in parallel. One rate-limited request becomes up to 8 more requests against the same quota, spending the user's own Cloudflare API budget, which their real tool calls then also run out of.Fix
temporarily_unavailablewith Cloudflare'sRetry-Afterstraight away.Retry-Afterwindow) so repeat requests from that token back off instead of re-probing.Related: about 24k/h of the probes are for tokens in workers-oauth-provider's own format (expired MCP access tokens), which the library hands to
resolveExternalToken. That's being fixed in the library (cloudflare/workers-oauth-provider, linked from here).