Skip to content

Operational: hard-timeout retry pass with 900s budget after next drain #63

Description

@splaice

Context

Of 1,991 failures from the Apr 2026 drain, 721 are `hard-timeout: killed after 300s`. Some fraction of these are borderline-slow rather than true hangs. Spike data showed 15/45 successes were in the 300–600s window; with a 900s budget more would land.

Proposal

After the next full drain finishes, run:

```
uv run maildb process_attachments retry --hard-timeouts-only --extract-timeout 900 --workers 1
```

The supervisor will catch the true hangs (now at 900s) and let borderline-slow docs complete. Each true-hang doc costs 900s of wall time (vs 300s) but only fires once per row per session.

Acceptance criteria

  • Pass executed against the 721 hard-timeout rows after next drain.
  • New successes recorded in DB; remaining hard-timeouts re-stamped with new `hard-timeout: killed after 900s` reason.
  • Result added to retrospective.

Estimated impact

200–360 additional successful extractions.

🤖 Generated with Claude Code

Activity

  1. added
    P1High: significant value, do soon
    extractionAttachment extraction pipeline
    opsOperational concerns: runbooks, monitoring, lifecycle
    on Apr 28, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    P1High: significant value, do soonextractionAttachment extraction pipelineopsOperational concerns: runbooks, monitoring, lifecycle

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions