Skip to content

Epic: Complete the attachment drain (Phase 2) #78

Description

@splaice

Context

The Apr 2026 drain ended with 9,546 extracted / 1,991 failed / 24,033 skipped / 0 pending. Full retrospective: docs/retrospectives/attachment-drain-retrospective.md.

The retrospective identifies ~1,900 of the 1,991 failures as recoverable, but only after specific gaps are closed. This epic gathers the relevant issues, sequences them, and notes opportunistic items worth pulling in alongside.

Status (2026-05-02)

All codable child issues merged. PRs:

Operational/external items remain by design:

Codebase is drain-ready end-to-end. Next step: kick off the drain.

Status (2026-07-12)

The drain ran. Two phases on 2026-05-03: Phase 1 (300s extract-timeout, 3h 52m) + Phase 2 (900s extract-timeout, 8h). Final state: 10,746 extracted / 790 failed / 24,034 skipped / 0 pending. The failed pool is dominated by one bucket — 681 slow-tail items (mid-size PDFs, oversized images, large xlsx) that exceed even the 900s budget. Full analysis: docs/retrospectives/attachment-residuals-2026-05-03.md.

Since then, the MarkItDown extraction leg (#84/#85, "Tier 4") recovered part of the intentionally-skipped pool (ical/csv/xls/json/xml/vcard/pages).

Remaining on this epic: the external/passive items only (#71 upstream watch, #75 file upstream issue) and a decision on whether the 681-item slow tail is worth further budget. Closing out the tracking below.

Recommended sequence

  1. ✅ Ship Tier 1 (Add Docling fallback for marker-failed office formats (DOCX/XLSX/PPTX) #61, Sanitize NUL bytes in extracted markdown before INSERT #62, NUL-strip _set_status reason field #67) as a single PR.
  2. ⏳ Flip remaining `failed → pending`.
  3. ⏳ Run the drain (`process_attachments run --workers 1 --extract-timeout 300 --max-runtime 86400`).
  4. ⏳ Run the hard-timeout retry pass (Operational: hard-timeout retry pass with 900s budget after next drain #63) at 900s.
  5. ✅ Tier 2 telemetry/safety items landed before the drain begins.

Tier 1 — Must do (recovers the failure pool)

The retrospective explicitly calls these out as "do these before the next phase":

# Title P Estimated yield Status
#61 Docling fallback for marker-failed office formats P0 ~1,170 ✅
#62 Sanitize NUL bytes in extracted markdown before INSERT P0 ~8 + prevents the entire class ✅
#63 Hard-timeout retry pass with 900s budget P1 ~200–360 ⏳ post-drain
#67 NUL-strip `_set_status` reason field P2 sister fix to #62, trivial ✅

Tier 2 — Opportunistic, ship alongside Tier 1

Cheap and tightly related to the drain. Skipping these means re-learning lessons from the retrospective.

# Title P Why now Status
#64 Document parallel-supervisor MPS hazard (code + runbook) P1 prevents the dual-supervisor disaster from recurring ✅
#66 Per-content-type yield dashboard in `maildb jobs` P1 surfaced the office-format gap immediately on smoke-test ✅
#65 Pre-extraction filter on tiny images P2 telemetry honesty ✅
#76 `just smoke-marker` sanity check P2 catches dep corruption before a drain ✅
#69 Orphan-process detector in `maildb jobs` P2 would have caught the dual-supervisor disaster ~5h earlier ✅
#70 `maildb jobs --kill-orphans` P3 builds on #69; operator ergonomics ✅

Tier 3 — Resilience extras

# Title P Status
#73 `--max-runtime` safety flag P3 ✅
#72 Persist supervisor logs to `~/.maildb/logs//` P3 ✅
#77 Standardize drain log location P3 (pair with #72) ✅
#75 File upstream surya issue for residual non-`.max()` MPS bugs P2 ⏳ external action
#71 Watch upstream surya for PR #493 release P1 (passive) ⏳ re-check next quarter

Explicitly excluded

Success criteria

  • Tier 1 PR(s) merged.
  • Tier 2 telemetry/safety items in place before the run begins.
  • Tier 3 resilience extras merged.
  • Drain run completes; remaining failures investigated (residuals report, 2026-05-03).
  • Hard-timeout retry pass at 900s completes (Phase 2 of the 2026-05-03 drain).
  • Final yield: >95% office formats (via Docling), >90% PDFs (Marker + 900s retry budget).

Tracking

Tier 1: #61 ✅, #62 ✅, #63 ⏳, #67 ✅
Tier 2: #64 ✅, #65 ✅, #66 ✅, #69 ✅, #70 ✅, #76 ✅
Tier 3: #71 ⏳, #72 ✅, #73 ✅, #75 ⏳, #77 ✅

🤖 Generated with Claude Code

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    P0Critical: blocks next phase or causes data lossepicEpic: tracks a multi-issue initiativeextractionAttachment extraction pipelineopsOperational concerns: runbooks, monitoring, lifecycle

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions