Repository navigation
Conversation
The artwork revision GC's dormant sweep re-verifies parked revisions so a reference that disappears through a surface without a displacement trigger still becomes collectible. It picked rows by updated_at and re-stamped updated_at on every row still referenced, so each parked row was rewritten once a day. updated_at is the key of the sweep's own index, so none of those updates could be HOT and each added entries to all seven indexes. On a server with about 110k parked revisions that was 258k row rewrites and 753 MB of WAL in three days, for rows that didn't change. The sweep now walks parked rows in id order from a cursor persisted in artwork_revision_gc_dormant_cursor and writes only the rows it requeues. It still skips rows changed within the recheck interval, and starts a new pass at most once per interval. The cursor only advances from the position a sweep read, so concurrent sweeps repeat work at worst. The (updated_at, id) index existed only for the old ordering and is dropped.
Silo Kody — review completeReview finished. Check the inline comments for findings and verify each suggestion against the code and tests. Reviewing changes in Silo
Review settingsReview OptionsThe following review options are enabled or disabled:
|
|
No actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info⚙️ Run configuration
📒 Files selected for processing (5)
Included review availability: This review used your included allowance. Your plan provides up to 4 included reviews per hour; 1 remain after this review. 📝 WalkthroughWalkthroughThe dormant artwork-revision sweep now processes parked candidates in bounded batches using a persisted ID cursor. A migration adds the cursor table and updates the dormant-candidate index. Tests cover cursor continuation, cycle timing, orphan requeueing, and migration behavior. ChangesDormant Revision Sweep
Priority: ➖ Normal Estimated code review effort: 3 (Moderate) | ~25 minutes Change: Refactor Sequence Diagram(s)sequenceDiagram
participant sweepDormant
participant cursor_table as artwork_revision_gc_dormant_cursor
participant database
sweepDormant->>cursor_table: Read cursor and cycle time
sweepDormant->>database: List eligible candidates after cursor ID
sweepDormant->>database: Check candidate references
sweepDormant->>database: Requeue unreferenced candidates
sweepDormant->>cursor_table: Advance or reset cursor
Suggested reviewers: Merge Risk: ⚪ Minimal · up to The cursor sweep is mergeable after normal checks; no concrete issue requiring a pre-merge fix was identified. 🚥 Pre-merge checks | ✅ 4 | ❌ 1❌ Failed checks (1 warning)
✅ Passed checks (4 passed)
Full details: Docstring CoverageExplanation Docstring coverage is 50.00% which is insufficient. The required threshold is 80.00%. Docstring coverage is scoped to functions touched by this diff. Analyzed 8 functions across 3 files. (2 skipped: 2 unsupported.)
✨ Finishing Touches🧪 Generate unit tests (beta)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
Problem
Related issue: N/A
Validation tasks: none affected. Artwork that GC keeps or deletes doesn't change, only how often parked rows are rewritten.
The artwork revision GC rewrites every parked revision that is still in use once a day, even though nothing about the row changes. Its dormant sweep re-verifies parked revisions (
next_attempt_at IS NULL), so a reference that disappears through a surface without a displacement trigger still becomes collectible. It picks rows whoseupdated_atis more than 24 hours old, then setsupdated_at = NOW()on every row that is still referenced.updated_atis the key of the sweep's own index (artwork_revision_gc_dormant_idx), so none of those updates can be HOT, and each adds an entry to all seven of the table's indexes.On my server (read-only statistics, about three days): 110,862 parked revisions. The touch UPDATE ran 41 times, rewrote 258,011 rows and wrote 753 MB of WAL.
n_tup_hot_updwas 0 of 327,314 updates. Theoriginal_pathunique index is 271 MB for 140k rows, and the heap 654 MB.This keeps the sweep's checks and stops it rewriting rows that didn't change.
Approach
artwork_revision_gc_dormant_cursorrow (after_id,cycle_started_at) holds the sweep's position. Each run reads up to 10,000 parked rows withid > after_idin id order, still skipping rows changed within the last 24 hours as before, and requeues the ones nothing references. It writes nothing to the rows that are still referenced. Then it moves the cursor to the last id it read.after_id = 0). A new cycle starts only oncecycle_started_atis 24 hours old, so each parked row is still checked about once a day.artwork_revision_gc_dormant_idx, which only the oldORDER BY updated_at, idused. The new query walks the primary key. Down recreates the index exactly as20260714120826defined it.A parked row is still checked about once a day, but in the worst case it now waits about two days instead of one (see Risks). With 10,000 rows a run and an hourly task, a cycle over my 110k parked rows takes about 11 runs.
Validation
TestArtworkRevisionGCDormantSweepnow also checks that the referenced row keeps its row version (ctid and xmin). It fails withmain'ssweepDormant(referenced dormant row was rewritten) and passes here.TestArtworkRevisionGCDormantSweepCyclescovers six things: a bounded sweep stops at its last row and the next resumes there; a row behind the cursor waits for the next cycle; the next cycle doesn't start before the recheck interval; once it does, it records its start and re-arms the unreferenced row; and a missing cursor row is recreated. Each test passes on its own on a fresh database. Both tests are added toscripts/ci/db-contracts.txt.TestArtworkRevisionGCDormantCursorMigrationChangesIndexConcurrentlychecks the migration isNO TRANSACTIONand that Down drops a failed build concurrently before recreating the index.go test ./internal/metadata/against a migrated PostgreSQL 18 database passes. Applying the migration, rolling back with--migrate-down-to, and applying again leaves the expected table and index.make migrate-validatepasses.gofmtandgo vetare clean. golangci-lint with gocritic on./internal/metadata/and./migrations/, filtered to changed lines, reports 0 issues. (make lint-changedlints the whole tree whenever a.sqlfile undermigrations/changes, which CI's Go lint job covers.)Benchmarks
Baseline
mainat ca186fe, changed this branch. PostgreSQL 18.6 with pgvector 0.8.7 on an Apple Silicon Mac,shared_buffers512 MB, autovacuum off for the table during the run.Workload: two databases, one migrated to
mainand one to this branch, each seeded the same way: 100,000 parked revisions still referenced by a movie poster and 1,000 that lost their reference, all last changed two days ago. Then 24 calls to the realsweepDormantwith the production batch size, one day of hourly runs, through a temporary test harness. WAL is thepg_current_wal_insert_lsn()difference, and updates arepg_stat_user_tablesafter the sweep connections closed. One repetition, because the first day's writes change the state for a second.mainBoth versions check every parked row once in the day and requeue the same 1,000 rows. On
mainthe same 100,000 rows are rewritten again the next day.Limitations: synthetic rows on one machine, so this local WAL per row is smaller than on my server, where rows are wider and the indexes bloated. Most of each run's time in production is the reference check (
referencedPaths, 3.3 to 4.0 s per 10,000-path batch on my server), which this doesn't change. A serial plan for that query measured 0.6 s; that's a separate change.Evidence
No user-visible change: the same artwork is kept and deleted; only the bookkeeping writes stop. The raw output behind Benchmarks and Validation:
My server, read-only statistics
Local benchmark, raw output
Tests
Risks
ORDER BY updated_at, idquery without the index, as a sequential scan of the candidates table, and still touches rows, which the new code doesn't mind.Checklist
AI Disclosure
CASE, the missing cursor row and long-lived databases weren't covered. I fixed the helper and added those cases: with theCASEbroken, the cycle test now fails. It also pointed out the longer worst-case recheck, which is under Risks.AI-assisted with Claude Opus. I directed the task and designed the work.