Skip to content

perf(catalog): skip episode catalog writes that change nothing - #2086

Open
blurbery wants to merge 1 commit into
Silo-Server:mainfrom
blurbery:perf/catalog-skip-unchanged-episode-entries
Open

blurbery wants to merge 1 commit into
Silo-Server:mainfrom
blurbery:perf/catalog-skip-unchanged-episode-entries

Conversation

@blurbery

@blurbery blurbery commented Oct 8, 2026 •

Copy link
Copy Markdown
Contributor

Problem

Related issue: N/A
Validation tasks: none affected. Catalog results don't change, only how often their rows are rewritten.

Episode catalog entries get rewritten even when nothing in them changed, and on my server those rewrites are about a quarter of all WAL. episode_catalog_entries is refreshed in place by triggers on episodes, media files, memberships and series. The refresh UPDATE assigned all 30 stored columns on every call with no change check. Each rewrite adds a row version plus entries in most of the table's 30 indexes (14% of its updates are HOT).

Two things make many of those rewrites no-ops:

  • The episodes trigger counts still_path and still_thumbhash as catalog changes, but entries store no still. Catalog reads join the still from episodes (episode_catalog_source.go). Every still cache write therefore refreshes the entry for nothing.
  • File and series changes that the trigger watches but the entry doesn't surface still rewrite the whole entry: a re-probe that rewrites video_tracks with the same HDR and codec facts, or a series runtime that an episode's own runtime masks.

This makes the refresh write only when a stored value differs, and stops still edits from triggering it.

Approach

One migration, which replaces the definitions from 20260930130212:

  • refresh_episode_catalog_entry's in-place UPDATE now has ROW(<30 stored columns>) IS DISTINCT FROM ROW(<src values>) in its WHERE. It still rewrites an entry whose search documents are NULL, so the BEFORE UPDATE trigger keeps repairing them. When the UPDATE finds nothing to change, the loop exits if the entry already holds the refreshed values, instead of falling through to the INSERT ... ON CONFLICT DO NOTHING, which would otherwise spin. A missing entry still gets inserted. An entry another session inserted with different values is still retried and updated.
  • episode_catalog_entries_episodes_trigger and trg_episode_catalog_entries_episodes drop still_path and still_thumbhash from the change check and from UPDATE OF.
  • Down restores the previous definitions verbatim.

updated_at on an entry now only moves when the entry changes. Nothing in the server reads it.

Validation

  • New TestSkipUnchangedEpisodeCatalogRefreshMigrationPostgres runs the real functions in an isolated schema. It applies 20260930130212, then this migration's Up, Down and Up again, and counts refresh calls and entry writes with transaction-local statistics. With the previous definitions (the predecessor and Down stages) a still edit refreshes and rewrites the entry, and so do an unchanged refresh and a masked series runtime. After Up they write nothing. A file edit, a series genre change, a deleted entry and a missing search document still write once each, and the entry matches its sources at every stage. Added to scripts/ci/db-pins.txt.
  • The existing TestEpisodeSearchRefreshWorkPostgres, TestEpisodeSearchRefreshMigrationPostgres and the other episode catalog and search tests in ./internal/catalog pass against a database migrated with this branch.
  • make migrate-validate passes. gofmt and go vet ./migrations/ are clean. golangci-lint with gocritic on ./migrations/, filtered to changed lines, reports 0 issues. (make lint-changed lints the whole tree whenever a .sql file under migrations/ changes, which CI's Go lint job covers.)

Benchmarks

Baseline main at ca186fe, changed this branch. PostgreSQL 18.6 with pgvector 0.8.7 on an Apple Silicon Mac, shared_buffers 512 MB, C locale. I built one database migrated to main and one to this branch, each seeded with the same library: 20 series of 500 episodes, one file each, so 10,000 entries. Each workload ran as one transaction. WAL is pg_current_wal_insert_lsn() before and after, entry updates are pg_stat_xact_user_tables.n_tup_upd for episode_catalog_entries, and time is wall time for the statement. There was a CHECKPOINT between runs, and each workload ran three times in the order shown.

Workload main: entry updates / WAL / time This branch: entry updates / WAL / time
New still on 10,000 episodes 10,000 / 59.0, 92.0, 97.2 MB / 1,290, 1,589, 1,538 ms 0 / 21.8, 32.8, 35.5 MB / 566, 604, 637 ms
Re-probe of 10,000 files (video_tracks rewritten, same facts) 10,000 / 96.0, 95.2, 100.6 MB / 1,226, 1,378, 1,356 ms 0 / 32.0, 30.2, 30.9 MB / 746, 797, 813 ms
Runtime change on 20 series whose episodes have their own 10,000 / 77.3, 71.0, 74.0 MB / 1,095, 989, 1,007 ms 0 / 6.8, 0.06, 0.02 MB / 505, 496, 496 ms

The WAL left on this branch is the episodes and media_files rows themselves and, for stills, the artwork GC queue. After those three workloads the entries table was 125 MB on main and 45 MB on this branch.

A change that does reach the entries costs the same. I ran it first, on freshly seeded databases, so neither table was bloated:

Genre change on 20 series (10,000 entries really change) main This branch
Entry updates 10,000 10,000
WAL 37.2, 59.4, 62.0 MB 37.2, 59.5, 61.8 MB
Time 671, 825, 808 ms 701, 799, 816 ms

On my server (read-only pg_stat_statements, about three days since a stats reset): the entry refresh UPDATE ran 176,249 times and wrote 3.1 GB of WAL, about a quarter of the cluster's 12 GB in that window. The still cache's UPDATE episodes SET still_path ... ran 26,513 times, and each one refreshed an entry. 527 series-level refreshes rewrote 125,723 entries. I can't tell from the statistics how many of those rewrites changed nothing.

Limitations: synthetic data on one machine. The production figures are the before side only, because measuring after would need a deploy.

Evidence

No user-visible change: catalog and search results are the same, and the tests check entries against their sources at every stage. The raw output behind Benchmarks and Validation:

My server, read-only statistics
pg_stat_statements (track = all), reset 2026-10-05 01:25 UTC; pg_stat_wal since the same reset

calls  | rows   | wal     | total_s | statement
176249 | 174940 | 3121 MB | 158.9   | UPDATE public.episode_catalog_entries SET series_id = src.series_id, ... (the refresh)
527    | 527    | 2437 MB | 142.4   | SELECT public.refresh_episode_catalog_entries_for_series(NEW.content_id)
527    | 125723 | 2437 MB | 142.1   | SELECT public.refresh_episode_catalog_entry(el.episode_id, el.media_folder_id) FROM ... (series fan-out)
24870  | 24870  | 385 MB  | 63.9    | SELECT public.refresh_episode_catalog_entry(NEW.episode_id, NEW.media_folder_id)
42069  | 42069  | 268 MB  | 46.9    | SELECT public.refresh_episode_catalog_entries_for_episode(NEW.content_id)
26513  | 26513  | 522 MB  | 97.3    | UPDATE episodes SET still_path = $3, still_source_path = $2, still_thumbhash = $4, updated_at = NOW ...

cluster WAL in the same window: 12 GB

episode_catalog_entries: n_live_tup 282045, n_tup_upd 175007, n_tup_hot_upd 24454 (14.0%), 30 indexes, 1467 MB total
Local benchmark, raw rows
PostgreSQL 18.6 (local). 20 series x 500 episodes, one file each: 10,000 entries per database.
Each run is one transaction; CHECKPOINT between runs.
Columns: workload, run, WAL bytes, entry n_tup_upd, entry n_tup_hot_upd, ms

== main
still_edit_10k	1	58973584	10000	8	1289.5
still_edit_10k	2	91980336	10000	17	1589.2
still_edit_10k	3	97219304	10000	25	1537.7
file_reprobe_10k	1	95986608	10000	30	1225.5
file_reprobe_10k	2	95166720	10000	35	1378.4
file_reprobe_10k	3	100597248	10000	45	1356.4
masked_series_runtime_20	1	77273416	10000	46	1094.5
masked_series_runtime_20	2	71009680	10000	48	988.7
masked_series_runtime_20	3	74002448	10000	50	1007.2
episode_catalog_entries total size after: 125 MB
== branch
still_edit_10k	1	21817648	0	0	565.6
still_edit_10k	2	32831608	0	0	604.0
still_edit_10k	3	35536928	0	0	636.7
file_reprobe_10k	1	32032224	0	0	745.8
file_reprobe_10k	2	30201144	0	0	797.1
file_reprobe_10k	3	30887512	0	0	812.9
masked_series_runtime_20	1	6830536	0	0	505.1
masked_series_runtime_20	2	61032	0	0	495.8
masked_series_runtime_20	3	15552	0	0	496.0
episode_catalog_entries total size after: 45 MB

Real change, run first on freshly seeded databases:
== main
series_genre_20	1	37227472	10000	0	671.2
series_genre_20	2	59359624	10000	0	825.1
series_genre_20	3	61968528	10000	0	807.8
== branch
series_genre_20	1	37227696	10000	0	700.5
series_genre_20	2	59508520	10000	0	798.7
series_genre_20	3	61811384	10000	0	815.6

Workloads:
still_edit_10k            UPDATE episodes SET still_path = <new original path>, still_thumbhash = <new> (10,000 rows)
file_reprobe_10k          UPDATE media_files SET video_tracks = jsonb_set(video_tracks, '{0,probe}', <random>) (10,000 rows)
masked_series_runtime_20  UPDATE media_items SET runtime = runtime + 1 (20 series whose episodes have their own runtime)
series_genre_20           UPDATE media_items SET genres = <toggle Drama/Comedy> (20 series)
Tests
== TestSkipUnchangedEpisodeCatalogRefreshMigrationPostgres and TestEpisodeSearchRefreshMigrationPostgres
--- PASS: TestEpisodeSearchRefreshMigrationPostgres (0.07s)
--- PASS: TestSkipUnchangedEpisodeCatalogRefreshMigrationPostgres (0.07s)
ok  	github.com/Silo-Server/silo-server/migrations	0.472s

== same test with the missing-document clause removed from the migration (mutation check)
    skip_unchanged_episode_catalog_refresh_test.go:107: up: missing document wrote 0 entry rows, want 1
FAIL

== go test -run 'EpisodeSearch|EpisodeCatalog|Episode' ./internal/catalog/ (database migrated with this branch)
ok  	github.com/Silo-Server/silo-server/internal/catalog	3.933s

Risks

  • This migration replaces a hot trigger function. The Up and Down are both covered by the migration test, and Down restores 20260930130212 exactly.
  • A refresh now compares 30 columns before writing. For a change that really reaches the entry that cost is within noise above. For a no-op it replaces a full row and index write.
  • Anything that wanted episode_catalog_entries.updated_at to move on every refresh would see it move less often. I found no reader of that column outside one test, which sets it directly.

Checklist

  • I read and can explain the complete diff.
  • This pull request addresses one concern.
  • The Evidence section shows every change a user can see, or says there is none.

AI Disclosure

  • Harness: Claude Code (desktop app)
  • Tool(s): Claude Code, with a read-only subagent for an adversarial review of the migration
  • Model(s): claude-opus-5-5
  • Involvement: AI-assisted
  • Adversarial review: a separate read-only agent reviewed the migration and test for loop termination under concurrent sessions, type and collation agreement in the row comparisons, readers of updated_at and the still columns, an exact Down, and whether the test fails on the old behaviour. It found no defects. On its suggestions I added a test stage for a refresh repairing a missing search document (checked by removing that clause, which makes the test fail), a statement timeout so a broken retry loop fails the test instead of hanging it, a comment tying the compared columns to the SET list, and a sentence in docs/architecture/postgres-search.md.

AI-assisted with Claude Opus. I directed the task and designed the work.

Episode catalog entries are refreshed in place by triggers on episodes,
media files, memberships and series. The refresh rewrote all 30 stored
columns on every call, even when the values matched, and each rewrite
adds a row version and an entry in most of the table's 30 indexes.

Two sources of those no-op rewrites:

- The episodes trigger counted still_path and still_thumbhash as
  catalog changes, but entries store no still: catalog reads join the
  still from episodes. Every still cache write refreshed an entry.
- Changes to watched file or series columns that the entry doesn't
  surface (a re-probed video track, a series runtime masked by the
  episode's own) still rewrote the whole entry.

The refresh UPDATE now only writes when a stored value differs, or a
search document is missing. When it finds nothing to change, the loop
exits if the entry already holds the refreshed values instead of falling
through to the INSERT. The episodes trigger no longer watches the still
columns. Down restores the previous definitions.
@silo-kody

silo-kody Bot commented Oct 8, 2026 •

Copy link
Copy Markdown

Silo Kody — review complete

Review finished. Check the inline comments for findings and verify each suggestion against the code and tests.

Reviewing changes in Silo
  • Include the related issue, expected behavior, and validation steps in the PR description.
  • For API changes, describe the effect on Apple and Android clients and Jellyfin compatibility.
  • For plugin changes, identify the affected SDK contract, plugin, and catalog entry.
  • Follow this repository's AGENTS.md and CONTRIBUTING.md.
  • Request another review with @kody start-review in a PR comment.
  • React with 👍 or 👎 to give feedback on individual suggestions.
Review settings
Review Options

The following review options are enabled or disabled:

Options Enabled
Bug ✅
Performance ✅
Security ✅
Business Logic ❌

@coderabbitai

coderabbitai Bot commented Oct 8, 2026

Copy link
Copy Markdown

Warning

Review limit reached

You've used all free OSS reviews for now. Wait for the free limit to reset to keep reviewing this public repository.

Next included review available in 13 minutes.

Check out review usage here.

View limit details

Limit details: You’ve used all 4 included reviews currently available.

Learn how review limits work.

Review configuration:

⚙️ Run configuration
  • Configuration used: Organization UI
  • Review profile: CHILL
  • Plan: Advanced
  • Run ID: 15088a8f-85a8-4327-88b4-13b5dd7da648
📥 Commits

Reviewing files that changed from the base of the PR and between ca186fe and b85ffd3.

📒 Files selected for processing (4)
  • docs/architecture/postgres-search.md
  • migrations/skip_unchanged_episode_catalog_refresh_test.go
  • migrations/sql/20261008034854_skip_unchanged_episode_catalog_refresh.sql
  • scripts/ci/db-pins.txt
  • Autopilot · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@Quick104 Quick104 added priority: P2 Limited scope, workaround exists, or polish impact: perf Unusable slowness on normal hardware labels Oct 8, 2026 — with Cursor
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

impact: perf Unusable slowness on normal hardware priority: P2 Limited scope, workaround exists, or polish

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants