Skip to content

Sync upstream silo-server main (~925 commits, API v2) - #183

Merged
JonahMMay merged 937 commits into
mainfrom
sync/upstream-2026-09-28
Sep 29, 2026
Merged

JonahMMay merged 937 commits into
mainfrom
sync/upstream-2026-09-28

Conversation

@JonahMMay

@JonahMMay JonahMMay commented Sep 29, 2026 •

Copy link
Copy Markdown

Summary

Syncs Prairie with upstream Silo-Server/silo-server main, about 930 commits up to 19e81846. The sync brings in /api/v2, which prairie-apple and prairie-android main already call. There are two real merge commits, both with upstream as a parent:

  • 08794f7b merges upstream f84d2dee (the bulk of the sync).
  • 0bad5e15 merges the 7 commits upstream landed while this was in progress.

Prairie branding is kept, and everything new from upstream is rebranded:

  • Go module paths and the plugin SDK import path.
  • cmd/silo references, which now point to cmd/prairie.
  • Web UI copy, Go string literals and example: tags.
  • Data directories, the Dockerfiles, the compose files and the macOS workflow.
  • The wire identifiers listed below.

The SDK is pinned to prairie-plugin-sdk main at c2f90523e166 (v0.12.1-0.20260928152428-c2f90523e166). go mod tidy has been run. mattn/go-sqlite3 drops out because nothing imports it any more (see the SQLite fix below).

Merge with 'Create a merge commit' — do not squash. A squash would repeat what went wrong with #174: upstream history would be lost, and the next sync would diff against the wrong base.

Wire renames (silo.* → prairie.*)

Before After
silo.playback-control.v2 prairie.playback-control.v2
silo.events.v2 prairie.events.v2
silo.room.v2 prairie.room.v2
silo.admin-logs.v2 prairie.admin-logs.v2
silo.ticket.<ticket> prairie.ticket.<ticket>
every X-Silo-* header (Device-Id/Name/Platform, Client, Client-Version/Build/Channel/Family, Stream-Token, Mutation-Id, Latest-Sequence, Idempotent-Replay, Restart-Required, Ingress-Token, Theme, User-/Profile-, Event/Webhook-Id/Delivery-Id/Timestamp/Signature, Tone-Map-, Transcode-, Capture-, Download-, Ebook-Conversion, and the rest) the same names as X-Prairie-*
OpenAPI header parameter location header.x-silo-device-id header.x-prairie-device-id
plugin gRPC service prefix silo.plugin.v1. (metrics label trim) prairie.plugin.v1. (matches the SDK proto package)
plugin repository source_kind silo (v2 enum and web) prairie (what the server already sends)
prairie:// invite and Watch Party deep links were silo:// in upstream's new code
theme-audio signing label silo.theme-audio.v2 prairie.theme-audio.v2

These are deliberately not renamed: the private-use language tag x-silo-original, which is a stored settings value that prairie-apple also uses; the OpenAPI vendor extensions x-silo-class and friends; and the silo-audiobooks ABS virtual library ID, which Prairie always used.

Client alignment: prairie-apple main already uses prairie.playback-control.v2, prairie.events.v2, prairie.room.v2 and prairie.ticket.. prairie-android main still offers the silo.* names, so Prairie-Server/prairie-android#31 must merge together with this PR. Without it, Android's playback-control, events and Watch Party sockets fail the subprotocol check. The SDK already uses X-Prairie-Ingress-Token, which this PR matches.

Conflict approach

  • Go: conflicts were resolved in the earlier WIP pass, then fixed up here until every package passes go vet. For files Prairie had only rebranded, the upstream version was taken and rebranded. Where Prairie's lint autofixes had deleted "unused" upstream code, or rewritten code in ways that broke upstream callers, upstream was restored. That covers catalogseed, playback_realtime, intromarkers/compare, opslog, the settings handlers and the episode-catalog SQL builders.
  • Web: hunks that differed only in formatting were resolved mechanically first. Prairie had lost its .prettierrc and drifted to 80 columns, so upstream's .prettierrc is restored and web/src is reformatted at upstream's 100 columns, which should make future syncs much smaller. Each remaining hunk was resolved by hand: Prairie features were kept (Live TV, trickplay, Quick Connect, AVIF artwork props, Prairie login layout, the prairie-dusk theme, button icons) on top of upstream's rewrites.
  • Where upstream redesigned, upstream won:
    • Setup wizard: the new step rail replaces Prairie's back/forward review navigation and its AuthBrandHero styling.
    • Theme system: upstream now ships one fixed theme. Prairie Dusk stays as that single base theme, with index.html set to data-theme="prairie-dusk" and Fraunces self-hosted in fonts.css. The other curated themes are gone.
    • API keys and the admin dashboard.
    • Autoscan inline picker: Prairie's cancel-restore was dropped because it conflicts with upstream's draft guard.
  • Prairie-owned files kept: README, TRADEMARK, CONTRIBUTING, AGENTS, DEVELOPMENT, docker.yml, discord-commits.yml and ci.yml.
  • CI changes:
    • A go-verify job with upstream's contract gates (settings bindings, playback fixtures, route inventory, migration ledger, scenario catalogs, offline routes, OpenAPI artifact, v2 fixtures). The base-diff breaking-change step is left out because it compares ~930 upstream commits against Prairie main.
    • The golangci timeout goes from 5 to 20 minutes, because a sync touches every package.
    • Catalog.test.tsx is excluded from the vitest gate. It is upstream's known failure (WEBTEST_KNOWN_FAILURES).
  • Generated contract artifacts: OpenAPI, web v2 types, v2 fixtures, the route inventory, the migration ledger (30 Live TV/Prairie rows seeded, plus a new livetv section), offline routes, route manifests and settings bindings. All are regenerated from the merged code.
  • v2 contract addition: Library gains trickplay_enabled and trickplay_supported in v2 (additive), so the web library editor keeps the trickplay toggle now that it goes through /api/v2.
  • chi pinned to v5.2.5 (upstream's version): chi 5.3 registers a QUERY method that upstream's route inventory does not model.

Prairie fixes lost in #174 (b5ca13b) and how each is handled

Restored in this PR:

  • ba9188f, Tizen HLS tags: -hls_flags temp_file, no #EXT-X-PLAYLIST-TYPE:VOD in the synthetic manifest, and RewriteManifestPaths strips #EXT-X-INDEPENDENT-SEGMENTS and …TYPE:VOD. EVENT is kept for upstream's remount timeline.
  • Web branding:
    • index.html had been reset to <title>Silo</title>, the arctic-frost theme and Google Fonts. It is restored.
    • Silo's favicon.ico, apple-touch-icon.png, maskable-icon-512.png and web-app-icon-192/512.png had replaced Prairie's (the same thing happened in android). They are restored, and the stray silo-icon-1024.png is removed.
    • assets/icon.png is restored.
    • The @fontsource-variable/fraunces dependency is back.
  • Docker: Dockerfile and Dockerfile.dev were upstream copies that built ./cmd/silo into /silo, so every Docker Image run on main has failed since feat: sync silo-server main upstream (~413 commits) #174. The last good image is 69ab5c1b. They are rebranded again, with Prairie's ldflags (versionOverride) and the Debian ffmpeg Prairie's AVIF encoder expects. The compose files are rebranded again, with the /dev/dri device, the books mount and the audiobook-covers mount put back.
  • Default paths back on /var/lib/prairie/...: userdb, the jellyfin-web install and artwork.
  • 37e2ae5, HEAD on transcode segments: both the route and the method-preserving node proxy.
  • e934c36, Redis restart-loop fix: an unreachable redis.url is rejected on save, and the log/realtime hub start logs a warning and continues instead of calling Fatal.
  • PRAIRIE_* env names: ebook/audiobook/manga enrich workers, scan workers, ebook backfill caps and rate-limit cooldown, metadata match score and push-relay development URL. Each reads PRAIRIE_* first and falls back to SILO_*.
  • AVIF/WebP artwork admin settings: defaults and validation are back. The config loader still read these settings, but the admin API could no longer set them.
  • Collection posters generate the w200 rung again. Card requests ask for it, and without it they got a 404.
  • Plugin IDs: prairie.theintrodb for the legacy IntroDB key migration, and prairie.audiobooks for the ABS install ID.
  • X-Prairie-* headers (httpheaders) had been reverted to X-Silo-* across the server, while both apps send X-Prairie-Device-Id. The rename above covers this.

Pre-existing Prairie bugs found and fixed along the way:

  • A misspell autofix changed calibre:series, calibre:series_index and calibre:isbn to caliber:*, so EPUB series and ISBN metadata were never read.
  • An ineffassign autofix dropped the argIdx++ in the episode catalog SQL builders, which reused a placeholder number.
  • userdb registers modernc as sqlite3, and an upstream package also imported mattn/go-sqlite3, which panics at init when both are linked. bridgeimport now uses modernc.

Not restored (listed so you can decide):

  • 291ec67, windowed encoded HLS resume at the start segment: not ported. Upstream reworked resume around -copyts, full-length synthetic manifests and seek restarts, and porting the window would undo that design. Its tests stay and fail: TestGenerateFullManifest_ResumeWindowsAtStartSegment, TestSegmentRecoveryDecision_BeforeStartNeverRestartsAtZero and TestSegmentRecoveryDecision_StartupProbeBeforeResumeDoesNotRestart in internal/playback. They are not in Prairie CI. This is a pre-existing regression on main that needs a Prairie-side fix if AVPlay still ignores EXT-X-START there.
  • CGO-free WebP backend (ffmpeg/WASM): feat: sync silo-server main upstream (~413 commits) #174 replaced it with upstream's bimg/libvips. The server stays CGO/libvips-based, as it has been on main since feat: sync silo-server main upstream (~413 commits) #174.
  • Rust/WASM fingerprint comparer: CompareFingerprints is now upstream's reworked Go algorithm. processor.go and the .wasm file remain but are unused.
  • internal/artworkstore (the Prairie local artwork store) was dead code after the sync, because upstream's blobstore filesystem backend replaces it, so it is removed. A local artwork cache written in the old layout may be re-cached.

Needs your judgement

  1. Image ladder: Prairie keeps its TV-tuned ladder: posters/profiles 500/300/200, stills 500/300, logo 500. Upstream moved to 780/500/300 with 1280 logos. Tests that the CI gates run were updated to the Prairie ladder. internal/metadata ladder tests and a few internal/api/handlers image-size tests still assume upstream's rungs and fail locally. These packages are not in Prairie CI.
  2. Merging main with the 291ec67 regression unresolved, and with prairie-android#31 required alongside.
  3. Env var names: new upstream variables keep the SILO_* prefix (stream telemetry, debug, scenario, test DB), and only Prairie's historical ones read PRAIRIE_* first. Decide whether you want a blanket PRAIRIE_* alias.
  4. Theme consolidation: only Prairie Dusk remains. The midnight-cinema block was removed so the high-contrast gate sees one base theme.
  5. Upstream logredact limitation: a fmt.Errorf("fetch %s: %w", rawURL, errors.Join(...)) wrapper still prints the raw URL, because the redactor only rewrites the *url.Error text. This was noticed while adding coverage tests for Prairie's gate and was not changed.

Testing

  • go vet passes on all 194 packages. golangci-lint (Prairie config, --new-from-merge-base=origin/main) shows 0 issues locally apart from one batch that the host OOM-killed; CI covers it.
  • go test, one package at a time locally, with libvips headers and libs extracted to a user prefix: the packages in Prairie's CI gates pass, including the coverage packages (lang and logredact got tests to stay ≥95%), jellycompat, imageutil, contract ledger, contract spec, route inventory, api and policy.
  • Known local failures outside Prairie CI:
    • the resume-window tests above;
    • the ladder tests in metadata and api/handlers;
    • environment-dependent tests: HDHomeRun LAN discovery finds a real tuner, plugin network-access tests time out, and the taskmanager suite times out;
    • three api/handlers tests (calendar poster batching, the first-frame metric, and v2 request labels in apiv2) that fail deterministically and were not investigated.
  • Web: tsc, prettier and eslint pass (0 errors). Vitest passes 685/685 files; one Watch Party test is timing-sensitive under load. Global coverage is 98.2/95.5/99.3/99.0.

Deploy note

Merging to main publishes an image. docker.yml runs on every push to main and pushes ghcr.io/prairie-server/prairie-server:latest plus a SHA tag. This PR fixes that workflow, which has failed on every main push since #174, so the first successful :latest since 69ab5c1b will be published on merge. Anything tracking :latest (keel, compose pulls) will roll forward. No release is cut: release.yml is dispatch-only. discord-commits.yml posts to the commits channel.

AI disclosure

Tool: Claude Code; Model: claude-opus-5-5; Involvement: AI-generated.

🤖 Generated with Claude Code

Quick104 and others added 30 commits September 23, 2026 12:03
…lear errors (Silo-Server#1329)

* fix(historyimport): import Emby watch dates, episode favorites, and clear errors

Emby list queries omit ProductionYear, UserData.PlayCount, and
UserData.LastPlayedDate unless Fields names them (verified on Emby 4.8.11,
4.9.5, and 4.10.0.40). Silo only asked for ProviderIds, so every Emby import
created no history rows, stamped progress at the Unix epoch, and could never
update progress on a re-import.

- Request the missing fields; page item lists and batch ID lookups.
- Import episode favorites; report season favorites, which Silo cannot hold.
- Mark every episode of a played multi-episode file watched.
- Keep importing when Emby show details cannot be read.
- Accept Emby accounts without a password.
- Report unreachable sources and invalid input with specific messages
  instead of a 500, and put the message in the v2 problem detail.
- Show safe per-cause summaries for run warnings and unmatched items, and
  count unmatched items by cause.
- Web: count skipped items once, ask for a new Emby Connect sign-in after a
  run consumes the session, and fix the Emby card copy.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(historyimport): keep v1 error decisions and fixed Emby warnings

The shared handler mapping also answers /api/v1, which is frozen. Restore its
decisions and map unreachable sources and invalid input in the v2 adapter,
which reads the cause the seam's APIError wraps.

Emby favorites and series-metadata warnings now store fixed text and log the
upstream error, since v1 returns stored warnings verbatim and Emby error
bodies can reach them.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…erver#1337)

* feat(admin): run full library metadata refreshes from the web

The library refresh button always queued a quick refresh, which skips
healthy matched items, so provider-side fixes never reached existing
libraries without API calls.

The button now opens a dialog offering the quick refresh or a full
refresh of every item. A new manual task, Refresh All Library Metadata,
runs a full refresh of each enabled library. It runs the library refresh
executor inline because admin jobs need a requesting account. A cluster
advisory lock allows one task run across servers, and the executor now
holds a per-library lock so a refresh job and the task never refresh
the same library at once.

The typed v2 request helper learns to send JSON bodies the contract
marks optional, which the refresh mode needs.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(adminjob): let a recovered library refresh wait for its lock

A library refresh job recovered after its worker stopped heartbeating
failed straight away on the per-library advisory lock, because the earlier
attempt can still hold it: its worker needs up to a heartbeat interval to
notice the new claim, and a server that disappears without closing its
connection keeps the lock until PostgreSQL drops the session. Recovery
used to rerun the refresh to completion.

A job on its second or later claim now retries the lock every five
seconds, reporting that it is waiting, until the holder lets go, the job
is cancelled, or the six-hour job limit runs out. First claims and the
full refresh task still fail fast on a held lock.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…Silo-Server#1331)

* fix(historyimport): import Jellyfin favorites and support Jellyfin 12

Jellyfin imports skipped favorites, and several requests misbehaved
against Jellyfin 12:

- A rejected password was retried under the /emby prefix, which
  Jellyfin 12 removed, so a credentials error surfaced as a 404. On
  10.11 the retry counted as a second failed login.
- Items came from /Users/{userId}/Items, which Jellyfin 12.1 no longer
  lists in its OpenAPI document.
- A source address ending in "/" produced //Items requests that 404.
- Items without a LastPlayedDate were stamped with the import time, so a
  re-import could overwrite newer Silo progress.

Fetch IsFavorite movies, series, and episodes as favorite-only records
and keep the IsFavorite flag on watched items. Sign in once at the
entered address, read items through /Items?UserId=, trim trailing
slashes, leave UpdatedAt zero without a last-played date, and request
only ProviderIds.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(historyimport): keep favorite series lookup failures non-fatal

A favorite episode added its series ID to the series lookup that the
watch history depends on, so a failed lookup discarded the whole run
even though favorites are otherwise best-effort. Look up series needed
only by favorites separately and report a failure as a run warning.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(historyimport): store fixed text for Jellyfin favorite warnings

Run warnings now hold fixed diagnostics that PublicWarning maps to
monitor text (Silo-Server#1329), because v1 returns them verbatim and upstream
errors can carry response bodies. Log the Jellyfin favorites and
favorite-series lookup errors and store fixed warnings with their own
public summaries, as Emby does.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(historyimport): say Jellyfin favorites may be incomplete after a failed query

Watched items that Jellyfin marks as favorites still import as
favorites when the favorites query fails, so the run summary must not
claim none were imported.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Requiring an open issue before a pull request cost more than it bought. It
did not prevent duplicate work — agents would find an open issue and race
the author to a PR — and it pushed contributors into filing throwaway issues
that only restated what the PR body already had to say.

An open issue is no longer a precondition. The pull request has to state the
problem itself: what breaks or is missing, who it affects, and why the change
is the right answer. `Related issue:` now links an issue when one covers the
work and reads `N/A` when none does, instead of needing a "narrow fix"
justification.

Opening an issue first for scope-changing work stays as advice, with the
tradeoff stated: it is the cheapest way to find out the work is already in
flight or out of scope, and a rejected pull request is the contributor's risk.

Adds the rule that actually addresses the collision: an open issue is not an
unclaimed one. Read its comments and linked pull requests before implementing
it, and agents should raise a likely collision rather than race the author.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…Silo-Server#1341)

* refactor(tasks): trim the Scheduled Tasks page to tasks admins manage

The admin Scheduled Tasks page listed about 38 tasks, most of them
internal queue workers and pollers that tick every 15 seconds to 5
minutes and need no administrator attention.

- Hide internal workers (matching, artwork caching and delivery checks,
  download preparation, request reconcile, collection syncs, autoscan
  poll, frequent log cleanups, provider-ID and watch-history repair).
  They still run and stay reachable by key; Server Activity still shows
  them while they run.
- List library-scoped tasks (audiobook, ebook, manga, podcast) only when
  a library of that kind exists.
- Run six retention sweeps as steps of one Database Maintenance task,
  daily at 05:00. A migration drops the old keys' saved schedules.
- Make the repair tools manual-only and group them under "On demand".
  Saved schedules for manual-only tasks are ignored at startup.
- Use plugin-supplied display names for plugin scheduled tasks.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(tasks): keep v1 task list unscoped and report maintenance steps

Review follow-ups for the Scheduled Tasks cleanup:

- Library scoping now applies only to the v2 default list through
  ListRelevantTasks. The frozen v1 list and the realtime tasks snapshot
  keep their previous signatures and behavior.
- v2 task history adds a steps array (key, name, status) for
  database_maintenance, and the task detail page names failed steps.
  Step errors are logged, since v2 does not return error text.
- Server Activity links each running task to its detail page, because
  hidden workers do not appear on the task list.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(tasks): run database maintenance on one server and keep canceled steps out

- Every API process fires the same 05:00 trigger, so Database
  Maintenance now takes a PostgreSQL advisory lock. One server runs the
  sweeps and the others finish without doing anything.
- A step interrupted by cancellation is left out of the step results
  instead of being recorded and logged as failed.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…ver#1313)

* feat(scanner): support flexible movie and episode filenames

* fix(scanner): correct flexible filename parsing regressions

Review of the flexible-naming change found cases where the new parser
wrote wrong episode numbers or types that the previous parser left unset.

- Read "Show - 05 - Part 2" from the first dash-delimited number instead of
  the last trailing number. A number completing the show folder title
  ("24 - 05", "9-1-1 - 01") yields to a later field-ending number, so
  titles and times ("24 - 12-00 AM") no longer replace the episode.
- Reject NxM coordinates from audio layouts (2.0x2, 5.1x264) and
  dimensions in year-numbered seasons (2048x1080).
- Keep "Show (2005) - S01E01" files in a show folder as series in mixed
  libraries.
- Pass library roots to IsMisplacedSeriesFile, the series refresh scope,
  and the scanner's fallback root and variant parsing so directories such
  as /mnt/s3 above a library are not read as season folders.
- Record a stale-identity series row without failing the running scan.
- Wake backed-off queue rows when a rescan changes a file's group key.
- Link episode-only numbers without a season only where absolute and
  per-season numbering agree.
- Do not cache a failed folder lookup as a disabled library.
- Keep the full undated title as a fallback search for loose titles ending
  in a number, and require a rescan when an anonymous sibling shares a
  stale root with a file-rooted show.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(scanner): address follow-up review comments

- Treat a leading NxM followed by a bracketed release year as a movie
  title (10x10 (2018)), not season and episode coordinates.
- Document that a number alone links an unseasoned episode only in the
  first regular season.
- Tighten two tests to check exact processing counts and the fallback
  release-year search.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(naming): use one movie-folder evidence rule for path classification

ResolvePathContext still used the older folder-evidence check, so a dated
show folder containing "Show (2005) - S01E01" files parsed as a movie while
root inference stored a series, and scans skipped the episode numbers. Use
the root-inference rule in both paths and remove the unused helper.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(naming): keep format words in dated titles and defer partial links

- An explicit "Title (Year)" boundary keeps title words that also name a
  format ("Opus", "UHD", "4K"). Resolution, source, and codec terms before
  the year still end the title.
- Treat a leading NxM followed by a dotted or spaced year (4x4.2019) as a
  movie title.
- Link an unseasoned number alone only to season 1, so a partial episode
  fetch cannot make a later season look like the first.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(metadata): share one provisional item across flat series episodes

Files directly in a series library root are their own observed roots, so
matching one show cannot relink its neighbors. Skeleton dedup keyed on the
observed root therefore created one provisional item per episode until the
show matched. Reuse the series item already linked to another file of the
same resolved scanner group, preferring a confirmed match. Ambiguous groups
keep separate items.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(naming): reject aspect ratios as episode coordinates

- An NxM inside a release tag ([16x9 1080p]) or a known aspect ratio
  directly before release details (Movie.16x9.1080p) is video metadata,
  not a season and episode.
- A resolution token no longer blocks a later dot- or space-separated
  episode marker (Show.1920x1080.E02). A dash-attached marker
  (1920x1080-E15) is still rejected.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(naming): let episode markers and title words outrank format tokens

- An aspect ratio followed by an explicit episode marker (Show.16x9.E02)
  yields to that marker and the season folder.
- A format word that appears before any number, such as "Opus" in a show
  title, no longer truncates the name before its episode number.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(naming): parse episode markers after audio layouts

- A rejected audio layout (AAC.2.0x2, DD5.1x264) no longer blocks a later
  dot- or space-separated episode marker, matching resolutions and aspect
  ratios.
- Treat a leading NxM followed by a dash-separated release year
  (10x10 - 2018) as a movie title.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(naming): handle decimal aspect ratios and dotted title years

- Reject decimal x-tokens with a one-digit count (2.35x1) and the 3x2
  ratio before an explicit episode marker, so Show.2.35x1.E02 keeps its
  season folder and episode.
- A bare release year also marks the title boundary when every format term
  before it is an ordinary title word (Mr.Hollands.Opus.1995).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(naming): reject small picture sizes as episode coordinates

An NxM token with two multi-digit sides (176x144) describes a picture
size. Reject it before the season range check so a later episode marker or
the movie identity wins.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(naming): keep dated titles and year seasons from misreading NxM

An NxM inside the episode title of a dated file ("Show - 2016-10-25 -
Salon Fail!; 2x4 Vandalism Victim") replaced the Season 21 folder and the
air date with S02E04. An NxM now counts after an air date only when it
directly follows the date ("Show - 2024-10-01 - 2024x246").

The picture-size rule also rejected year-numbered seasons with a
three-digit day count (2024x246). A year-shaped season with at most 366
episodes now falls through to the year-season check.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(naming): keep bare aspect ratios from classifying movies as series

A trailing aspect-ratio tag in a mixed library (Movie.Name.16x9.mkv) was
read as S16E09 and sent the movie down the series path. A recognized
aspect ratio now counts as a coordinate only in a season folder or a
declared series library.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(metadata): converge flat-series provisional items across nodes

The dedup lock that lets episodes of a flat library-root show share one
provisional series is per process. Two nodes matching episodes of the
same show at once could each miss the other's link and mint a separate
local item keyed on its own file path. A resolved flat group now anchors
its provisional content ID on the group key, so concurrent creators
write the same item.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(metadata): link to an existing flat-group item instead of resetting it

With a shared provisional ID, a node that passed the group check before
another node linked its episode would Upsert the item and overwrite
every field, resetting an already matched series to an empty skeleton.
A flat group now links to the item when it already exists and creates
it only when it is absent. The test fake now reports a missing item
with catalog.ErrItemNotFound, as the real repository does.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(metadata): create flat-group provisional items atomically

Loading the group item before upserting it left a window: a node that
stalled between the two steps while another node matched the item would
still overwrite it with an empty skeleton. The item repository now has
InsertIfAbsent, which uses ON CONFLICT DO NOTHING and reports whether it
wrote the row. Flat-group skeletons use it; a creator that loses the
insert links to the existing item.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(metadata): keep shared skeletons when a creator's cleanup runs

Skeleton cleanup after a failed link or membership write used the
unconditional item delete. With a shared content_id, the node that won
the insert could fail to link after another node had linked and enriched
the same item, and its cleanup would delete that item.
ItemRepository.DeleteIfUnreferenced deletes one item only while nothing
references it, using the orphan sweep's safety conditions, and skeleton
cleanup now uses it.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(naming): keep numbers in show titles from becoming episodes

The trailing-number fallback read a number that completes the show
folder's title as the episode: "The 100/Season 01/The 100.mkv" became
S01E100. Its non-overlapping search also consumed the separator after
the title number, so "The 100 05" never considered 05. Trailing
candidates now resume after each number, and a number whose prefix
matches the show folder's title is skipped.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(naming): keep format-word titles and low-resolution sizes apart

An undated folder and file sharing a title such as "Mr Holland's Opus"
or "The UHD Journey" lost the format-like word and everything after it.
When the folder title matches and every format term in it can also be a
title word, the full title is kept.

A picture size with a two-digit height (128x96) was read as S128E96 and
hid a later E02 marker. A three-digit width with a multi-digit height is
now a picture size, and an episode marker after it keeps the season
folder.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* feat(storage): fence S3 clients and artwork stores during transitions

Co-Authored-By: Codex <gpt-5>

* feat(storage): add verified transitions and resumable recovery

Co-Authored-By: Codex <gpt-5>

* feat(api): expose storage transition status and cancellation

Co-Authored-By: Codex <gpt-5>

* feat(web): guide managed artwork storage transitions

Co-Authored-By: Codex <gpt-5>

* docs: explain managed artwork storage transitions

Co-Authored-By: Codex <gpt-5>

* fix(storage): address transition review findings

Prevent duplicate active transition jobs, expose capability discovery to clients, and correct recovery field documentation. Reject stale execution after settings commit and cover the behavior with focused and PostgreSQL-backed tests.

Co-Authored-By: Codex <gpt-5>

* fix(storage): harden transition failure handling

Keep cancellation recoverable when staged-state cleanup fails, distinguish request validation from infrastructure errors, and preserve unrelated settings edits during managed transitions. Bound synchronization tests and make AWS smoke configuration portable.

Co-Authored-By: Codex <gpt-5>

* test(api): refresh media route manifest

Record the additive storage-transition capability endpoint in the checked-in media route classification manifest.

Co-Authored-By: Codex <gpt-5>

* fix(storage): preserve transition state and private artifacts

* fix(storage): detect aliases between source namespaces

* fix(storage): move operational blobs with managed transitions

Since blob storage became backend-neutral, a local root holds subtitles,
diagnostic bundles, job artifacts, and avatars beside artwork. The transition
still treated operational data as private-S3-only, so it:

- never copied local diagnostics or job artifacts to private S3, and left
  their "local" bucket rows pointing at a bucket that does not exist;
- never copied private S3 diagnostics or artifacts into a local root, and left
  their rows on the old bucket where a later repoint could not find them;
- skipped subtitles on every local target;
- disabled diagnostic uploads on commit to local storage.

Model a transition as two moves, derived the way blobstore.Open selects
stores: the assets store and the operational store (private bucket, local
root, or none). The assets copy excludes operational namespaces; the
operational copy runs when that location changes and uses explicit prefixes
whenever a local root is involved. Restart recovery repoints rows between the
old and new operational bucket, treating a local root as "local".

Disabling S3 still moves everything to local disk. A local install keeps its
private bucket, including a legacy operational alias, unless the transition
changes it, which makes a private-only move on a local backend possible.
Fence each distinct store once: a local root is both stores, and fencing it
twice would never return.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(storage): lock and transition a local install's private bucket

A private bucket owns diagnostics, job artifacts, and avatars on either
backend. On a local install that has stored data, adding one moved those
reads to an empty bucket and stranded the files in the local root, and the
settings API allowed it because the lock only covered private keys on an S3
backend.

Lock the private endpoint, bucket, and key prefix on both backends. On a
local install, editing them now opens the managed transition as a
private-only change: the dialog shows the new private location (or local
disk when the bucket is removed), omits the artwork path, checks the old
bucket's reachability before a copy policy, and sends the private keys with
the request. Transition copy now says subtitles move in every direction and
no longer claims local mode cannot read diagnostics or artifacts.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* test(settings): give the layout form mock getPersistedValue

The storage page now reads the persisted private bucket on every render to
decide whether a local install's copy source needs a reachability check. The
real settings form provides getPersistedValue; the admin layout test's partial
mock did not, which only surfaced once the call stopped short-circuiting on
non-S3 sources.

Also name the private endpoint setting once in the transition service, which
its new location summary made a repeated literal.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(storage): keep avatars out of an artwork-only preserve preflight

When only the public location changes, the operational store stays where it
is and preserve_uploads copies no avatars. The preflight said it did.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(storage): address review findings in managed transitions

- Commit writes only the location keys the copy verified. Other storage
  settings keep their current value, so a credential rotated or a read
  endpoint changed while a transition ran is no longer reverted at commit.
- Start applies the settings API's per-key rules. An unknown backend or a
  relative local path can no longer be committed and stop the next boot.
- Start rejects a target whose normalized identities equal the active
  stores before creating a job, and builds the preflight from those same
  identities, so the preflight matches what the copy does.
  blobstore.LocalIdentity computes a local identity without creating the
  root.
- Clearing the private bucket clears the rest of the private location, so
  removing it no longer fails validation on a leftover endpoint.
- A local-to-local transition keeps saved public S3 settings and
  credentials instead of blanking them.
- Restart recovery repoints rows naming the "local" bucket under every
  policy. An S3 reader took it as a real bucket name, presigned downloads
  against it, and could never expire those rows.
- The final pass fences only the stores being copied from, so removing a
  local install's private bucket no longer blocks artwork writes until
  restart.
- The preflight no longer promises avatar copies when no operational
  location moves.

Tests cover each case, including a PostgreSQL test for the repoint that
refuses to run against a database already holding "local" rows, and a
test that blobstore.Open fences the shared local root.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(settings): keep private-storage edits from dead-ending

- The settings lock covers the private endpoint and key prefix only while
  a private bucket is configured, and compares the prefix after trimming
  slashes. A leftover prefix or an equivalent edit no longer needs a
  transition that could only fail as a no-op.
- The storage page compares a dirty location field with its stored value,
  so retyping a bucket or adding a trailing space is a plain save.
- The transition dialog's text follows the actual move: removing an S3
  install's private bucket says private data becomes unavailable, Start
  fresh when disabling S3 says nothing is copied, avatar copies are
  mentioned only when the private location moves, and a private bucket
  without an endpoint cannot be queued.
- The setup wizard makes locked location fields read-only instead of
  saving into a 409.
- docs/settings-api.md lists the new lock triggers.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(storage): close legacy, lock, checksum, and throughput gaps

Legacy aliases: Start now builds its target from canonical keys and drops
the s3.operational_* rows first. A request that cleared a canonical key was
refilled from its alias, and an auto backend whose public bucket existed
only as an alias resolved to local, so a private-only API request became a
move to local disk.

Lock before artwork: private writes never recorded a storage identity, so a
private bucket holding avatars, diagnostics, or job artifacts could be
repointed before the first artwork write. The private S3 client now records
storage.operational_identity on its first successful write, through a
write observer every private writer shares. The settings lock engages on
either row and, with only the private row, locks only the private keys.
A transition that moves the private location rewrites or clears the row.

Checksums: objects up to 8 MiB are copied with Put, which records the
silo-sha256 metadata S3 compares in Matches. Streamed copies had none, so
the image cache re-uploaded every migrated variant it touched.

Throughput: each page of 250 keys copies through eight workers, with
receipts, counters, and progress under one lock and the page flushed before
its cursor advances. Progress reports at most every two seconds instead of
writing the job row and publishing an event per object. The fenced pass
accepts an object this run already verified once its source digest still
matches, without downloading the target copy again.

The architecture, settings API, and admin wiki docs describe the managed
transitions and the new identity row; the wiki still described a manual
database procedure.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(storage): address review of the lock and identity changes

- The settings lock and the storage page compare locations the way stores
  name them: endpoint scheme and host and bucket names are
  case-insensitive, and key prefixes ignore slashes. The page also treats a
  leftover private endpoint or prefix with no bucket as a plain save. These
  edits previously dead-ended between a 409 on save and a transition that
  failed as a no-op.
- artwork_storage.locked again means the artwork location is recorded.
  /api/v2 adds private_locked for the private bucket alone, so a private
  bucket holding data no longer disables the artwork path in the settings
  page or the public fields in the setup wizard. The frozen /api/v1
  response is unchanged.
- Recording the private bucket's identity can no longer fail a write that
  already stored its object. A failed recording is logged and retried on
  the next write, with a context detached from the request.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* test(settings): type private_locked in the storage page's status mock

The production build typechecks tests with tsc -b; the mock's hand-written
status type did not know the new field.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(storage): tolerate scalar restart receipts during finalization

* fix(storage): reject partial orphan and probe deletions

* fix(storage): guard transition review on capability discovery

* docs(storage): describe managed transition policies accurately

* fix(storage): reject ambiguous source namespace copies

* fix(storage): route local path changes and preserve endpoint identity

* fix(storage): expire recovered transition job receipts

* fix(storage): preserve committed transition receipts during cancellation

* fix(apiv2): honor storage transition cancellation contract

* chore(contracts): refresh combined storage transition digest

* test(storage): wait for mutation fence contention

* fix(storage): bind private bucket identity at startup

* fix(storage): validate transition destinations before commit

* docs(storage): describe private startup lock and local path changes

* fix(storage): keep empty artwork identity on private move

* chore(contracts): refresh storage identity digest

* fix(storage): route automatic S3 switch through transition

* fix(storage): wait for lock status before editing settings

* fix(setup): wait for storage lock status

* fix(settings): avoid passing click event to save

* fix(storage): finalize recovery after transient stage read

* fix(storage): keep config watcher running with unreadable stage

* fix(storage): warn when sidecar artwork needs refresh

* fix(storage): scope transition validation to storage settings

* fix(storage): tolerate vanished objects during bulk copy

* fix(settings): exclude recovery state from admin settings ETag

* fix(storage): condition artwork reconcile certification on active identity

* fix(storage): show safe transition progress in admin settings

* fix(storage): expose safe transition progress in admin jobs

* fix(storage): reject reconcile from stale API nodes

* fix(storage): sanitize storage jobs in admin streams

* fix(storage): coordinate API nodes with storage transitions

* fix(storage): hide stale progress while transition resumes

* test(storage): use structured transition progress in admission checks

* fix(storage): preserve restart receipt for resumed claims

* fix(storage): reuse transition phase constants

* fix(storage): reuse artwork path setting key

* fix(admin): clean up storage transition lint findings

* fix(storage): reuse backend setting key in reconcile guard

* fix(storage): count missing S3 objects as deleted

* fix(storage): bind local root after private data copy

* fix(storage): explain private location mismatch recovery

* fix(storage): report when lock state is known

* fix(api): keep v1 status response unchanged

* fix(settings): guard storage transition review

* fix(setup): guard storage edits when lock status is unavailable

* fix(storage): rejoin node admission after a database restart

Any failed probe of the admission session stopped the process, so a
PostgreSQL restart, a failover, or a slow ping took down every API node
even with no transition running. A node holding only the shared lock now
rejoins on a new session with a non-blocking shared acquire. Its blob
writes pause while the database cannot be reached and resume once it
rejoins. It still stops if a transition took the exclusive lock while it
was out, or if it owned the transition whose session was lost.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(storage): hold committed transition claims until restart

After a commit, or a commit the process could not confirm, the process
keeps its source fences until it restarts. If a database outage outlasted
the stale window, recovery requeued the job, and the next claim either
failed a committed transition or ran it again against fences the process
still held. A late cancellation also left a committed job reported as
canceling. Such claims now record the restart receipt with cancellation
cleared and keep it alive, and each process requests its restart once.
Boot recovery completes the job after the restart.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(storage): check the recorded location before a node rejoins

A node cut off from PostgreSQL for another node's whole transition found
no exclusive holder once that owner restarted, rejoined, and kept writing
to the stores it opened before the commit. Under the new shared lock, a
rejoin now compares the recorded storage identities with this process's
stores and stops the node when a transition moved them.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(storage): wait for admission rejoin before failing a transition

The runner can claim a queued transition as soon as the database answers,
up to one admission probe before the node has rejoined. Execute now waits
a few seconds for the rejoin instead of failing the job.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(settings): compare storage edits with the saved location

After a committed backend switch and before the restart, the page
compared drafts with the running backend, so every save became a
transition that Start rejects. It now compares with the saved location.
With lock status unknown, a location edit also holds back the storage
credentials that belong with it.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(setup): explain storage locks with the settings page's name

The wizard said files were stored in a public bucket that Automatic
storage on local disk leaves empty, and pointed to a page named
Infrastructure that the admin settings call Storage & Database.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(storage): order node rejoin checks after in-flight commits

A node could rejoin in the moment between a transition owner losing its
admission session and its commit becoming visible, then keep writing to
the old stores. The owner now confirms its session inside the commit's
settings transaction, and the rejoin check reads the identity rows under
the same settings lock.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(settings): use saved storage only while a commit awaits restart

Comparing drafts with the saved backend broke an install whose saved
settings changed while unlocked and whose first write then recorded the
running store: edits saved directly into a 409. The page now uses the
saved location only while the latest transition awaits restart, and the
running one otherwise. The wizard's footnote and backend lock text name
the Storage & Database page.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Quick104 <31828688+Quick104@users.noreply.github.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Bring the Jellyfin-compatible API in line with Jellyfin 12.1 and fix two
problems found while validating it with Jellyfin Web.

- Jellyfin 12 API: /Items/{id}/Collections, audio and subtitle language
  filters and Filters2 facets, root recursion for filtered requests,
  episode parent posters, OriginalLanguage, /Persons name and parent
  filters, display-preference defaults, Jellyfin 12 HLS codec strings and
  the Dolby Vision variant, LocalizedLanguage/LocalizedOriginal.
- New installs report 12.1.0 and install Jellyfin Web 12.1; configured
  servers keep their versions.
- AudioLanguagePreference "OriginalLanguage" maps to x-silo-original
  (settings contract revision 9); SubtitleMode maps onto the profile's
  subtitle settings, and PlaybackInfo picks the default subtitle with
  Jellyfin's selection rules.
- Embedded text subtitles are extracted in one pass per file, and in the
  background when Jellyfin Web starts playback, so switching tracks no
  longer waits a minute or more per track on large remuxes.
- Jellyfin Web device IDs longer than 256 bytes are accepted.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* feat(playback): configure profile-wide seek intervals

* fix(player): track only accepted seek targets

* fix(player): discard seek targets after failed reanchors

---------

Co-authored-by: Quick104 <31828688+Quick104@users.noreply.github.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…ilo-Server#1342)

* fix(intromarkers): stop re-refining unchanged silence backfill files

A chapter silence refinement that finds no better intro end, or fails,
leaves the chapter:v1 marker unchanged, so intro_markers_detected_at never
moves and the backfill picks the same files ahead of everything else on
every run.

Record each such attempt in intro_silence_refinement_attempts with the file
identity, the refined range, and a hash of the silence settings. The backfill
skips a file while its attempt still matches: until the inputs change after a
clean no-improvement result, and until retry_after (12h doubling, capped at
7 days) after a failure. Files never attempted are picked first.

Fixes Silo-Server#1272

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(intromarkers): keep silence retry backoff for failures inside the window

Scheduled tasks run per process, so every API server runs the nightly
marker job over the same files. A failure recorded while an earlier failure
is still backing off is a duplicate run (or a forced episode analysis), not
a retry, and no longer escalates the failure count or moves retry_after.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(intromarkers): re-run silence refinement when chapters change

The refined intro segment comes from the file's chapters, and a re-probe
can rewrite them without changing the file identity or the stored marker.
Record a SHA-256 of the chapters loaded with the candidate and require it
to match mf.chapters before skipping a file.

The column lands in its own migration because a development deployment
already applied the table migration. Existing rows get an empty hash, so
those files are refined once more.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(intromarkers): scope silence refinement failures to the recording server

Every server runs the nightly marker job against its own ffmpeg and media
mounts. A failure recorded by one server (missing ffmpeg, lost mount) no
longer defers the file on other servers: the attempt records the host that
wrote it, the backfill honors retry_after only for its own failures, and
backoff escalates only on the same server. No-improvement results still
apply everywhere.

The unapplied follow-up migration now adds recorded_by next to
chapters_hash and is renamed to cover both.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(intromarkers): store silence attempt file IDs as bigint

media_files.id is bigint, so an int4 media_file_id would reject attempts
for files past the int32 range and leave them in the backfill loop. The
table migration now creates the column as bigint, and the follow-up
migration widens it on a database that already applied the integer
version.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…el (Silo-Server#1325)

* fix(playback): fall back to CPU decode and software under auto hw accel

With playback.hw_accel=auto, a GPU that can encode a source but cannot
decode it made FFmpeg exit before its first manifest, and playback failed
instead of trying a safer path.

An automatic video transcode that resolves to a hardware backend now tries
GPU decode and encode, then CPU decode with GPU encode, then software. It
moves on only when FFmpeg exits before its first manifest, never while a
process is still running, and briefly remembers the path that worked for
that file. Explicit acceleration settings, copy/remux, and tone-map
recipes keep their existing behavior.

The fallback covers native local and remote starts, transcode-node starts
(keeping replacement isolation), Jellyfin-compatible local and remote
starts, and reconstruction on the API server and nodes. The NVENC mixed
path uploads CPU-decoded frames to the configured CUDA device. Nodes
report software_video_decode so recipe cards record the executed path.

Fixes Silo-Server#918

Co-authored-by: MrDoudou <22004704+MrDoudou@users.noreply.github.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(playback): address lint findings and refresh contract digest fixture

Check session.Close errors in the node fallback tests, name the transport
startup "ready" outcome, split a double-advance assertion, and regenerate
get_system_info_ok.json for the updated v2 contract digest.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(playback): size auto remote start budget from the request

Native remote starts extended their deadline for the hw_accel=auto
fallback only when the node's stored capability report resolved to
hardware. A missing or stale report still let the node walk every
fallback path under the single-wait deadline, so the API could cancel a
fallback that was about to succeed. The node resolves auto against live
hardware, so size the budget from the dispatched request instead.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(playback): judge manifest identity from the file that was read

getManifest read the playlist by path and then stat'd the path again. If
FFmpeg's rename landed between the two calls, the inherited playlist's
bytes were paired with the new file's identity and accepted as current.
Read and stat the same open handle.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(transcodenode): let the node decide auto fallback readiness

Jellyfin-compat remote starts asked the node to wait for readiness only
when the node's stored capability report resolved to hardware. A missing
or stale report either kept the fallback from a node that now has a GPU,
or forced a software-only node to close a slow start at the readiness
deadline.

Add auto_fallback_ready to the start request: the node waits for the
first manifest only when its own live hw_accel=auto pipeline is enabled,
and otherwise starts unwaited. Jellyfin-compat sends it for auto video
transcodes; older nodes ignore it and keep the unwaited start.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: MrDoudou <22004704+MrDoudou@users.noreply.github.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…o-Server#1347)

* fix(watch-party): stop warning on routine room socket reconnects

The server ends every room socket after five minutes and the web client
reconnects in about half a second. The player showed "Reconnecting to
room" on each rotation and never cleared it, so both viewers saw a false
warning that stayed up for eight seconds.

Wait two seconds of lost connection before warning, and clear the
warning when the socket reconnects. Expire notices from player state
instead of hiding them inside the overlay, so a later identical notice
shows again rather than staying invisible.

Fixes Silo-Server#1346

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(watch-party): keep notices raised while the player is minimized

Notice expiry now lives in player state, so a notice raised while the
player is minimized would expire before anyone saw it, including the
one-time "Lower quality" offer. Start the expiry only once the player
is shown again.

Advance timers in the displacement test so it still proves the
replacement path suppresses the delayed reconnect warning.

Refs Silo-Server#1346

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(watch-party): keep newer notices and hold the warning for an outage

The delayed reconnect warning replaced whatever notice was showing when
it fired, so an admin message or server-restart notice that arrived
during the two-second delay was lost. Only replace the notice that was
showing when the outage began, or an empty one.

The warning also expired after eight seconds while the room was still
disconnected, leaving controls unavailable with no indication. Keep it
until the socket reconnects, the room closes, or the viewer is
displaced.

Refs Silo-Server#1346

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(watch-party): restore the reconnect warning after a newer notice

A notice raised during the reconnect delay keeps its place, but once it
expired the room could still be disconnected with no warning on screen.
Remember that the outage has outlasted the delay, and let an expiring
notice hand back to the reconnect warning until the socket reconnects.

Refs Silo-Server#1346

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(watch-party): preserve notice expiry through socket rotation

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-authored-by: Quick <31828688+Quick104@users.noreply.github.com>
…ilo-Server#1348)

* fix(subtitles): keep SRT {\anN} placement in the WebVTT conversion

The SRT to WebVTT conversion copied cue text verbatim, so ASS override
blocks such as {\an8} reached WebVTT players as literal text and the cue
fell back to the bottom of the frame. The cue's \an alignment now becomes
WebVTT cue settings on the timing line, and override blocks are removed
from the text. The Jellyfin-compatible windowed .srt response writes the
tag back.

The converter also no longer merges cues when an SRT omits the blank
line between them, keeps a stray CR out of the timing line, and
recognizes timing lines that use a period before the milliseconds.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(subtitles): split CR-only SRT line endings

A classic-Mac SRT uses a lone CR between lines, so the converter saw one
line, recognized no cue, and left {\anN} blocks in the text. Normalize
\r\r\n, then \r\n, then a lone \r to LF.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(subtitles): require real SRT timestamps on timing lines

The converter ends a cue at the next timing line, and the timing-line
check only looked for an arrow after a colon and a separator. Cue text
such as "Meet at 10:30. --> go now" split the cue, and Jellyfin windowing
then failed to parse it with a 500. Both sides must now be SRT
timestamps.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…o-Server#1405)

* fix(jellycompat): serve Jellyfin Web preference discovery routes

* fix(jellycompat): preserve preference IDs during route normalization
…Rip (Silo-Server#1358)

* feat(playback): serve original SRT sidecars to clients that parse SubRip

Clients that send subrip_sidecar_v1 receive external and downloaded SRT
tracks as the original .srt file (with original=1) instead of the WebVTT
conversion, so SRT features WebVTT cannot express survive. A bare .srt
request keeps its historical WebVTT response for the frozen v1 route.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(playback): keep subrip_sidecar_v1 off the frozen /api/v1 surface

The playback and stream handlers are shared with /api/v1, so the new
feature was advertised, negotiated and served there too. Requests through
/api/v2 now carry a context mark: only they advertise and negotiate
subrip_sidecar_v1 and receive original SRT bytes for .srt?original=1. A v1
start drops the feature, which keeps it out of the attempt's replans, and
the v1 route keeps answering .srt with WebVTT.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(playback): keep /api/v2 SRT attempts from continuing on /api/v1

An attempt that negotiated subrip_sidecar_v1 on /api/v2 could still be
replayed or replanned through /api/v1, which published its
.srt?original=1 plan on a route that serves WebVTT for those URLs. v1 now
refuses to continue such an attempt with the existing 409
playback_attempt_reused error.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(playback): label original SRT as UTF-8 only when it is

Original SRT bytes were always served as charset=utf-8, which mislabels a
legacy-encoded file. The charset is now declared only for valid UTF-8, and
HEAD, which does not read the bytes, declares none. The protocol doc's
feature-list note names plan_source_duration_v1 instead of "the last one".

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(playback): refuse a v2 retry of a start that /api/v1 negotiated

A start first sent to /api/v1 with subrip_sidecar_v1 stores WebVTT after v1
drops the feature, but its request digest still matches the same body
retried through /api/v2. That retry replayed the WebVTT plan to a client the
protocol promises original SRT. It now gets the existing 409
playback_attempt_reused, the same binding the v1 direction already has.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(playback): decide the SRT representation from what the attempt published

A server that predates subrip_sidecar_v1 stored the token verbatim beside
WebVTT URLs. Trusting the stored token let the new cross-surface guard
refuse such a legacy v1 attempt mid-playback, and let every replan other
than a seek reanchor switch its SRT URLs to .srt?original=1. Both now read
the representation the attempt's current plan actually published, and
every replan keeps it.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* docs(playback): every replan keeps the published SRT representation

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(playback): keep the negotiated SRT surface before any SRT is published

With no SRT track in the current plan, the published representation says
nothing, so a v1 replan could continue a v2-negotiated attempt and publish
.srt?original=1 for a track that appears later. The stored token now
decides in that case: this server drops it from every /api/v1 start, so it
marks only attempts negotiated through /api/v2. The protocol doc also notes
that a convert artifact stays WebVTT.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(playback): record the /api/v2 offer of subrip_sidecar_v1 on the attempt

Falling back to the stored token when no SRT was published misclassified a
legacy /api/v1 attempt whose old server stored the token verbatim, so its
next v1 retry or replan got a 409. /api/v2 start and replan decisions now
advertise the v2 feature list, which the attempt persists as its
StartResponse. Before any SRT is published, an attempt counts as
negotiated only when the client sent the token and this server offered it,
and replans and realtime events use the same rule.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(playback): replay a terminal start before the surface check

A terminal start is persisted without the /api/v2 feature marker, so an
identical opted-in retry through /api/v2 failed the surface check with a
409 instead of replaying the terminal idempotently. A terminal publishes
no subtitle URLs, so it now replays before the check runs.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(playback): omit a realtime subtitle track when its representation is unknown

A failed attempt lookup made the subtitle-ready notifier fall back to the
default features, which could publish a .vtt URL to a session that
negotiated original SRT. The resolver now returns the error, and the event
omits the track so the client refetches its plan. A session with no v3
attempt still uses the defaults.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…ilo-Server#1350)

* fix(watchsync): classify Trakt and Simkl rate limits and pace writes

Trakt and Simkl answered every non-2xx the same way, so a 429 never
became a RateLimitedError and the sync service never deferred the
account. Both providers now map 429 (and Simkl's documented
rate_limit/RATE_LIMIT bodies) to RateLimitedError with the provider's
Retry-After, retry short waits in place, and pace authenticated writes
to one per second per access token as the providers document.

Retry-After parsing and the per-credential write limiter move into
internal/watchsync so MDBList, Trakt and Simkl share them.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(apiv2): answer provider rate limits with 429 and Retry-After

A provider 429 during a watch-provider call, such as Trakt's slow_down
answer to a device-code poll, now reaches clients as a rate-limited
problem carrying the provider's wait instead of an internal error.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(watchsync): defer limiter refusals and Trakt device-code 429s

A local limiter refuses at once when the next slot lies past the
context deadline. That was an ordinary error, so watched exports that
were never sent were marked failed. The refusal is now a
RateLimitedError, which leaves the work pending for a later run; a
cancelled context still returns its own error. MDBList's limiter gets
the same treatment.

Starting Trakt device authorization now reports a 429 as a rate limit
with the provider's Retry-After, like polling and token refresh.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…er#1351)

* fix(watchsync): classify Trakt and Simkl rate limits and pace writes

Trakt and Simkl answered every non-2xx the same way, so a 429 never
became a RateLimitedError and the sync service never deferred the
account. Both providers now map 429 (and Simkl's documented
rate_limit/RATE_LIMIT bodies) to RateLimitedError with the provider's
Retry-After, retry short waits in place, and pace authenticated writes
to one per second per access token as the providers document.

Retry-After parsing and the per-credential write limiter move into
internal/watchsync so MDBList, Trakt and Simkl share them.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(watchsync): keep the MDBList API key out of error text

The MDBList watch-sync provider puts the API key in the request query
string. Transport errors wrap *url.Error, whose message includes that
URL, and those errors reach the connection's last_error, sync run
errors shown in the web UI, and logs. Build the keyed URL only inside
the single request helper, sanitize transport and request-build errors
with logredact.SanitizeURLError, and mask the key (raw or escaped) in
error-body excerpts. Cancellation and timeout classification is kept.

Same bug class as Silo-Server#692, which covers the separate internal/mdblist
discovery client.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(apiv2): answer provider rate limits with 429 and Retry-After

A provider 429 during a watch-provider call, such as Trakt's slow_down
answer to a device-code poll, now reaches clients as a rate-limited
problem carrying the provider's wait instead of an internal error.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(watchsync): defer limiter refusals and Trakt device-code 429s

A local limiter refuses at once when the next slot lies past the
context deadline. That was an ordinary error, so watched exports that
were never sent were marked failed. The refusal is now a
RateLimitedError, which leaves the work pending for a later run; a
cancelled context still returns its own error. MDBList's limiter gets
the same treatment.

Starting Trakt device authorization now reports a 429 as a rate limit
with the provider's Retry-After, like polling and token refresh.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(watchsync): keep error classification when masking the MDBList key

When a transport error quoted the keyed URL in its own message, the
masking fallback replaced it with a plain error, so errors.Is and
errors.As no longer matched cancellations, deadlines, or net.Error
timeouts. The masked error now answers Is and As from the original,
without an Unwrap that would expose its message.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…rver#1363)

* fix(logging): keep Redis off the request path for log writes

With Redis configured, every captured slog record and every request
through the activity middleware ran a synchronous RPUSH on the caller's
goroutine. When the push failed, the writer logged a warning through the
slog default, which is the operational log handler, so the warning was
pushed through the same writer. While Redis refused connections one
slog.Info never returned and every request that logged stalled.

Every node now uses the path single-node installs already had: operational
and activity entries go into a bounded in-memory buffer
(logstream.Buffer), and a consumer on the same node inserts them into
Postgres in batches and publishes each row to the admin live tail.
Writers never block or log. A full buffer or a failed batch insert is
counted in silo_log_writer_dropped_total{stream,reason}.

The opslog consumer's own warnings go to stderr and OTLP only, so a
failing batch cannot queue entries about itself. Cross-node publishes for
a batch share a two-second deadline, so an unreachable Redis cannot stall
the consumer row by row. The consumers outlive appCtx and flush from
deferred calls in main, so records logged during graceful shutdown are
persisted. The Redis writers, the Redis drain loops and their dedicated
Redis clients are removed.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(logging): retry failed log inserts and publish to other nodes off the drain

Review of the first revision found that a Redis outage still throttled
Postgres persistence and that a Postgres blip dropped whole batches.

The consumer published each inserted row to other nodes inline, under a
two-second budget per batch. With Redis refusing connections or
unreachable, one consumer persisted only 25 to 35 entries per second, so
any busy node overflowed its buffer. logstream.Hub now hands rows to local
subscribers inline and queues them (1,024 rows) for its own goroutine,
which publishes to the event bus with a two-second deadline per publish.
The consumer never waits on Redis. Rows that miss other nodes' live tails
are counted in silo_log_tail_publish_dropped_total{stream,reason}, and the
hub logs only when publishing starts failing and when it recovers. main
closes the hub after the log consumers flush and before the event bus
closes.

A failed batch insert was dropped on the first attempt. logstream.Drain
now retries connection and availability errors with backoff from one to
ten seconds while the node runs, so the buffer holds entries through a
Postgres restart or failover. A batch the server rejects is dropped at
once, a stopping node makes one attempt per batch, and every attempt has
a ten-second deadline.

The metric help and the observability doc no longer claim that stderr
and OTLP receive every record: audit entries have no other copy. The doc
also covers the lost log.Fatalf record on exit and the leftover Redis
list keys from earlier versions.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(logging): keep log rows beside a value Postgres rejects

Without the Redis JSON round trip, raw strings reached the multi-row
INSERT, so one request with a 0xFF byte in its path or User-Agent made
Postgres reject the whole batch and dropped up to 100 other audit
entries. Both consumers now replace invalid UTF-8 and NUL bytes in text
fields, and \u0000 escapes in attrs JSON, with U+FFFD. When Postgres
still rejects a value (SQLSTATE class 22, such as a client address that
is not an inet), the drain inserts that batch one row at a time so only
the rejected rows are dropped and counted.

opslog.Handler now encodes non-scalar attribute values when the record
is logged. The consumer encoded them seconds later on its own goroutine,
so a caller that changed a logged map could race that encode.

A write refused as read-only (25006) after a switchover is now retried.
The request pipeline test serves 100 requests.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…lo-Server#1400)

TestControlSocketReconnectResumesOnlySameOwnerAndInstallation flakes in
CI with "second connection did not receive the command". The client's
dial returns once the upgrade response arrives, before the server
registers the new connection with the realtime hub. The hello helper
waits for HasRealtimeConnection, which the first connection already
set, so it returns at once and the command can still be routed to the
first connection's lane.

Wait until the handler holds a new lane for the session before sending.
With a 100 ms sleep added before Register to widen the window, the old
test failed 5 of 5 runs and this one passes 5 of 5.

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…Silo-Server#1364)

Every v1 and v2 progress sample called SyncNow, which reconciles the whole
node's live-session snapshot: one transaction of N+5 statements for N local
sessions, then a sessions.replaced push to every admin socket and a
playback-sessions invalidation on every replica. With the web and Android
clients reporting progress every 10 seconds, the shared admin view was
rewritten and its caches dropped once per heartbeat per stream.

Position does not need that. The periodic reconcile tick already carries
it, and jellycompat progress never triggered a sync. A progress sample now
syncs only when it flips the pause state, so the admin view still shows
pause and resume at once. Start, replan, stop and abort keep syncing from
their own call sites.

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…Silo-Server#1365)

* perf(playback): refresh the taste profile only when a play changes it

Every successful progress heartbeat on the native v1/v2 playback path and
the Jellyfin progress report marked the viewer's taste profile stale and
queued a full profile rebuild. A client reporting every 10 seconds
produced 360 stale-mark UPDATEs and up to 360 rebuild requests per
viewer-hour. Mid-play, a heartbeat can at most move the item between
watch-weight bands, and the stop that ends the play refreshes the profile
with the final position anyway.

Progress writes now report whether they flipped the item to watched
(userstore.UpdateProgressReportingCompletion reads the prior row only for
samples past the watched threshold). Native playback refreshes on that
crossing and, as before, on stop; the Jellyfin report refreshes on the
crossing and on the Stopped report. Ratings, favorites, watchlist and
watched-state changes keep their existing triggers, and expired sessions
still refresh through the stop finalizer.

MarkProfileStale now writes only when the profile is not already pending
(stale_at IS NULL OR stale_at <= updated_at), so a repeat mark before the
sweep consumes the first one is a no-op instead of a row rewrite.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(jellycompat): refresh the taste profile on a Stopped report without a position

The Stopped-report refresh sat inside the progress-persist block, which
runs only when the report carries a positive PositionTicks. Some clients
send Stopped without a position (or with 0), and Jellyfin teardown stops
the native session without running its stop finalizer, so such a play
got no taste-profile refresh at all once heartbeats stopped refreshing.

The Stopped report now refreshes whether or not it writes progress; the
completion crossing still refreshes from inside the persist block, and a
report that does both refreshes once.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…hing (Silo-Server#1366)

* perf(progress): index Continue Watching's resume-point predicate

The in-progress listings behind Continue Watching, Next Up, Jellyfin
Resume, and v2 GET /progress select rows with position_seconds > 0.
Completed rows hold position 0 since the reset_completed_progress_position
migration, and a rewatch keeps completed = TRUE with a live position, so
that is the right predicate. The partial index built for these lists is
still WHERE completed = FALSE, which the predicate does not imply. The
planner cannot use it, so every page read all of the profile's progress
rows and sorted the survivors.

Add idx_uwp_profile_resumable on (user_id, profile_id, updated_at DESC)
WHERE position_seconds > 0, built concurrently, and drop
idx_uwp_profile_in_progress. The only queries whose filter implies
completed = FALSE either look up a row by primary key or also require
position_seconds > 0, which the new index serves with fewer rows.

On a profile with 19,700 progress rows (600 resumable), a Continue
Watching page went from 19,700 rows read, 1,759 shared buffers, and a
sort to 100 rows read and 202 buffers via an ordered index scan.

A DB-backed test EXPLAINs the SQL each in-progress listing sends and
fails if a page reads anything but the profile's resume points.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* perf(progress): re-measure the resumable index and guard its rollback

Re-ran the before and after measurement on a fresh synthetic profile
whose rows are scattered through the heap the way interleaved writes
leave them, and replaced the migration comment's figures with those: a
Continue Watching page goes from reading 20,752 progress rows and 1,850
shared buffers to 102 rows and 310 buffers.

The comment now also lists the episode catalog's in_progress rule and
Next Up's resumable list among the readers, says that only the ordered
Continue Watching page stops at its LIMIT, and states why nothing can
use idx_uwp_profile_in_progress: every WHERE clause that implies
completed = FALSE is a primary-key lookup or an ON CONFLICT guard.

Down gets the same invalid-index guard as Up. Without it, an
interrupted concurrent rebuild of idx_uwp_profile_in_progress survives
a retry as an invalid index that IF NOT EXISTS skips.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* perf(progress): correct why the in-progress index can go

The migration comment said no query could use idx_uwp_profile_in_progress.
The Jellyfin IsResumable browse filter does: when it plans as a join over
the profile's rows, it reads the profile's unfinished rows through that
index. The filter also requires position_seconds > 0, so after the
migration it reads the new resumable index instead. On the benchmark data
that is 509 rows instead of 709 and 2.0 ms instead of 8.1 ms for a movie
browse page. Say so in the comment, and name the Audiobookshelf
COALESCE(completed, FALSE) predicate, which the planner cannot match to
the old index.

The guard test's plan walker only took index names from nodes that carry
the relation name. A Bitmap Heap Scan carries only the relation, and its
Bitmap Index Scan child carries only the index, so a bitmap read of the
new index counted as no index and would fail the guard. Collect names
from the bitmap children too, and cover it with a plan-JSON unit test.

Refresh the Continue Watching figures in the comment from a fresh
measurement on the same synthetic profile shape.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…ilo-Server#1367)

* fix(downloads): stop series monitors re-registering deleted episodes

Deleting a managed episode hard-deletes its downloads row, and monitor sync
only skipped episodes whose row still existed. The next sync registered the
deleted episode again under a new ID. iOS delete_watched retention deletes
watched episodes, so each one came back on a later monitoring run; capped
monitors spent their storage budget on episodes the user had deleted; and
manual deletes of monitored episodes did not stick.

Deleting an episode of a series the device monitors now records a
(monitor, episode) exclusion in the same statement. Sync reads the device's
existing entries and then its exclusions inside the monitor lock, before the
storage cap is applied, so a deleted episode neither returns nor consumes the
budget. An explicit episode, season or series download clears the exclusion,
and deleting the monitor drops its exclusions by cascade.

delete_watched monitors also skip episodes the profile has completed, using
the per-user progress store's batch lookup, so the device never fetches a
file its retention pass is about to delete.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(downloads): forget monitor deletions when the monitor is created again

A create for a series the device already monitors returns the existing
monitor. Clients patch a monitor they know, so a create that finds one means
the client lost track of it. The iOS app does this after an upgrade: it
deletes the downloads an earlier version registered, which now records
exclusions under that version's monitor, and hides the monitor. When the user
monitors the series again, the create answered with the old monitor and its
exclusions kept the back catalog from registering.

The native create and the bridge re-monitor now drop the monitor's
exclusions, so re-monitoring a series starts from a clean slate, as a new
monitor does.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* docs(downloads): correct the watched-progress lookup chunk comment

The bundled SQLite allows 32766 bound parameters, not 999. Chunking stays, as
it does for the catalog's playable targets, because the SQLite user store
binds one parameter per episode and a bridge sync passes a whole series.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(downloads): keep managed deletes working while a monitor is deleted

A managed delete of a monitored episode joined the monitor from its
snapshot and inserted the exclusion. If the monitor was deleted while the
insert's foreign-key check waited on the monitor lock, the whole
statement failed with SQLSTATE 23503 and the download row stayed. The
statement now locks the monitor FOR KEY SHARE: a monitor deleted
meanwhile yields no row, so the delete records nothing and succeeds. It
still deletes the downloads row first, the order a device delete's
cascade takes.

Registration now reads the device's existing entries and the monitor's
exclusions in one statement. The delete commits the row removal and the
exclusion together, so one snapshot sees one or the other, and a monitor
held at its storage cap no longer pays an extra statement on every sync.

The delete_watched filter fails open: if the progress lookup fails, sync
logs a warning and registers without it, as the storage gate does. The
retention pass deletes any finished episode registered that way, and the
exclusion keeps it from coming back.

The subscription test helper defines the exclusions table inline instead
of reading the migration file, and the migration comment now says that
re-creating a monitor clears its exclusions.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* docs(apiv2): say that re-creating a monitor clears its recorded deletions

createDownloadSubscription returns an existing monitor with its options
unchanged, but since the monitor exclusions landed it also forgets the
episodes the device deleted under that monitor. The published summary
still said the monitor is returned "without changing it". Update it and
regenerate the OpenAPI document, the web contract types and the fixture
that carries the contract digest.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* test(downloads): poll for the monitor lock wait on a ticker

waitForLockWait polled pg_stat_activity back to back, so a lock wait that
never appeared would query the shared test database nonstop for the full
10 s timeout. Check every 5 ms instead and stop on the context deadline.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* docs(apiv2): shorten the createDownloadSubscription summary

Generated clients and docs render an operation summary as a one-line title,
and the create summary had grown to several clauses. Keep the summary to what
the operation does and move the existing-monitor behavior and the resend
guidance into the operation description. Regenerated the OpenAPI document,
web types and the system-info fixture's contract digest.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* docs(downloads): say when a managed delete can deadlock with a device delete

The DeleteManaged comment said a device delete's cascade normally takes the
same row order. Postgres fires the cascade triggers in name order, and their
names end in OIDs compared as text, so the order depends on the database, not
on whether it was migrated in order. Say so, and that the loser gets a
retryable deadlock error.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(downloads): keep a delete's exclusion when a re-download clears it late

An explicit download of a deleted episode registers its row, then clears
the monitor's exclusion for the episode in a separate statement. A delete
of the new row that committed in between left the device with neither the
row nor the exclusion, so the next monitor sync registered the episode
again even though the user deleted it last.

The clear now forgets an episode only while the device still holds its
row, checked in the same statement. It locks that row FOR KEY SHARE SKIP
LOCKED: a delete that has run its statement but not committed still holds
the row, and a plain read would see the row and erase that delete's
exclusion.

The clear also skips an exclusion another transaction holds, so it never
waits on a row lock. It is best-effort cleanup after the download
succeeded, and waiting could hold the download behind a delete that is
itself waiting on the monitor lock behind a sync. A skip keeps the
exclusion, the direction the clear already tolerates.

DeleteManaged's statement moves into a constant so the test can run it in
an open transaction and hold a delete between its statement and commit.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…eads (Silo-Server#1371)

* perf(plugins): stop re-hashing a verified plugin binary on manifest reads

Every manifest read went through ArchiveCache.Ensure, which reads and
SHA-256 hashes the whole plugin binary under the cache's global mutex.
The admin plugin list reads the manifest of every installation on each
load, and a plugin asset request reads it twice, so each navigation
hashed tens of megabytes serially per installed plugin.

Add ArchiveCache.Manifest for callers that only read the manifest or
serve packaged assets. It skips the hash while the binary is the same
file, with the same size and mtime, that this process last hashed
against the checksum the manifest still names. Any change, a new
release path, or a failed check sends the read through the full check
and repair. Launch paths (Start, the connection test, preload) keep
calling Ensure, which still hashes the binary on every call, so the
tamper check in front of exec is unchanged.

The hash now streams the file instead of loading it whole, which drops
a launch check's allocation from the binary's size to about 37 KB.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* test(plugins): cover the file identity check on manifest reads

Review follow-up for the manifest read change.

The archive cache test now renames a different binary with the
verified size and mtime over the installed one and expects Manifest to
hash and repair it. Without the os.SameFile comparison the test fails;
before, nothing covered it.

Rename the service helper installedManifest to loadForManifestRead so
it no longer shares a name with a local variable in doEnsureClient, and
rename the benchmark case launch to ensure, since it measures
ArchiveCache.Ensure and starts no process.

Correction to the previous commit message: startup preload checks and
rehydrates files without executing them. Launches go through start or
the connection test, and both call Ensure.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…er event (Silo-Server#1372)

* perf(plugins): index event subscribers instead of reading the store per event

The plugin event dispatcher listed enabled installations and then read
each installation's capabilities for every event it saw, on every
replica: 1+N store reads per event even when no plugin subscribes to
that event. Progress, playback, and admin events all pass through it, so
busy servers paid hundreds of plugin-table reads per second per replica
for events that no plugin receives.

The dispatcher now keeps an in-memory index from event name to
subscribers, built from the store on first use. The index is dropped by
a lifecycle hook that SetEventDispatcher registers, so local install,
enable, disable, upgrade, and uninstall take effect on the next event.
Other replicas drop theirs when cache.EventPluginsChanged arrives on the
admin channel, and an index older than DefaultLifecyclePollInterval is
rebuilt so a missed publish costs at most one interval. A generation
guard, the same one the installation cache uses, keeps a rebuild that
races an invalidation from being stored. An index missing an
installation whose capabilities could not be read serves that event but
is not kept.

Delivery semantics are unchanged: the first matching event_consumer.v1
capability per installation receives the event, builtin installations
are skipped, and TargetPluginID narrows delivery to one subscriber.
Payloads are decoded only when an event has a subscriber.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* test(plugins): measure dispatcher reads with the real user-state envelope

The hub measurement in TestDispatcherStoreReadsPerEvent used an invented
progress_updated event. Use user_state.changed with the payload and scope
fields a progress sync publishes, so the measured envelope matches what
reaches the dispatcher in production. Also tighten the index field comment.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(plugins): drop the dispatcher index before other lifecycle hooks

SetEventDispatcher appended its invalidation hook after the
plugins_changed publish hook. When this replica's own publish came back
before that hook ran, the dispatcher rebuilt its index on the echo and
the hook then dropped it, so one local change cost two rebuilds. The
hook now runs first: a local change with an echoing bus costs one
rebuild (5 store reads with the test fixture, down from 10), and events
dispatched while resident reconcile and provider reloads run already
see the change.

The tests now assert store reads for a remote plugins_changed and for a
local change with and without an event bus. The racing-invalidation
test is renamed to say what it checks, and the startup
OnLifecycleChange comment in main.go states why the call remains.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…ilo-Server#1373)

* perf(catalog): read series and season user_data from the SQL rollup

The native series detail, season detail, and season list requests loaded
every available episode row of the parent and then read the viewer's
progress and completed history in 500-ID chunks, only to count watched and
in-progress episodes. On a 1,200-episode series that is 1,200 episode rows
plus six progress and history statements per request.

The Postgres user store already computes the same counts in one aggregate
(SeriesEpisodeWatchCounts), and jellycompat uses it for series. Add a
season-number variant for the season list and a season-row variant for the
season detail, sharing the existing query, and use them in the native
paths when the viewer's store implements the rollup. Stores without it
(the SQLite user store) and rollup query errors keep the episode fold.
The season list also takes each season's episode count from the rollup,
so it no longer lists episodes at all.

Measured with a pgx tracer on a 1,200-episode fixture: series detail goes
from 19 statements / 2,739 rows to 13 / 1,365, season detail from 13 / 334
to 11 / 212, and the season list from 13 / 2,749 to 7 / 1,387. A
DB-backed test proves value-for-value parity with the fold across watched,
partially watched, history-only, hidden, unwatched, unavailable, and
other-profile episodes; a unit test pins the rollup path in CI.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* perf(catalog): roll up season-by-number user_data in SQL

GET /catalog/series/{id}/seasons/{num} still folded per-episode progress
after the series, season detail and season list paths moved to the SQL
rollup. When the season's episodes are linked to its row, read user_data
from SeasonEpisodeWatchCounts; the by-number fallback keeps the fold.
The parity test now covers this request as well.

A canceled request no longer logs a rollup warning, and the season list
returns the error instead of retrying with the fold.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…#1374)

* ci: run the DB-backed query-budget pins against Postgres

make test-go runs without SILO_TEST_DATABASE_URL, so every DB-backed
test skips and go test reports the skip as a pass. The statement-count
and query-plan pins that guard item detail, listings, metadata refresh,
home sections, settings resolution and book scans have never run in CI,
so a change that turns a pinned budget into an N+1 merges green.

Add a Go DB pins job that migrates a fresh pgvector/pgvector:pg18
database with cmd/silo --migrate-only and runs the tests listed in
scripts/ci/db-pins.txt through make test-db-pins. The runner in
scripts/ci/dbpins reads go test -json and fails when a listed test is
missing, skipped or failing, and on any skipped subtest: without a
database, two of the pinned tests pass at the top level because only
their subtests skip.

The job is separate from go-test so it adds nothing to that job's time,
and because the full schema needs pgvector, which go-test's image lacks.
CONTRIBUTING.md lists the new target in the pre-PR gate.

Refs Silo-Server#599.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* ci: scope the DB pins in CONTRIBUTING and pin card play targets

CONTRIBUTING.md listed make test-db-pins in the mandatory pre-PR gate,
which would ask every contributor, even on a web-only change, for a
migrated pgvector database. Move it out of the gate: CI runs it for
every pull request, and contributors run it when they change database or
query code or add a pin. Say how to create and migrate the database with
the Compose defaults.

Add TestPlayableTargetResolverProfileStateAvailabilityAndAccess to the
list: leaf cards and a validated anchor hint must resolve without a
watch-progress query. It passes on a freshly migrated database.

The runner now drains go test's output if reading the event stream
fails, so a go test still writing cannot block on a full pipe before
Wait returns.

Refs Silo-Server#599.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…o-Server#1376)

* perf(artwork): keep revisioned artwork URLs stable for a UTC day

Clients and CDNs cache artwork by full URL, so every new URL for unchanged
bytes costs a download the client already has. Local signed URLs rolled every
15 minutes. S3 presigned URLs carried the signing second and Cloudflare token
URLs the issue second, so they changed on every resolve and differed between
replicas.

A revisioned key names immutable bytes: its path is the identity and its
query only authorizes. At the default lifetime such a URL now holds for its
UTC day. The local signer uses a day-long issuance bucket, and S3 presigns at
the start of the day with X-Amz-Expires set to the day plus the TTL,
shortened to fit the SigV4 seven-day limit. The WAF rule fixes a token's
lifetime from its timestamp, so token timestamps are truncated to a quarter
of the token TTL instead. Every replica mints the same URL for the same
window. Mutable keys, short-lived capabilities such as avatars and chapter
thumbnails, and job artifact downloads keep their current buckets.

Over a simulated day of per-minute resolves through the production
resolvers, distinct URLs for one revisioned key drop from 97 to 2 (local),
1440 to 2 (S3 presigned), and 1440 to 33 (three-hour Cloudflare token), and
a second replica resolving 20 seconds later now always agrees.

A leaked revisioned URL now works for up to a day plus the TTL instead of
the TTL plus 15 minutes.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* docs(artwork): state URL stability limits from review

Record that replica-identical presigned URLs rely on the shared static
access key, that lowering s3.metadata_presign_expiry shortens the TTL but
not the day window, and how the Cloudflare Token TTL sets how often token
URLs change. Give the revisioned cache-policy test a few seconds of slack
so a run that crosses midnight UTC cannot fail it.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Silo-Server#1377)

* perf(access): stop re-reading legacy settings on every profile request

Every profile-scoped request resolved hidden libraries from the canonical
ui.disabled_library_ids row and, when the profile had none (the default),
also read the frozen legacy disabled_library_ids account row. Home section
requests did the same for ui.next_up_mode, up to three times per request.
The legacy settings endpoint stopped accepting both keys at the settings
cutover, so those reads return values that can no longer change.

A Go migration on Postgres and SQLite schema v26 copy each legacy value
onto every profile that has no canonical row. They use the settings planner,
so each stored value matches what the fallback read produced. Stored rows
win, and the legacy rows stay in place. With the values materialized, the
fallback reads go away. ui.next_up_mode now resolves in the same settings
read as the rest of the viewer scope and rides on access.Scope (excluded
from its JSON, so access fingerprints do not change), and section fetchers
take it from the request scope.

Measured with a MaxConns=1 pgx tracer: the auth prelude for a default
profile drops from 6 to 5 statements (7 to 6 for an access-group member),
and a NextUpMode call inside a request drops from 2 statements to 0.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* docs(middleware): correct the RequireAuth session cache comment

RequireAuth checks session validity with the SessionValidator on every
request, and the production validator is an uncached query. The comments
claimed an in-memory cache, which hides the per-request cost and suggests a
revocation might lag.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* test(sections): pin zero next-up reads on Home section requests

Measure the Home half of the identity-prelude change through
SectionHandler.HomeSectionItems instead of deriving it from per-call
counts. With a scope from the production viewer resolver, the Continue
Watching and Next Up section requests make no next-up mode reads for
either mode. On the base commit the same requests made 2 to 4.

Review follow-ups on the retired-fallback migration:
- read both legacy keys in one pass with key = ANY($1)
- log unconvertible values on SQLite, as Postgres already does
- correct the comment that said user_settings will be dropped; other
  keys still use it

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* refactor(database): run the retired-fallback backfill as a Goose SQL migration

New database changes belong in timestamped Goose SQL migrations under
migrations/sql, but the retired-fallback backfill was registered as a Go
migration. It now lives in
20260924052112_materialize_retired_settings_fallbacks.sql, created with
make migrate-create, and the Go migration is removed.

The SQL reproduces settingsmigrate.PlanRetiredFallback, which the SQLite
user store still runs in Go at schema v26. A hidden-library list converts
only in the shape Go's encoding/json decodes into []int; nulls and
non-positive ids are dropped, repeats are removed in first-seen order, and
a list of more than 512 ids is skipped. A next-up mode is trimmed of Go's
Unicode whitespace and kept only when it is "combined" or "separate".
Stored rows still win, legacy rows stay, and a re-run changes nothing.

TestRetiredSettingsFallbacksMigrationMatchesPlanner seeds 53 accounts
covering every shape the planner handles and fails if the migration
writes anything other than what the planner plans.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(settings): keep padded legacy next-up modes at the default

The retired fallback returned the legacy next_up_mode string unchanged, and
the section fetchers compared it with "separate" exactly, so a padded value
such as " separate " never enabled the separate Next Up row. The backfill
trimmed before validating and would have written "separate" for it,
turning the row on for that profile.

Only an exact member is converted now. The planner rejects a padded value
before the generic write normalization trims it, so the SQLite v26 path
skips it, and the SQL migration matches the two members exactly. A skipped
value leaves the profile at the default, which is what the fallback gave it.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* docs(settings): describe what unconvertible legacy values did before

The comments said a padded legacy next-up mode behaved as the default. It
did not: the section fetchers matched the raw value exactly against each
member, so a padded or unknown value showed next-up in neither place. The
contract has no such state, so those profiles now get the default. Say so,
and say that a hidden-library list over the contract's 512-id limit is not
copied, which makes those libraries visible.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…ilo-Server#1378)

* feat(jellycompat): add request metrics and a server span per request

The Jellyfin-compatible listener runs on its own http.Server, so the
native request metrics never see it. Its only per-request signal was an
INFO log line, which gives no p95 by route, no error rate and no split by
client. It also started no server span, so every Postgres, Redis and S3
dependency span in a compat request became its own root trace.

observeCompatRequest runs right after the request ID middleware and
records each request once, whatever answered it:

- silo_jellycompat_requests_total{route,method,status_class,client}
- silo_jellycompat_request_duration_seconds{route,method}

route is the chi template read after routing (or "unmatched"), method is
folded into the standard set, status_class adds "hijacked" for the
session socket, and client is a fixed family parsed from the MediaBrowser
Client field and then the User-Agent. The playback and transfer media
routes and /socket are counted but not timed, because their duration is
the client's viewing or connection time. HLS manifests stay timed.

Each request also opens a fresh server span, so dependency spans join
one trace. Like native v2, a caller's trace context cannot choose the
trace ID or the sampling decision.

The request logger's writer becomes statusResponseWriter, shared with
the new middleware, and hijacks through http.ResponseController so the
socket upgrade reaches the connection through every wrapper.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(jellycompat): time requests per client and bound socket traces

Resolve the review findings on the compat request metrics:

- Add silo_jellycompat_client_request_duration_seconds{client} over the
  timed routes, so operators can see which client is slow without a route
  by client histogram (about 200 series instead of about 26,000).
- Run the session socket's periodic checks outside the request's server
  span. A socket left open all day no longer grows one trace without
  bound; the check before the upgrade still joins the request trace.
- Match client families without allocating: fold ASCII case into a stack
  buffer and read at most 512 bytes of the Client field and User-Agent.
  An 8 KB User-Agent no longer costs an 8 KB copy per request.
- Add the moonfin family.
- Document the client histogram, that media and socket requests are
  counted when they finish, and the socket trace boundary.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
rangoDJ and others added 20 commits September 28, 2026 01:34
Deleting a group revoked each moved member's sign-ins with three
statements per member while holding the group-writer lock. The new
RevokeSignInsForUsersInTransaction does it in three statements total,
and the single-user form now delegates to it. DeleteMovingMembers also
wraps its Begin error with context.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…s-new-content-only

fix(notifications): skip releases for content added long before it matched
Docs in this repository are for people and agents changing the code.
Guides for installing, configuring, and running Silo belong in the user
manual on siloserver.org. CONTRIBUTING.md and AGENTS.md now say so, the
wiki index takes no new pages or sections, and the README's
documentation links point to the manual.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…ves-members-1402

fix(access): move a deleted group's members into the default group
…ver#1516)

Log migration status, per-version progress, and rollback candidates and results. Keep heartbeats active during lock waits and name registered Go migrations.

Verified migration and rollback behavior in an isolated sandbox; focused race checks and CI passed.

AI review: gpt-6-astra in T3 Code using the Codex provider.
…ides-to-manual

docs: send operator guides to the siloserver.org manual
Silo-Server#1627)

* fix(notifications): skip second library copies in server channel posts

Release events are per library, so a title added to a second library (a
4K library beside an HD one) produced a second "new" event and the
server-wide channel announced it again.

Server channel sweeps now drop events whose episode (series and episode
key) or item a different library made available before the event.
Profile fanout is unchanged: it already deduplicates per episode across
libraries and still reaches profiles that can see only the newer library.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(notifications): count only existing libraries as earlier copies

Availability rows have no foreign key to their library and outlive a
deleted one, so a title from a deleted library kept counting as an
earlier copy and its next arrival was never posted. The repeat lookup
now ignores rows whose library no longer exists.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Silo-Server#1634)

Remove scripts that nothing in the Makefile, CI, hooks, or docs uses:
the native rsync + air dev setup and the root .air.toml, one-time data
fixes and backfills, ad hoc SQL and benchmark scripts, a duplicate of
make jellyfin-web, jellycompat-diff.sh, the k6 load test, and an unwired
silo-profile test.

Also remove the Continuum-to-Silo migration script, its guide, and the
migrate-continuum-check target, and drop the links to the guide.

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
)

* fix(catalog): match name_prefix on the sort title only

The alphabetical jump matched name_prefix against the sort key or the raw
title, so a movie titled "The Hobbit" with sort title "Hobbit, The" was
listed under both H and T. Title sorting already orders by the sort key,
and Jellyfin's NameStartsWith compares SortName, so the letter rail now
matches only the sort key, which falls back to title when sort_title is
empty. Browse, favorites, and query-source previews share one helper.

* fix(catalog): apply sort-title prefix matching to every catalog path

Review follow-up. The in-memory prefix filter (relevance search,
exact-order collections, offset pages) and the watch-history source still
matched the raw title, so offset and cursor pages disagreed and history
results kept "The Hobbit" under T while history facets did not. Both now
use the sort key. The helper reuses sortTitleKeyExpr, the episode entry
executor drops its redundant title arm, test comments no longer claim the
prefix LIKE is index-backed, and docs/catalog-api.md states the rule.

* fix(catalog): trim sort titles like SQL BTRIM in the in-memory prefix filter

Review follow-up. strings.TrimSpace also removed tabs, which BTRIM keeps,
so a tab-led sort title could land under a different letter in memory
than in SQL. docs/catalog-api.md notes that recently added TV also
matches episode titles.

* docs(apiv2): name the sort-title rule in name_prefix descriptions

Review follow-up. The GET /api/v2/catalog parameter and the catalog
query body field now say name_prefix matches the sort title, or the
title when none is set. Regenerated the OpenAPI artifact, web types, and
the system-info fixture's contract digest.
…ed (Silo-Server#1630)

* fix(downloads): pick the version the profile watches when none is named

A download that omits media_file_id (season, series and monitored
episodes, and any client that leaves Auto to the server) always got the
highest-resolution file, with ties going to database order. A user who
watches a different version, such as a 1080p file with a sidecar .srt
next to a 2160p file without one, got the other file offline.

Native requests now pick from the profile's watch history: the file last
played for that movie or episode, then for an episode the version closest
to the series' most recent play (edition, resolution, HDR, video codec)
within the profile's latest 200 progress rows, then the highest
resolution with the lowest file id breaking ties. The v1 bridge keeps its
highest-resolution default, so a repeated v1 request never swaps files a
device already has. Items with one file skip the history reads.

* fix(downloads): only pick versions the profile may play

An automatic pick could register a version the profile's library access
or quality ceiling forbids, leaving an entry that looks ready but can
never be served. Candidates are now narrowed with the request's access
filter first; when none qualify the full list is kept and serving reports
the refusal as before.

* docs(downloads): describe the fallback when no version is playable

An automatic pick narrows candidates to versions the profile may play but
keeps all of them when none qualify, so the create or the later file
request can still be refused. The doc claimed that could not happen.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Quick <31828688+Quick104@users.noreply.github.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…1635)

* docs: move operator guides out of the server repository

Delete the operator guides that the siloserver.org manual now owns:
docs/wiki/, the S3 setup guide, and the monitoring and profiling
runbooks. The contributor material they held moves into
docs/architecture/: profiling bounds, node resource sampling, and
observability validation into observability.md; the workload metric
catalog; new collection-templates.md and media-naming.md; and NFO known
limitations. docs/README.md sends operators to the manual.

docs/update-to-1.0.md stays until 1.0 ships.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* docs: state the node mount cap per process and what the poster tests check

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* docs: name the node health route as /api/v1/health

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…-Server#1625)

* fix(notifications): stop notifying for series removed from Home

New-episode interest ignored Home removals. A series removed from Continue
Watching or Next Up kept matching the continue_watching and next_up
reasons, because the interest recompute only read favorites, watchlist,
progress and watch history, and nothing queued a recompute on removal.

The recompute now follows Home: an active series drop clears both
reasons, a per-card Continue Watching dismissal clears continue_watching
while the dismissed progress is unchanged, and a per-card Next Up
dismissal clears next_up while the dismissed episode is still next.
Favorites and watchlist still notify. Dismissal writes, drop changes
(Home handler and watch-provider sync) and the first progress write of a
new watch session queue a recompute.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(notifications): recheck Home removals after the session gap

A resume within ten minutes of the previous progress write was not a new
watch session, so it queued no recompute even when it lifted a series
drop or a Continue Watching dismissal made in between. Interest stayed
cleared until the next state transition or the daily rebuild.

A removal now queues a second recompute after the session gap. A resume
later than that starts a new session, which queues its own recompute.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(notifications): match Home's Next Up card and read dismissals once per rebuild

A per-card Next Up dismissal compared against the next catalog episode
after the progression cursor. Home shows the first episode after its
anchor that has a present file and has not been started, so a dismissal
of that card could fail to match and next_up kept notifying. The
recompute now looks up the same episode, only when the series has a
Next Up dismissal.

The interest rebuild read the profile's full dismissal lists once per
series. It now shares one read per profile across its series.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(notifications): read a series' dismissals directly and anchor Next Up like Home

The interest rebuild shared one snapshot of a profile's dismissals across
all of its series, so a card restored mid-rebuild could be overwritten
from the stale snapshot. Each recompute now reads only the dismissals of
the series' own episodes; the Postgres store gets a targeted
ListHomeDismissalsForItems, and other stores list the surface and filter.

Home anchors Next Up on the most recently completed episode, not the
highest. The card lookup now uses the same anchor, so a dismissal still
matches after a rewatch of an earlier episode.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(notifications): drop superseded progress and recompute on every applied import

Home's Continue Watching hides an in-progress episode once a later
episode was completed more recently; the recompute now applies the same
rule, so a stale entry no longer keeps continue_watching set.

Timestamped progress writes (imports, watch sync, offline-queued client
events) now queue a recompute whenever they apply. A late import can move
the stamp a Continue Watching dismissal holds for without a state change
or a new watch session.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(notifications): keep stamped playback ticks free and recheck restores

625d085 queued a recompute for every applied timestamped progress
write, but SyncProgress routes any write carrying updated_at through
SetProgressIfNewer, so a client that stamps live ticks would recompute
its series every flush. Timestamped writes now use the session test
again, extended to the clock: a write queues when it lands more than the
session gap after the stored stamp by its own stamp or by wall time, which
still covers a late import while leaving ticks free.

Restoring a dismissal or undropping a series now also queues the second,
deferred recompute, which repairs a concurrent recompute (such as the
interest rebuild) that read the state from before the restore.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(notifications): count started episodes like Home and keep rewatch ticks free

Home's Next Up treats any progress row that is completed or has a
position as started, including rows a history hide covers. The card
lookup built "started" from visible progress, so a started-then-hidden
episode could be chosen as the card; it now uses Home's test in SQL.

Timestamped writes keep completion sticky in both stores, so a stamped
rewatch tick of a finished episode is no longer read as a transition.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(notifications): lift Home removals on the first write after them

A resume inside the session gap right after a removal waited for the
removal's deferred recompute, and a release processed in that window was
decided on the stale interest. Each node now remembers a profile's last
Home change for the session gap, and a progress write whose row predates
it queues a recompute at once; later ticks stay free.

The Next Up card lookup also skips episodes the profile's own store
reports as started, since Home's Postgres progress test sees nothing for
profiles whose progress lives elsewhere.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(notifications): keep sub-second Home change times and document node limits

Truncating the Home-change time to the second could miss a progress row
written earlier in the same second, so the lift waited for the deferred
recompute. The marker keeps the full time again; a tick in the same second
may queue one extra, coalesced recompute.

The notifications doc now states that the marker and the deferred
recompute are per node. New-test cleanups report their errors.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…rver#1326)

Collection artwork was stored under fixed keys, so a replacement kept the
same URL and CDN and browser caches kept showing the old image. Keys now
carry a content revision in the shared artworkkey format
({imageType}/original.{revision}.webp), so new bytes get a new URL.

Every upload path (admin artwork upload and source URL, v1 create/update
artwork inputs, collage generation, template posters, and personal
posters) now uploads and commits the replacement first, then removes only
the revision it replaced, including legacy fixed keys. Cleanup outlives a
canceled request.

Closes Silo-Server#1258

Co-authored-by: Quick104 <31828688+Quick104@users.noreply.github.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…-Server#1647)

* fix(apiv2): keep sendfile for file responses and log body bytes

The v2 status recorder hid io.ReaderFrom from io.Copy, so every file
response went through a 32 KB user-space buffer instead of sendfile.
The recorder now forwards ReadFrom and Flush, and counts the body bytes
the handler wrote. The v2 request log reports them as body_bytes.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(apiv2): keep error statuses and flush errors through the recorder

Forwarding ReadFrom let chi's wrapper commit 200 before an asset copy
read its first byte, so artwork whose store failed at once went out as
an empty, cacheable 200. The artwork and subtitle routes now hide
ReadFrom, and the recorder sets 200 only once a copy writes bytes.

The recorder's Flush also hid the transport's flush error from
http.ResponseController; it now implements FlushError.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* refactor(apiv2): trim test wrappers and explain the recorder's Flush

iotest.ErrReader already offers only Read, so the tests don't need to
hide its type. Flush's comment now says why it matters: chi offers
ReadFrom only over a writer that can also flush.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…-Server#1648)

* fix(downloads): answer 503 when the offline artwork store fails

A transport failure fetching offline artwork answered v2 clients with a
500, and an upstream 5xx or 429 with a 404, so clients treated a store
outage as a server bug or as missing artwork. Both now return
dependency_unavailable (503, Retry-After: 5) on v2. The frozen v1 route
keeps its status codes.

Artwork fetches without an injected client use a 30 s timeout, and the
error no longer repeats the presigned URL.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(downloads): keep artwork failures safe to log and retry correctly

- Log only the upstream status or timeout for a failing artwork store;
  a malformed redirect's error quoted the presigned URL.
- Treat an artwork URL without a host as broken, not retryable.
- Keep the 30 s artwork client when the server wires no client;
  SetOfflineDeps had replaced nil with http.DefaultClient.
- Declare Retry-After on the artwork route's 503 and document it.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* refactor(downloads): simplify artwork failure handling

- Declare the artwork 503 and its Retry-After before registering the
  route instead of patching the registered operation.
- Use errors.AsType and fold the two broken-URL test cases into a loop.
- Treat a store that stops sending mid-body as unavailable, so v2
  answers 503 before the first byte, as it does for other store
  failures.
- Assert that an upstream 403 or 404 is never retryable.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(downloads): blame only store read failures on the artwork store

A failed write to the client was classified as the store being
unavailable, logging a false store outage. Only a read from the store's
response now counts. Also regenerate the web v2 types for the artwork
503's Retry-After header.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* test(downloads): assert both halves of each artwork upstream status

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(downloads): time out only the artwork store's waits, not the copy

http.Client.Timeout covered the whole exchange, including time spent
writing to a slow client, so a slow download of a large image could be
cut off and logged as a store outage. The artwork client now bounds
only the wait for response headers, and each read of the image body
gets its own stall timeout while the copy waits on the store.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(downloads): drop image headers when artwork fails before its first byte

With the stall timeout, a store that sends headers and then nothing ends
the copy while the response is still uncommitted. The image's length and
immutable cache policy were left on the writer, so v1's error response
went out with them. They are now removed whenever nothing was written.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…Silo-Server#1655)

* fix(watch-party): keep room members out of post-roll at the end of an item

A web member of a Watch Party entered post-roll in the last 30 seconds of
an episode. When the room then returned to its lobby, the player's room
exit saw a non-foreground mode, tore playback down without navigating, and
left the member on a black player page with no room socket.

Skip the early post-roll trigger inside a room. The room decides what
follows the item, and the member stays in the player until the room's
lobby snapshot sends it back to the room page.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(watch-party): read the room from the active request

The room is part of the request key the callback already depends on, so a
ref adds a render of lag and nothing else. Match the other room branches
in checking both the room id and its token.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…wal (Silo-Server#1657)

* fix(watch-party): drop a room command with the socket that delivered it

The web room connection kept the last transport command when its socket
closed. On reconnect the player saw the connection come back with a
command it had not applied on this connection and ran it again, projecting
the old command's position forward. Every planned five-minute socket
rotation made a web member correct twice: once to the stale projection,
then to the fresh sync the server sends after the member reattaches. If
the room had re-anchored in between, the first seek went the wrong way.

Clear the command when its socket closes, as a replaced connection already
does. The server sends a fresh command once the member attaches again.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(watchtogether): stop resending the room command after a reconnect sync

Connect clears a member's last delivered command for every new socket, so
the reconciler resent the room's last command within two seconds of each
socket renewal, with its original id and execution time. The member had
already been synced by the attach, and applied the old command on top of
it, projected from when it first ran.

Record the room's command as delivered to a member whenever the member is
synced to the room.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(watch-party): drop the room command when the connection effect resets

A changed authority, room token, rejoin, or replacement scope re-runs the
connection effect. Its cleanup lets go of the socket before closing it, so
the close handler returns early and the command survived onto the next
connection. Clear it in the cleanup too.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…ilo-Server#1624)

The Playback settings page described Chapter thumbnail workers as
"Parallel extraction jobs per library scan". The value sets how many
general workers the chapter thumbnail service starts in each server
process, each handling one file at a time. Library scans have no per-scan
job count. The service also starts one worker that takes only playback
requests, which the general workers pick up first too, so a title being
played does not wait behind queued work. Describe the setting the way the
other worker settings are described, and mention that extra worker.

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
@coderabbitai

coderabbitai Bot commented Sep 29, 2026 •

Copy link
Copy Markdown

Walkthrough

Warning

Review details and warnings were omitted to fit the comment limit.

Syncs Silo-Server/silo-server main at f84d2de, which brings /api/v2 that
prairie-apple and prairie-android main already call.

Prairie branding is kept and everything new from upstream is rebranded:
module paths, cmd/prairie, UI copy, Go string literals, data directories,
Docker and compose files. Wire identifiers move to prairie.*: WebSocket
subprotocols and ticket prefix, every X-Silo-* header, the plugin proto
package prefix and the plugin source kind.

Also restores Prairie code the squashed sync #174 dropped: the Tizen HLS
tag strip, web branding assets, HEAD on transcode segments, the Redis save
guard and non-fatal hub start, PRAIRIE_* env names, AVIF admin settings,
the w200 collection poster rung, calibre metadata keys and the episode
catalog placeholder numbering.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@JonahMMay
JonahMMay force-pushed the sync/upstream-2026-09-28 branch from b045747 to 08794f7 Compare September 29, 2026 05:59
Brings the sync up to upstream 19e8184: advisory age poster badge, stable
collection artwork URLs, v2 sendfile, offline artwork 503, and two Watch
Party fixes. New upstream code is rebranded as in the parent merge, and the
contract artifacts (OpenAPI, web types, fixtures, route inventory, ledger,
settings bindings) are regenerated.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@JonahMMay
JonahMMay marked this pull request as ready for review September 29, 2026 06:29
@JonahMMay
JonahMMay merged commit 12d9337 into main Sep 29, 2026
16 checks passed
@JonahMMay
JonahMMay deleted the sync/upstream-2026-09-28 branch September 29, 2026 12:58
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.