Sync upstream silo-server main (~925 commits, API v2) - #183
Merged
Merged
Conversation
…lear errors (Silo-Server#1329) * fix(historyimport): import Emby watch dates, episode favorites, and clear errors Emby list queries omit ProductionYear, UserData.PlayCount, and UserData.LastPlayedDate unless Fields names them (verified on Emby 4.8.11, 4.9.5, and 4.10.0.40). Silo only asked for ProviderIds, so every Emby import created no history rows, stamped progress at the Unix epoch, and could never update progress on a re-import. - Request the missing fields; page item lists and batch ID lookups. - Import episode favorites; report season favorites, which Silo cannot hold. - Mark every episode of a played multi-episode file watched. - Keep importing when Emby show details cannot be read. - Accept Emby accounts without a password. - Report unreachable sources and invalid input with specific messages instead of a 500, and put the message in the v2 problem detail. - Show safe per-cause summaries for run warnings and unmatched items, and count unmatched items by cause. - Web: count skipped items once, ask for a new Emby Connect sign-in after a run consumes the session, and fix the Emby card copy. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(historyimport): keep v1 error decisions and fixed Emby warnings The shared handler mapping also answers /api/v1, which is frozen. Restore its decisions and map unreachable sources and invalid input in the v2 adapter, which reads the cause the seam's APIError wraps. Emby favorites and series-metadata warnings now store fixed text and log the upstream error, since v1 returns stored warnings verbatim and Emby error bodies can reach them. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…erver#1337) * feat(admin): run full library metadata refreshes from the web The library refresh button always queued a quick refresh, which skips healthy matched items, so provider-side fixes never reached existing libraries without API calls. The button now opens a dialog offering the quick refresh or a full refresh of every item. A new manual task, Refresh All Library Metadata, runs a full refresh of each enabled library. It runs the library refresh executor inline because admin jobs need a requesting account. A cluster advisory lock allows one task run across servers, and the executor now holds a per-library lock so a refresh job and the task never refresh the same library at once. The typed v2 request helper learns to send JSON bodies the contract marks optional, which the refresh mode needs. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(adminjob): let a recovered library refresh wait for its lock A library refresh job recovered after its worker stopped heartbeating failed straight away on the per-library advisory lock, because the earlier attempt can still hold it: its worker needs up to a heartbeat interval to notice the new claim, and a server that disappears without closing its connection keeps the lock until PostgreSQL drops the session. Recovery used to rerun the refresh to completion. A job on its second or later claim now retries the lock every five seconds, reporting that it is waiting, until the holder lets go, the job is cancelled, or the six-hour job limit runs out. First claims and the full refresh task still fail fast on a held lock. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…Silo-Server#1331) * fix(historyimport): import Jellyfin favorites and support Jellyfin 12 Jellyfin imports skipped favorites, and several requests misbehaved against Jellyfin 12: - A rejected password was retried under the /emby prefix, which Jellyfin 12 removed, so a credentials error surfaced as a 404. On 10.11 the retry counted as a second failed login. - Items came from /Users/{userId}/Items, which Jellyfin 12.1 no longer lists in its OpenAPI document. - A source address ending in "/" produced //Items requests that 404. - Items without a LastPlayedDate were stamped with the import time, so a re-import could overwrite newer Silo progress. Fetch IsFavorite movies, series, and episodes as favorite-only records and keep the IsFavorite flag on watched items. Sign in once at the entered address, read items through /Items?UserId=, trim trailing slashes, leave UpdatedAt zero without a last-played date, and request only ProviderIds. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(historyimport): keep favorite series lookup failures non-fatal A favorite episode added its series ID to the series lookup that the watch history depends on, so a failed lookup discarded the whole run even though favorites are otherwise best-effort. Look up series needed only by favorites separately and report a failure as a run warning. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(historyimport): store fixed text for Jellyfin favorite warnings Run warnings now hold fixed diagnostics that PublicWarning maps to monitor text (Silo-Server#1329), because v1 returns them verbatim and upstream errors can carry response bodies. Log the Jellyfin favorites and favorite-series lookup errors and store fixed warnings with their own public summaries, as Emby does. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(historyimport): say Jellyfin favorites may be incomplete after a failed query Watched items that Jellyfin marks as favorites still import as favorites when the favorites query fails, so the run summary must not claim none were imported. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Requiring an open issue before a pull request cost more than it bought. It did not prevent duplicate work — agents would find an open issue and race the author to a PR — and it pushed contributors into filing throwaway issues that only restated what the PR body already had to say. An open issue is no longer a precondition. The pull request has to state the problem itself: what breaks or is missing, who it affects, and why the change is the right answer. `Related issue:` now links an issue when one covers the work and reads `N/A` when none does, instead of needing a "narrow fix" justification. Opening an issue first for scope-changing work stays as advice, with the tradeoff stated: it is the cheapest way to find out the work is already in flight or out of scope, and a rejected pull request is the contributor's risk. Adds the rule that actually addresses the collision: an open issue is not an unclaimed one. Read its comments and linked pull requests before implementing it, and agents should raise a likely collision rather than race the author. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…Silo-Server#1341) * refactor(tasks): trim the Scheduled Tasks page to tasks admins manage The admin Scheduled Tasks page listed about 38 tasks, most of them internal queue workers and pollers that tick every 15 seconds to 5 minutes and need no administrator attention. - Hide internal workers (matching, artwork caching and delivery checks, download preparation, request reconcile, collection syncs, autoscan poll, frequent log cleanups, provider-ID and watch-history repair). They still run and stay reachable by key; Server Activity still shows them while they run. - List library-scoped tasks (audiobook, ebook, manga, podcast) only when a library of that kind exists. - Run six retention sweeps as steps of one Database Maintenance task, daily at 05:00. A migration drops the old keys' saved schedules. - Make the repair tools manual-only and group them under "On demand". Saved schedules for manual-only tasks are ignored at startup. - Use plugin-supplied display names for plugin scheduled tasks. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(tasks): keep v1 task list unscoped and report maintenance steps Review follow-ups for the Scheduled Tasks cleanup: - Library scoping now applies only to the v2 default list through ListRelevantTasks. The frozen v1 list and the realtime tasks snapshot keep their previous signatures and behavior. - v2 task history adds a steps array (key, name, status) for database_maintenance, and the task detail page names failed steps. Step errors are logged, since v2 does not return error text. - Server Activity links each running task to its detail page, because hidden workers do not appear on the task list. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(tasks): run database maintenance on one server and keep canceled steps out - Every API process fires the same 05:00 trigger, so Database Maintenance now takes a PostgreSQL advisory lock. One server runs the sweeps and the others finish without doing anything. - A step interrupted by cancellation is left out of the step results instead of being recorded and logged as failed. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…ver#1313) * feat(scanner): support flexible movie and episode filenames * fix(scanner): correct flexible filename parsing regressions Review of the flexible-naming change found cases where the new parser wrote wrong episode numbers or types that the previous parser left unset. - Read "Show - 05 - Part 2" from the first dash-delimited number instead of the last trailing number. A number completing the show folder title ("24 - 05", "9-1-1 - 01") yields to a later field-ending number, so titles and times ("24 - 12-00 AM") no longer replace the episode. - Reject NxM coordinates from audio layouts (2.0x2, 5.1x264) and dimensions in year-numbered seasons (2048x1080). - Keep "Show (2005) - S01E01" files in a show folder as series in mixed libraries. - Pass library roots to IsMisplacedSeriesFile, the series refresh scope, and the scanner's fallback root and variant parsing so directories such as /mnt/s3 above a library are not read as season folders. - Record a stale-identity series row without failing the running scan. - Wake backed-off queue rows when a rescan changes a file's group key. - Link episode-only numbers without a season only where absolute and per-season numbering agree. - Do not cache a failed folder lookup as a disabled library. - Keep the full undated title as a fallback search for loose titles ending in a number, and require a rescan when an anonymous sibling shares a stale root with a file-rooted show. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(scanner): address follow-up review comments - Treat a leading NxM followed by a bracketed release year as a movie title (10x10 (2018)), not season and episode coordinates. - Document that a number alone links an unseasoned episode only in the first regular season. - Tighten two tests to check exact processing counts and the fallback release-year search. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(naming): use one movie-folder evidence rule for path classification ResolvePathContext still used the older folder-evidence check, so a dated show folder containing "Show (2005) - S01E01" files parsed as a movie while root inference stored a series, and scans skipped the episode numbers. Use the root-inference rule in both paths and remove the unused helper. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(naming): keep format words in dated titles and defer partial links - An explicit "Title (Year)" boundary keeps title words that also name a format ("Opus", "UHD", "4K"). Resolution, source, and codec terms before the year still end the title. - Treat a leading NxM followed by a dotted or spaced year (4x4.2019) as a movie title. - Link an unseasoned number alone only to season 1, so a partial episode fetch cannot make a later season look like the first. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(metadata): share one provisional item across flat series episodes Files directly in a series library root are their own observed roots, so matching one show cannot relink its neighbors. Skeleton dedup keyed on the observed root therefore created one provisional item per episode until the show matched. Reuse the series item already linked to another file of the same resolved scanner group, preferring a confirmed match. Ambiguous groups keep separate items. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(naming): reject aspect ratios as episode coordinates - An NxM inside a release tag ([16x9 1080p]) or a known aspect ratio directly before release details (Movie.16x9.1080p) is video metadata, not a season and episode. - A resolution token no longer blocks a later dot- or space-separated episode marker (Show.1920x1080.E02). A dash-attached marker (1920x1080-E15) is still rejected. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(naming): let episode markers and title words outrank format tokens - An aspect ratio followed by an explicit episode marker (Show.16x9.E02) yields to that marker and the season folder. - A format word that appears before any number, such as "Opus" in a show title, no longer truncates the name before its episode number. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(naming): parse episode markers after audio layouts - A rejected audio layout (AAC.2.0x2, DD5.1x264) no longer blocks a later dot- or space-separated episode marker, matching resolutions and aspect ratios. - Treat a leading NxM followed by a dash-separated release year (10x10 - 2018) as a movie title. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(naming): handle decimal aspect ratios and dotted title years - Reject decimal x-tokens with a one-digit count (2.35x1) and the 3x2 ratio before an explicit episode marker, so Show.2.35x1.E02 keeps its season folder and episode. - A bare release year also marks the title boundary when every format term before it is an ordinary title word (Mr.Hollands.Opus.1995). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(naming): reject small picture sizes as episode coordinates An NxM token with two multi-digit sides (176x144) describes a picture size. Reject it before the season range check so a later episode marker or the movie identity wins. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(naming): keep dated titles and year seasons from misreading NxM An NxM inside the episode title of a dated file ("Show - 2016-10-25 - Salon Fail!; 2x4 Vandalism Victim") replaced the Season 21 folder and the air date with S02E04. An NxM now counts after an air date only when it directly follows the date ("Show - 2024-10-01 - 2024x246"). The picture-size rule also rejected year-numbered seasons with a three-digit day count (2024x246). A year-shaped season with at most 366 episodes now falls through to the year-season check. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(naming): keep bare aspect ratios from classifying movies as series A trailing aspect-ratio tag in a mixed library (Movie.Name.16x9.mkv) was read as S16E09 and sent the movie down the series path. A recognized aspect ratio now counts as a coordinate only in a season folder or a declared series library. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(metadata): converge flat-series provisional items across nodes The dedup lock that lets episodes of a flat library-root show share one provisional series is per process. Two nodes matching episodes of the same show at once could each miss the other's link and mint a separate local item keyed on its own file path. A resolved flat group now anchors its provisional content ID on the group key, so concurrent creators write the same item. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(metadata): link to an existing flat-group item instead of resetting it With a shared provisional ID, a node that passed the group check before another node linked its episode would Upsert the item and overwrite every field, resetting an already matched series to an empty skeleton. A flat group now links to the item when it already exists and creates it only when it is absent. The test fake now reports a missing item with catalog.ErrItemNotFound, as the real repository does. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(metadata): create flat-group provisional items atomically Loading the group item before upserting it left a window: a node that stalled between the two steps while another node matched the item would still overwrite it with an empty skeleton. The item repository now has InsertIfAbsent, which uses ON CONFLICT DO NOTHING and reports whether it wrote the row. Flat-group skeletons use it; a creator that loses the insert links to the existing item. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(metadata): keep shared skeletons when a creator's cleanup runs Skeleton cleanup after a failed link or membership write used the unconditional item delete. With a shared content_id, the node that won the insert could fail to link after another node had linked and enriched the same item, and its cleanup would delete that item. ItemRepository.DeleteIfUnreferenced deletes one item only while nothing references it, using the orphan sweep's safety conditions, and skeleton cleanup now uses it. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(naming): keep numbers in show titles from becoming episodes The trailing-number fallback read a number that completes the show folder's title as the episode: "The 100/Season 01/The 100.mkv" became S01E100. Its non-overlapping search also consumed the separator after the title number, so "The 100 05" never considered 05. Trailing candidates now resume after each number, and a number whose prefix matches the show folder's title is skipped. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(naming): keep format-word titles and low-resolution sizes apart An undated folder and file sharing a title such as "Mr Holland's Opus" or "The UHD Journey" lost the format-like word and everything after it. When the folder title matches and every format term in it can also be a title word, the full title is kept. A picture size with a two-digit height (128x96) was read as S128E96 and hid a later E02 marker. A three-digit width with a multi-digit height is now a picture size, and an episode marker after it keeps the season folder. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* feat(storage): fence S3 clients and artwork stores during transitions Co-Authored-By: Codex <gpt-5> * feat(storage): add verified transitions and resumable recovery Co-Authored-By: Codex <gpt-5> * feat(api): expose storage transition status and cancellation Co-Authored-By: Codex <gpt-5> * feat(web): guide managed artwork storage transitions Co-Authored-By: Codex <gpt-5> * docs: explain managed artwork storage transitions Co-Authored-By: Codex <gpt-5> * fix(storage): address transition review findings Prevent duplicate active transition jobs, expose capability discovery to clients, and correct recovery field documentation. Reject stale execution after settings commit and cover the behavior with focused and PostgreSQL-backed tests. Co-Authored-By: Codex <gpt-5> * fix(storage): harden transition failure handling Keep cancellation recoverable when staged-state cleanup fails, distinguish request validation from infrastructure errors, and preserve unrelated settings edits during managed transitions. Bound synchronization tests and make AWS smoke configuration portable. Co-Authored-By: Codex <gpt-5> * test(api): refresh media route manifest Record the additive storage-transition capability endpoint in the checked-in media route classification manifest. Co-Authored-By: Codex <gpt-5> * fix(storage): preserve transition state and private artifacts * fix(storage): detect aliases between source namespaces * fix(storage): move operational blobs with managed transitions Since blob storage became backend-neutral, a local root holds subtitles, diagnostic bundles, job artifacts, and avatars beside artwork. The transition still treated operational data as private-S3-only, so it: - never copied local diagnostics or job artifacts to private S3, and left their "local" bucket rows pointing at a bucket that does not exist; - never copied private S3 diagnostics or artifacts into a local root, and left their rows on the old bucket where a later repoint could not find them; - skipped subtitles on every local target; - disabled diagnostic uploads on commit to local storage. Model a transition as two moves, derived the way blobstore.Open selects stores: the assets store and the operational store (private bucket, local root, or none). The assets copy excludes operational namespaces; the operational copy runs when that location changes and uses explicit prefixes whenever a local root is involved. Restart recovery repoints rows between the old and new operational bucket, treating a local root as "local". Disabling S3 still moves everything to local disk. A local install keeps its private bucket, including a legacy operational alias, unless the transition changes it, which makes a private-only move on a local backend possible. Fence each distinct store once: a local root is both stores, and fencing it twice would never return. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(storage): lock and transition a local install's private bucket A private bucket owns diagnostics, job artifacts, and avatars on either backend. On a local install that has stored data, adding one moved those reads to an empty bucket and stranded the files in the local root, and the settings API allowed it because the lock only covered private keys on an S3 backend. Lock the private endpoint, bucket, and key prefix on both backends. On a local install, editing them now opens the managed transition as a private-only change: the dialog shows the new private location (or local disk when the bucket is removed), omits the artwork path, checks the old bucket's reachability before a copy policy, and sends the private keys with the request. Transition copy now says subtitles move in every direction and no longer claims local mode cannot read diagnostics or artifacts. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * test(settings): give the layout form mock getPersistedValue The storage page now reads the persisted private bucket on every render to decide whether a local install's copy source needs a reachability check. The real settings form provides getPersistedValue; the admin layout test's partial mock did not, which only surfaced once the call stopped short-circuiting on non-S3 sources. Also name the private endpoint setting once in the transition service, which its new location summary made a repeated literal. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(storage): keep avatars out of an artwork-only preserve preflight When only the public location changes, the operational store stays where it is and preserve_uploads copies no avatars. The preflight said it did. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(storage): address review findings in managed transitions - Commit writes only the location keys the copy verified. Other storage settings keep their current value, so a credential rotated or a read endpoint changed while a transition ran is no longer reverted at commit. - Start applies the settings API's per-key rules. An unknown backend or a relative local path can no longer be committed and stop the next boot. - Start rejects a target whose normalized identities equal the active stores before creating a job, and builds the preflight from those same identities, so the preflight matches what the copy does. blobstore.LocalIdentity computes a local identity without creating the root. - Clearing the private bucket clears the rest of the private location, so removing it no longer fails validation on a leftover endpoint. - A local-to-local transition keeps saved public S3 settings and credentials instead of blanking them. - Restart recovery repoints rows naming the "local" bucket under every policy. An S3 reader took it as a real bucket name, presigned downloads against it, and could never expire those rows. - The final pass fences only the stores being copied from, so removing a local install's private bucket no longer blocks artwork writes until restart. - The preflight no longer promises avatar copies when no operational location moves. Tests cover each case, including a PostgreSQL test for the repoint that refuses to run against a database already holding "local" rows, and a test that blobstore.Open fences the shared local root. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(settings): keep private-storage edits from dead-ending - The settings lock covers the private endpoint and key prefix only while a private bucket is configured, and compares the prefix after trimming slashes. A leftover prefix or an equivalent edit no longer needs a transition that could only fail as a no-op. - The storage page compares a dirty location field with its stored value, so retyping a bucket or adding a trailing space is a plain save. - The transition dialog's text follows the actual move: removing an S3 install's private bucket says private data becomes unavailable, Start fresh when disabling S3 says nothing is copied, avatar copies are mentioned only when the private location moves, and a private bucket without an endpoint cannot be queued. - The setup wizard makes locked location fields read-only instead of saving into a 409. - docs/settings-api.md lists the new lock triggers. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(storage): close legacy, lock, checksum, and throughput gaps Legacy aliases: Start now builds its target from canonical keys and drops the s3.operational_* rows first. A request that cleared a canonical key was refilled from its alias, and an auto backend whose public bucket existed only as an alias resolved to local, so a private-only API request became a move to local disk. Lock before artwork: private writes never recorded a storage identity, so a private bucket holding avatars, diagnostics, or job artifacts could be repointed before the first artwork write. The private S3 client now records storage.operational_identity on its first successful write, through a write observer every private writer shares. The settings lock engages on either row and, with only the private row, locks only the private keys. A transition that moves the private location rewrites or clears the row. Checksums: objects up to 8 MiB are copied with Put, which records the silo-sha256 metadata S3 compares in Matches. Streamed copies had none, so the image cache re-uploaded every migrated variant it touched. Throughput: each page of 250 keys copies through eight workers, with receipts, counters, and progress under one lock and the page flushed before its cursor advances. Progress reports at most every two seconds instead of writing the job row and publishing an event per object. The fenced pass accepts an object this run already verified once its source digest still matches, without downloading the target copy again. The architecture, settings API, and admin wiki docs describe the managed transitions and the new identity row; the wiki still described a manual database procedure. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(storage): address review of the lock and identity changes - The settings lock and the storage page compare locations the way stores name them: endpoint scheme and host and bucket names are case-insensitive, and key prefixes ignore slashes. The page also treats a leftover private endpoint or prefix with no bucket as a plain save. These edits previously dead-ended between a 409 on save and a transition that failed as a no-op. - artwork_storage.locked again means the artwork location is recorded. /api/v2 adds private_locked for the private bucket alone, so a private bucket holding data no longer disables the artwork path in the settings page or the public fields in the setup wizard. The frozen /api/v1 response is unchanged. - Recording the private bucket's identity can no longer fail a write that already stored its object. A failed recording is logged and retried on the next write, with a context detached from the request. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * test(settings): type private_locked in the storage page's status mock The production build typechecks tests with tsc -b; the mock's hand-written status type did not know the new field. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(storage): tolerate scalar restart receipts during finalization * fix(storage): reject partial orphan and probe deletions * fix(storage): guard transition review on capability discovery * docs(storage): describe managed transition policies accurately * fix(storage): reject ambiguous source namespace copies * fix(storage): route local path changes and preserve endpoint identity * fix(storage): expire recovered transition job receipts * fix(storage): preserve committed transition receipts during cancellation * fix(apiv2): honor storage transition cancellation contract * chore(contracts): refresh combined storage transition digest * test(storage): wait for mutation fence contention * fix(storage): bind private bucket identity at startup * fix(storage): validate transition destinations before commit * docs(storage): describe private startup lock and local path changes * fix(storage): keep empty artwork identity on private move * chore(contracts): refresh storage identity digest * fix(storage): route automatic S3 switch through transition * fix(storage): wait for lock status before editing settings * fix(setup): wait for storage lock status * fix(settings): avoid passing click event to save * fix(storage): finalize recovery after transient stage read * fix(storage): keep config watcher running with unreadable stage * fix(storage): warn when sidecar artwork needs refresh * fix(storage): scope transition validation to storage settings * fix(storage): tolerate vanished objects during bulk copy * fix(settings): exclude recovery state from admin settings ETag * fix(storage): condition artwork reconcile certification on active identity * fix(storage): show safe transition progress in admin settings * fix(storage): expose safe transition progress in admin jobs * fix(storage): reject reconcile from stale API nodes * fix(storage): sanitize storage jobs in admin streams * fix(storage): coordinate API nodes with storage transitions * fix(storage): hide stale progress while transition resumes * test(storage): use structured transition progress in admission checks * fix(storage): preserve restart receipt for resumed claims * fix(storage): reuse transition phase constants * fix(storage): reuse artwork path setting key * fix(admin): clean up storage transition lint findings * fix(storage): reuse backend setting key in reconcile guard * fix(storage): count missing S3 objects as deleted * fix(storage): bind local root after private data copy * fix(storage): explain private location mismatch recovery * fix(storage): report when lock state is known * fix(api): keep v1 status response unchanged * fix(settings): guard storage transition review * fix(setup): guard storage edits when lock status is unavailable * fix(storage): rejoin node admission after a database restart Any failed probe of the admission session stopped the process, so a PostgreSQL restart, a failover, or a slow ping took down every API node even with no transition running. A node holding only the shared lock now rejoins on a new session with a non-blocking shared acquire. Its blob writes pause while the database cannot be reached and resume once it rejoins. It still stops if a transition took the exclusive lock while it was out, or if it owned the transition whose session was lost. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(storage): hold committed transition claims until restart After a commit, or a commit the process could not confirm, the process keeps its source fences until it restarts. If a database outage outlasted the stale window, recovery requeued the job, and the next claim either failed a committed transition or ran it again against fences the process still held. A late cancellation also left a committed job reported as canceling. Such claims now record the restart receipt with cancellation cleared and keep it alive, and each process requests its restart once. Boot recovery completes the job after the restart. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(storage): check the recorded location before a node rejoins A node cut off from PostgreSQL for another node's whole transition found no exclusive holder once that owner restarted, rejoined, and kept writing to the stores it opened before the commit. Under the new shared lock, a rejoin now compares the recorded storage identities with this process's stores and stops the node when a transition moved them. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(storage): wait for admission rejoin before failing a transition The runner can claim a queued transition as soon as the database answers, up to one admission probe before the node has rejoined. Execute now waits a few seconds for the rejoin instead of failing the job. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(settings): compare storage edits with the saved location After a committed backend switch and before the restart, the page compared drafts with the running backend, so every save became a transition that Start rejects. It now compares with the saved location. With lock status unknown, a location edit also holds back the storage credentials that belong with it. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(setup): explain storage locks with the settings page's name The wizard said files were stored in a public bucket that Automatic storage on local disk leaves empty, and pointed to a page named Infrastructure that the admin settings call Storage & Database. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(storage): order node rejoin checks after in-flight commits A node could rejoin in the moment between a transition owner losing its admission session and its commit becoming visible, then keep writing to the old stores. The owner now confirms its session inside the commit's settings transaction, and the rejoin check reads the identity rows under the same settings lock. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(settings): use saved storage only while a commit awaits restart Comparing drafts with the saved backend broke an install whose saved settings changed while unlocked and whose first write then recorded the running store: edits saved directly into a 409. The page now uses the saved location only while the latest transition awaits restart, and the running one otherwise. The wizard's footnote and backend lock text name the Storage & Database page. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Quick104 <31828688+Quick104@users.noreply.github.com> Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Bring the Jellyfin-compatible API in line with Jellyfin 12.1 and fix two
problems found while validating it with Jellyfin Web.
- Jellyfin 12 API: /Items/{id}/Collections, audio and subtitle language
filters and Filters2 facets, root recursion for filtered requests,
episode parent posters, OriginalLanguage, /Persons name and parent
filters, display-preference defaults, Jellyfin 12 HLS codec strings and
the Dolby Vision variant, LocalizedLanguage/LocalizedOriginal.
- New installs report 12.1.0 and install Jellyfin Web 12.1; configured
servers keep their versions.
- AudioLanguagePreference "OriginalLanguage" maps to x-silo-original
(settings contract revision 9); SubtitleMode maps onto the profile's
subtitle settings, and PlaybackInfo picks the default subtitle with
Jellyfin's selection rules.
- Embedded text subtitles are extracted in one pass per file, and in the
background when Jellyfin Web starts playback, so switching tracks no
longer waits a minute or more per track on large remuxes.
- Jellyfin Web device IDs longer than 256 bytes are accepted.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* feat(playback): configure profile-wide seek intervals * fix(player): track only accepted seek targets * fix(player): discard seek targets after failed reanchors --------- Co-authored-by: Quick104 <31828688+Quick104@users.noreply.github.com> Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…ilo-Server#1342) * fix(intromarkers): stop re-refining unchanged silence backfill files A chapter silence refinement that finds no better intro end, or fails, leaves the chapter:v1 marker unchanged, so intro_markers_detected_at never moves and the backfill picks the same files ahead of everything else on every run. Record each such attempt in intro_silence_refinement_attempts with the file identity, the refined range, and a hash of the silence settings. The backfill skips a file while its attempt still matches: until the inputs change after a clean no-improvement result, and until retry_after (12h doubling, capped at 7 days) after a failure. Files never attempted are picked first. Fixes Silo-Server#1272 Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(intromarkers): keep silence retry backoff for failures inside the window Scheduled tasks run per process, so every API server runs the nightly marker job over the same files. A failure recorded while an earlier failure is still backing off is a duplicate run (or a forced episode analysis), not a retry, and no longer escalates the failure count or moves retry_after. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(intromarkers): re-run silence refinement when chapters change The refined intro segment comes from the file's chapters, and a re-probe can rewrite them without changing the file identity or the stored marker. Record a SHA-256 of the chapters loaded with the candidate and require it to match mf.chapters before skipping a file. The column lands in its own migration because a development deployment already applied the table migration. Existing rows get an empty hash, so those files are refined once more. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(intromarkers): scope silence refinement failures to the recording server Every server runs the nightly marker job against its own ffmpeg and media mounts. A failure recorded by one server (missing ffmpeg, lost mount) no longer defers the file on other servers: the attempt records the host that wrote it, the backfill honors retry_after only for its own failures, and backoff escalates only on the same server. No-improvement results still apply everywhere. The unapplied follow-up migration now adds recorded_by next to chapters_hash and is renamed to cover both. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(intromarkers): store silence attempt file IDs as bigint media_files.id is bigint, so an int4 media_file_id would reject attempts for files past the int32 range and leave them in the backfill loop. The table migration now creates the column as bigint, and the follow-up migration widens it on a database that already applied the integer version. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…el (Silo-Server#1325) * fix(playback): fall back to CPU decode and software under auto hw accel With playback.hw_accel=auto, a GPU that can encode a source but cannot decode it made FFmpeg exit before its first manifest, and playback failed instead of trying a safer path. An automatic video transcode that resolves to a hardware backend now tries GPU decode and encode, then CPU decode with GPU encode, then software. It moves on only when FFmpeg exits before its first manifest, never while a process is still running, and briefly remembers the path that worked for that file. Explicit acceleration settings, copy/remux, and tone-map recipes keep their existing behavior. The fallback covers native local and remote starts, transcode-node starts (keeping replacement isolation), Jellyfin-compatible local and remote starts, and reconstruction on the API server and nodes. The NVENC mixed path uploads CPU-decoded frames to the configured CUDA device. Nodes report software_video_decode so recipe cards record the executed path. Fixes Silo-Server#918 Co-authored-by: MrDoudou <22004704+MrDoudou@users.noreply.github.com> Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(playback): address lint findings and refresh contract digest fixture Check session.Close errors in the node fallback tests, name the transport startup "ready" outcome, split a double-advance assertion, and regenerate get_system_info_ok.json for the updated v2 contract digest. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(playback): size auto remote start budget from the request Native remote starts extended their deadline for the hw_accel=auto fallback only when the node's stored capability report resolved to hardware. A missing or stale report still let the node walk every fallback path under the single-wait deadline, so the API could cancel a fallback that was about to succeed. The node resolves auto against live hardware, so size the budget from the dispatched request instead. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(playback): judge manifest identity from the file that was read getManifest read the playlist by path and then stat'd the path again. If FFmpeg's rename landed between the two calls, the inherited playlist's bytes were paired with the new file's identity and accepted as current. Read and stat the same open handle. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(transcodenode): let the node decide auto fallback readiness Jellyfin-compat remote starts asked the node to wait for readiness only when the node's stored capability report resolved to hardware. A missing or stale report either kept the fallback from a node that now has a GPU, or forced a software-only node to close a slow start at the readiness deadline. Add auto_fallback_ready to the start request: the node waits for the first manifest only when its own live hw_accel=auto pipeline is enabled, and otherwise starts unwaited. Jellyfin-compat sends it for auto video transcodes; older nodes ignore it and keep the unwaited start. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: MrDoudou <22004704+MrDoudou@users.noreply.github.com> Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…o-Server#1347) * fix(watch-party): stop warning on routine room socket reconnects The server ends every room socket after five minutes and the web client reconnects in about half a second. The player showed "Reconnecting to room" on each rotation and never cleared it, so both viewers saw a false warning that stayed up for eight seconds. Wait two seconds of lost connection before warning, and clear the warning when the socket reconnects. Expire notices from player state instead of hiding them inside the overlay, so a later identical notice shows again rather than staying invisible. Fixes Silo-Server#1346 Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(watch-party): keep notices raised while the player is minimized Notice expiry now lives in player state, so a notice raised while the player is minimized would expire before anyone saw it, including the one-time "Lower quality" offer. Start the expiry only once the player is shown again. Advance timers in the displacement test so it still proves the replacement path suppresses the delayed reconnect warning. Refs Silo-Server#1346 Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(watch-party): keep newer notices and hold the warning for an outage The delayed reconnect warning replaced whatever notice was showing when it fired, so an admin message or server-restart notice that arrived during the two-second delay was lost. Only replace the notice that was showing when the outage began, or an empty one. The warning also expired after eight seconds while the room was still disconnected, leaving controls unavailable with no indication. Keep it until the socket reconnects, the room closes, or the viewer is displaced. Refs Silo-Server#1346 Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(watch-party): restore the reconnect warning after a newer notice A notice raised during the reconnect delay keeps its place, but once it expired the room could still be disconnected with no warning on screen. Remember that the outage has outlasted the delay, and let an expiring notice hand back to the reconnect warning until the socket reconnects. Refs Silo-Server#1346 Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(watch-party): preserve notice expiry through socket rotation --------- Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Co-authored-by: Quick <31828688+Quick104@users.noreply.github.com>
…ilo-Server#1348) * fix(subtitles): keep SRT {\anN} placement in the WebVTT conversion The SRT to WebVTT conversion copied cue text verbatim, so ASS override blocks such as {\an8} reached WebVTT players as literal text and the cue fell back to the bottom of the frame. The cue's \an alignment now becomes WebVTT cue settings on the timing line, and override blocks are removed from the text. The Jellyfin-compatible windowed .srt response writes the tag back. The converter also no longer merges cues when an SRT omits the blank line between them, keeps a stray CR out of the timing line, and recognizes timing lines that use a period before the milliseconds. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(subtitles): split CR-only SRT line endings A classic-Mac SRT uses a lone CR between lines, so the converter saw one line, recognized no cue, and left {\anN} blocks in the text. Normalize \r\r\n, then \r\n, then a lone \r to LF. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(subtitles): require real SRT timestamps on timing lines The converter ends a cue at the next timing line, and the timing-line check only looked for an arrow after a colon and a separator. Cue text such as "Meet at 10:30. --> go now" split the cue, and Jellyfin windowing then failed to parse it with a 500. Both sides must now be SRT timestamps. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…o-Server#1405) * fix(jellycompat): serve Jellyfin Web preference discovery routes * fix(jellycompat): preserve preference IDs during route normalization
…Rip (Silo-Server#1358) * feat(playback): serve original SRT sidecars to clients that parse SubRip Clients that send subrip_sidecar_v1 receive external and downloaded SRT tracks as the original .srt file (with original=1) instead of the WebVTT conversion, so SRT features WebVTT cannot express survive. A bare .srt request keeps its historical WebVTT response for the frozen v1 route. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(playback): keep subrip_sidecar_v1 off the frozen /api/v1 surface The playback and stream handlers are shared with /api/v1, so the new feature was advertised, negotiated and served there too. Requests through /api/v2 now carry a context mark: only they advertise and negotiate subrip_sidecar_v1 and receive original SRT bytes for .srt?original=1. A v1 start drops the feature, which keeps it out of the attempt's replans, and the v1 route keeps answering .srt with WebVTT. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(playback): keep /api/v2 SRT attempts from continuing on /api/v1 An attempt that negotiated subrip_sidecar_v1 on /api/v2 could still be replayed or replanned through /api/v1, which published its .srt?original=1 plan on a route that serves WebVTT for those URLs. v1 now refuses to continue such an attempt with the existing 409 playback_attempt_reused error. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(playback): label original SRT as UTF-8 only when it is Original SRT bytes were always served as charset=utf-8, which mislabels a legacy-encoded file. The charset is now declared only for valid UTF-8, and HEAD, which does not read the bytes, declares none. The protocol doc's feature-list note names plan_source_duration_v1 instead of "the last one". Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(playback): refuse a v2 retry of a start that /api/v1 negotiated A start first sent to /api/v1 with subrip_sidecar_v1 stores WebVTT after v1 drops the feature, but its request digest still matches the same body retried through /api/v2. That retry replayed the WebVTT plan to a client the protocol promises original SRT. It now gets the existing 409 playback_attempt_reused, the same binding the v1 direction already has. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(playback): decide the SRT representation from what the attempt published A server that predates subrip_sidecar_v1 stored the token verbatim beside WebVTT URLs. Trusting the stored token let the new cross-surface guard refuse such a legacy v1 attempt mid-playback, and let every replan other than a seek reanchor switch its SRT URLs to .srt?original=1. Both now read the representation the attempt's current plan actually published, and every replan keeps it. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * docs(playback): every replan keeps the published SRT representation Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(playback): keep the negotiated SRT surface before any SRT is published With no SRT track in the current plan, the published representation says nothing, so a v1 replan could continue a v2-negotiated attempt and publish .srt?original=1 for a track that appears later. The stored token now decides in that case: this server drops it from every /api/v1 start, so it marks only attempts negotiated through /api/v2. The protocol doc also notes that a convert artifact stays WebVTT. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(playback): record the /api/v2 offer of subrip_sidecar_v1 on the attempt Falling back to the stored token when no SRT was published misclassified a legacy /api/v1 attempt whose old server stored the token verbatim, so its next v1 retry or replan got a 409. /api/v2 start and replan decisions now advertise the v2 feature list, which the attempt persists as its StartResponse. Before any SRT is published, an attempt counts as negotiated only when the client sent the token and this server offered it, and replans and realtime events use the same rule. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(playback): replay a terminal start before the surface check A terminal start is persisted without the /api/v2 feature marker, so an identical opted-in retry through /api/v2 failed the surface check with a 409 instead of replaying the terminal idempotently. A terminal publishes no subtitle URLs, so it now replays before the check runs. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(playback): omit a realtime subtitle track when its representation is unknown A failed attempt lookup made the subtitle-ready notifier fall back to the default features, which could publish a .vtt URL to a session that negotiated original SRT. The resolver now returns the error, and the event omits the track so the client refetches its plan. A session with no v3 attempt still uses the defaults. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…ilo-Server#1350) * fix(watchsync): classify Trakt and Simkl rate limits and pace writes Trakt and Simkl answered every non-2xx the same way, so a 429 never became a RateLimitedError and the sync service never deferred the account. Both providers now map 429 (and Simkl's documented rate_limit/RATE_LIMIT bodies) to RateLimitedError with the provider's Retry-After, retry short waits in place, and pace authenticated writes to one per second per access token as the providers document. Retry-After parsing and the per-credential write limiter move into internal/watchsync so MDBList, Trakt and Simkl share them. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(apiv2): answer provider rate limits with 429 and Retry-After A provider 429 during a watch-provider call, such as Trakt's slow_down answer to a device-code poll, now reaches clients as a rate-limited problem carrying the provider's wait instead of an internal error. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(watchsync): defer limiter refusals and Trakt device-code 429s A local limiter refuses at once when the next slot lies past the context deadline. That was an ordinary error, so watched exports that were never sent were marked failed. The refusal is now a RateLimitedError, which leaves the work pending for a later run; a cancelled context still returns its own error. MDBList's limiter gets the same treatment. Starting Trakt device authorization now reports a 429 as a rate limit with the provider's Retry-After, like polling and token refresh. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…er#1351) * fix(watchsync): classify Trakt and Simkl rate limits and pace writes Trakt and Simkl answered every non-2xx the same way, so a 429 never became a RateLimitedError and the sync service never deferred the account. Both providers now map 429 (and Simkl's documented rate_limit/RATE_LIMIT bodies) to RateLimitedError with the provider's Retry-After, retry short waits in place, and pace authenticated writes to one per second per access token as the providers document. Retry-After parsing and the per-credential write limiter move into internal/watchsync so MDBList, Trakt and Simkl share them. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(watchsync): keep the MDBList API key out of error text The MDBList watch-sync provider puts the API key in the request query string. Transport errors wrap *url.Error, whose message includes that URL, and those errors reach the connection's last_error, sync run errors shown in the web UI, and logs. Build the keyed URL only inside the single request helper, sanitize transport and request-build errors with logredact.SanitizeURLError, and mask the key (raw or escaped) in error-body excerpts. Cancellation and timeout classification is kept. Same bug class as Silo-Server#692, which covers the separate internal/mdblist discovery client. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(apiv2): answer provider rate limits with 429 and Retry-After A provider 429 during a watch-provider call, such as Trakt's slow_down answer to a device-code poll, now reaches clients as a rate-limited problem carrying the provider's wait instead of an internal error. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(watchsync): defer limiter refusals and Trakt device-code 429s A local limiter refuses at once when the next slot lies past the context deadline. That was an ordinary error, so watched exports that were never sent were marked failed. The refusal is now a RateLimitedError, which leaves the work pending for a later run; a cancelled context still returns its own error. MDBList's limiter gets the same treatment. Starting Trakt device authorization now reports a 429 as a rate limit with the provider's Retry-After, like polling and token refresh. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(watchsync): keep error classification when masking the MDBList key When a transport error quoted the keyed URL in its own message, the masking fallback replaced it with a plain error, so errors.Is and errors.As no longer matched cancellations, deadlines, or net.Error timeouts. The masked error now answers Is and As from the original, without an Unwrap that would expose its message. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…rver#1363) * fix(logging): keep Redis off the request path for log writes With Redis configured, every captured slog record and every request through the activity middleware ran a synchronous RPUSH on the caller's goroutine. When the push failed, the writer logged a warning through the slog default, which is the operational log handler, so the warning was pushed through the same writer. While Redis refused connections one slog.Info never returned and every request that logged stalled. Every node now uses the path single-node installs already had: operational and activity entries go into a bounded in-memory buffer (logstream.Buffer), and a consumer on the same node inserts them into Postgres in batches and publishes each row to the admin live tail. Writers never block or log. A full buffer or a failed batch insert is counted in silo_log_writer_dropped_total{stream,reason}. The opslog consumer's own warnings go to stderr and OTLP only, so a failing batch cannot queue entries about itself. Cross-node publishes for a batch share a two-second deadline, so an unreachable Redis cannot stall the consumer row by row. The consumers outlive appCtx and flush from deferred calls in main, so records logged during graceful shutdown are persisted. The Redis writers, the Redis drain loops and their dedicated Redis clients are removed. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(logging): retry failed log inserts and publish to other nodes off the drain Review of the first revision found that a Redis outage still throttled Postgres persistence and that a Postgres blip dropped whole batches. The consumer published each inserted row to other nodes inline, under a two-second budget per batch. With Redis refusing connections or unreachable, one consumer persisted only 25 to 35 entries per second, so any busy node overflowed its buffer. logstream.Hub now hands rows to local subscribers inline and queues them (1,024 rows) for its own goroutine, which publishes to the event bus with a two-second deadline per publish. The consumer never waits on Redis. Rows that miss other nodes' live tails are counted in silo_log_tail_publish_dropped_total{stream,reason}, and the hub logs only when publishing starts failing and when it recovers. main closes the hub after the log consumers flush and before the event bus closes. A failed batch insert was dropped on the first attempt. logstream.Drain now retries connection and availability errors with backoff from one to ten seconds while the node runs, so the buffer holds entries through a Postgres restart or failover. A batch the server rejects is dropped at once, a stopping node makes one attempt per batch, and every attempt has a ten-second deadline. The metric help and the observability doc no longer claim that stderr and OTLP receive every record: audit entries have no other copy. The doc also covers the lost log.Fatalf record on exit and the leftover Redis list keys from earlier versions. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(logging): keep log rows beside a value Postgres rejects Without the Redis JSON round trip, raw strings reached the multi-row INSERT, so one request with a 0xFF byte in its path or User-Agent made Postgres reject the whole batch and dropped up to 100 other audit entries. Both consumers now replace invalid UTF-8 and NUL bytes in text fields, and \u0000 escapes in attrs JSON, with U+FFFD. When Postgres still rejects a value (SQLSTATE class 22, such as a client address that is not an inet), the drain inserts that batch one row at a time so only the rejected rows are dropped and counted. opslog.Handler now encodes non-scalar attribute values when the record is logged. The consumer encoded them seconds later on its own goroutine, so a caller that changed a logged map could race that encode. A write refused as read-only (25006) after a switchover is now retried. The request pipeline test serves 100 requests. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…lo-Server#1400) TestControlSocketReconnectResumesOnlySameOwnerAndInstallation flakes in CI with "second connection did not receive the command". The client's dial returns once the upgrade response arrives, before the server registers the new connection with the realtime hub. The hello helper waits for HasRealtimeConnection, which the first connection already set, so it returns at once and the command can still be routed to the first connection's lane. Wait until the handler holds a new lane for the session before sending. With a 100 ms sleep added before Register to widen the window, the old test failed 5 of 5 runs and this one passes 5 of 5. Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…Silo-Server#1364) Every v1 and v2 progress sample called SyncNow, which reconciles the whole node's live-session snapshot: one transaction of N+5 statements for N local sessions, then a sessions.replaced push to every admin socket and a playback-sessions invalidation on every replica. With the web and Android clients reporting progress every 10 seconds, the shared admin view was rewritten and its caches dropped once per heartbeat per stream. Position does not need that. The periodic reconcile tick already carries it, and jellycompat progress never triggered a sync. A progress sample now syncs only when it flips the pause state, so the admin view still shows pause and resume at once. Start, replan, stop and abort keep syncing from their own call sites. Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…Silo-Server#1365) * perf(playback): refresh the taste profile only when a play changes it Every successful progress heartbeat on the native v1/v2 playback path and the Jellyfin progress report marked the viewer's taste profile stale and queued a full profile rebuild. A client reporting every 10 seconds produced 360 stale-mark UPDATEs and up to 360 rebuild requests per viewer-hour. Mid-play, a heartbeat can at most move the item between watch-weight bands, and the stop that ends the play refreshes the profile with the final position anyway. Progress writes now report whether they flipped the item to watched (userstore.UpdateProgressReportingCompletion reads the prior row only for samples past the watched threshold). Native playback refreshes on that crossing and, as before, on stop; the Jellyfin report refreshes on the crossing and on the Stopped report. Ratings, favorites, watchlist and watched-state changes keep their existing triggers, and expired sessions still refresh through the stop finalizer. MarkProfileStale now writes only when the profile is not already pending (stale_at IS NULL OR stale_at <= updated_at), so a repeat mark before the sweep consumes the first one is a no-op instead of a row rewrite. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(jellycompat): refresh the taste profile on a Stopped report without a position The Stopped-report refresh sat inside the progress-persist block, which runs only when the report carries a positive PositionTicks. Some clients send Stopped without a position (or with 0), and Jellyfin teardown stops the native session without running its stop finalizer, so such a play got no taste-profile refresh at all once heartbeats stopped refreshing. The Stopped report now refreshes whether or not it writes progress; the completion crossing still refreshes from inside the persist block, and a report that does both refreshes once. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…hing (Silo-Server#1366) * perf(progress): index Continue Watching's resume-point predicate The in-progress listings behind Continue Watching, Next Up, Jellyfin Resume, and v2 GET /progress select rows with position_seconds > 0. Completed rows hold position 0 since the reset_completed_progress_position migration, and a rewatch keeps completed = TRUE with a live position, so that is the right predicate. The partial index built for these lists is still WHERE completed = FALSE, which the predicate does not imply. The planner cannot use it, so every page read all of the profile's progress rows and sorted the survivors. Add idx_uwp_profile_resumable on (user_id, profile_id, updated_at DESC) WHERE position_seconds > 0, built concurrently, and drop idx_uwp_profile_in_progress. The only queries whose filter implies completed = FALSE either look up a row by primary key or also require position_seconds > 0, which the new index serves with fewer rows. On a profile with 19,700 progress rows (600 resumable), a Continue Watching page went from 19,700 rows read, 1,759 shared buffers, and a sort to 100 rows read and 202 buffers via an ordered index scan. A DB-backed test EXPLAINs the SQL each in-progress listing sends and fails if a page reads anything but the profile's resume points. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * perf(progress): re-measure the resumable index and guard its rollback Re-ran the before and after measurement on a fresh synthetic profile whose rows are scattered through the heap the way interleaved writes leave them, and replaced the migration comment's figures with those: a Continue Watching page goes from reading 20,752 progress rows and 1,850 shared buffers to 102 rows and 310 buffers. The comment now also lists the episode catalog's in_progress rule and Next Up's resumable list among the readers, says that only the ordered Continue Watching page stops at its LIMIT, and states why nothing can use idx_uwp_profile_in_progress: every WHERE clause that implies completed = FALSE is a primary-key lookup or an ON CONFLICT guard. Down gets the same invalid-index guard as Up. Without it, an interrupted concurrent rebuild of idx_uwp_profile_in_progress survives a retry as an invalid index that IF NOT EXISTS skips. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * perf(progress): correct why the in-progress index can go The migration comment said no query could use idx_uwp_profile_in_progress. The Jellyfin IsResumable browse filter does: when it plans as a join over the profile's rows, it reads the profile's unfinished rows through that index. The filter also requires position_seconds > 0, so after the migration it reads the new resumable index instead. On the benchmark data that is 509 rows instead of 709 and 2.0 ms instead of 8.1 ms for a movie browse page. Say so in the comment, and name the Audiobookshelf COALESCE(completed, FALSE) predicate, which the planner cannot match to the old index. The guard test's plan walker only took index names from nodes that carry the relation name. A Bitmap Heap Scan carries only the relation, and its Bitmap Index Scan child carries only the index, so a bitmap read of the new index counted as no index and would fail the guard. Collect names from the bitmap children too, and cover it with a plan-JSON unit test. Refresh the Continue Watching figures in the comment from a fresh measurement on the same synthetic profile shape. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…ilo-Server#1367) * fix(downloads): stop series monitors re-registering deleted episodes Deleting a managed episode hard-deletes its downloads row, and monitor sync only skipped episodes whose row still existed. The next sync registered the deleted episode again under a new ID. iOS delete_watched retention deletes watched episodes, so each one came back on a later monitoring run; capped monitors spent their storage budget on episodes the user had deleted; and manual deletes of monitored episodes did not stick. Deleting an episode of a series the device monitors now records a (monitor, episode) exclusion in the same statement. Sync reads the device's existing entries and then its exclusions inside the monitor lock, before the storage cap is applied, so a deleted episode neither returns nor consumes the budget. An explicit episode, season or series download clears the exclusion, and deleting the monitor drops its exclusions by cascade. delete_watched monitors also skip episodes the profile has completed, using the per-user progress store's batch lookup, so the device never fetches a file its retention pass is about to delete. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(downloads): forget monitor deletions when the monitor is created again A create for a series the device already monitors returns the existing monitor. Clients patch a monitor they know, so a create that finds one means the client lost track of it. The iOS app does this after an upgrade: it deletes the downloads an earlier version registered, which now records exclusions under that version's monitor, and hides the monitor. When the user monitors the series again, the create answered with the old monitor and its exclusions kept the back catalog from registering. The native create and the bridge re-monitor now drop the monitor's exclusions, so re-monitoring a series starts from a clean slate, as a new monitor does. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * docs(downloads): correct the watched-progress lookup chunk comment The bundled SQLite allows 32766 bound parameters, not 999. Chunking stays, as it does for the catalog's playable targets, because the SQLite user store binds one parameter per episode and a bridge sync passes a whole series. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(downloads): keep managed deletes working while a monitor is deleted A managed delete of a monitored episode joined the monitor from its snapshot and inserted the exclusion. If the monitor was deleted while the insert's foreign-key check waited on the monitor lock, the whole statement failed with SQLSTATE 23503 and the download row stayed. The statement now locks the monitor FOR KEY SHARE: a monitor deleted meanwhile yields no row, so the delete records nothing and succeeds. It still deletes the downloads row first, the order a device delete's cascade takes. Registration now reads the device's existing entries and the monitor's exclusions in one statement. The delete commits the row removal and the exclusion together, so one snapshot sees one or the other, and a monitor held at its storage cap no longer pays an extra statement on every sync. The delete_watched filter fails open: if the progress lookup fails, sync logs a warning and registers without it, as the storage gate does. The retention pass deletes any finished episode registered that way, and the exclusion keeps it from coming back. The subscription test helper defines the exclusions table inline instead of reading the migration file, and the migration comment now says that re-creating a monitor clears its exclusions. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * docs(apiv2): say that re-creating a monitor clears its recorded deletions createDownloadSubscription returns an existing monitor with its options unchanged, but since the monitor exclusions landed it also forgets the episodes the device deleted under that monitor. The published summary still said the monitor is returned "without changing it". Update it and regenerate the OpenAPI document, the web contract types and the fixture that carries the contract digest. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * test(downloads): poll for the monitor lock wait on a ticker waitForLockWait polled pg_stat_activity back to back, so a lock wait that never appeared would query the shared test database nonstop for the full 10 s timeout. Check every 5 ms instead and stop on the context deadline. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * docs(apiv2): shorten the createDownloadSubscription summary Generated clients and docs render an operation summary as a one-line title, and the create summary had grown to several clauses. Keep the summary to what the operation does and move the existing-monitor behavior and the resend guidance into the operation description. Regenerated the OpenAPI document, web types and the system-info fixture's contract digest. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * docs(downloads): say when a managed delete can deadlock with a device delete The DeleteManaged comment said a device delete's cascade normally takes the same row order. Postgres fires the cascade triggers in name order, and their names end in OIDs compared as text, so the order depends on the database, not on whether it was migrated in order. Say so, and that the loser gets a retryable deadlock error. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(downloads): keep a delete's exclusion when a re-download clears it late An explicit download of a deleted episode registers its row, then clears the monitor's exclusion for the episode in a separate statement. A delete of the new row that committed in between left the device with neither the row nor the exclusion, so the next monitor sync registered the episode again even though the user deleted it last. The clear now forgets an episode only while the device still holds its row, checked in the same statement. It locks that row FOR KEY SHARE SKIP LOCKED: a delete that has run its statement but not committed still holds the row, and a plain read would see the row and erase that delete's exclusion. The clear also skips an exclusion another transaction holds, so it never waits on a row lock. It is best-effort cleanup after the download succeeded, and waiting could hold the download behind a delete that is itself waiting on the monitor lock behind a sync. A skip keeps the exclusion, the direction the clear already tolerates. DeleteManaged's statement moves into a constant so the test can run it in an open transaction and hold a delete between its statement and commit. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…eads (Silo-Server#1371) * perf(plugins): stop re-hashing a verified plugin binary on manifest reads Every manifest read went through ArchiveCache.Ensure, which reads and SHA-256 hashes the whole plugin binary under the cache's global mutex. The admin plugin list reads the manifest of every installation on each load, and a plugin asset request reads it twice, so each navigation hashed tens of megabytes serially per installed plugin. Add ArchiveCache.Manifest for callers that only read the manifest or serve packaged assets. It skips the hash while the binary is the same file, with the same size and mtime, that this process last hashed against the checksum the manifest still names. Any change, a new release path, or a failed check sends the read through the full check and repair. Launch paths (Start, the connection test, preload) keep calling Ensure, which still hashes the binary on every call, so the tamper check in front of exec is unchanged. The hash now streams the file instead of loading it whole, which drops a launch check's allocation from the binary's size to about 37 KB. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * test(plugins): cover the file identity check on manifest reads Review follow-up for the manifest read change. The archive cache test now renames a different binary with the verified size and mtime over the installed one and expects Manifest to hash and repair it. Without the os.SameFile comparison the test fails; before, nothing covered it. Rename the service helper installedManifest to loadForManifestRead so it no longer shares a name with a local variable in doEnsureClient, and rename the benchmark case launch to ensure, since it measures ArchiveCache.Ensure and starts no process. Correction to the previous commit message: startup preload checks and rehydrates files without executing them. Launches go through start or the connection test, and both call Ensure. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…er event (Silo-Server#1372) * perf(plugins): index event subscribers instead of reading the store per event The plugin event dispatcher listed enabled installations and then read each installation's capabilities for every event it saw, on every replica: 1+N store reads per event even when no plugin subscribes to that event. Progress, playback, and admin events all pass through it, so busy servers paid hundreds of plugin-table reads per second per replica for events that no plugin receives. The dispatcher now keeps an in-memory index from event name to subscribers, built from the store on first use. The index is dropped by a lifecycle hook that SetEventDispatcher registers, so local install, enable, disable, upgrade, and uninstall take effect on the next event. Other replicas drop theirs when cache.EventPluginsChanged arrives on the admin channel, and an index older than DefaultLifecyclePollInterval is rebuilt so a missed publish costs at most one interval. A generation guard, the same one the installation cache uses, keeps a rebuild that races an invalidation from being stored. An index missing an installation whose capabilities could not be read serves that event but is not kept. Delivery semantics are unchanged: the first matching event_consumer.v1 capability per installation receives the event, builtin installations are skipped, and TargetPluginID narrows delivery to one subscriber. Payloads are decoded only when an event has a subscriber. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * test(plugins): measure dispatcher reads with the real user-state envelope The hub measurement in TestDispatcherStoreReadsPerEvent used an invented progress_updated event. Use user_state.changed with the payload and scope fields a progress sync publishes, so the measured envelope matches what reaches the dispatcher in production. Also tighten the index field comment. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(plugins): drop the dispatcher index before other lifecycle hooks SetEventDispatcher appended its invalidation hook after the plugins_changed publish hook. When this replica's own publish came back before that hook ran, the dispatcher rebuilt its index on the echo and the hook then dropped it, so one local change cost two rebuilds. The hook now runs first: a local change with an echoing bus costs one rebuild (5 store reads with the test fixture, down from 10), and events dispatched while resident reconcile and provider reloads run already see the change. The tests now assert store reads for a remote plugins_changed and for a local change with and without an event bus. The racing-invalidation test is renamed to say what it checks, and the startup OnLifecycleChange comment in main.go states why the call remains. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…ilo-Server#1373) * perf(catalog): read series and season user_data from the SQL rollup The native series detail, season detail, and season list requests loaded every available episode row of the parent and then read the viewer's progress and completed history in 500-ID chunks, only to count watched and in-progress episodes. On a 1,200-episode series that is 1,200 episode rows plus six progress and history statements per request. The Postgres user store already computes the same counts in one aggregate (SeriesEpisodeWatchCounts), and jellycompat uses it for series. Add a season-number variant for the season list and a season-row variant for the season detail, sharing the existing query, and use them in the native paths when the viewer's store implements the rollup. Stores without it (the SQLite user store) and rollup query errors keep the episode fold. The season list also takes each season's episode count from the rollup, so it no longer lists episodes at all. Measured with a pgx tracer on a 1,200-episode fixture: series detail goes from 19 statements / 2,739 rows to 13 / 1,365, season detail from 13 / 334 to 11 / 212, and the season list from 13 / 2,749 to 7 / 1,387. A DB-backed test proves value-for-value parity with the fold across watched, partially watched, history-only, hidden, unwatched, unavailable, and other-profile episodes; a unit test pins the rollup path in CI. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * perf(catalog): roll up season-by-number user_data in SQL GET /catalog/series/{id}/seasons/{num} still folded per-episode progress after the series, season detail and season list paths moved to the SQL rollup. When the season's episodes are linked to its row, read user_data from SeasonEpisodeWatchCounts; the by-number fallback keeps the fold. The parity test now covers this request as well. A canceled request no longer logs a rollup warning, and the season list returns the error instead of retrying with the fold. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…#1374) * ci: run the DB-backed query-budget pins against Postgres make test-go runs without SILO_TEST_DATABASE_URL, so every DB-backed test skips and go test reports the skip as a pass. The statement-count and query-plan pins that guard item detail, listings, metadata refresh, home sections, settings resolution and book scans have never run in CI, so a change that turns a pinned budget into an N+1 merges green. Add a Go DB pins job that migrates a fresh pgvector/pgvector:pg18 database with cmd/silo --migrate-only and runs the tests listed in scripts/ci/db-pins.txt through make test-db-pins. The runner in scripts/ci/dbpins reads go test -json and fails when a listed test is missing, skipped or failing, and on any skipped subtest: without a database, two of the pinned tests pass at the top level because only their subtests skip. The job is separate from go-test so it adds nothing to that job's time, and because the full schema needs pgvector, which go-test's image lacks. CONTRIBUTING.md lists the new target in the pre-PR gate. Refs Silo-Server#599. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * ci: scope the DB pins in CONTRIBUTING and pin card play targets CONTRIBUTING.md listed make test-db-pins in the mandatory pre-PR gate, which would ask every contributor, even on a web-only change, for a migrated pgvector database. Move it out of the gate: CI runs it for every pull request, and contributors run it when they change database or query code or add a pin. Say how to create and migrate the database with the Compose defaults. Add TestPlayableTargetResolverProfileStateAvailabilityAndAccess to the list: leaf cards and a validated anchor hint must resolve without a watch-progress query. It passes on a freshly migrated database. The runner now drains go test's output if reading the event stream fails, so a go test still writing cannot block on a full pipe before Wait returns. Refs Silo-Server#599. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…o-Server#1376) * perf(artwork): keep revisioned artwork URLs stable for a UTC day Clients and CDNs cache artwork by full URL, so every new URL for unchanged bytes costs a download the client already has. Local signed URLs rolled every 15 minutes. S3 presigned URLs carried the signing second and Cloudflare token URLs the issue second, so they changed on every resolve and differed between replicas. A revisioned key names immutable bytes: its path is the identity and its query only authorizes. At the default lifetime such a URL now holds for its UTC day. The local signer uses a day-long issuance bucket, and S3 presigns at the start of the day with X-Amz-Expires set to the day plus the TTL, shortened to fit the SigV4 seven-day limit. The WAF rule fixes a token's lifetime from its timestamp, so token timestamps are truncated to a quarter of the token TTL instead. Every replica mints the same URL for the same window. Mutable keys, short-lived capabilities such as avatars and chapter thumbnails, and job artifact downloads keep their current buckets. Over a simulated day of per-minute resolves through the production resolvers, distinct URLs for one revisioned key drop from 97 to 2 (local), 1440 to 2 (S3 presigned), and 1440 to 33 (three-hour Cloudflare token), and a second replica resolving 20 seconds later now always agrees. A leaked revisioned URL now works for up to a day plus the TTL instead of the TTL plus 15 minutes. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * docs(artwork): state URL stability limits from review Record that replica-identical presigned URLs rely on the shared static access key, that lowering s3.metadata_presign_expiry shortens the TTL but not the day window, and how the Cloudflare Token TTL sets how often token URLs change. Give the revisioned cache-policy test a few seconds of slack so a run that crosses midnight UTC cannot fail it. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Silo-Server#1377) * perf(access): stop re-reading legacy settings on every profile request Every profile-scoped request resolved hidden libraries from the canonical ui.disabled_library_ids row and, when the profile had none (the default), also read the frozen legacy disabled_library_ids account row. Home section requests did the same for ui.next_up_mode, up to three times per request. The legacy settings endpoint stopped accepting both keys at the settings cutover, so those reads return values that can no longer change. A Go migration on Postgres and SQLite schema v26 copy each legacy value onto every profile that has no canonical row. They use the settings planner, so each stored value matches what the fallback read produced. Stored rows win, and the legacy rows stay in place. With the values materialized, the fallback reads go away. ui.next_up_mode now resolves in the same settings read as the rest of the viewer scope and rides on access.Scope (excluded from its JSON, so access fingerprints do not change), and section fetchers take it from the request scope. Measured with a MaxConns=1 pgx tracer: the auth prelude for a default profile drops from 6 to 5 statements (7 to 6 for an access-group member), and a NextUpMode call inside a request drops from 2 statements to 0. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * docs(middleware): correct the RequireAuth session cache comment RequireAuth checks session validity with the SessionValidator on every request, and the production validator is an uncached query. The comments claimed an in-memory cache, which hides the per-request cost and suggests a revocation might lag. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * test(sections): pin zero next-up reads on Home section requests Measure the Home half of the identity-prelude change through SectionHandler.HomeSectionItems instead of deriving it from per-call counts. With a scope from the production viewer resolver, the Continue Watching and Next Up section requests make no next-up mode reads for either mode. On the base commit the same requests made 2 to 4. Review follow-ups on the retired-fallback migration: - read both legacy keys in one pass with key = ANY($1) - log unconvertible values on SQLite, as Postgres already does - correct the comment that said user_settings will be dropped; other keys still use it Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * refactor(database): run the retired-fallback backfill as a Goose SQL migration New database changes belong in timestamped Goose SQL migrations under migrations/sql, but the retired-fallback backfill was registered as a Go migration. It now lives in 20260924052112_materialize_retired_settings_fallbacks.sql, created with make migrate-create, and the Go migration is removed. The SQL reproduces settingsmigrate.PlanRetiredFallback, which the SQLite user store still runs in Go at schema v26. A hidden-library list converts only in the shape Go's encoding/json decodes into []int; nulls and non-positive ids are dropped, repeats are removed in first-seen order, and a list of more than 512 ids is skipped. A next-up mode is trimmed of Go's Unicode whitespace and kept only when it is "combined" or "separate". Stored rows still win, legacy rows stay, and a re-run changes nothing. TestRetiredSettingsFallbacksMigrationMatchesPlanner seeds 53 accounts covering every shape the planner handles and fails if the migration writes anything other than what the planner plans. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(settings): keep padded legacy next-up modes at the default The retired fallback returned the legacy next_up_mode string unchanged, and the section fetchers compared it with "separate" exactly, so a padded value such as " separate " never enabled the separate Next Up row. The backfill trimmed before validating and would have written "separate" for it, turning the row on for that profile. Only an exact member is converted now. The planner rejects a padded value before the generic write normalization trims it, so the SQLite v26 path skips it, and the SQL migration matches the two members exactly. A skipped value leaves the profile at the default, which is what the fallback gave it. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * docs(settings): describe what unconvertible legacy values did before The comments said a padded legacy next-up mode behaved as the default. It did not: the section fetchers matched the raw value exactly against each member, so a padded or unknown value showed next-up in neither place. The contract has no such state, so those profiles now get the default. Say so, and say that a hidden-library list over the contract's 512-id limit is not copied, which makes those libraries visible. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…ilo-Server#1378) * feat(jellycompat): add request metrics and a server span per request The Jellyfin-compatible listener runs on its own http.Server, so the native request metrics never see it. Its only per-request signal was an INFO log line, which gives no p95 by route, no error rate and no split by client. It also started no server span, so every Postgres, Redis and S3 dependency span in a compat request became its own root trace. observeCompatRequest runs right after the request ID middleware and records each request once, whatever answered it: - silo_jellycompat_requests_total{route,method,status_class,client} - silo_jellycompat_request_duration_seconds{route,method} route is the chi template read after routing (or "unmatched"), method is folded into the standard set, status_class adds "hijacked" for the session socket, and client is a fixed family parsed from the MediaBrowser Client field and then the User-Agent. The playback and transfer media routes and /socket are counted but not timed, because their duration is the client's viewing or connection time. HLS manifests stay timed. Each request also opens a fresh server span, so dependency spans join one trace. Like native v2, a caller's trace context cannot choose the trace ID or the sampling decision. The request logger's writer becomes statusResponseWriter, shared with the new middleware, and hijacks through http.ResponseController so the socket upgrade reaches the connection through every wrapper. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(jellycompat): time requests per client and bound socket traces Resolve the review findings on the compat request metrics: - Add silo_jellycompat_client_request_duration_seconds{client} over the timed routes, so operators can see which client is slow without a route by client histogram (about 200 series instead of about 26,000). - Run the session socket's periodic checks outside the request's server span. A socket left open all day no longer grows one trace without bound; the check before the upgrade still joins the request trace. - Match client families without allocating: fold ASCII case into a stack buffer and read at most 512 bytes of the Client field and User-Agent. An 8 KB User-Agent no longer costs an 8 KB copy per request. - Add the moonfin family. - Document the client histogram, that media and socket requests are counted when they finish, and the socket trace boundary. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Deleting a group revoked each moved member's sign-ins with three statements per member while holding the group-writer lock. The new RevokeSignInsForUsersInTransaction does it in three statements total, and the single-user form now delegates to it. DeleteMovingMembers also wraps its Begin error with context. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…s-new-content-only fix(notifications): skip releases for content added long before it matched
Docs in this repository are for people and agents changing the code. Guides for installing, configuring, and running Silo belong in the user manual on siloserver.org. CONTRIBUTING.md and AGENTS.md now say so, the wiki index takes no new pages or sections, and the README's documentation links point to the manual. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…ves-members-1402 fix(access): move a deleted group's members into the default group
…ver#1516) Log migration status, per-version progress, and rollback candidates and results. Keep heartbeats active during lock waits and name registered Go migrations. Verified migration and rollback behavior in an isolated sandbox; focused race checks and CI passed. AI review: gpt-6-astra in T3 Code using the Codex provider.
…ides-to-manual docs: send operator guides to the siloserver.org manual
Silo-Server#1627) * fix(notifications): skip second library copies in server channel posts Release events are per library, so a title added to a second library (a 4K library beside an HD one) produced a second "new" event and the server-wide channel announced it again. Server channel sweeps now drop events whose episode (series and episode key) or item a different library made available before the event. Profile fanout is unchanged: it already deduplicates per episode across libraries and still reaches profiles that can see only the newer library. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(notifications): count only existing libraries as earlier copies Availability rows have no foreign key to their library and outlive a deleted one, so a title from a deleted library kept counting as an earlier copy and its next arrival was never posted. The repeat lookup now ignores rows whose library no longer exists. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Silo-Server#1634) Remove scripts that nothing in the Makefile, CI, hooks, or docs uses: the native rsync + air dev setup and the root .air.toml, one-time data fixes and backfills, ad hoc SQL and benchmark scripts, a duplicate of make jellyfin-web, jellycompat-diff.sh, the k6 load test, and an unwired silo-profile test. Also remove the Continuum-to-Silo migration script, its guide, and the migrate-continuum-check target, and drop the links to the guide. Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
) * fix(catalog): match name_prefix on the sort title only The alphabetical jump matched name_prefix against the sort key or the raw title, so a movie titled "The Hobbit" with sort title "Hobbit, The" was listed under both H and T. Title sorting already orders by the sort key, and Jellyfin's NameStartsWith compares SortName, so the letter rail now matches only the sort key, which falls back to title when sort_title is empty. Browse, favorites, and query-source previews share one helper. * fix(catalog): apply sort-title prefix matching to every catalog path Review follow-up. The in-memory prefix filter (relevance search, exact-order collections, offset pages) and the watch-history source still matched the raw title, so offset and cursor pages disagreed and history results kept "The Hobbit" under T while history facets did not. Both now use the sort key. The helper reuses sortTitleKeyExpr, the episode entry executor drops its redundant title arm, test comments no longer claim the prefix LIKE is index-backed, and docs/catalog-api.md states the rule. * fix(catalog): trim sort titles like SQL BTRIM in the in-memory prefix filter Review follow-up. strings.TrimSpace also removed tabs, which BTRIM keeps, so a tab-led sort title could land under a different letter in memory than in SQL. docs/catalog-api.md notes that recently added TV also matches episode titles. * docs(apiv2): name the sort-title rule in name_prefix descriptions Review follow-up. The GET /api/v2/catalog parameter and the catalog query body field now say name_prefix matches the sort title, or the title when none is set. Regenerated the OpenAPI artifact, web types, and the system-info fixture's contract digest.
…ed (Silo-Server#1630) * fix(downloads): pick the version the profile watches when none is named A download that omits media_file_id (season, series and monitored episodes, and any client that leaves Auto to the server) always got the highest-resolution file, with ties going to database order. A user who watches a different version, such as a 1080p file with a sidecar .srt next to a 2160p file without one, got the other file offline. Native requests now pick from the profile's watch history: the file last played for that movie or episode, then for an episode the version closest to the series' most recent play (edition, resolution, HDR, video codec) within the profile's latest 200 progress rows, then the highest resolution with the lowest file id breaking ties. The v1 bridge keeps its highest-resolution default, so a repeated v1 request never swaps files a device already has. Items with one file skip the history reads. * fix(downloads): only pick versions the profile may play An automatic pick could register a version the profile's library access or quality ceiling forbids, leaving an entry that looks ready but can never be served. Candidates are now narrowed with the request's access filter first; when none qualify the full list is kept and serving reports the refusal as before. * docs(downloads): describe the fallback when no version is playable An automatic pick narrows candidates to versions the profile may play but keeps all of them when none qualify, so the create or the later file request can still be refused. The doc claimed that could not happen. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Quick <31828688+Quick104@users.noreply.github.com> Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…1635) * docs: move operator guides out of the server repository Delete the operator guides that the siloserver.org manual now owns: docs/wiki/, the S3 setup guide, and the monitoring and profiling runbooks. The contributor material they held moves into docs/architecture/: profiling bounds, node resource sampling, and observability validation into observability.md; the workload metric catalog; new collection-templates.md and media-naming.md; and NFO known limitations. docs/README.md sends operators to the manual. docs/update-to-1.0.md stays until 1.0 ships. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * docs: state the node mount cap per process and what the poster tests check Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * docs: name the node health route as /api/v1/health Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…-Server#1625) * fix(notifications): stop notifying for series removed from Home New-episode interest ignored Home removals. A series removed from Continue Watching or Next Up kept matching the continue_watching and next_up reasons, because the interest recompute only read favorites, watchlist, progress and watch history, and nothing queued a recompute on removal. The recompute now follows Home: an active series drop clears both reasons, a per-card Continue Watching dismissal clears continue_watching while the dismissed progress is unchanged, and a per-card Next Up dismissal clears next_up while the dismissed episode is still next. Favorites and watchlist still notify. Dismissal writes, drop changes (Home handler and watch-provider sync) and the first progress write of a new watch session queue a recompute. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(notifications): recheck Home removals after the session gap A resume within ten minutes of the previous progress write was not a new watch session, so it queued no recompute even when it lifted a series drop or a Continue Watching dismissal made in between. Interest stayed cleared until the next state transition or the daily rebuild. A removal now queues a second recompute after the session gap. A resume later than that starts a new session, which queues its own recompute. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(notifications): match Home's Next Up card and read dismissals once per rebuild A per-card Next Up dismissal compared against the next catalog episode after the progression cursor. Home shows the first episode after its anchor that has a present file and has not been started, so a dismissal of that card could fail to match and next_up kept notifying. The recompute now looks up the same episode, only when the series has a Next Up dismissal. The interest rebuild read the profile's full dismissal lists once per series. It now shares one read per profile across its series. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(notifications): read a series' dismissals directly and anchor Next Up like Home The interest rebuild shared one snapshot of a profile's dismissals across all of its series, so a card restored mid-rebuild could be overwritten from the stale snapshot. Each recompute now reads only the dismissals of the series' own episodes; the Postgres store gets a targeted ListHomeDismissalsForItems, and other stores list the surface and filter. Home anchors Next Up on the most recently completed episode, not the highest. The card lookup now uses the same anchor, so a dismissal still matches after a rewatch of an earlier episode. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(notifications): drop superseded progress and recompute on every applied import Home's Continue Watching hides an in-progress episode once a later episode was completed more recently; the recompute now applies the same rule, so a stale entry no longer keeps continue_watching set. Timestamped progress writes (imports, watch sync, offline-queued client events) now queue a recompute whenever they apply. A late import can move the stamp a Continue Watching dismissal holds for without a state change or a new watch session. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(notifications): keep stamped playback ticks free and recheck restores 625d085 queued a recompute for every applied timestamped progress write, but SyncProgress routes any write carrying updated_at through SetProgressIfNewer, so a client that stamps live ticks would recompute its series every flush. Timestamped writes now use the session test again, extended to the clock: a write queues when it lands more than the session gap after the stored stamp by its own stamp or by wall time, which still covers a late import while leaving ticks free. Restoring a dismissal or undropping a series now also queues the second, deferred recompute, which repairs a concurrent recompute (such as the interest rebuild) that read the state from before the restore. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(notifications): count started episodes like Home and keep rewatch ticks free Home's Next Up treats any progress row that is completed or has a position as started, including rows a history hide covers. The card lookup built "started" from visible progress, so a started-then-hidden episode could be chosen as the card; it now uses Home's test in SQL. Timestamped writes keep completion sticky in both stores, so a stamped rewatch tick of a finished episode is no longer read as a transition. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(notifications): lift Home removals on the first write after them A resume inside the session gap right after a removal waited for the removal's deferred recompute, and a release processed in that window was decided on the stale interest. Each node now remembers a profile's last Home change for the session gap, and a progress write whose row predates it queues a recompute at once; later ticks stay free. The Next Up card lookup also skips episodes the profile's own store reports as started, since Home's Postgres progress test sees nothing for profiles whose progress lives elsewhere. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(notifications): keep sub-second Home change times and document node limits Truncating the Home-change time to the second could miss a progress row written earlier in the same second, so the lift waited for the deferred recompute. The marker keeps the full time again; a tick in the same second may queue one extra, coalesced recompute. The notifications doc now states that the marker and the deferred recompute are per node. New-test cleanups report their errors. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…rver#1326) Collection artwork was stored under fixed keys, so a replacement kept the same URL and CDN and browser caches kept showing the old image. Keys now carry a content revision in the shared artworkkey format ({imageType}/original.{revision}.webp), so new bytes get a new URL. Every upload path (admin artwork upload and source URL, v1 create/update artwork inputs, collage generation, template posters, and personal posters) now uploads and commits the replacement first, then removes only the revision it replaced, including legacy fixed keys. Cleanup outlives a canceled request. Closes Silo-Server#1258 Co-authored-by: Quick104 <31828688+Quick104@users.noreply.github.com> Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…-Server#1647) * fix(apiv2): keep sendfile for file responses and log body bytes The v2 status recorder hid io.ReaderFrom from io.Copy, so every file response went through a 32 KB user-space buffer instead of sendfile. The recorder now forwards ReadFrom and Flush, and counts the body bytes the handler wrote. The v2 request log reports them as body_bytes. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(apiv2): keep error statuses and flush errors through the recorder Forwarding ReadFrom let chi's wrapper commit 200 before an asset copy read its first byte, so artwork whose store failed at once went out as an empty, cacheable 200. The artwork and subtitle routes now hide ReadFrom, and the recorder sets 200 only once a copy writes bytes. The recorder's Flush also hid the transport's flush error from http.ResponseController; it now implements FlushError. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * refactor(apiv2): trim test wrappers and explain the recorder's Flush iotest.ErrReader already offers only Read, so the tests don't need to hide its type. Flush's comment now says why it matters: chi offers ReadFrom only over a writer that can also flush. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…-Server#1648) * fix(downloads): answer 503 when the offline artwork store fails A transport failure fetching offline artwork answered v2 clients with a 500, and an upstream 5xx or 429 with a 404, so clients treated a store outage as a server bug or as missing artwork. Both now return dependency_unavailable (503, Retry-After: 5) on v2. The frozen v1 route keeps its status codes. Artwork fetches without an injected client use a 30 s timeout, and the error no longer repeats the presigned URL. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(downloads): keep artwork failures safe to log and retry correctly - Log only the upstream status or timeout for a failing artwork store; a malformed redirect's error quoted the presigned URL. - Treat an artwork URL without a host as broken, not retryable. - Keep the 30 s artwork client when the server wires no client; SetOfflineDeps had replaced nil with http.DefaultClient. - Declare Retry-After on the artwork route's 503 and document it. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * refactor(downloads): simplify artwork failure handling - Declare the artwork 503 and its Retry-After before registering the route instead of patching the registered operation. - Use errors.AsType and fold the two broken-URL test cases into a loop. - Treat a store that stops sending mid-body as unavailable, so v2 answers 503 before the first byte, as it does for other store failures. - Assert that an upstream 403 or 404 is never retryable. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(downloads): blame only store read failures on the artwork store A failed write to the client was classified as the store being unavailable, logging a false store outage. Only a read from the store's response now counts. Also regenerate the web v2 types for the artwork 503's Retry-After header. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * test(downloads): assert both halves of each artwork upstream status Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(downloads): time out only the artwork store's waits, not the copy http.Client.Timeout covered the whole exchange, including time spent writing to a slow client, so a slow download of a large image could be cut off and logged as a store outage. The artwork client now bounds only the wait for response headers, and each read of the image body gets its own stall timeout while the copy waits on the store. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(downloads): drop image headers when artwork fails before its first byte With the stall timeout, a store that sends headers and then nothing ends the copy while the response is still uncommitted. The image's length and immutable cache policy were left on the writer, so v1's error response went out with them. They are now removed whenever nothing was written. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…Silo-Server#1655) * fix(watch-party): keep room members out of post-roll at the end of an item A web member of a Watch Party entered post-roll in the last 30 seconds of an episode. When the room then returned to its lobby, the player's room exit saw a non-foreground mode, tore playback down without navigating, and left the member on a black player page with no room socket. Skip the early post-roll trigger inside a room. The room decides what follows the item, and the member stays in the player until the room's lobby snapshot sends it back to the room page. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(watch-party): read the room from the active request The room is part of the request key the callback already depends on, so a ref adds a render of lag and nothing else. Match the other room branches in checking both the room id and its token. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…wal (Silo-Server#1657) * fix(watch-party): drop a room command with the socket that delivered it The web room connection kept the last transport command when its socket closed. On reconnect the player saw the connection come back with a command it had not applied on this connection and ran it again, projecting the old command's position forward. Every planned five-minute socket rotation made a web member correct twice: once to the stale projection, then to the fresh sync the server sends after the member reattaches. If the room had re-anchored in between, the first seek went the wrong way. Clear the command when its socket closes, as a replaced connection already does. The server sends a fresh command once the member attaches again. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(watchtogether): stop resending the room command after a reconnect sync Connect clears a member's last delivered command for every new socket, so the reconciler resent the room's last command within two seconds of each socket renewal, with its original id and execution time. The member had already been synced by the attach, and applied the old command on top of it, projected from when it first ran. Record the room's command as delivered to a member whenever the member is synced to the room. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(watch-party): drop the room command when the connection effect resets A changed authority, room token, rejoin, or replacement scope re-runs the connection effect. Its cleanup lets go of the socket before closing it, so the close handler returns early and the command survived onto the next connection. Clear it in the cleanup too. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…ilo-Server#1624) The Playback settings page described Chapter thumbnail workers as "Parallel extraction jobs per library scan". The value sets how many general workers the chapter thumbnail service starts in each server process, each handling one file at a time. Library scans have no per-scan job count. The service also starts one worker that takes only playback requests, which the general workers pick up first too, so a title being played does not wait behind queued work. Describe the setting the way the other worker settings are described, and mention that extra worker. Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
WalkthroughWarning Review details and warnings were omitted to fit the comment limit. |
JonahMMay
force-pushed
the
sync/upstream-2026-09-28
branch
2 times, most recently
from
September 29, 2026 05:13
0bec150 to
b045747
Compare
Syncs Silo-Server/silo-server main at f84d2de, which brings /api/v2 that prairie-apple and prairie-android main already call. Prairie branding is kept and everything new from upstream is rebranded: module paths, cmd/prairie, UI copy, Go string literals, data directories, Docker and compose files. Wire identifiers move to prairie.*: WebSocket subprotocols and ticket prefix, every X-Silo-* header, the plugin proto package prefix and the plugin source kind. Also restores Prairie code the squashed sync #174 dropped: the Tizen HLS tag strip, web branding assets, HEAD on transcode segments, the Redis save guard and non-fatal hub start, PRAIRIE_* env names, AVIF admin settings, the w200 collection poster rung, calibre metadata keys and the episode catalog placeholder numbering. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
JonahMMay
force-pushed
the
sync/upstream-2026-09-28
branch
from
September 29, 2026 05:59
b045747 to
08794f7
Compare
Brings the sync up to upstream 19e8184: advisory age poster badge, stable collection artwork URLs, v2 sendfile, offline artwork 503, and two Watch Party fixes. New upstream code is rebranded as in the parent merge, and the contract artifacts (OpenAPI, web types, fixtures, route inventory, ledger, settings bindings) are regenerated. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Syncs Prairie with upstream
Silo-Server/silo-servermain, about 930 commits up to19e81846. The sync brings in/api/v2, which prairie-apple and prairie-androidmainalready call. There are two real merge commits, both with upstream as a parent:08794f7bmerges upstreamf84d2dee(the bulk of the sync).0bad5e15merges the 7 commits upstream landed while this was in progress.Prairie branding is kept, and everything new from upstream is rebranded:
cmd/siloreferences, which now point tocmd/prairie.example:tags.The SDK is pinned to prairie-plugin-sdk
mainatc2f90523e166(v0.12.1-0.20260928152428-c2f90523e166).go mod tidyhas been run.mattn/go-sqlite3drops out because nothing imports it any more (see the SQLite fix below).Merge with 'Create a merge commit' — do not squash. A squash would repeat what went wrong with #174: upstream history would be lost, and the next sync would diff against the wrong base.
Wire renames (
silo.*→prairie.*)silo.playback-control.v2prairie.playback-control.v2silo.events.v2prairie.events.v2silo.room.v2prairie.room.v2silo.admin-logs.v2prairie.admin-logs.v2silo.ticket.<ticket>prairie.ticket.<ticket>X-Silo-*header (Device-Id/Name/Platform, Client, Client-Version/Build/Channel/Family, Stream-Token, Mutation-Id, Latest-Sequence, Idempotent-Replay, Restart-Required, Ingress-Token, Theme, User-/Profile-, Event/Webhook-Id/Delivery-Id/Timestamp/Signature, Tone-Map-, Transcode-, Capture-, Download-, Ebook-Conversion, and the rest)X-Prairie-*header.x-silo-device-idheader.x-prairie-device-idsilo.plugin.v1.(metrics label trim)prairie.plugin.v1.(matches the SDK proto package)source_kindsilo(v2 enum and web)prairie(what the server already sends)prairie://invite and Watch Party deep linkssilo://in upstream's new codesilo.theme-audio.v2prairie.theme-audio.v2These are deliberately not renamed: the private-use language tag
x-silo-original, which is a stored settings value that prairie-apple also uses; the OpenAPI vendor extensionsx-silo-classand friends; and thesilo-audiobooksABS virtual library ID, which Prairie always used.Client alignment: prairie-apple
mainalready usesprairie.playback-control.v2,prairie.events.v2,prairie.room.v2andprairie.ticket.. prairie-androidmainstill offers thesilo.*names, so Prairie-Server/prairie-android#31 must merge together with this PR. Without it, Android's playback-control, events and Watch Party sockets fail the subprotocol check. The SDK already usesX-Prairie-Ingress-Token, which this PR matches.Conflict approach
go vet. For files Prairie had only rebranded, the upstream version was taken and rebranded. Where Prairie's lint autofixes had deleted "unused" upstream code, or rewritten code in ways that broke upstream callers, upstream was restored. That coverscatalogseed,playback_realtime,intromarkers/compare,opslog, the settings handlers and the episode-catalog SQL builders..prettierrcand drifted to 80 columns, so upstream's.prettierrcis restored andweb/srcis reformatted at upstream's 100 columns, which should make future syncs much smaller. Each remaining hunk was resolved by hand: Prairie features were kept (Live TV, trickplay, Quick Connect, AVIF artwork props, Prairie login layout, the prairie-dusk theme, button icons) on top of upstream's rewrites.index.htmlset todata-theme="prairie-dusk"and Fraunces self-hosted infonts.css. The other curated themes are gone.docker.yml,discord-commits.ymlandci.yml.go-verifyjob with upstream's contract gates (settings bindings, playback fixtures, route inventory, migration ledger, scenario catalogs, offline routes, OpenAPI artifact, v2 fixtures). The base-diff breaking-change step is left out because it compares ~930 upstream commits against Prairiemain.Catalog.test.tsxis excluded from the vitest gate. It is upstream's known failure (WEBTEST_KNOWN_FAILURES).livetvsection), offline routes, route manifests and settings bindings. All are regenerated from the merged code.trickplay_enabledandtrickplay_supportedin v2 (additive), so the web library editor keeps the trickplay toggle now that it goes through/api/v2.QUERYmethod that upstream's route inventory does not model.Prairie fixes lost in #174 (
b5ca13b) and how each is handledRestored in this PR:
-hls_flags temp_file, no#EXT-X-PLAYLIST-TYPE:VODin the synthetic manifest, andRewriteManifestPathsstrips#EXT-X-INDEPENDENT-SEGMENTSand…TYPE:VOD.EVENTis kept for upstream's remount timeline.index.htmlhad been reset to<title>Silo</title>, the arctic-frost theme and Google Fonts. It is restored.favicon.ico,apple-touch-icon.png,maskable-icon-512.pngandweb-app-icon-192/512.pnghad replaced Prairie's (the same thing happened in android). They are restored, and the straysilo-icon-1024.pngis removed.assets/icon.pngis restored.@fontsource-variable/frauncesdependency is back.DockerfileandDockerfile.devwere upstream copies that built./cmd/silointo/silo, so every Docker Image run onmainhas failed since feat: sync silo-server main upstream (~413 commits) #174. The last good image is69ab5c1b. They are rebranded again, with Prairie's ldflags (versionOverride) and the DebianffmpegPrairie's AVIF encoder expects. The compose files are rebranded again, with the/dev/dridevice, the books mount and the audiobook-covers mount put back./var/lib/prairie/...: userdb, the jellyfin-web install and artwork.redis.urlis rejected on save, and the log/realtime hub start logs a warning and continues instead of callingFatal.PRAIRIE_*env names: ebook/audiobook/manga enrich workers, scan workers, ebook backfill caps and rate-limit cooldown, metadata match score and push-relay development URL. Each readsPRAIRIE_*first and falls back toSILO_*.prairie.theintrodbfor the legacy IntroDB key migration, andprairie.audiobooksfor the ABS install ID.X-Prairie-*headers (httpheaders) had been reverted toX-Silo-*across the server, while both apps sendX-Prairie-Device-Id. The rename above covers this.Pre-existing Prairie bugs found and fixed along the way:
calibre:series,calibre:series_indexandcalibre:isbntocaliber:*, so EPUB series and ISBN metadata were never read.argIdx++in the episode catalog SQL builders, which reused a placeholder number.userdbregisters modernc assqlite3, and an upstream package also importedmattn/go-sqlite3, which panics at init when both are linked.bridgeimportnow uses modernc.Not restored (listed so you can decide):
-copyts, full-length synthetic manifests and seek restarts, and porting the window would undo that design. Its tests stay and fail:TestGenerateFullManifest_ResumeWindowsAtStartSegment,TestSegmentRecoveryDecision_BeforeStartNeverRestartsAtZeroandTestSegmentRecoveryDecision_StartupProbeBeforeResumeDoesNotRestartininternal/playback. They are not in Prairie CI. This is a pre-existing regression onmainthat needs a Prairie-side fix if AVPlay still ignoresEXT-X-STARTthere.mainsince feat: sync silo-server main upstream (~413 commits) #174.CompareFingerprintsis now upstream's reworked Go algorithm.processor.goand the.wasmfile remain but are unused.internal/artworkstore(the Prairie local artwork store) was dead code after the sync, because upstream'sblobstorefilesystem backend replaces it, so it is removed. A local artwork cache written in the old layout may be re-cached.Needs your judgement
internal/metadataladder tests and a fewinternal/api/handlersimage-size tests still assume upstream's rungs and fail locally. These packages are not in Prairie CI.mainwith the 291ec67 regression unresolved, and with prairie-android#31 required alongside.SILO_*prefix (stream telemetry, debug, scenario, test DB), and only Prairie's historical ones readPRAIRIE_*first. Decide whether you want a blanketPRAIRIE_*alias.logredactlimitation: afmt.Errorf("fetch %s: %w", rawURL, errors.Join(...))wrapper still prints the raw URL, because the redactor only rewrites the*url.Errortext. This was noticed while adding coverage tests for Prairie's gate and was not changed.Testing
go vetpasses on all 194 packages. golangci-lint (Prairie config,--new-from-merge-base=origin/main) shows 0 issues locally apart from one batch that the host OOM-killed; CI covers it.go test, one package at a time locally, with libvips headers and libs extracted to a user prefix: the packages in Prairie's CI gates pass, including the coverage packages (langandlogredactgot tests to stay ≥95%), jellycompat, imageutil, contract ledger, contract spec, route inventory, api and policy.metadataandapi/handlers;taskmanagersuite times out;api/handlerstests (calendar poster batching, the first-frame metric, and v2 request labels inapiv2) that fail deterministically and were not investigated.Deploy note
Merging to
mainpublishes an image.docker.ymlruns on every push tomainand pushesghcr.io/prairie-server/prairie-server:latestplus a SHA tag. This PR fixes that workflow, which has failed on everymainpush since #174, so the first successful:latestsince69ab5c1bwill be published on merge. Anything tracking:latest(keel, compose pulls) will roll forward. No release is cut:release.ymlis dispatch-only.discord-commits.ymlposts to the commits channel.AI disclosure
Tool: Claude Code; Model: claude-opus-5-5; Involvement: AI-generated.
🤖 Generated with Claude Code