Skip to content

SubContainerLazy.eager() caches a rejected materialization forever — one transient mount failure permanently wedges the daemon (projects/start-sdk/lib/util/SubContainer.ts:1006) #4024

Description

@stupleb

Environment

  • @start9labs/start-sdk 2.0.9. The defect is in every published 2.0.x tag (start-sdk/v2.0.0 … v2.0.9) and on Start9Labs/start-technologies master b90fbbab7 (verified 2026-09-19).
  • Not package-specific. Found while packaging ytptube-startos (master fd443c8, SDK 2.0.9), which works around it; worked example of an exposed package below: btcpayserver-startos (master 872cf9e, SDK 2.0.9). Node v22.22.3.
  • Area: service runtime — SubContainer / Daemons. Found by reading code, so there is no StartOS runtime capture.

What happens

SubContainerLazy.eager() memoizes its materialization promise with ??=:

// projects/start-sdk/lib/util/SubContainer.ts:1006
eager(): Promise<SubContainerEager<Manifest, Effects>> {
  return (this.materialized ??= SubContainerEager._of<Manifest, Effects>(...)
    .then(async eager => { ... return eager }))
}

The promise is stored at call time, before it's known to fulfil. If _of rejects — e.g. its mount step throws because a mounted dependency volume is momentarily absent — the rejected promise stays in this.materialized. Every later .eager() / .exec() / .writeFile() / .launch() on that handle re-returns the same rejection and never re-attempts createFs/mount, even after the underlying cause clears.

SubContainerEager._of itself is correct — on failure it tears down its own partial container (_destroyImmediate() + rethrow, SubContainer.ts:527-529). It's the lazy handle that poisons.

This turns a transient failure permanent. The daemon restart loop (projects/start-sdk/lib/mainFn/Daemon.ts:141, while (!signal.aborted)) re-execs on the same retained subcontainer with backoff — it does not re-invoke setupMain. Because a lazy handle first materializes inside that loop (on first exec), a first-mount rejection is caught by the loop and retried against the already-poisoned handle. The daemon then spins on the cached rejection until a full service restart rebuilds main with a fresh handle.

What should happen instead

A failed materialization must be retryable. A subsequent call (or the next restart-loop iteration) should re-attempt createFs/mount rather than replay a cached rejection. Only a successful materialization should be memoized.

Root cause / code pointers

Found by reading the SDK source, not reproduced at runtime.

  • projects/start-sdk/lib/util/SubContainer.ts:1006 — this.materialized ??= … caches the promise regardless of outcome.
  • projects/start-sdk/lib/mainFn/Daemon.ts:141 — restart loop re-execs on the retained subcontainer (startCommand closes over it, Daemon.ts:88-90), so it re-awaits the poisoned handle rather than rebuilding it.
  • projects/start-sdk/lib/mainFn/index.ts:26 — setupMain only re-runs fn (fresh handle) on a reactive .const() change or a full restart, not on daemon crash-restart.

How a package reaches it:

  1. In setupMain, build a daemon whose subcontainer mounts a dependency's volume via mountDependency(...) (at least 18 fleet packages call it), and hand the handle to the daemon without touching it in main first.
  2. The dependency's volume is unavailable at the moment the daemon first execs (dependency not installed yet, mid-install / mid-restore / mid-migration, or a transient unmount).
  3. The first .eager() rejects; the handle caches the rejection.
  4. The daemon's restart loop retries forever on the cached rejection — it never recovers even once the volume is present. Only stopping and starting the service clears it.

Blast radius widened with #3961 (4989ed8d8, 2026-09-16, not yet in a tagged StartOS release). A read-only dependency mount already failed when the dependency's volume root was missing, but a read-write one did not: Bind::pre_mount created the missing source and the mount resolved against an empty directory (wrong, but non-rejecting). Since #3961 the read-write mount fails too (volume-missing, NotFound) — which is precisely the rejection that poisons the handle.

Suggested fix — memoize only on success:

eager(): Promise<SubContainerEager<Manifest, Effects>> {
  if (this.materialized) return this.materialized
  const p = SubContainerEager._of<Manifest, Effects>(
    this.effects,
    { imageId: this.imageId, sharedRun: this.sharedRun },
    this.mounts, this.name, this.identity,
  ).then(async eager => {
    if (this.destroyPending) await eager.destroy()
    else if (this.detachPending) eager.detach()
    return eager
  })
  this.materialized = p
  // Don't cache a failed materialization: clear the slot so the next call retries.
  // Guard on identity so a concurrent retry that already replaced it isn't clobbered.
  p.catch(() => { if (this.materialized === p) this.materialized = null })
  return p
}

This preserves success-caching, idempotency and the destroy/detach-pending handling; the hold()/destroy()/detach() paths already treat a null materialized as never-materialized, which is now accurate after a failure.

Logs & evidence

Worked example — btcpayserver-startos (master 872cf9e), startos/main.ts:

  • :195-213 — nbxSub = sdk.SubContainer.of(…) mounts the Bitcoin dependency's main volume read-write (readonly: false): exactly the case fix(start-core): refuse a dependency mount whose volume does not exist #3961 turned from a silent empty directory into a rejection.
  • :269 the nbxplorer daemon and :289 the reset-start-height oneshot both take nbxSub, and the utxo-sync health check reads nbxSub.rootfs (:305). Nothing in main touches the handle first, so its first materialization happens inside the daemon loop.
  • If Bitcoin's volume root does not exist at that moment, the first exec rejects, nbxSub caches the rejection and NBXplorer never starts. The btcpay daemon requires: ['nbxplorer', 'postgres'], so BTCPay never starts either — even after Bitcoin's volume appears. Only restarting BTCPay Server clears it.

Same shape, verified in source: lightning-terminal-startos (litSub, LND's main volume, read-only) and mempool-startos (backendSub, the lightning node's volume, read-only). Packages that read .rootfs on the handle inside main before building their daemons (cln, lnd, eclair, nextcloud) fail main outright on the same mount error instead, which does not go through the cached-rejection path.

Found in ytptube-startos (stupleb/ytptube-startos master fd443c8), which mounts a File Browser / NextExplorer data volume read-write and had to defend itself — startos/main.ts: it materializes the handle inside main under a try (:115-116, await remote.eager()), and on failure (:120) discards that handle, falls back to a fresh SubContainer.of(...) on local storage (:131), and forces a later fresh main pass with a timer-backed .const() (:127) — so it never re-awaits the poisoned handle. None of the three packages above have such a guard.

No activity

Activity on this issue will appear here.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Labels

Type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions