Skip to content

feat(cluster): simplify live deployment and delete managed graphs - #878

Merged
aaltshuler merged 13 commits into
ModernRelay:mainfrom
aaltshuler:codex/cluster-lifecycle-completion
Oct 6, 2026
Merged

aaltshuler merged 13 commits into
ModernRelay:mainfrom
aaltshuler:codex/cluster-lifecycle-completion

Conversation

@aaltshuler

@aaltshuler aaltshuler commented Oct 5, 2026 •

Copy link
Copy Markdown
Contributor

What & why

Cluster deployment now previews changes through a running server and keeps long-running apply independent of the HTTP connection. Ordinary apply creates, updates and physically deletes managed graphs, including schema, query, policy, embedding-provider and Blob-binding changes without a server restart.

  • cluster plan --server returns an observational diff, schema migrations and exact roots scheduled for deletion. It uses retained accepted schema views without closing graph admission or acquiring schema gates; apply revalidates physical eligibility after drain.
  • Served apply submits once. HTTP 202 follows durable acceptance, and the existing owned executor continues through activation. The CLI polls the exact deployment ID by default; --no-wait returns after acceptance, --timeout bounds only caller waiting, and cluster status --deployment-id ID --wait resumes observation. Lost responses never trigger POST replay. Both outstanding and completed receipts must match the submitted input digest.
  • cluster uses explicit --managed to select the managed service; without it, folder context never redirects self-hosted commands. cluster operation --managed ID observes lifecycle work separately from run status. Wrong-mode arguments and competing targets refuse before context or external access. The top-level managed command is removed without an alias.
  • Removing a graph declaration closes affected admission, drains admitted requests and streams, and purges the exact managed root and retained history. Deletion intent is recorded first; interrupted deletion resumes under its original ID and verifies root absence. Peer roots and external Blob sources are preserved. Policy authorizes deletion for storage-owner callers too.
  • One existing ledger owns durable deployment authority. Immutable runtime bindings and process activation install together; unrelated blocked graphs do not invalidate independent live changes. The drain deadline cannot strand a durably applied deployment before activation. Original authenticated initiators retain only exact-receipt access after management-policy handoff.
  • Remove lifecycle files, adoption/recreation/drift-repair overrides, obsolete CLI/HTTP aliases, the unused HTTP schema-apply route and version-1 managed credentials. Recognized obsolete completed receipts require explicit ledger conversion; normal reads stay strict. Update the accepted RFC, guides, distributed CLI skill, OpenAPI and release notes.

Backing RFC

Implements the accepted Server runtime and online deployment RFC, including its simplification, physical deletion, served CLI and explicit managed-mode amendments.

Validation

  • Inventory follow-up: remove all eight vocabulary records for the intentionally retired LegacyReadOutput type. All 52 vocabulary-guard tests and the exact-base OpenAPI inventory check pass. The complete canonical workspace/failpoint test graph, including documentation tests, passed in a clean worktree with an explicitly built server binary; no extra tests were skipped.
  • Final merge fixes: all 28 server startup tests passed (1 ignored), the merge/plan ownership test passed 100 repeated runs with alternating 2/8-worker runtimes, and all 35 architectural source guards passed. Strict Clippy for both affected test owners and formatting passed. The CI ownership race is fixed by waiting for completed read producers to settle before comparing full ownership snapshots; the root-marker delete is explicitly registered after auditing its existing admission and durable-intent callers.
  • Explicit managed-mode follow-up: the full CLI suite passed (439 passed, 4 ignored); strict CLI all-target Clippy and current docs/static checks passed. Managed-service coverage uses local HTTP fixtures.
  • CLI, cluster and server test suites with cluster failpoints. The real CLI/server/proxy deployment matrix covers caller timeout and socket loss before acceptance and lost delivery after acceptance, requiring one ledger result, one schema publication, eventual exact-ID completion, and the same server PID. The required CI cell rejects removal or skipping of that journey.
  • Deterministic server tests cover owner start/finish/turnover during status reads, bounded observation refusal, exact-ID isolation, and schema preview while a merge owns the graph gate. CLI fixtures cover 429/503/truncated GET retries without POST replay and immediate malformed-receipt refusal. The truncated-body case reproduced the missing retry before the fix.
  • Existing process journeys cover live schema/query changes, policy handoff, providers/Blob bindings and deletion preview → acceptance → exact-ID waiting. The original submitter observes its exact receipt after management revocation and real restart, while aggregate status and new apply remain denied. Managed status/history scope and filters, direct apply under managed folder context, and timeout recovery guidance are covered in their existing owners. Explicit-mode regressions cover both flag positions and 33 refused argument combinations with no HTTP, storage, or context effects.
  • Existing fake HTTP owner reproduces changed-input/old-ID false success before the fix and passes for both wait modes afterward. Existing cluster owner covers deletion planning without a ready handle.
  • Engine schema owner: 30 tests passed, including pure planning against an exact accepted contract. No engine publication behavior changes in the served-planning addition.
  • Strict workspace all-target Clippy, OpenAPI regeneration/drift, formatting, docs, agent-map, required-test-name, dependency-source, spelling and diff checks.
  • Earlier storage and GQT evidence for the same PR: 50 storage tests and 473 GQT tests passed. This CLI follow-up does not change storage mechanisms or query semantics.

Local validation had no live S3/Azure credentials or managed control-plane service; skipped cloud wrappers and local managed HTTP fixtures do not qualify those live integrations. Configured RustFS and Azurite suites, storage-upgrade compatibility, GQT and DST passed in CI on the preceding commit. Benchmarks were not rerun. CI must verify the final head before merge.

Operator changes

  • Deletion is destructive: omission plus apply deletes managed storage and history. Cloud object-version retention remains provider-controlled. Writer admission and request drainage do not establish full native-I/O settlement or fence arbitrary embedded writers. Uncertain deletion remains outstanding and closed to serving; an already-missing runtime predecessor requires restart to recapture inventory.
  • Deploy matching CLI/server builds. Exact served lookup now returns {deployment, active, in_progress}; aggregate status returns {status, active, in_progress}. Caller wait expiry exits 5 and preserves the original ID. Managed commands use cluster <command> --managed; standalone schema apply remains direct-storage only. Version-1 credentials must be replaced explicitly.
  • Finish outstanding deployments with their originating build. Recognized obsolete completed-result fields require omnigraph --cluster ROOT cluster upgrade-ledger --writers-stopped after stopping writers and establishing prior-I/O quiescence. Conversion preserves data, identity, history and achieved receipts.
  • Provider changes do not re-embed stored vectors. Storage-root migration, graph-format conversion and credential/trust replacement remain outside apply. Empty-directory creation retry does not authorize adoption or general root cleanup.

ragnorc added a commit to ragnorc/omnigraph that referenced this pull request Oct 6, 2026
v0.12.0 is tagged and its snapshot generated, but release.json still named
it, so the documentation check compared every note added since with that
published snapshot and failed. The release procedure's post-publication
step moves the configuration to the next version on the published base;
this is the same change ModernRelay#871 and ModernRelay#878 carry.
@aaltshuler aaltshuler changed the title feat(cluster): restore live configuration and explicit graph lifecycle feat(cluster): simplify live deployment and delete managed graphs Oct 6, 2026
Qualify submit-once behavior with real CLI/server disconnect and timeout cases, deterministic owner races, held-merge planning, and receipt access across policy handoff and restart. Extend managed command coverage and fix its timeout recovery hint.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant