Skip to content

fix(service-storage): chunks sent in parallel to one chunked upload each record their part (#22332) - #22348

Merged
objectstack-fleet[bot] merged 6 commits into
mainfrom
claude/issue-22332-chunk-part-atomic
Oct 8, 2026
Merged

objectstack-fleet[bot] merged 6 commits into
mainfrom
claude/issue-22332-chunk-part-atomic

Conversation

@objectstack-fleet

Copy link
Copy Markdown
Contributor

Fixes #22332

Clause-②: no

What was wrong

The chunk door (PUT /storage/upload/chunked/:uploadId/chunk/:chunkIndex) read the session row, merged its part into the parts list in memory (recordChunk) and wrote the whole record back with an unconditional by-id update. Two PUTs to one upload at once both read the same record, and the later write erased the earlier part: both answered 200, sys_upload_session.parts kept one part, uploaded_chunks / uploaded_size counted one chunk, and the completion door refused the upload 409 naming the lost chunk.

The atomic route taken, and why

A compare-and-set on the row's three progress columns, with a bounded re-read-and-merge retry. This is the route the store already supports; it needs no new object, no new field and no packages/spec change.

  • On a wired engine it is the engine's own conditional update: engine.update('sys_upload_session', progress, { where: { parts, uploaded_chunks, uploaded_size, id }, multi: true, context }). The update-dispatch module (resolveEngineUpdateDispatch, @objectstack/metadata-core) names exactly this shape, an id in where beside further keys with multi: true, as "the compare-and-set spelling". It routes to driver.updateMany, which evaluates the guard in the same statement that writes and answers the matched-row count. driver-sql issues one UPDATE … WHERE with the tenant scope applied; driver-memory matches and writes in one synchronous step. Every driver's updateMany (sql, memory, turso, mongodb) resolves a count.
  • On the engine-absent stand-in it is one synchronous compare-and-set on the Map.
  • Why these columns are the guard: they are exactly the values the merge read, so the comparison is exact. updated_at as a version has millisecond grain, so two writes in one millisecond would still match. A per-part row would be a new object, a schema change, and was not needed.
  • ⛔ No process-local lock: a second server process would defeat it. ⛔ The completion door's guard is untouched.

The store side is updateSessionProgressIfUnchanged in metadata-store.ts. It is a module export that is not re-exported from the package entry, following the organizationOutOfWriteReach pattern, so StorageMetadataStore's public face does not grow.

The door merges and writes up to CHUNK_RECORD_ATTEMPTS (16) times. A conditional write loses only to a write that landed between its read and itself, so each lost attempt is another chunk recorded.

  • On exhaustion the door answers 409 RESOURCE_CONFLICT with error.details: { chunkIndex, attempts }. The chunk's bytes are stored but the record does not hold them, and re-sending the chunk replaces its slot. It never answers a silent 200.
  • When the conditional write cannot reach the row (it misses, and the row still holds the progress the write was conditioned on), the door answers 500 at once. It does not loop into a 409 that would tell the uploader to retry something that cannot succeed.

Measured: what the store supports (H2)

A probe against the real ObjectQL over SqlDriver (better-sqlite3, :memory:) at e36ee5351b:

conditional update answer
guard as read 1
stale guard 0
guard on an 8-part parts string 1
guard of NULL columns (lowered to IS NULL) 1
row stamped org_B, acting organization org_A 0, no throw
the same row, acting organization org_B 1
missing row 0

Measured: the reproduction, before and after

pnpm --filter @objectstack/service-storage exec vitest run --maxWorkers=2 src/chunk-part-record-concurrency.test.ts

Before, the pins on main at e36ee5351b (which carries the completion guard): 8 failed and 2 passed; the 2 that passed are the sequential controls.

  • The stand-in and the real engine alike lost parts: every concurrently sent chunk is in the record: expected [ 1 ] to deeply equal [ +0, 1 ] (2 concurrent), and expected [ +0 ] to deeply equal [ +0, 1, 2, 3, 4, 5, 6, 7 ] (8 concurrent).
  • The parallel upload's completion was refused 409 with missingChunks: [1].
  • With a competing write landing between the read and the write, the door answered {"success":true,…} (expected 200 to be 409), and the record held [ +0 ]: the competitor's part was erased.

After, at 006344eb50: 24 passed (24). Both stores pass the 2 and 8 concurrent pins (every part, every eTag and size, uploaded_chunks, uploaded_size, and the progress door), the parallel completion 200 with the whole file, the sequential control (the same answers and record, one session write per chunk), and the exhaustion pin (409 RESOURCE_CONFLICT after exactly 16 reads, with every competitor's part kept). The store-level pins also pass on both stores: the write lands on an unchanged row, misses once any one progress column moved, and misses on a gone row. On the real engine, the call shape is pinned (dispatch verdict multi, the same { tenantId, isSystem } context), another organization's row is not reached, a non-count answer is refused loudly, and an unreachable row answers 500 once.

Over HTTP, a real boot (bootStack, sqlite-wasm): packages/qa/dogfood/test/storage-chunked-parallel-parts.dogfood.test.ts passes 2 of 2 (2 concurrent PUTs, then progress and completion 200; 8 concurrent PUTs, then progress). With the sibling storage-chunked-resume-integrity.dogfood.test.ts, 6 passed (6) at 006344eb50.

Ablations

Each mutation went through scripts/ablation-replace.mjs: the anchor hit 1 to 0, the blob changed, and the restore was proven with blob == HEAD and an empty git diff HEAD. The runner also carried its own EXIT/INT/TERM restore over absolute paths. The control run with no mutation passed 24 of 24.

mutation pins turned red
M1: the stand-in's comparison removed 7, all stand-in: 2 and 8 concurrent, the parallel completion, exhaustion, and 3 column-moved pins
M2: the engine guard replaced by an unchanging column 8, all engine: 2 and 8 concurrent, the parallel completion, exhaustion, 3 column-moved pins, and the where-shape pin
M3: the door treats a lost write as landed 9: on both stores, 2 and 8 concurrent, the parallel completion and exhaustion; plus the unreachable-row 500
M4: exhaustion falls through to 200 2: exhaustion, on both stores
M5: a non-count engine answer accepted 1
M6: the unreachable-row check removed 1
M7: the stand-in lands on a missing row 1
M8: the organization scope dropped from the conditional write 2

The dogfood pin reads @objectstack/service-storage from dist/, so its ablation included a rebuild.

  • Mutation leg: M3, then a rebuild. ablation-dist-preflight found the marker in dist/index.js and dist/index.cjs, and 2 of 2 failed (the progress fell short of uploadedChunks).
  • Restore leg: a rebuild, after which the marker was absent from all 6 built files and the tree was clean against HEAD. 2 of 2 passed.

Tests and gates (HEAD 006344eb50)

  • @objectstack/service-storage: 46 files and 773 tests passed. typecheck passed (tsc, the scripts project, and check:test-typecheck: OK). --listFiles shows the new test file is in the test-layer program.
  • @objectstack/dogfood: typecheck passed.
  • Two fake engines now answer a declared predicate update with its matched-row count, as the real engine does: tenant-audit-update-delete-half-repairs.test.ts and storage-routes.metadata-outage.test.ts. The chunk-door pin in the first now expects the conditional where with multi: true, under the same context.
  • check:tenant-audit-census: the new conditional write is one more engine write call site (236 to 237). The census was regenerated (node scripts/tenant-audit-census.mjs --write), and the page's hand-written prose figures were moved with it. The gate and its self-test pass.
  • Lint, narrowed and stated as a measurement. ① Population: the 6 touched TypeScript files, all inside the eslint.config.mjs glob **/*.{ts,tsx,mts,cts,js,jsx,mjs,cjs}. ② --format json linted 6 files with 0 errors and 0 warnings. ③ Invariance: the config sets no parserOptions.project (--print-config gives {"ecmaVersion":"latest","sourceType":"module"}), so linting is not type-aware and this diff cannot move a verdict on an untouched file. The repo-wide pnpm lint is CI's.
  • The derived gate families (dispatch-gates --commands, derived at 006344eb50) were run after the last commit, with exit codes captured before any pipe. dispatch-gates --ran reads 97 derived, 97 run, 0 NOT-MEASURED, 0 UNRUN (a derived zero: every family recorded exit 0). pnpm check:error-status-conformance was run by hand and exited 0: ✓ every derivable runtime status is documented, and every documented status is reachable. A few verdict lines:
    • check-engine-double-contract: OK — 983 pinned, 129 in the DEBT ledger, 3 exempt.
    • check-nul-bytes: OK (scanned 10310 text file(s) …; no raw ASCII control bytes).
    • ✓ check-tenant-audit-census: OK -- 237 write call sites certified …
    • check-test-source-alias OK — 73 packages with tests scanned …
    • ✓ check:dual-build-cjs-loads — 106 published require entry point(s) across 66 package(s) load …. Its first run answered PREREQUISITE NOT MET (no dist/ for 8 packages), which is not a measurement. It passed after a full cache-backed turbo run build.
    • ✓ This diff introduces no major bump. (check-changeset-no-major)

Acceptance notes

  • Clause-② stays no, as the claim declared. No accept set, export or public signature moves. The one new answer, 409 after 16 lost record writes, replaces a 200 that had recorded nothing.
  • The progress write is now a predicate update. On an engine-backed store, the chunk door's write publishes the bulk data.records.updated event (a count) where it published data.record.updated, and its after-hooks go through the bulk per-row path. Nothing in the tree registers an update hook on sys_upload_session or subscribes to its record events, and the config-change audit excludes the object.
  • Observation, read-only and not measured, not filed (承接者:无). The chunk door's upload-limit check judges the record as it was read. Concurrent PUTs can therefore each pass it against the same stale total, and the backend can briefly hold more bytes than maxUploadBytes. The completion door still refuses any upload whose held bytes differ from its declared size, and the init door bounds that size.

Generated by Claude Code

claude added 6 commits October 8, 2026 18:54
…rt (red on main)

Claude-Session: https://claude.ai/code/session_01WkL6Eijt432S1Y7ekb6ovQ
Co-authored-by: Claude <noreply@anthropic.com>
…and-set, so concurrent chunk PUTs record every part

Claude-Session: https://claude.ai/code/session_01WkL6Eijt432S1Y7ekb6ovQ
Co-authored-by: Claude <noreply@anthropic.com>
…pdate with its matched-row count

Claude-Session: https://claude.ai/code/session_01WkL6Eijt432S1Y7ekb6ovQ
Co-authored-by: Claude <noreply@anthropic.com>
…and-set progress write

Claude-Session: https://claude.ai/code/session_01WkL6Eijt432S1Y7ekb6ovQ
Co-authored-by: Claude <noreply@anthropic.com>
…nditional progress write

Claude-Session: https://claude.ai/code/session_01WkL6Eijt432S1Y7ekb6ovQ
Co-authored-by: Claude <noreply@anthropic.com>
@github-actions github-actions Bot added size/l documentation Improvements or additions to documentation tests tooling labels Oct 8, 2026
@github-actions

github-actions Bot commented Oct 8, 2026

Copy link
Copy Markdown
Contributor

📓 Docs Drift Check

This PR changes 1 package(s): @objectstack/service-storage, touching 16 documentable anchor(s).

4 hand-written doc(s) NAME something this change touched and may need an implementation-accuracy re-verification:

  • content/docs/api/error-catalog.mdx (via RESOURCE_CONFLICT (literal, a string literal in CHUNK_NOT_RECORDED_CODE))
  • content/docs/api/error-handling-server.mdx (via RESOURCE_CONFLICT (literal, a string literal in CHUNK_NOT_RECORDED_CODE))
  • content/docs/permissions/attachments-access.mdx (via sys_upload_session (literal, a string literal in updateSessionProgressIfUnchanged))
  • content/docs/ui/translations.mdx (via sys_upload_session (literal, a string literal in updateSessionProgressIfUnchanged))

⛔ 4 release-owned page(s) also name something this change touched. These are read-only:

  • content/docs/releases/v15.mdx (via sys_upload_session (literal, a string literal in updateSessionProgressIfUnchanged))
  • content/docs/releases/v16.mdx (via sys_upload_session (literal, a string literal in updateSessionProgressIfUnchanged))
  • content/docs/releases/v17/17-6.mdx (via RESOURCE_CONFLICT (literal, a string literal in CHUNK_NOT_RECORDED_CODE))
  • content/docs/releases/v17/17-7.mdx (via RESOURCE_CONFLICT (literal, a string literal in CHUNK_NOT_RECORDED_CODE))

content/docs/releases/ is RELEASE-OWNED (AGENTS.md "Documentation Guardrails"): release
notes are written centrally at release time, and a code PR that edits them is the exact PR
that guardrail exists to stop. They are still audited — read-only. If one of them is actually
wrong, file an issue or open a dedicated docs-only PR; do not edit it here.

What this run could not see
  • 4 name(s) were too generic to anchor anything (single lowercase words)
  • the SDK route bridge reached 54 of 206 client-bound route-ledger rows — the other 152 have no registrar path: tail to select them, so pages documenting THEIR client methods cannot appear above, on this or any run. Of those 152: 0 are remediable by widening that discovery convention (an in-repo file declares the path; the convention did not scan it); 55 are structural — on a ledger where NOT ONE row is declared in-repo, so no discovery change reaches them at any price; 97 are undecided (no in-repo declaration, on a ledger that has other in-repo registrars — absence and an unreadable spelling are not distinguishable here). The rows themselves: node scripts/docs-audit/affected-docs.mjs --bridge-coverage
  • a page that states a rule by its inputs shares no identifier with the emitter that implements the rule, so an emitter-only diff cannot list it — not on this run and not on any run. Measured on fix(driver-sql): emit varchar(maxLength) for a text field a declared index keys on #11430: content/docs/protocol/objectql/types.mdx documents the text-family column mapping by the ObjectQL type names it maps FROM (text / textarea / html) while the diff changed createColumn; it went unlisted, and it was the page that diff falsified, in four places. No shared token exists to detect this on, so a rule your change carries has to be re-read by hand in the pages that restate it.
  • a key NAME is not a key, so the hand re-read the line above prescribes can land on the wrong schema. The same spelling is authorable on one governed type and a [REMOVED] tombstone on another for each of active, aria, joins, objects, template, tools and version (censused on [finding] tools is a key on BOTH AgentSchema (tombstoned, dead) and SkillSchema (live, cloud-attested), so a name-based search attributes skill examples to the agent key — it produced a false stop-the-line alarm on PR #19059 #19093 over the liveness ledger's governed types, top-level keys); nothing in a search result distinguishes the two, so a grep hit on a LIVE example reads as evidence about the DEAD key. Measured on fix(spec): the agent.tools liveness row says dead — it claimed live on a key the schema tombstoned #19059: content/docs/ai/agents.mdx was reported as contradicting the agent.tools tombstone over its tools: example at :161, which is inside the defineSkill({ block opened at :155 — the page was already correct. Settle ownership by PARSING the value against both schemas, never by the name: that literal PASSES SkillSchema, and as an AgentSchema it FAILS at tools with the tombstone prescription. ⛔ These names are not the whole class — a key retired through a .strict() guidance map leaves no tombstone in the walked shape and none of them here (tool.category, live as AIToolDefinition.category).

Coarse fallback — 8 page(s) merely mention a changed package (the pre-#9192 predicate, kept for the deliberately-wide backstop): node scripts/docs-audit/affected-docs.mjs --json 5ff7cbe364f939a55633a04891650163c9e8884e → packageMentionDocs.

Which tree this was computed on

This run read content/docs from ab4870b39f83791115b20a3cb15308471b476085 — the merge of head 006344eb50aad472b817b8fdc56f72392c3569db into base 5ff7cbe364f939a55633a04891650163c9e8884e, which is what actions/checkout gives a pull_request run. Not the PR head.

A worktree cut from an older main holds a different content/docs, so re-deriving there can legitimately return a different list — that is a different tree, not a wrong row. To answer on the same tree:

# while this PR is open — GitHub drops the merge commit once it closes
git fetch origin ab4870b39f83791115b20a3cb15308471b476085 && git checkout ab4870b39f83791115b20a3cb15308471b476085
# afterwards, rebuild it from the two parents, which stay fetchable
git fetch origin 5ff7cbe364f939a55633a04891650163c9e8884e 006344eb50aad472b817b8fdc56f72392c3569db && git checkout -B drift-repro 5ff7cbe364f939a55633a04891650163c9e8884e && git merge --no-ff 006344eb50aad472b817b8fdc56f72392c3569db

node scripts/docs-audit/affected-docs.mjs --json 5ff7cbe364f939a55633a04891650163c9e8884e

⚠️ That checkout carried uncommitted changes, so the commit above does not fully identify what was read.

Advisory only, and a precision-first one (#9192): a page is listed because it names a
symbol, wire route or SDK method this diff touched — not because it mentions a changed
package. Each row says which anchor put it there, so a wrong row is reportable rather than
merely annoying. To re-verify, run the docs-accuracy-audit workflow scoped to these files:
node scripts/docs-audit/affected-docs.mjs 5ff7cbe364f939a55633a04891650163c9e8884e → pass the list as
args.docs, on the commit named under Which tree this was computed on.

@objectstack-fleet
objectstack-fleet Bot marked this pull request as ready for review October 8, 2026 20:20
@objectstack-fleet
objectstack-fleet Bot enabled auto-merge October 8, 2026 20:20
@objectstack-fleet
objectstack-fleet Bot added this pull request to the merge queue Oct 8, 2026
Merged via the queue into main with commit 54c3ce1 Oct 8, 2026
37 checks passed
@objectstack-fleet
objectstack-fleet Bot deleted the claude/issue-22332-chunk-part-atomic branch October 8, 2026 20:51
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

documentation Improvements or additions to documentation size/l tests tooling

Projects

None yet

2 participants