Skip to content

feat(formal): close portable behavior and native evidence gaps - #193

Draft
lan17 wants to merge 33 commits into
claude/fault-mapfrom
codex/formal-interactions
Draft

lan17 wants to merge 33 commits into
claude/fault-mapfrom
codex/formal-interactions

Conversation

@lan17

@lan17 lan17 commented Sep 21, 2026

Copy link
Copy Markdown
Owner

DialCache’s formal checks could credit a rule without showing that its fault changed an observable result in the real libraries. This change connects model rules to exact generated histories and vector inputs, and requires the declared consequence in completed clean and faulty runs in both TypeScript and Go. TypeScript remains the executable reference; Quint defines the portable behavior and expectations that language implementations must satisfy.

The expanded portable coverage includes:

  • Dark work across request, local and shared layers, with captured policy, invalidation, instance isolation and clock changes.
  • Caller, source, whole-job and independent C0/C1 read deadlines, including cancellation and capacity retained by unfinished raw work.
  • Public coalescing inspection: leaders, followers, oldest-leader age, cleanup and exclusions.
  • Exact stale-frame expiry, refill and later reuse; local millisecond expiry; recovery acquisition and age checks; inherited policy; tracked publication and fence boundaries.
  • Key, frame, envelope and compression boundaries. Every native mutation shard also runs all 288 generated invalidation vectors through production Lua on its own Redis server.

Native fault credit requires the selected public consequence, a matching clean baseline, completed recordings, current source/corpus fingerprints and the correct language binding. Crashes, incomplete histories, unrelated earlier counters and disabled model histories cannot substitute for that evidence. All 79 model challenges have deterministic reproducers. All 71 native mappings have exact boundaries: 63 history mappings and eight vector mappings. Six model-only and two unobservable challenges retain their explicit classifications; both historical mapping backlogs are closed.

The setup keeps independent model properties while sharing rules with identical meaning. The cleanup removes catalog mirror arrays, 363 redundant structural-exclusion explanations, closed migration escape paths and unused assertion-log reconstruction. formal/WALKTHROUGH.md follows one contract through the model and both ports; formal/AUTHORING.md and formal/PORTING.md define the extension path.

Go replay uses built-in parallel testing with a separate coordinator per profile. A controlled 467-history race benchmark improved from 15.94s to 10.01s (37% less elapsed time). Boundary recordings remain serial. Mutation cohorts use -parallel=1 because an intentional process-global fault must not connect independent virtual-time histories; whole mutation shards still run concurrently. M20 now completes its generated cohort with 448 assertion detections and confirms both cross-instance boundaries.

Exploration no longer repeats the identical pinned model-fault campaign; full validation still requires that campaign exactly once. A same-seed comparison with three Quint workers improved from 65m44s to 13m02s while preserving every one of the 6,022 behavioral histories, all 497 required witness labels/counts and both 7,759-case native inventories. This measures the combined changes under differing machine load, not an isolated or whole-CI speedup. Quint output is captured through private regular files because its immediate exit can truncate piped failure diagnostics; complete-output checks on macOS/Linux and a real faulty model verify the fix without relaxing classification.

Validation at 8e2bc79: local acceptance is complete; the full hosted run with fresh-seed exploration has passed all native replay, mutation, symbolic and exploration lanes; its final model gate and aggregate remain pending.

Evidence Result
Full native replay 7,759 required cases per language over the same 6,022 histories
Model faults 79/79 detected; all 1,173 profile classifications complete; zero measurement errors
Native mutations 63/63 required detections and 71/71 exact boundaries in each port, with 55 clean baseline histories
Behavioral inventory 69 contracts; 244 behavioral and 22 protocol cases; 12 feature families
Corpus comparison 4,993 forward / 5,003 reverse histories across 13 comparable profiles; zero disagreements

Ordinary local checks and all standard hosted PR checks pass. Separate local symbolic and integration results remain applicable through verified unchanged inputs: seven finite rule properties, 819 TypeScript integration tests and 1,019 Go integration leaves. Local acceptance combines completed lanes and that source-verified evidence; it is not a claim that one uninterrupted make ci command ran at this commit.

The first local fresh-seed exploration (0x5ca6fd8d33dc235f) remains a recorded sampling-quality failure: both languages passed all cases and every required label was reached, but one shadow logging label had 10 sampled hits against a threshold of 10.02175 (baseline 33). The gate was unchanged. The fixed-seed performance comparison does not replace that result; the independent hosted fresh-seed exploration passed.

Coverage is finite and explicitly scoped. The counts measure registered cases and required evidence, not arbitrary schedules, all input combinations, liveness or a universal correctness percentage. Fixed scenarios retain handwritten expectations; native API/value/codec/exporter and real-server obligations remain separate. The two new profiles and changed effects observation schema are explicit exclusions from the old-corpus comparison. Production TypeScript and Go code is unchanged.

Stacked on #188 (84fccb1), targeting claude/fault-map. Retarget to main after #188 merges. No merge or auto-merge is enabled.

Replay named histories past complete observation mismatches and record each divergent field at the declared checkpoint. Derive evidence from public Quint regressions, preserving explicit gaps and refusing partial projections or incomplete recordings. Measure clean and mutated histories in both ports outside the existing detection cohorts, and retain evidence through shard merges. Report boundary results before enabling the stricter gate.
…ions

Reject held and local-fault state combinations whose assigned transitions do not define their lifecycle. Schedule atomic payload, single-source and admission clock bounds, and pin enabled failed-key bypass behavior in the kernel fixtures. Measure shared model-fault reachability from import closures, require listed profiles to detect it, and check reaching exclusions rather than relying on dated prose. Preserve inconclusive mutant histories separately from actual assertion failures.
Add 18 public regressions and 18 invariants for held dark work across request and local layers, source deadlines, instance isolation, captured fills and tracked fences. Correct source-failure diagnostic attribution against the TypeScript reference, with matching Go replay. Register five portable challenge reproducers and measured cross-profile fault partitions, retaining explicit native scope limits for M01, M30, M35 and M37. Both ports replay all 274 dark histories, and the witness controls require the distinguishing public consequences.
Gate every reproduced native mapping on a complete clean baseline and a consequential comparison at the declared checkpoint. Repair the histories and shared publication fault anchors whose old witnesses were inert or could not finish under mutation. Keep typed semantic monitor failures recordable while infrastructure and malformed observations remain incomplete evidence. Preserve explicit vector and unreproduced states instead of granting boundary credit from cohort detection alone.
… evidence

Require standalone boundary gating to match the current catalog, source, configuration and corpus fingerprints while retaining historical inspection. Refresh the integrated fixture and inventory snapshots and the command expectations for the new invariants and reproduced histories. Document the current profile composition and allow hosted model partition checks a provisional sixty-minute budget pending measurement.
Remove the earlier private-state expectation from the dark read fence reproducer so its clean prefix remains executable under both inclusive-fence model faults. The public serialization and decode consequence at step seven remains the required boundary. Exact before/through probes pass for all five dark-profile challenge mappings, and the exported history states are unchanged.
The scope cleanup fault is detected by scope and recovery-read, while layers currently lacks a preserving-memo assertion. Record that bounded gap and correct the dark-layers exclusion rationale. Finish independent challenge measurements before rejecting a failed campaign, preserving every original error and incomplete report status. Focused checker and execution tests pass, and the source audit remains current.
Observe the source-versus-decode and fill-versus-decode decisions before the next settlement command. The future-timestamp fault now fails public assertions in effects and shadow instead of disabling a later history step. Clean named suites pass and the fault produces genuine assertion failures in both profiles. Regenerated fixtures retain their contents and update only the model fingerprints; source audits and 59 focused tests pass.
Pin outside-call deadlines, independent recovery values, and source-budget origins to existing exported histories. Require complete caller-result divergences in both ports and retain the independent model property checks. Measure every importing profile before declaring detection or exclusion, and remove the three reproduced challenges from the grandfathered backlog. Update the execution ledger and inventory expectations without changing model transitions or production code.
Add an explicit wall-clock input and portable last-fresh and exact-expiry histories that cover source fallback, refill, reuse and the expired diagnostic. Check the category with an independent acquisition receipt and a compiling model fault mapped to native mutant M57. Version the expanded effects input domain and add consequential witness controls for both language replays.
Resolve named generated inputs into input-only native operations and recompute each boundary mismatch from typed actual results. Require matching clean baseline, artifact and input fingerprints, and report language for every mapped vector challenge. Add deterministic model cases and exact native witnesses for fence, key, envelope, and compression tie faults.
Replay every generated invalidation vector through each production Lua script in the native mutation campaigns. Record exact input-only Redis operations for cutoff monotonicity and the inclusive maximum buffer, with clean baselines and measured server time for TTL comparisons. Require Docker for these campaigns, isolate server ownership, and reject transport or malformed-output failures as detection evidence. Update shard fixtures to exercise typed vector evidence.
Include the separate shadow read profile in differential inventory controls. Restore compact formatting for unchanged fixture recipes, preserving all parsed values and their order. Refresh only the recipe input fingerprint; all 36 generated artifacts remain identical.
@lan17 lan17 changed the title feat(formal): validate dark-layer interactions at native fault boundaries feat(formal): close portable behavior and native evidence gaps Sep 21, 2026
Record the complete combined corpus after the shared rule connections. All existing sampled witness counts are unchanged; the seven additional boundary labels remain anchored by required exported histories. Preserve the seed, density thresholds and every existing required count.
Infer structural profile exclusions from imports while retaining every measured importing-profile claim. Require reproducible evidence directly now that the migration backlogs are closed, and remove catalog snapshot arrays and unused assertion-log reconstruction. Correct three partition declarations exposed by the complete model challenge campaign. All 1,173 partition decisions are preserved by the metadata simplification; focused validation and the source audit pass.
Record the completed cleanup paragraph after removal of historical assertion-log reconstruction. The guide now directs readers to completed native recordings while retaining raw assertion diagnostics. Refresh its audit and dependent parity digest; no executable behavior changes.
The in-memory serializer test can cross the default real-time read deadline when coverage workers contend for CPU. Use its suite’s existing fake-clock convention so the serializer assertions depend on controlled operations. Keep deadline-specific tests unchanged and refresh the reviewed native-test evidence.
Give each feature profile its own coordinator and let the Go test runner bound concurrency. Preserve leaf names, action coverage, virtual clocks, settlement checks, and ordered serial boundary recordings. Aggregate executed counts after parallel children finish so empty selections still fail. Paired race-enabled replay of the same 467 histories averaged 37 percent less elapsed time with two workers; selector, recording, completion, and native controls passed.
The hosted exploration exhausted its old one-hour budget while checking model fault partitions, before generating or replaying the alternate corpus. Give model checking slow-runner headroom and leave exploration additional time for generation and both ports. Preserve every model, history, partition, and per-command hang bound. Existing workflow validation passes all 23 checks.
Serialize mutation cohorts because deliberate global-state faults can join work across otherwise independent synctest histories. Keep clean profile replay parallel and retain concurrent mutation shards. Remove copied inventory totals from the formal guides so the executable catalogs remain authoritative.
Capture evaluator output through private regular files so immediate CLI exits cannot drop piped diagnostics. Keep strict failure classification, and make the pinned model fault campaign an explicit acceptance stage so fresh-seed exploration does not repeat it.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant