Revealed Uncollapsed Manifold Instrument
A local-first instrument for discovering latent affordances — features your codebase already almost contains, that your users are already trying to collapse into existence.
RUMI does not tell you what happened. It lets you see something you could not previously perceive: the gap between what a system selected, what humans keep correcting it toward, what the code can already support, and what is actually being used.
Registered as DCC-2026-001 in the Displacement Code Challenge. RUMI is the instrument. Rectifier Seed is the correction-field core inside it.
Every system selection displaces alternatives. Some of those displaced alternatives become visible when users correct, override, rephrase, retry, or repair the system. RUMI reads three fields over a shared index of capabilities:
| Field | Symbol | Question |
|---|---|---|
| Correction | C(x) |
What do humans keep pushing the system toward? |
| Capacity | K(x) |
What does the codebase structurally already support? |
| Utilization | U(x) |
What is actually being used? |
The interesting region is high C, high K, low U — strong correction pressure meeting latent, unused capacity. RUMI's original observable multiplied the three fields directly:
Collapse Potential CP(x) = C(x) · K(x) · (1 − U(x)) (retained as a component)
But a causal benchmark ([docs/worldforge.md] / WorldForge v2 — synthetic worlds with a hidden causal truth RUMI never sees, scored by simulated counterfactual collapse) showed that gating on capacity as a magnitude multiplier is net-harmful in every noise regime tested: text-presence capacity is too noisy a proxy for real buildability, and multiplying by it crushes true positives. The benchmark-validated primary ranking is:
Collapse Score S(x) = C(x) · confidence(x) · (1 − U(x))
Capacity isn't discarded — it survives inside confidence (a capability with no matching code has capacityConfidence = 0, so unsupported wishes still score zero). It just no longer acts as a magnitude gate. Across the full noise space swept, S weakly dominates C·K·(1−U): it is never worse (the two tie only when capacity is a perfect, noise-free proxy — the gate doesn't help even then) and strictly better under any realistic noise, by the widest margin exactly where capacity is a poor proxy for buildability. A learned combiner does better still, so the long arc is to learn the weighting. This is synthetic-validated; validation on a real correction stream is pending — CP is retained alongside S so the change is transparent and reversible.
The qualitative regions are unchanged — high corrections alone is an unsupported wish, high capacity alone is dormant code, low usage alone is nobody-wants-it — and the intersection is still where an uncollapsed feature lives.
npm install
npm run build
# scan the bundled example (the "enterprise renewal risk" demo)
npm run scan
# or directly:
node dist/index.js scan --repo <path-to-repo> --data <path-to-data-dir>
# explore everything in the browser observatory (Scan / Discover / Reflect)
npm run dashboard # http://localhost:4317 — computed live, fully localThe observatory has three views: Scan (Level-1 candidates with the ripe/deep quadrant and confidence), Discover (emergent capabilities proposed from the correction field), and Reflect (the Level-2/3 recursion and its split fixed point). Everything is computed on the local machine — nothing is uploaded.
Running against examples/sample-repo surfaces, as the top candidate:
▸ Enterprise Renewal Risk Review [enterprise-renewal-risk]
Collapse Score : 0.626 ← primary ranking (C · confidence · (1−U))
confidence : 0.811
(Collapse Potential: 0.667 — legacy C·K·(1−U))
C correction : 0.811 (5 events, coherence 1)
K capacity : 0.865 (6 signals, 4 files)
U utilization : 0.049 (1 uses)
D integration : 1.000 DEEP — pieces scattered with no shared dependency path
→ Uncollapsed feature: strong correction pressure meets latent capacity that is barely used.
Each field is scored on a fixed, scan-independent scale — a capability's CP
depends only on its own evidence, never on which other capabilities share the
scan — so readings are comparable across scans and experiment compare is a
real before/after test. Every reading also carries a confidence in [0,1]:
CP says how strong the collapse signal is; confidence says how much to trust
that number. A high CP with low confidence (e.g. correction pressure but
unknown utilization) is a lead to verify, not a conclusion to act on — absent
telemetry is treated as unknown, never as "confirmed unused".
The pieces (segmentation, renewalDate, riskScore, reportBuilder, ownerRouting, blockers) all already exist in the repo, scattered across four files. No workflow composes them. Users keep correcting generic reports toward renewal-risk reviews. RUMI names the gap.
Contrast with the other capabilities the same scan classifies:
- Bulk Data Export → unsupported wish (users want it, no code supports it)
- Dark Mode → already collapsed (realized and in active use)
rumi scan --repo <dir> --data <dir> [--json] [--top N]
rumi discover --repo <dir> --data <dir> [--json] [--top N]
rumi reflect --repo <dir> --data <dir> [--json] [--top N]
rumi dashboard [--port 4317]
rumi experiment baseline --repo <dir> --data <dir>
rumi experiment compare --repo <dir> --data <dir>
experiment baseline snapshots the field; after you ship a change, experiment compare checks whether correction pressure actually decayed and utilization rose — i.e. whether the latent affordance collapsed into a real, used workflow. The instrument verifies collapse; it does not stop at discovery.
Collapse Potential tells you a latent feature is wanted and possible. It does
not tell you how much work it is. Integration distance D(x) is a second,
independent observable: given the files where a capability's pieces were found,
it measures how far apart they are in the repo's import graph.
Collapse Potential → how strongly the system wants this (is it latent?)
Integration distance → how far apart the pieces are (how hard to build?)
These are orthogonal, so they sort candidates into actionable quadrants:
- high CP, low D → RIPE. Strongly wanted, pieces already connected. A quick wire-up — build it now.
- high CP, high D → DEEP. Strongly wanted, but pieces scattered with no shared dependency path. Real composition work.
On the example, the strongest candidate by CP is deep, while a slightly weaker one is ripe — exactly the prioritization a backlog can't give you:
▸ Enterprise Renewal Risk Review CP 0.667 D 1.000 DEEP (4 unconnected files)
▸ Weekly Digest CP 0.421 D 0.333 RIPE (2 files, already importing)
D prefers a symbol reference graph — distance between the actual
definitions a capability resolves to (does buildDigest really reference
sendNotification?), so two functions in the same file that never call each
other read as deep, not falsely co-located. When a capability's signals don't
resolve to ≥2 defined symbols (other languages, or signals that map only to
uses) it falls back to the file import graph (JS/TS and Python imports
resolved). Like every RUMI field, D is scan-independent.
scan measures Collapse Potential over capabilities you declared. discover removes that scaffolding — it is handed no capabilities.json and proposes the capabilities itself, purely from the correction field. It clusters corrections by what users keep pushing toward, derives each cluster's capacity signals from the tokens that recur across it, scans the repo for them, and runs the ordinary collapse engine. A dense, coherent knot of corrections that no one named is exactly where an emergent feature first becomes visible — before there is a word for it.
Clustering is distributional, not just lexical: terms that keep appearing in the same context are treated as related, so corrections that share no surface words still group when they are about the same thing. On the example, "export everything to CSV" and "download all results" merge into a single emergent Csv / Download / Result capability despite having no word in common.
npm run discoverOn the bundled example, given 10 raw corrections and zero declared capabilities, RUMI's top emergent proposal is:
▸ Renewal / Risk / Blocker (×5) [emergent-renewal-risk-blocker]
proposed signals : renewal, risk, blocker, owner, report
Collapse Potential : 0.315
confidence : 0.162 ⚠ usage unverified
capacity in : risk-score.ts, report-builder.ts, owner-routing.ts
→ Candidate uncollapsed feature — BUT utilization is unknown: confirm it isn't already in use before acting.
It reconstructs the renewal-risk capability on its own and points at the real files — with no capability declared. Because an auto-proposed capability has no usage record, its utilization is unknown by construction, so every emergent candidate arrives flagged as a lead to verify, never a conclusion. This is the line between divining rod and horoscope: the proposal is only confirmed if naming and building it makes the correction pressure actually decay (experiment compare).
The clustering is local and dependency-free (corrections never leave the machine): distributional co-occurrence over the corpus itself. A pretrained-embedding backend — for true outside-world synonymy ("cancel my plan" ≈ "downgrade subscription") — is a planned optional add-on; it would keep data local but introduce a model download, so it is opt-in rather than default.
RUMI measures a target system's displacement field — a selection displaces alternatives, and the displaced field reveals latent features. But RUMI is itself a system that makes a selection: it picks the top candidates, and that choice depends on RUMI's own arbitrary configuration (the saturation scales, the unknown-usage prior, the classification thresholds). So RUMI can be turned on itself.
reflect sweeps RUMI's own parameter manifold (729 configurations), re-runs the
engine at each point, and measures collapse stability — does a discovery
survive RUMI displacing its own parameters?
npm run reflectHeadline: RUMI's top selection [enterprise-renewal-risk] holds rank #1 in 100%
of its own plausible self-configurations.
▸ Enterprise Renewal Risk Review [enterprise-renewal-risk]
stays rank #1 : 100%
classed a feature : 70%
most sensitive to : cHalf (feature 89% at cHalf=3 vs 33% at cHalf=6)
→ consistently RUMI's #1 pick; MOSTLY ROBUST — a feature across most of
RUMI's configuration space, threshold-sensitive at the margin.
This separates a robust discovery (survives RUMI's self-displacement) from a
fragile one (an artifact of the current tuning), and attributes any fragility
to a specific knob — here, whether renewal-risk reads as a feature depends most
on cHalf, how much correction volume RUMI requires to call pressure "high".
It is RUMI's own thesis — examine the field a selection displaces — applied to
RUMI's own selection, yielding a meta-confidence no single reading can.
The Level-2 verdict itself depends on how the reflection was configured — how
widely the knobs are perturbed, how fine the grid. reflect --level 3 sweeps
that reflection-design space (perturbation radius × granularity) and asks the
terminus question: does the stability estimate hold, or is it an artifact too?
npm run reflect -- --level 3On the example the recursion reaches a split fixed point:
RANKING CONVERGES — renewal-risk is RUMI's #1 pick in 100% of every
reflection design. A true fixed point.
CLASSIFICATION does NOT converge — its 'feature' label spans 44%–100% as RUMI's
self-perturbation widens (Level-2 reported a single 70%).
That is the honest terminus: the trustworthy invariant is the ranking; the single "70%" feature-stability figure was itself partly an artifact of how RUMI was set to reflect. Level-3 is precisely where that becomes visible — and where the recursion stops, because a Level-4 sweep would only keep re-measuring an already-proven-unstable quantity. Looking was the only way to know.
RUMI reads a data directory containing:
capabilities.json— the coordinates: each capability's id, label, and the codesignalsthat indicate capacity for it. Required forscan;discoverdoes not use it (it proposes capabilities itself).corrections.json— intent-gap signals tagged by capability. A correction (before→after) is the gold standard, but each event may set akind(see below) to admit other behaviour.usage.json— optional telemetry: how much each capability is actually exercised.
Plus a target repo, scanned locally for capacity signals. Nothing is uploaded. The instrument runs where the code lives.
A correction is just the cleanest evidence of the real thing — the gap between
what people are trying to do and what the system lets them do. Lots of behaviour
reveals that gap; the difference is whether it carries an arrow (it tells you
which way to build) or only heat (it tells you there's friction, but not the
destination). Each signal sets a kind:
| kind | weight | carries an arrow? |
|---|---|---|
correction |
1.0 | yes — the gold standard (before → after) |
request |
1.0 | yes — the ask names the destination |
repetition |
0.8 | yes — a repeated manual sequence implies the missing composite |
workaround |
0.6 | yes, if it names what was wanted |
abandonment |
0.5 | no — heat |
retry |
0.4 | no — heat |
Heat raises demand (RUMI sees that something is wanted) but earns no confidence (it can't say what). So a feature with only abandonment signals shows real correction pressure yet near-zero confidence and a "direction uncertain" flag — honest about the difference between "people struggle here" and "here's what to build." Confidence is what keeps the loud-but-vague signals from masquerading as actionable ones.
1.1.0 — working instrument: three-field engine, scan-independent Collapse Potential with per-reading confidence (unknown utilization is never mistaken for confirmed-unused), code-aware capacity across many languages (JS/TS via the TypeScript compiler; Python, Go, Ruby, Java, Rust, PHP, C# via tree-sitter; a comment/string/keyword-aware text analyzer as the fallback — a signal in a comment never counts as code), integration distance (a second observable ranking candidates ripe vs. deep from the symbol reference graph, with the file import graph as fallback), emergent capability discovery (discover) with distributional clustering — proposing undeclared capabilities from the correction field alone, grouping demand that shares no surface words — and recursive reflection (reflect, reflect --level 3), RUMI turning its instrument on itself to tell robust discoveries from artifacts of its own tuning — and testing whether that judgement itself converges. Plus a CLI, a three-view browser observatory (Scan / Discover / Reflect, computed live and fully local), and baseline/compare.
All parsing is local: tree-sitter runs on prebuilt wasm grammars shipped on disk — nothing touches the network at run time.
See docs/ARCHITECTURE.md for the design and roadmap. Next depth: meaning-based clustering for discover, import resolution for more languages, correction-capture SDK, VS Code panel, recursive / Level-3 analysis.