Skip to content

Latest commit

 

History

History
468 lines (402 loc) · 28.3 KB

File metadata and controls

468 lines (402 loc) · 28.3 KB

Roadmap

ThreadMesh uses milestone exit criteria rather than date promises. Priorities may change as the safety and adapter contracts become clearer. See the current status and active acceptance for the evidence-backed snapshot, current product decision, and ordered workstreams.

The roadmap now optimizes for one outcome: parallel agent sessions hand off completion, blockers, review findings, and dependency-ready state without the user acting as their message bus. ThreadMesh is the attention and admission policy layer; A2A, Cotal, ACP, or harness-native APIs may supply transport.

Active priority — existing desktop clients (2026-09-08)

Current order, confirmed by the user: independent first use → visible local status and stop → demonstrated relay savings → a real user case → DeepSeek and quota handoff. Work happens in original local agent conversations. The website helper is removed, including its code, browser-only tests and deployment workflow. Install once → select allowed tasks → ordinary work → useful receiver result is the first-use target. Manual prompt setup is a fallback, not completion.

Order User outcome Remaining acceptance
1 Install ThreadMesh and get a useful result inside Codex without maintainer help A supported installation path, readable activation and explicit task selection, with no website/form, source checkout, hand-edited JSON or long setup-prompt relay. Then verify an independent user's pair and B's own correct result. Existing-task adoption must be tested separately from new-task pickup.
2 Know what happened and stop safely In-task readiness/pending/result/stop explanations; never confuse idle with configured or sent with done. Guidance is being improved; persistent receipt/control and simultaneous-input safety remain open (#135/#136).
3 Less relaying than native-only use A small matched case counts setup actions, manual relays, correct edits and unwanted contact. No new benchmark framework; keep optional guidance small if it adds no benefit.
4 Others can understand and reproduce the value A consented real native recording and independent case, followed by relevant community sharing. No recreated conversation or promised star count.
5 Extend a proven useful workflow Provider-configured DeepSeek initiative, then continuation from an actual quota-blocked task's already saved checkpoint. Neither has a completed live acceptance.

Orders 1 and 2 form the immediate product work: fix observed local usability while seeking an independent participant, without pretending an internal test is that participant. The old #91/#93 multi-role loop and #7 formal review remain separate open tracks, not new prerequisites for this desktop-first alpha.

Completed evidence — do not repeat as a new gate

Retired experiment: the website generated setup text but did not meet the installation-first requirement. Its source and deployment are removed; the historical API/client result remains valid bounded evidence, not a current website feature.

Completed this slice: public workflow + chat-link pairing, both original tasks enabled, ordinary request triggering A's chosen advice, original B's correct own edit with prior decisions preserved, and both stopped. The bilingual no-terminal entry uses the official copied local-chat link, not a global task list or maintainer-local skill path. This is a controlled maintainer link-input pass; manually copying links in the GUI is not yet verified.

Native-value checkpoint: Codex already supplies the tested native messaging and continuation. The skill is optional guidance, not new transport. Read the responsibility map and when not to install. The retained desktop evidence shows feasibility, not improvement over native Codex alone.

For community growth, first make one independent Codex user's own pair succeed; fix their first blocker, then prepare a consented real recording and a concise case study. A consolidated evidence-backed reply has been posted; further updates need new results or actionable feedback. No fabricated video, unsolicited promotion or guaranteed star count.

Primary audience: Codex users, especially separate existing conversations inside Codex desktop. Pi is a supported option, not a prerequisite or replacement. A Codex-only first-use entry must reuse the user's own Codex login/model, never silently switch harnesses, and keep old-desktop-session acceptance separate. Success in a newly created App Server thread pair is not native desktop adoption.

The Codex-first installed-package case now passes with the default command in 272.604 seconds; alpha.3 is published and its public install was checked. Earlier timeout and outdated-runtime failures remain recorded. This is new-session CLI acceptance, not the primary existing-desktop gate.

Completed native slice: the skill-only workflow uses task tools already exposed by Codex, with an explicitly selected pair. It needs no Node/MCP/hook setup. One controlled opted-in desktop pair passed, including prior context, B's own edit and busy/stop checks. The subsequent failed title-based entry and empty-preview diagnosis are retained. The September 8 deep-link run above completes selected-task lookup and the full controlled public-source handoff without a list operation. Independent manual onboarding and normal plugin activation remain open. Sending is not race-free. External adapter/hook adoption is a separate portability route, not a prerequisite for trying native guidance.

The desktop-first plan supersedes the ordering below. CLI integration is not no-terminal first use.

Priority: same-product sessions first, multiple related workstreams second, cross-product interoperability and quota recovery third. The first useful experience must not require a second agent product or account. The Pi-pair CLI evidence is a developer baseline, not a pass for existing desktop chats. The fresh same-product copy run confirmed model-selected handoff and same-session continuation, but failed the free-plan meaning check. Keep content correctness open alongside entry. The bounded repair follow-up subsequently passed one same-product copy run and a fresh no-contact control. Broader quality and the earlier cross-product copy failure remain open; return to existing-session desktop entry instead of further prompt tuning.

  • Review official Codex/ZCode entry points; inspect local ZCode settings.
  • Prepare a developer hook probe; no messaging or native pass claimed.
  • Add dual-host MCP/hook identity correlation diagnostics and passive Codex endpoint checks; keep native receipt and desktop ownership unverified.
  • Verify plugin loading, native identity and adoption of a prior conversation. Native attempt: installation succeeded, but neither prior conversation exposed the diagnostic; keep open.
  • Controlled native-skill pair: two original same-client tasks with explicit links and scope, without a shared path or JSON setup. Independent user setup and the separate external-plugin adapter are not covered by this pass.
  • Prove one controlled model-selected native advice and same-receiver edit; source attribution verified in turn data, not a rendered UI recording.
  • Verify full business constraints, unrelated silence and user-input priority.
  • Package without developer prerequisites; observe an independent GUI user.

The external desktop adapter and independent-user entry remain unresolved. The controlled skill route does not close those gates. Do not substitute a new CLI session, unscoped remote control or a promotional UI for existing-conversation acceptance. Existing quality, quota and DeepSeek live gaps remain open.

Earlier delivery checkpoint — community feedback

The independent first-use report #158 was submitted on September 5 and acknowledged on September 7. It found a roughly five-minute, mostly quiet install and a Codex quota block before a model turn. Its public-API harness check is real external evidence, not a live agent collaboration pass. One report is not a community popularity ranking.

The following delivered work addressed that first failed user journey; the current ordered acceptance above supersedes its implementation sequence:

Implementation checkpoint: the packaged try --live entry, bounded failure handling and bilingual guides are implemented. Real copy and installed-package API cases passed; records and limits are retained. Independent first-user success and existing GUI conversations remain open; do not count maintainer samples as either gate.

  1. Package a self-contained real same-agent example through the public install. No repository checkout, custom harness code, user-created fixture or second agent account. Show setup progress and distinguish model/account failures from successful installation. Label simulation separately; never fall back to it while claiming live success.
  2. Show the model choosing a relevant peer, the same receiver continuing with its earlier constraints, and a useful artifact change with visible source. Measure setup steps, time to useful result and first failure. Existing Pi users can validate this developer entry; it does not satisfy GUI acceptance.
  3. Keep existing same-client desktop adoption as the primary integration gate. Use one supported host entry; if unavailable, record the exact missing host capability instead of repeating diagnostics or advertising desktop support.
  4. Ship and test the version users will install, then invite two more willing first users. Credit #158 without counting its quota-blocked run as a live pass. Fix their first blocking step before wider promotion; publish one short, truthful A-to-B case once the entry works independently.

Do not add another protocol, harness matrix or promotional redesign to this checkpoint. Cross-product quota recovery remains an extension, not a second subscription requirement. No deadline or star count is guaranteed.

Retained first-use alpha ledger (2026-09-05)

The first-use plan supersedes the older harness-expansion freeze below. Existing M0/M5 acceptance gaps remain open; they are not prerequisites for a usable, honestly labelled local alpha.

  • Shared local workspace with names/goals, persistent inbox and four tools.
  • Invocation-scoped Codex/Pi launch, project-scoped Kimi MCP configuration.
  • Official DeepSeek Harness MCP plugin integration and native runtime check.
  • Explicit portable checkpoints and cross-harness continuation command.
  • Practical API, preferences and quota previews, labelled as simulated.
  • First-run feedback form accepts unsuccessful installs and silent agents.
  • Real Pi-pair ordinary-task advice → idle wake → file update → independent business assertion; real Pi continuation from an explicit saved checkpoint.
  • Second real task family: an approved name/free-tier change reaches the website session, which updates copy and leaves unrelated price/data work alone. Evidence and limits.
  • Native Codex task-time context plus one ordinary Codex → Pi API pass: same receiver session, model-selected advice and verified receiver file edit.
  • Fresh unrelated-change control with real Codex discovery/inbox calls: zero source send attempts and no same-receiver follow-up.
  • Preserve the complete business meaning in the second cross-harness task: the copy run delivered and edited, but lost the free-plan qualifier. Keep this failed assertion visible; a successful receipt is not verified completion.
  • Native resume/attachment with a real earlier project context, not merely a newly named workspace or seeded checkpoint.
  • Recover useful work after an actual quota-blocked session using a prior saved checkpoint; distinguish retained context from unavailable history.
  • Native busy-receiver/queued-user-input experiment; current guard evidence is deterministic, not a real typing-race result.
  • Close live ordinary-task collaboration evidence, including DeepSeek with an available provider credential; publish failures alongside successes.
  • Get three independent first-run reports and fix the first blocking step.
  • Measure useful messages, irrelevant contact and checkpoint recoverability across at least three different everyday tasks, not only pagination.
  • Publish one concise evidence-backed community demo and one relevant harness integration submission; target 100 legitimate stars, not paid growth.

Do not claim arbitrary-host idle wake, lossless chat migration, production security, reliable speedups, or independent adoption from maintainer tests. The next focused product acceptance is tracked in #156, including the retained copy-quality failure and prior-session adoption gap.

Historical milestone ledger

The following sections preserve prior accounting. Where execution order conflicts, the active priority above governs; incomplete evidence stays open.

M0 — Foundation and protocol draft

  • Define the problem, scope, and core terminology.
  • Separate notify, suggest, steer, and interrupt.
  • Establish context sovereignty and least-authority principles.
  • Publish draft envelope and capability schemas.
  • Resolve the initial design questions captured by ADRs 0004–0007.
  • Run distributed-systems, safety, and adapter internal review lanes.
  • Publish authenticated authority and executable operation bindings (#15, #17).
  • Define crash-safe receipts, unknown-outcome reconciliation, disposition CAS, and durable harness-idempotency gating (#19).
  • Define typed interruption results and authenticated verification attestations (#16).
  • Enforce summary, relationship, disposition, and capability coherence (#18).
  • Accept two independent design reviews.

Current accounting: 10 milestone issues closed and 1 open. The internal reviews approved the conservative experimental prototype after fixes, but they do not satisfy #7.

Exit: a reader can implement a compatible prototype without relying on undocumented assumptions, and two independent reviews have accepted the safety and distributed-systems boundaries.

M1 — Local reference coordinator

  • Versioned SQLite storage, migration, rollback, retention, and deletion contract.
  • Complete durable task registry, mailbox, and scoped audit API.
  • Complete relationship- and intent-based policy engine with stable reasons.
  • Complete freshness, idempotency, expiry, receipts, and reconciliation.
  • Local event stream and provenance inspector.
  • Two-profile mock-harness conformance kit.
  • Retention-driven sensitive-content purge.

M1 is merged and its GitHub milestone is closed. This is experimental reference runtime evidence, not a production deployment claim.

Exit: two mock harnesses can discover, notify, suggest, accept, reject, defer, and explicitly decline unsupported steer/interrupt behavior with a complete audit trail. Successful interruption is not an M1 exit requirement unless the typed cancellation contract is implemented.

M2 — First real adapters

  • Codex App Server adapter.
  • Generic subprocess/JSON-RPC adapter.
  • Minimal installable harness SDK and short integration example.
  • One real non-Codex harness pass through the shared coordinator path.
  • Adapter capability negotiation and graceful degradation.

Kimi Code 0.38.0 now passes a real accepted suggestion through ACP with exact binary/capability evidence, context admission, and delete-plus-absence cleanup. The Codex App Server path has a real receiver pass, a model-selected A-to-B case, and the first scored control/relevant/irrelevant comparison with exact cleanup. Repetition and interference-budget evidence remain open. Gemini CLI headless stream-json is selected as the materially different third harness. Its pinned official package and no-model capability preflight pass; the checklist remains open until an explicitly authorized real model executes the shared scenario. The same receiver-accepted suggestion passes real Codex App Server and Kimi ACP products, plus deterministic ACP, Codex, and Gemini fakes. Gemini live remains optional rather than a competing mainline.

Exit: the same scenario runs across at least two different harness families.

M3 — Proactive dependency discovery

  • Explicit dependency graph.
  • Privacy-preserving task summaries.
  • Bounded model-selected relationship lookup and send experiment.
  • First no-contact and irrelevant control conditions.
  • Receiver decision and interference-cost budget.
  • Evaluation suite for useful versus harmful coordination.

The initial repetition matrix rejected default enablement. The shorter outcome-bearing benchmark and two-stage policy subsequently passed relevant 3/3, while fresh control used no tool and fresh irrelevant performed one read-only lookup without sending or activating B. A real Codex-to-Kimi case then passed with exact cleanup. The bounded profile is therefore eligible for explicit experimental opt-in, while repository-wide default enablement remains off during pre-alpha.

Exit: proactive coordination improves task outcomes in a benchmark without exceeding the defined interference budget.

M4 — Reusable harness integration kit

  • Export a transport-agnostic proactive tool bridge from the package.
  • Bound relationship discovery and suggestion budgets per model turn.
  • Publish a runnable sender-plus-receiver harness example.
  • Verify the packed package from an external consumer project.
  • Collect the first independent harness-integration feedback.

Exit: a harness can add bounded proactive discovery and suggestion without importing coordinator, adapter, or validation internals.

M5 — Attention and handoff router MVP

The deterministic vertical slice passes locally and from a packed consumer, and M5.1 has a real two-session Codex pass. The fixture closes the full lifecycle chain after one user kickoff: A → R → same-A → V → dependent, with zero fixture-runner activation dispatches, phase/business prompts, manual relay, or polling. The pump starts protected receiver decision and admitted business turns. Trusted finalization precedes the dependent turn, the irrelevant control starts no turn, and exact cleanup passes.

The retained foundation includes:

  • coordinator-owned decision/admission activation plumbing in #118 at 2a0d8550abc1a8c5dcebceb86d0372ea8d337b4d;
  • the in-process autonomous event pump in #119 at d37cb428ea84b0683dac24787889e259a0a18c71;
  • verifier finalization, dependent gating, and exact preverified provenance in #120 at 3b91dcff82622a0fed936e8295b77905777c6ada.
  • durable per-dispatch selection and publication recovery in #122 at 711da6606ac8b0c326f199a96d1713bc7a6de68c, including publication leasing, fencing, and committed-orphan recovery;
  • protected exact multi-tool receiver turns in #124, and the operator-run event-pump gate plus strict Codex evidence boundaries in #125#127.

The deterministic chain remains fixture evidence: liveProductEvidence=false, deterministicPolicyOracle=true, externalIndependentVerifier=false, the signer is a fixture-owned ephemeral key, and there is no global cross-dispatch selection chain. Five real Codex event-pump attempts then failed closed at successively narrower product boundaries. A sixth attempt completed the full real proactive A -> R -> same-A -> V -> dependent chain with one kickoff, nine bound native turns, zero later runner prompts or direct activations, an irrelevant zero-turn control, and exact cleanup.

This exposed an execution-order imbalance rather than a change in product direction. The behavioral checkpoint is passed. The bounded Git worktrees and process-isolated child verifier are now wired into the correlated path by #133, with deterministic positive and wrong-finding negative coverage. Attempt 16 then completed the merged real-effects live path. A measured operator-triggered control completed with nine actions and exact cleanup; its same-condition ThreadMesh arm failed closed at reviewer admission. The deterministic M5.3 relevant 3/3, irrelevant, stale/unverified, restart/replay, and failure-cleanup matrix passes. Product token cost remains unavailable. New substrate, generalized recovery, cross-harness, transport, and protocol expansion remains frozen. No partial integration attempt is promoted to M5.2 evidence.

  • Ship a one-command local demo with generated identities, grants, example sessions, and an inspector (#89).
  • Make completed, blocked, needs-input, review-failed, artifact-ready, and dependency-satisfied the primary product events (#90).
  • Route accepted and sufficiently verified events to eligible dependent sessions without treating receipt as verification or authority (#90).
  • Demonstrate a real Codex-first implementation/review/fix loop with zero manual relay and zero polling turns (#91):
    • M5.1: prove the real Codex dependency wake/unlock seam using durable cursor reconciliation; the adapter remains idleWake: false.
    • M5.2 fixture: prove the no-plan single-kickoff A/R/same-A/V/dependent chain with trusted pre-turn finalization, zero irrelevant turns, and exact cleanup; persist each dispatch through selection, turn settlement, and publication recovery. This does not satisfy the real-product M5.2 gate.
    • Real-chain checkpoint: retain one fresh Codex event-pump A/R/same-A/V/dependent run with one kickoff, zero runner phase/business prompts or direct activations, exact real session/turn/dispatch bindings, dependent ordering, an irrelevant zero-turn control, and exact cleanup. The earlier behavioral run's simulated effects remain explicitly labeled.
    • Reuse the existing bounded Git topology and process-isolated child verifier in the correlated event-pump implementation, with exact cleanup and no new coordinator or verifier subsystem.
    • Retain one successful live Codex traversal of that real-effects path; attempt 16 completed with certificate-verified connectivity and exact cleanup.
    • Add executable manual workflow accounting: one kickoff plus four checks plus four relays is a nine-action lower bound, versus one ThreadMesh kickoff. The real operator-triggered control later measured all nine actions and elapsed time; product tokens remain unavailable.
    • Add the active-receiver negative: a completion remains pending at a checkpoint while B stays running, with zero steer, interrupt, or native-turn starts.
    • M5.2 closure: combine the successful correlated real-effects run with a complete measured two-arm baseline while keeping raw product data out of public output. The control arm passes; the ThreadMesh arm failed closed.
    • M5.3: deterministic relevant 3/3, irrelevant, stale/unverified, restart/replay, and cleanup pass. Real Codex relevant 3/3 and the complete live baseline remain open.
  • Add one bounded, read-only correlated handoff state vector after the current live-value gate, without adding a scheduler or transport (#135).
  • Prove a real busy receiver is not silently steered and preserve admitted context across restart/compaction through a durable attention inbox (#136).
  • Repeat the loop across Codex and one ACP-compatible harness (#93).
  • Publish the bounded inspector and reproducible deterministic evidence record (#92).
  • Publish a 76-second evidence walkthrough generated from fresh executable demo output, with retained real Codex evidence and honest claim boundaries.

The executable closure gates for the real-agent phases are in the M5 real Codex loop plan. A local verifier simulation proves plumbing only and must not be represented as an independent external service.

Exit: a new operator can run and understand the closed loop in under 15 minutes; the bounded scenario has zero manual relay, irrelevant wakes, and incorrect dependency unlocks.

M6 — Independent adoption and ecosystem bridges

  • Collect three independent setup attempts and one completed real workflow.
  • Publish the 15-minute operator challenge and structured report template.
  • Close #79 with independent harness-author feedback.
  • Make ACP the preferred multi-harness gateway.
  • Map ThreadMesh lifecycle, evidence, and admission semantics to A2A without duplicating A2A transport.
  • Prototype a Cotal transport bridge only after the local loop passes.
  • Receive one external connector contribution or equivalent clean-room integration.
  • After #91 and #93 each have real evidence, add durable workflow budgets, recursive stop, and an owner-visible circuit breaker (#137).

Exit: an operator outside the maintainer organization completes a useful loop without maintainer intervention, and the integration contract is ready for a versioned 0.1 release candidate.

M7 — Evidence-driven production hardening

  • Prioritize claimant leases, crash recovery, authentication, isolation, and remote transport from observed operator failures.
  • Define service-level expectations only after a real persistent deployment.
  • Add wake, steer, or interruption capabilities only for a validated workflow that cannot use checkpoint admission.

Exit: production claims are backed by real deployment evidence rather than prototype inference.

Explicitly deferred

  • Public agent discovery across trust domains.
  • Payments, markets, or autonomous contracting.
  • Cross-user coordination without an identity and consent design.
  • A hosted multi-tenant control plane.
  • A general-purpose orchestrator, DAG engine, or agent-team framework.
  • New protocol intentions that are not required by the M5 closed loop.
  • Gemini live validation as a competing product mainline.
  • A global cross-dispatch chain, full OS-kill matrix, and long-turn heartbeat until the real-chain checkpoint shows they block the user-visible behavior.
  • Kimi flagship-loop parity, new harnesses, external verifier service design, and further Git-evidence generalization until the real Codex chain is retained. Existing foundations remain available for reuse after that gate.
  • Inspector and README presentation polish that does not record new product evidence.