diff --git a/.agents/skills/augustus/SKILL.md b/.agents/skills/augustus/SKILL.md index 4c52665..3d4e50c 100644 --- a/.agents/skills/augustus/SKILL.md +++ b/.agents/skills/augustus/SKILL.md @@ -1,6 +1,6 @@ --- name: augustus -description: "Use when placing typed probabilistic judgment (Jev-class System One / decision models) with mathematical, logical, or algorithmic mental models — in AI, software, business, knowledge work, or life, not only SWE; deciding where a fast cheap categorization/classification/scoring model belongs versus generation, exact policy/code, or proof; applying expected utility, selective classification/abstention, calibration, cost-sensitive thresholds, value of information, MCDA, signal detection, search/control substitutions, or Leveson-style org/safety; using NATM/snap-fit/Norman as design intuition; designing mixed architecture (decision model + LLM writing); auditing an existing system, PR, workflow, or non-software practice for judgment-shaped holes and code smells; debugging a question that hovers near 0.5, clusters mid-scale, or hides two judgments; placing agent self-supervision gates (pre-action, output judge, done-check, stuck-detector, context sieve); coupling a typed judge as an optimizer metric (Ax, DSPy); choosing among TypeSafe Jev, open heads (Laya, kev, openjev-lm, Nimble, encoder DeBERTa, LoRA distill, open multimodal RLCD / blackwood), announced open decision-model (Watch — still not landed), constrained-AR (TypeAR, pcdServer, decision-token LoRA), diffusion structured reads, GLiNER/GLiClass/GLiGuard encoder family (locate vs categorize vs safety-schema classify vs local multi-head), listwise rankers, or vision scorers; placing judgment beside TLA+/Alloy/Apalache/Dafny/DST (Antithesis, Resonate, PufferLib) without laundering a Noul as a proof; answering \"it's just classification\", \"is Jev probabilistic programming\" (marginals vs joint, not a PPL), \"low/medium/high entropy\" (allocator, not a meter), \"perception specialist then judgment vs shared multimodal System One\", \"wait for Archer vs open multimodal RLCD\", \"screenshot/DOM candidates → typed Choice\", \"eval path\", \"jevals\", \"Harbor taskset\", \"shared bake-off ECE/NLL/Brier\", \"LLM-as-judge is not the System One score\", \"pipeline / measure / hill-climb perception into a decision\", \"Ax vs DSPy\", \"held-out\", \"correctness is not confidence\", \"is this only for software?\", Alloy vs Apalache, GLiNER vs Jev, \"is GLiGuard Jev?\", LLM-as-judge, paraphrase brittleness, allowlist then judge (allowlist *proves*; fail-open cannot block), \"missing other → confident wrong Choice\", \"lint the Jev request\", \"training confronts Choice other / none-of-the-above\", \"S1 reflex keeps control / optional S2 one-use advice\", \"soft AGENTS.md rules vs linter\" (Abide / jev-pref), \"edit-phase vs turn-phase observation window\", banded confidence fail-open preference lint, \"extractive selection / pointer-not-generator\", \"encoder GLiNER compaction vs Jev Score compaction (same job; pointer not summarizer)\", \"fail-closed keep_full under mutation envelope\", \"CI flaky-vs-real merge gate\", \"fail-open VOI wake/resume (Horvitz)\", \"claim/evidence Stop integrity\", \"S1 extract + escalate-S2 indexer\", \"Harbor on/off routing\", \"policy-as-judgment PR marshal\", \"shadow-mode compaction rollout\", \"Jev Ultrafast vs GLiNER Ultrafast (observe-score-act backend-agnostic)\", \"hybrid local decide + remote fill\", \"DONE ≠ verified success\", \"observed a11y/DOM candidates vs screenshot multimodal\", \"evidence-preserving stdout prune (not summarize)\", \"hard token/format envelope then soft Noul\", \"fail-safe keep original on prune failure\", \"stdout prune vs session compaction\", \"specialist S1 computer-use (Cua-S1 form-v0; not TypeSafe Jev)\", \"plan ≠ execute / dry-run default\", \"observed-element option head (fill/check/click/skip)\", \"local /v1/systemone drop-in (stub until hf scorer)\", \"dataframe-native semantic columns\", \"route≠memory\", \"advisory sidecar receipts\", \"structure induction over bags\", \"AST ∩ semantic lint\", \"extractable-from-state / retrieve first\", \"decision-model vs constrained-LLM bake-off\", \"dual-process S1 decide / S2 generate\", \"combinatorial grid ≠ extractive\", TOCTOU-of-Noul, vacuous specs, open weights vs constrained decoding vs encoder vs LoRA vs kev, whether a decision needs a model at all (meta-VOI), env-break vs policy-break, sqlite-jev / in-engine vs CLI store index, hard safety envelope (Jev proposes, code clamps), host-adapter routing (not MCP), distill-to-device memory gate, \"uncalibrated local likelihoods vs Noul / CUDA replica\", \"decision-native RAG retrieve wide then decide then evidence set\", \"classify-first MCP / read selectively\", \"living applied-mappings atlas / class patterns not a 342 hit list\", \"draft-gate silence as safer / heartbeat\", \"robotics text-state vs pixels\", \"verbatim session ledger / scored recall\", \"judgment as language primitive / English-as-config\", \"pre-registered AMBIGUOUS eval / cascade sign-flip\", \"healthcare Harbor-shaped S1+S2\", \"pre-exec tool gate allow/block/review\", \"productized public primitive / judgment wall\", \"meaning-search without embeddings\", \"attention≠correctness PR review\", \"skills→oxlint / AST prove ∩ remainder\", \"session-sticky first-prompt routing\", \"measured RAG rerank vs generative rerank\", \"Stagehand extract pick-and-copy / judge\", \"harness observe-score-act productization\", or \"formally verify with Jev\", \"capability kernel / secrets never in the agent\", \"Jev is SENSOR not policy\", \"type-safe ≠ correct\", \"typed control plane around DSPy\", \"native-probability calibration / Brier/ECE arena\", \"fan-out as measurement economics\", \"engine owns truth / Jev owns judgment\", \"human-confirmed kill gate\", \"train specialist when downstream reads p vs few-shot hosted when only argmax\", \"decide→policy→LLM leftover cascade\", \"Noul 0.5 cannot-tell never rounded\", \"calibration ≠ sortable / ORDER BY over Jev probs\", \"pairwise inversion / Score ordinality / two-decimal ties\", \"wire-compat self-hosted /v1/systemone GLiFormer\", \"class-backend economics\", \"loopback gateway hosted + local OpenJev\", \"do not distill Jev as teacher of record\", \"active-learning triage / training-data VOI\", \"index-once ask-many / citable evidence packets\", \"meaning-grep AND/OR/NOT line Nouls\", \"closed-vote-only computer-use / no planner LLM\", \"Jev vs local MLX PCD Harbor\", \"PCD O(1) speed ≠ calibrated Noul\", \"host-owned handlers × System One\", \"OMP/pi fail-open acceptance gate\", \"permission vs probability / operator owns the safety bar\", \"judgment ≠ permission / Jev never grants access\", \"eval integrity / instrument not score / dinostomp jev-as-if\", \"constrained optimizer + S1 features / never sole hot-path gate\", \"privilege ≠ verdict / effect contracts not tokens\", \"attention filter / VOI for human review / never blocks / never green unless sure\", \"measurement owns endorsement / evidence-gated question packs\", \"Jev supplies evidence / code owns authority\", \"ranking ≠ calibration / never hard-threshold raw p as frequency\", \"hot-click CU / indexed element table / S1 on click path\", \"Jev judges relevance / code decides structure / never rewrite\", \"local rules first then remainder / never auto-train on model's own hides\", \"combinators / System One as control plane / not chat turns\", \"receipts not leaderboard / type-safe ≠ correct jaggedness\", \"VOI over skill library / skillranker abstention\", \"OOD calibration / AUC ≠ ECE / sign of miscalibration by type\", \"Jev vs thinking-budget small models / frontier-100\", \"turnstile / replayable evidence≠authority\", \"MLX one-pass schema→JSON / Apple Silicon replica economics\", \"memory leases ended by new evidence\", \"never confidently wrong / TLA+ compose with judgment / escalate instead of hard-gate\", \"no seal no advance / coverage ledger / mint ≠ product brain\", \"skill-broker sibling turnstile/skillranker / judgment ≠ permission\", \"sureness / CERTAIN|CONFIDENT|LEANING|TORN|CLUELESS / max_prob is generous\", \"JevBench Harbor/jevals practice / calibration not in Main Score\", \"CI typed gate before expensive review / ci-gatekeeper\", \"Codex MCP host adapter / jev_select_capability\", \"judgment as attention redirect not merge blocker / jev-preflight\", \"compress-before-first-send / dizk jev-lens vs rashed attention filter\", \"tools≠use / SessionStart over hoping the model recalls\", \"observational memory / keep-kind verbatim / pi-om\", \"open-Jev class / openvons / JevPick menu decode\", \"physical-world System One / HA-Jev / not for locks\", \"judgment outside the store / jevql CLI\", \"landed-script trust / headless≠auto-approve\", \"digital-design combinators / extended Router Loop Retry Fallback Memory\", \"VOI cache admission / same-intent skip LLM\", \"BM25 vs Jev skill routing Harbor harness\", \"zeroshot vs BERT / contamination DiD / label-equivalence\", \"typed escalate continue abort baton / inverted loop\", \"worth-your-attention VOI / ThinkyMiner Winnow vs kevinpita winnow\", \"Jev WHETHER Python HOW LLM WHAT\", \"conflict vs ignorance / named Choice escape\", \"Playwright executes Jev chooses / sample-from-distribution\", \"OpenJev /v1/decide not TypeSafe drop-in\", \"SemIf wire-compat runoff\", \"decision-as-memory flywheel\", \"record/replay CI / jevassert\", \"failure-finding arena / jevarena ≠ jev-arena\", \"BBQ stereotype/uncertainty/cost\", \"decider≠executor / jeffrey\", \"sentence-as-rule lint / jevlint ≠ JevLint\", \"VOI hunk prune / prune-review\", \"whole-repo intent VERIFIED/VIOLATION/UNKNOWN\", \"GLiNER2 System One spec ≠ replica\", \"Rust/WebGPU grande / Clojure Laya byte parity / CPU SemIf\", \"ONNX ModernBERT local-jev measured not equivalent\", \"persist constraints across compaction / pi-heed\", \"calibration+cost as first-class gates\", \"Harbor-shaped Jev vs schema-guided LLM-as-judge / jev-judge-bench ≠ jevarena ≠ jevbench\", \"hand no-text steps to Jev / jev-use / Vercel drops confidence / margin fallback\", \"Pi System-One control plane / pi-jev-control\", \"generation as tree of Choices / never free-generates / jev-gpt\", \"OpenRouter recipe atlas / samples not benches / jev-cookbook\", \"personal history feed / no social graph / jevfeed\", \"competing NAR claims / dual-channel ECE / claim-verification / openJev-verdict ≠ OpenJev\", \"empty compaction-proxy skip / IPECTER\", \"throughput ≠ latency / like-for-like ECE\", \"1-token logprob endpoint ≠ Noul / coverage ≠ correctness / chakuho\", \"open replica engine / jevinf / argmax-parity ≠ ECE\", \"unofficial Elixir SDK ≠ OTP peer / dannote/jev\", \"jevex rename + n=16 SWE VOI / files-to-read\", \"commit pre-review attention≠verdict / middle band never rounded / commitjev\", \"Hermes plugin is Agnes not TypeSafe\", \"pi-jev-compact ≠ pi-jev-compaction / verbatim summarizer replacement\", \"empty Codex-proxy skip / IPECTER runway\", \"decision-native inbox / mailordinal / humans own ambiguity\", \"unofficial jev-cli not ready / ≠ jevql\", \"laya-multilingual / English checkpoint confident-wrong OOD / ships uncalibrated\", \"schema-conditioned DeBERTa scorer / peaked ranking ≠ calibration\", \"HF 401 access / GitHub 404 Hub-only\", \"productized System One HTTP / classifier.dev / label+confidence public contract\", \"escalate-under-threshold / smart tier 0.7 / multi-label ignores tier\", \"silent-fallback FALLBACK marker / granite 0.546 vs advertised 0.800\", \"vs_jev tracked JSON not transcription / read eval/README before quoting\", \"choxos/jev-reviewer ≠ egma-ai / systematic-review pointer-not-generator\", \"two-pass Choice+Noul / relative which-line + absolute does-this-line\", \"not-found is an answer / no paraphrase invent\", \"human check as productized judgment / checked never overwritten\", \"githubnext/localjev ≠ kunchenguid/local-jev / prompted JSON ≠ structured logit read\", \"wire-compat ≠ logit-equiv / self-reported probs / entropy confidence\", \"institutional open-replica / GitHub Next /v1/systemone\", \"Harbor-shaped bake-off AG News BoolQ SST-5 / 1200-request caveats\", \"LM Studio runner gap / structured-read primitives for OpenJev parity\", \"NandhaKishorM/laya packaging ≠ Hub-only / Router script-before-p\", \"post-T ECE ≠ raw ECE / Banking77 token-budget / 0.85 still soft / not TypeSafe drop-in / external census ≠ scored bake-off / GLiNER2+routers class-boundary / incomplete vs watch / Harbor honesty watch / JevBench v1.2 geometric-mean I/C/S/K / cal now ON rank / weight sensitivity / option-order 72→21 / instruction models class-boundary / ×2 latency assumption / est. costs / Laya absent gap / Qwen3.8 27B ≠ Archer\", \"hourly already-folded watch / apply-the-five / skip thin noise\", \"hard-gate Noul as PR gate is soundness theater / totally-tim/jev-gate ≠ jev-gateway\", \"S1 never stalls waiting / S2 one-use advisory\", \"purple telemetry = consumed not arrived\", \"Local controller ≠ githubnext/localjev\", \"seed = geometry not async replay\", \"20% starting gate still soft / schema-safe ≠ correct\", \"no pixels to either provider / confidence ≠ selected probability\", \"experimental viz not a flight controller / S2 never grants\", \"OCR+AX observe-score-act / typesafe-computer-use\", \"never send screenshot to frontier for the decision\", \"overlapping CU options = false low confidence\", \"split kind/item/site / offscreen\", \"writer/decider split + post-type Noul still soft\", \"155× one-screenshot Harbor-shaped ≠ taskset\", \"AX never sole / Spotify 0\", \"decision ≠ answer-reader capture\", \"typesafe-computer-use ≠ jev-ultrafast ≠ cua-s1 ≠ camoufox\", \"ASR observe-score-act / jev-voice-browser\", \"partial-speech VOI / complete Noul / free-text waits\", \"spoken confirm ≠ hard auth\", \"numbered overlay disambiguate without another model\", \"moritzkremb/jev-voice-browser ≠ jev-voice-control ≠ nikolas-j\", \"wrap-as-execution / AgentGhost ALLOW ASK DENY\", \"rules first then Jev remainder / ASK throws / fail-closed\", \"reddpy/AgentGhost ≠ jwen5419807/agentghost ≠ vventirozos\", \"JP genre atlas / studio_yebisu / stars ephemeral ≠ eval\", \"Jev Clearly Explained / akshay_pachaar / LLM hammer\", \"schema-safe ≠ correct / 200× 400× TypeSafe ceiling\", \"questions-as-code / shadow first / not a TypeSafe how-to\", \"proposition ≠ embedding / contrast-set refund\", \"boolean composition of soft Nouls / AND OR NOT after threshold\", \"uehaj/jev-semgrep ≠ semgrep.dev\", \"meaning-grep dedicated fold / not a gate\", \"decision-validated UI / Jev never authors text / gram-render\", \"decision-as-assert / jevtest ambiguous band\", \"typed decisions drive UI / jev2ui\", \"hybrid S1 closed verb menu / anima3 / jeff confidently flat\", \"pointer-not-generator search / JevFind\", \"jev-frontier-bench ≠ frontier-100 / ChaosNLI JS\", \"product bakeoff ≠ architecture duel / jev-gliclass-bench\", \"four engines same questions / majority floor / calibration ≠ discrimination\", \"authorship named escape / not courtroom evidence\", \"ha-switchboard HA remains execution / ≠ HA-Jev\", \"n8n classify/route/score / Low Confidence abstention\", \"fast-jev-compaction-pi ≠ pi-jev-compact ≠ pi-jev-compaction\", \"jevloop full-distribution optimizer / no LLM in the loop\", \"laya-vision SmolVLM / score untrained / ≠ blackwood ≠ Archer\", \"Cerebellum-2B /v1/decide ≠ TypeSafe / wire-compat vs agent-routing\", \"laya-grounded not drop-in / phishing regress / Platt not temperature\", \"GestaltLabs/Jeff-1 ≠ logan-markewich/jeff / acc vs ECE n=9730\", \"stanley-code empty findings ≠ approval / human promote\", \"findme ≠ JevFind / NL memory beam-search FS\", \"jevsubrouter price workers not conversation / counts ≠ dollars\", \"feelings .feels() default 0.5 is Noul-0.5-never-rounded / ≠ hunch ≠ Probably\", \"apa-agent-harness ≠ AntonioCoppe/jev-harness / unpublished npm\", \"grok-bot-jev skill cannot force a bot that ignores it / A/B proxies not tokens\", \"Essentiel-Jev never authority / human every action\", \"enzo-mcp independently falsifiable claims / ≠ jev-sift\", \"pigeonhole OTHER skip / decision-as-filing\", \"jev-reliability Nothing about accuracy\", \"clduab11/jev-test ≠ realZachi/jevtest / Nothing runs yet\", \"jev-rag-benchmark Jev wins is not an assumption\", \"dairui1/jev-lab ≠ BrendanH18/jev-lab\", \"jevmail gmail.readonly / mailjay archive/trash\", \"ZHUBoer/ego-jev reserved __none__\", \"runWorkflow completed ≠ success\", \"jsort scores are relative\", \"Noul not Choice for scale\", \"groundedness-judge-bench native vs schema-guided\", \"implicit_true included in yes\", \"jev_playground 0 promotions\", \"routing-backtest 0.0447%\", \"yuyang2230/jev-agent-skill jev-1.13-free\", \"jev-techstack-classifier stack_config.json\", \"s1_ruby collapse late\", \"undecided? abstain\", \"2389-research/judgement license null\", \"confidence ≠ winner p\", \"typesafeai-sdk-community not a new species\", \"tpellet/hunch exit 3\", \"never-execute list\", \"jev-file-search scores not calibrated accuracy\", \"jev-linkmap Jev never sees S2 prose\", \"muhammedilyasy/jev-mail metadata only\", \"tidy none-of-folders stay\", \"tab-bouncer pinned/audio/current never closed\", \"lkclean Show fail-open\", \"jev-yt-time-saver Show anyway\", \"ORIGIN pause-if-no-Jev\", \"validResponse sums-to-1\", \"jev-crawlers risk bands never raw boolean\", \"jevbrain AUTO_ACT is not a Noul\", \"judgekit YAML classify/score/route/verify\", \"typed-judge-kit verdict-in-code\", \"alsoleg89/decide packing VOI\", \"0.8 ≠ 80% accuracy\", \"Jev-Calibration Platt ECE 0.117→0.052\", \"jev-calibration-arena never acts\", \"ctmx/openrouter-jev-mcp Decision-as-Plugin\", \"FrancoisChastel/jev-code ≠ npm jev-code\", \"claudecode-jev-marketplace fail-open not hot path\", \"pedroknigge/mcp_jev packs not ask_jev\", \"cyrusasco/typesafe-mcp noul deadband 0.35–0.65\", \"codaaiteam/jev-skill jevtypesafeai.com ≠ TypeSafe\", \"hermes-switchyard ≠ hermes-jev-router ≠ hermes-plugin-jev\", \"nanoprune 2.8MB ECE 2.58%\", \"smartdio/jev-browser-agent ≠ ZHUBoer/ego-jev\", \"Dakai/omp-jev-web DONE ≠ proof\", \"hari007sh/jev ≠ dannote/jev\", \"0thernet/system-one-skills deterministic verify\", \"typed-gate band [0.40,0.60] is refusal\", \"pi-jev-gate fail-closed; choice is the verdict\", \"Foq ~25ms/2.2GB local\", \"rev prefill-only + HF jev-0.5b\", \"robfrase/jev planning memo\", \"typesafe_agent_gates 27/27 / 31/31\", \"EpicEric/safe-sh static remainder\", \"pastepilot Confirm before act\", \"Jev-Reranker live Jev not yet measured\", \"sessionwise opt-in relevance\", \"jev-search pointer sieve\", \"400ms Salesforce WebMCP\", \"typesafe-scheduler-diagnostics advisory\", \"droidjev screenshot-free\", \"Tewoto1 jevcu planner still writes\", \"ha-conversation-jev Jev→Grok\", \"dsh-jev can only gate\", \"jev-classification-benchmark specified not run\", \"jev-luna-pagerduty p≥0.50\", \"meldltd/meldecision laya-go ONNX\", \"laya-doom never pixels\", \"logixism/laya-api empty README\", \"akpsahan/laya ≠ Archer\", \"choxos/jevchess engine owns truth\", \"jev-drive sim not AV\", \"story-arc Jev never authors\", \"jev-hs-assistant HS6\", \"golergka/jev-plays-starcraft-2 UI-verified ≠ API Victory\", \"awesome-jev-use-cases catalog\", \"Nibir1/typesafe-go ≠ official\", \"fingerprint after redact\", \"recall vs decide\", \"publish fingerprints+answers\", \"CI replay as Harbor cousin\", \"Cache hit ≠ correctness\", \"hyperspaceai/jevcache ≠ kushals256/jevcache\", \"human labels only\", \"score never auto-accepts\", \"production capture flywheel\", \"sutro-sh/jev-align ≠ caiovicentino/jev-align\", \"guidance ≠ hook\", \"catalysts ≠ summaries\", \"compile-time System One\", \"unofficial ≠ TypeSafe\", \"format_version modernbert-jev/1\", \"Argos1111/jev_local ≠ us/jev-local ≠ kunchenguid/local-jev\", \"LFM default ≠ ModernBERT backend\", \"Nemotron ≠ TypeSafe Jev\", \"not a calibrated replacement\", \"djev-dev complements djev-spark\", \"images as Choice options\", \"Laya essay numbers *theirs*\", \"Router/OOD confidence\", \"hosted bootstrap ≠ silent TypeSafe\", \"difficulty + policy thresholds + JSONL trace\", \"jev-codex-pilot model + reasoning depth\", \"keep/shadow/hybrid/reject\", \"quarry evidence projection\", \"Frank-ZY-Dou/awesome-jev robotics/3D/control\", \"one-dollar-tahoe TypeSafe Jev defense eval\", \"jevguard calibrator/cache/escape\", \"jev-ci-selector CI shadow mode\", \"llama-jev llama.cpp replica\", \"petercr/jev-orchestrator ≠ FleeexCorp/jev-orchestrator\", \"seb4ez/jevguard ≠ AseemPrasad/JevGuard ≠ pablozr/JevGuard\", \"webNeat/llama-jev ≠ WiktorB2004/llama-index-jev\", \"OpenCode jev-pruner context sieve\", \"observe→score-candidates→prune\", \"jev-zen / jev-1.13-free\", \"zen-chat ≠ Noul\", \"fail-open original\", \"keepScore >0.1 floor\", \"host port of tamaratran/jev-pruner\", \"indiejoseph/opencode-jev-pruner ≠ nrdz-labs/fast-jev-opencode\", \"jev-webagent-bench empty stub\", \"Kiln-AI/jev_jsonschema noul_threshold 0.5\", \"NSStudent/JevSwiftSDK unofficial\", \"GLiNER2 native Apple path\", \"unofficial Swift/Core ML GLiNER 2.5-small\", \"entity spans + confidence\", \"not Choice/Score/Noul\", \"not TypeSafe\", \"label descriptions as schema\", \"on-device ANE economics\", \"honesty locks\", \"shershah1024/gliner-native-runtime ≠ Fastino\", \"≠ gliner25-compaction ≠ gliner2-ultrafast ≠ Eran-BA/Jev_from_GLiNER2 ≠ NSStudent/JevSwiftSDK ≠ jevmlx\", \"default threshold 0.1 still soft\", \"Decision Graph Protocol frame→assess→commit\", \"app retains permissions/effects\", \"Jev-first assessor-neutral\", \"guarded commit / receipt/next frame\", \"assessment batching\", \"hard-gating DGP as safety theater\", \"numerous-com/dgp ≠ TypeSafe official\", \"jegrep calibrated path+range Nouls\", \"no embeddings/index/daemon\", \"~$0.01–0.03 typical\", \"agent --json\", \"can1357/jegrep ≠ Bentlybro/jevgrep ≠ uehaj/jev-semgrep\", \"Archer-arch fidelity\", \"kev family OOD 0.76–0.77 vs Jev 0.86\", \"block-causal isolation\", \"pointer/readout CE-trained\", \"/v1/systemone drop-in\", \"replica honesty\", \"cost-sensitive decision theory × System One probabilities → control flow\", \"thresholds derived from costs not hard-coded\", \"YES / NO / UNSURE from cost_false_yes / cost_false_no / cost_human\", \"auto-batching same-object questions\", \"Kungie/gut ≠ tpellet/hunch ≠ carldaws/hunch\", \"judgment vs generation\", \"deterministic execution after probabilistic judgment\", \"exactly one app-owned callback\", \"explicit uncertain branch\", \"Illusion47586/judge ≠ lexingtonhibiki/judgekit ≠ Ascurse/typed-judge-kit\", \"variable-N option scoring as the trainable object\", \"dynamic candidate bags not fixed label sets\", \"zwliJay/jev-forge ≠ NanoJev\", \"open replica economics / latency vs closed Jev\", \"NAR local drop-in\", \"wfzyx/von late-catch HIGH\", \"competing NAR claims / replica honesty\", \"typed judgments vs chat judges on guardrailing\", \"ishaannk/llm-vs-jev cross-note only\", \"deeper integrity fold is rh-guard\", \"nothing wins outright\", \"can be argued out of guarding\", \"Jev IS the if-statement\", \"judgments/probabilities drive branches\", \"text model only writes prose\", \"interpreter owns variables/loops/budgets/replay\", \"otherwise maybe / confidence gate\", \"chaos samples after the gate\", \"southpolesteve/probably ≠ carldaws/hunch ≠ feelings ≠ Kungie/gut ≠ Illusion47586/judge ≠ tidymodels/probably\", \"133★ / forks 10 live\", \"build calibrated classifiers from human feedback\", \"retrieve by relevance not resemblance\", \"one calibrated yes/no per memory in one request\", \"pointer mode 17/18 19/20 *theirs*\", \"embedding resemblance misses the allergy\", \"samdotmak/jev-recall ≠ jev-search ≠ jev-sift ≠ carryforward ≠ chopratejas/invalidate\", \"memory leases ended by new evidence\", \"six Nouls then fixed rules in code\", \"0 of 157 false invalidations\", \"questions/plans/directives are not evidence\", \"unsure → review queue\", \"host keeps the store\", \"name↔body / comment truth / test-claims\", \"mizchi/jev-lint is mizchi/jevlint rename\", \"no shipped rule has severity error\", \"~1 in 5 findings wrong *theirs*\", \"mizchi/jev-lint ≠ huntedman/JevLint ≠ MichitoSugawara/jev-lint\", \"JSON Schema → typed JSON via Jev\", \"noul_threshold 0.5 decoder not a proof\", \"IncompatibleSchemaError lists every bad property\", \"on-device Laya CoreML ANE\", \"~5 ms P50 short decisions\", \"189/189 FP16 checkpoint parity\", \"10× not achieved\", \"mizorewww/laya-coreml ≠ gliner-native-runtime ≠ jevmlx ≠ NandhaKishorM/laya\", \"softmax over allowed tokens ≠ Noul\", \"question-first cache\", \"Micha0827/snapjudge ≠ githubnext/localjev ≠ jevmlx ≠ cendress/SnapJudge\", \"Jev-first Pi agent loop\", \"slow-LLM fallback\", \"explicit action menu / CandidateSource unimplemented\", \"62 tests wiring not quality\", \"direwolfiy/JevPi ≠ standardagents/jevpilot ≠ pi-jev-control\", \"resume-screening bias audit methodology\", \"name×resume factorial independent Nouls\", \"callback determined by resume quality\", \"mean-probability name gaps operationally negligible\", \"natemoo-re/bias-bench ≠ BBQ\", \"Plan/PRD panel → code-owned pass|review|block\", \"cheerleading out of scope\", \"austindixson/planalyzer ≠ single-goodness Noul\", \"cost-aware multi-model routing/escalation\", \"decide vs do\", \"successful-task cost\", \"cannacre8ive/switchboard-ai ≠ ha-switchboard ≠ hermes-switchyard\", \"frozen-protocol zero-shot bench\", \"TypeSafe Jev vs PrismNLI vs Laya\", \"contamination caveat\", \"elcronos/jev-vs-open-decision-models ≠ JevBench ≠ DMB\", \"context-window admission control\", \"VOI gate which tokens are worth the expensive model\", \"fail polarity per lens\", \"on small inputs lenses lose money\", \"cvsgireesh/jevusher ≠ jev-sift ≠ winnow\", \"typed decision control plane\", \"receipt ≠ authorization\", \"historical-v0 zero retained cases\", \"MokiMeow/jev-fabric ≠ jev-forge ≠ dgp\", \"live 15-dim typed rubric re-score per pause\", \"scoring economics exemplar\", \"OpenJev/Codiv ≠ TypeSafe hosted\", \"jose-troche/live-rubric ~$0.000004 desc / ~$0.000006 README\", \"adversarial pre-registered Jev eval\", \"28 predictions before data\", \"123,805 requests\", \"confidence does not track ignorance\", \"polite injection 65% / crude 0%\", \"willkelly/jev-evaluation ≠ jevals ≠ jev-baselines-eval\", \"provider-neutral Elixir/BEAM Noul/Choice/Score SDK\", \"class infrastructure\", \"nshkrdotcom/system_one_sdk ≠ typesafe_sdk ≠ dannote/jev\", \"question-linting of Jev questions themselves\", \"nine jaggedness rules, no API key, no labelled data\", \"static lint ≠ measured separation\", \"yodablocks/jevq ≠ tenbin ≠ JevLint ≠ commitjev\", \"open-weights Laya as class exemplar (binding)\", \"Nx/Bumblebee runtime\", \"host chooses backend\", \"ChristianAlexander/laya_ex ≠ system_one_sdk ≠ dannote/jev ≠ NandhaKishorM/laya\", \"on-chain/edge Laya deploy\", \"parity_verified stays false\", \"model output never grants Tx\", \"humandebri/IC-Laya ≠ laya_ex\", \"auditable weekend replica\", \"Jev outputs never used for training\", \"soft human-vote distributions\", \"unpaired 0.577 vs 0.727\", \"agilabs-ai/jev48 ≠ JevBench ≠ Mapika/decider\", \"adversarial dual-judge / framing attack surface\", \"comparative framing is the usable judgment\", \"prior injection crowds out evidence\", \"copyleftdev/ember ≠ ember.js\", \"Laya specialist fine-tune pipeline\", \"training still GPU-pending\", \"PIXELZX0/XERON ≠ convaiinnovations/laya\", \"Hub Laya replica drop\", \"daliborsb/laya ≠ convaiinnovations/laya ≠ NandhaKishorM/laya\", \"System One student distillation corpus\", \"gold is programmatic\", \"teacher is closed-API clone\", \"do not distill Jev as teacher of record\", \"MagaBitmex/jev-4b-distill-data ≠ missing student checkpoint\", \"non-LLM VIN System One\", \"planning depth not chat\", \"lewislululu/jevon ≠ douglance/jevon\", \"source-bound evidence checks\", \"local quote mismatch needs no API\", \"exit 0 ≠ claim truth\", \"WaynezProg/jev-kit ≠ jonathanavis96/jev-kit (Airlock) ≠ jev-use ≠ jev-mcp\", \"independent System One evidence catalog\", \"scores not one leaderboard\", \"no external record currently reproduced\", \"TokenTrim no-Jev matched hybrid 62.4%\", \"reachjalil/system-one-bench ≠ mallahyari/system-one-benchmark\", \"21 tasks · 134 items · 208 questions\", \"scenes from public GitHub contracts, not production logs\", \"SivletLabs/jev-eval ≠ willkelly/jev-evaluation ≠ 4esv/jev-eval ≠ xxkuboxx/jev-eval ≠ onlyoneaman/jev-eval ≠ dayhaysoos/jevals\", \"option isolation (sibling-blind)\", \"permutation-equivariant\", \"Hub OWNER not published\", \"nafisazizir/hev ≠ jaredpalmer/kev\", \"frozen local LLM logits, no trained decision head\", \"residual-head 9,222-param decreased 73/96→67/96\", \"confidence = 1−normalized entropy, not P(correct)\", \"yuki-oshio/mini-jev ≠ r-ms/mini-jev\", \"Jev classifier as autoregressive next-token predictor\", \"ChatJev-style soundness theater\", \"erik-dunteman/ChatJev ≠ dannote/jev ≠ jev-gpt\", \"calibrated decision head × AlphaProof value head\", \"implementation-layer isomorphism, semantic difference\", \"timeout = censoring\", \"do not launder Noul as proof\", \"parallel rank-prediction vs serial selection\", \"independent questions can conflict\", \"zzzzzec/jevsort ≠ keltokhy/jsort\", \"curated open System One ecosystem catalog\", \"rupeshpoojary9/awesome-open-system-one ≠ AnotiaWang/awesome-jev\", \"arXiv paper radar with Jev relevance scoring\", \"ranking ≠ calibration / 0.5 still soft\", \"fail-open failed evals not marked seen\", \"train calibrated ~27M from scratch\", \"typed Q→prob dist / one forward pass / no LLM decode\", \"hyusi2003/MiniSystemOne ≠ Colvin0315/MiniSystemOne\", \"description-only stub / size 5\", \"ESCI hard probe fails four of six\", \"jev_bool ECE 0.242 inversion 0.255\", \"do not re-fold §60 six-gates as new\", \"jobbyjev one-request-per-company from batch-size result\", \"find/design/evaluate TypeSafe Jev decision loops\", \"karanb192/jev-architect ≠ samtay32/jev-system-architect\", \"Jairik/jev-distiller size 1\", \"distill-Jev UI stub / do not distill Jev as teacher of record\", \"post-launch scored use-case map / Jev self-scores then human curation\", \"licensedsaucer9-web/jev-opportunities\", \"Jev-inize a use case into classifier/router\", \"gavinHuang/jevinize → simple-jev not TypeSafe\", \"featherless-ai/simple-jev\", \"compare saved decisions / same label can still change the branch\", \"VihaanAgarwal/jev-diff ≠ Saik0s/diffusiongemma-jev-macos\", \"not tested with a live Jev API key\", \"constrained logprob + temp/Platt ≠ Noul\", \"OpenJevPro pastes openjev-sglang JevBench as own\", \"zhangcy122/OpenJevPro ≠ IamBusy/OpenJev ≠ ekzhang/openjev-sglang\", \"PolyForm Noncommercial\", \"SmolLM-135M / sub-70ms / 0 output tokens\", \"demo P(True) 0.5052 / Choice conf 0.2872 / Score conf 0.0055\", \"README claims MIT / GitHub license null / no LICENSE file\", \"patelvishwa112/jev-system-one-rlcd ≠ arnabgho/rlcd-lite ≠ blackwood-rlcd\", \"source-backed Awesome Jev radar / 306+ commit-pinned\", \"logicrw/awesome-jev-projects ≠ AnotiaWang/awesome-jev ≠ yibie/awesome-jev ≠ cobanov/awesome-jev ≠ rupeshpoojary9/awesome-open-system-one\", \"auto GitHub sync / Issue-only submissions\", \"hashed n-gram encoder / rival-aware attention\", \"olanotolu/jevbetter vs jevlike starter\", \"synthetic hard menus top-1 0.916 vs 0.873 / ECE 0.0182 vs 0.0367 / 40 vs 4608 menus/sec\", \"shuffled-context control 0.335\", \"structured probability readouts\", \"distribution > argmax\", \"Noul 0.5 midpoint\", \"score is expectation not integer\", \"bare HTTP not SDK\", \"Arohtea/jev-readout\", \"Jev-style Choice/Score/Noul from ordinary models\", \"optional DSH plugin\", \"schema-valid ≠ calibrated\", \"gulagala001/jevify ≠ Mintzs/jevify\", \"Laya RLCD benchmark\", \"40.3% below constant-answer\", \"open-weight measurement\", \"mourad-ghafiri/laya-rlcd-benchmark ≠ yibie/laya-jev-lab\", \"cheap fail-open semantic edge\", \"second signal not sole\", \"FastLoopError catch\", \"SupremeDreamZ/jev-fastloop ≠ jev-ultrafast\", \"asking more questions in one call\", \"0.980 at every N\", \"nearly not fully deterministic\", \"TheWebDevel/jev-fanout\", \"Qwen3-VL perception + Jev decisions train RL\", \"0 model calls at deployment\", \"VLM alone 1.7 vs +Jev 4.4\", \"harneet2512/reflexrl ≠ khordoo/jev-reflex-autonomy-lab\", \"independent Jev API vs Laya\", \"cascade 0.60 matches 78% at 1.8×\", \"noul facts not judgements\", \"yibie/laya-jev-lab ≠ dairui1/jev-lab ≠ BrendanH18/jev-lab\", \"GLiNER vs GLiFormer vs Laya vs Jev\", \"extractors ≠ decision engines\", \"Laya dict-instructions collapse 58.3%\", \"umstek/zero-shot-ie-bench\", \"decisions-per-minute & cost\", \"204 moves vs 73\", \"throughput not intelligence\", \"angelgalvisc/snake-arena-jev-vs-llms ≠ vtrivedy/jev-plays-games\", \"behavioral contracts\", \"pin expectations eval upgrades\", \"raw 0.94 is not a release\", \"sathariels/jevcheck ≠ dayhaysoos/jevals ≠ SivletLabs/jev-eval\", \"evidence-linked dependency upgrade\", \"Jev never generates filenames\", \"no_direct_evidence ≠ safe to merge\", \"GaneshVG18/upgrade-radar ≠ LYchoon/paper-radar-jev\", \"discography theme/mood/complexity\", \"five atomic questions one call\", \"lirantal/discoprint\", \"Turn any open LLM into System-One Jev\", \"uspraveen/Jevify ≠ Mintzs/jevify ≠ gulagala001/jevify\", \"Jevify-any-LLM architecture probe\", \"description-only stub / size 0\", \"Train encoder-only calibrated decision models from a task sentence\", \"Exu is a toolkit, not a method\", \"strictly proper scoring rule\", \"Pre-alpha\", \"Ruivalim/exu-base\", \"scratch-trained calibrated decision model\", \"typed Q → probability dists\", \"Colvin0315/MiniSystemOne ≠ hyusi2003/MiniSystemOne\", \"no published weights download URL\", \"90.5 seconds / 29.2% pipeline evidence\", \"p_i/p_j independent of other candidates\", \"Recipe for calibrated decision models — small model out\", \"init → synth → train → eval → serve\", \"91.1 % / ECE 0.022 *theirs*\", \"Jev zero-shot 75.1\", \"scienthoon/luce\", \"Put Jev's three headline claims on trial\", \"0.5B local GPU\", \"46x speedup / accuracy identical\", \"ECE 0.624 sentiment catastrophe\", \"bigger model worse calibration\", \"RichardoMrMu/jev-mini ≠ yuki-oshio/mini-jev ≠ r-ms/mini-jev\", \"System-1 decision engine for local LLMs\", \"structured choices only\", \"JSON parse of generated text ≠ Noul\", \"TypefAI JEV / Journal Entry Voucher\", \"tapsin/jev-local ≠ us/jev-local ≠ Argos1111/jev_local\", \"Jev 1.13 reward-model eval across 8 benchmark tracks\", \"40,940 examples / 0 API errors\", \"RewardBench v1 92.58%\", \"Precise IF 50.63%\", \"goya4140/jev-reward-model-evaluation\", \"Scaffolding in progress\", \"Jev vs LLM support-ticket routing\", \"static + live decision bench\", \"TypeSafe's own published benchmark\", \"illustrative simulations, not live API calls\", \"JevBench v1 — smart/cheap/fast/reliable\", \"I/C/S/K 25% geometric mean\", \"classifier.dev fast tier 84.8 is Jev behind its own API\", \"do not re-fold §78 v1.2 board as new\", \"Laya (421M) 70.1 now on board\", \"Zero-shot/few-shot LLM routing\", \"hard budget filter before Jev\", \"Jev never asked to perform budget arithmetic\", \"Jev judges the next state, XState enforces transitions\", \"simulation uses synthetic keyword fixtures\", \"catalog gravity\", \"v-modal/awesome-jev-tools\", \"★339 live REST\", \"curation is not endorsement\", \"crawler-maintained directory\", \"Daily GitHub + npm sweep, human-merged\", \"RadRebelSam/awesome-jev ≠ AnotiaWang ≠ yibie ≠ cobanov ≠ logicrw ≠ v-modal\", \"HF peft SPLADE/BGE reranker\", \"rdxtremity/jev-reranking ≠ carlaiau/jev-reranking\", \"query-side encoders, not a Jev replica\", \"ONNX System One Qwen3.5-4B scorer\", \"source:pngwn/system-one-qwen3.5-4b-scorer\", \"CC-BY-NC-4.0\", \"temperature 1.75\", \"transformers.js AutoModel cannot load this graph\", \"Consistency benchmark Space\", \"This Space contains no benchmark result yet\", \"12-case plumbing fixture\", \"Benchmark-driven Jev router and judge\", \"cheap alone is not success\", \"Jev does not write, sum prices, or claim accuracy %\", \"Sol 94.2 / Luna 83.9 / Jev path 89.7\", \"19.2% Sol / 62.3% cost save / 4.5pp miss of 2pp non-inferiority\", \"p50 latency worse than Sol due to routing overhead\", \"erendikmenn/jev-llm-router-benchmark ≠ jev-rag-benchmark ≠ ryantsai/jev-llm-router\", \"Express + node:sqlite\", \"mock and Jev decision engines\", \"previous_ticket_count >= 3 is code\", \"MIN_CONFIDENCE 0.6 still soft\", \"substring false positives\", \"aesaganda/jev-ticket-router ≠ SarathChandraBellam/jev-vs-llm-ticket-router\", \"Universal Figure & Diagram Router\", \"confidence ≥ 0.85 hard-gate is theater\", \"generative AI banned from scientific plots\", \"six visual branches\", \"hoangngochuong24947-gif/jev-figure-router\", \"human-labeled (state, question, label)\", \"166,054 rows / 22 configs\", \"soft_label for human uncertainty\", \"Praveenrajus/jev-bench ≠ fstandhartinger/jevbench\", \"ternary bonsai System One GGUF\", \"openjev's mechanism, Bonsai's weights\", \"Hub does not ship weights\", \"100/100 easy T/F is not Harbor\", \"label_mass ≠ correctness\", \"stock llama.cpp Q2_0 silently gibberish\", \"NicolaiMTLassen/open-bonzi-jev ≠ NicolaiLassen\", \"transformers.js DeBERTa ONNX\", \"source:com-kotobalabs/open-jev-deberta-v3-large\", \"temperature 1.05\", \"AutoModel from_pretrained works\", \"onnx-community/open-jev-deberta-v3-large-ONNX ≠ system-one-qwen3.5-4b-scorer-ONNX\", \"107★ densify\", \"GH 151M vs README 149.6M\", \"PR #1 now closed unmerged\", \"do not re-fold §71 claim-audit as a beat\", \"typed decisions, RLCD, confidence-gated routing\", \"structured ≠ correct\", \"mock not live API\", \"26 tests\", \"wjdjdakf17/jev-study ≠ baekenough/jev-study\", \"bonzi-27b-v2 / ternary-8b / 27b-v1 GGUF family densify\", \"Hub still does not ship weights\", \"WANLI-256 74.6% / 65.2% / 71.1% *theirs*\", \"Bonsai 1 27B Q1_0 runs on stock llama.cpp\", \"ternary still needs PrismML fork\", \"hf:heman10x/openJev-verdict-2.0 twin tokenizer-only\", \"OpenJev Vision image classification + uncertainty\", \"CLEVR-4 held-out joint 0%\", \"hfdataset:IamBusy/OpenJev-Vision-Research-v0.1 12,832\", \"294,912 derived targets not independent samples\", \"Laya multilingual ONNX WebGPU typed-decisions port\", \"63/63 selected answers / 5.1e-4 CPU / 1.2e-2 WebGPU\", \"UpHash-Network/mini-jev is yuki-oshio transfer\", \"jev-injection-bench 11,900 labelled prompts\", \"Jev best ranking / Haiku better ECE 0.021 vs 0.058\", \"0.5–0.9 band is where Jev's numbers do not mean what they say\", \"Prompt wording moves panic 28%\", \"manojlds/jev-dspy-bench ≠ dspachos/jev-dspy ≠ jmanhype/jev-dspy-lab\", \"Jev agreement is similarity, never ground truth\", \"no aggregate quality grade or merge gate\", \"AbstentionBench-on-Jev rank 1 of 20 vs 2025 field\", \"question-asymmetry\", \"forward-looking 0.465 never extreme\", \"openkev calibration layer not a runtime\", \"ECE vs coverage independent\", \"select_threshold returns inf\", \"escalation catches uncertainty not ignorance\", \"misakaikato/openkev ≠ jaredpalmer/kev\", \"pdf-race Docling→Jev vs Gemini\", \"parser owns the wall clock\", \"12/12 tie is a tie\", \"titles selected not generated\", \"flopcheck 16 calibrated tweet judgments\", \"mechanical tells in code\", \"ZeroX-01/jev-atlas ≠ Zaious/jev-capability-atlas ≠ gorock007/jev-atlas\", \"catalog not endorsement\", \"Laya calibration lab Gradio MCP\", \"T never changes argmax\", \"confidence ≠ top-label p\", \"easy probe set refused\", \"40–48 rows too small to ship T\", \"Gemma-4 26B-A4B jevify classification+calibration\", \"LoRA adapter twin not independent eval\", \"Gemma-4 E4B jevify\", \"E4B LoRA stub card\", \"kushalpatil/jevify-gemma4 ≠ Mintzs/jevify ≠ gulagala001/jevify ≠ uspraveen/Jevify\", \"GH kushalpatil07/jevify 404\", \"PAWS 0.580/ece 0.288 is the weak cell\", \"smaller E4B slightly better OOD ECE than 26B-A4B\", \"Hub jevify merged LoRA ships weights\", \"bonzi Bonsai-8B v1 GGUF densify\", \"Bonsai-1.7B v1\", \"Bonsai-4B v1\", \"WANLI-256 64.5% / 60.2% / 52.0% *theirs*\", \"rank #4 / #5 / #6 of 6\", \"JulesHuisman/jev-eval scaffolding / README SHA c356a584 (was empty e69de29b)\", \"JulesHuisman/jev-eval ≠ SivletLabs/jev-eval ≠ willkelly/jev-evaluation ≠ 4esv ≠ xxkuboxx ≠ onlyoneaman ≠ dayhaysoos/jevals\", \"7 bands 6/10 vs 40 bands 0/10\", \"source receipts + confidence slider re-policy without re-inference\", \"32/32 synthetic is smoke not production\", \"classify HF datasets across typed semantic dimensions\", \"roadus2 watch misspelling\", \"lock roadius2/ultra_laya\", \"ultra_laya REVIEW defects\", \"default branch claude/laya-jev-review-gg5ppo\", \"XNLI EN 88.3% ECE 0.032 → RU 77.3% ECE 0.096\", \"Δ −11.0 pp [−14.2,−7.8]\", \"ECE +0.063\", \"MASSIVE no detectable difference at n=600\", \"confidence is function of p_max (r=1.000)\", \"pointer-not-generator 400 human-authored responses\", \"proposed ≠ authorized\", \"FewRel 160: Jev 85.0% vs lexical 13.125%\", \"gated 100% (95/95) coverage 59.375%\", \"J++ composable semantic computation language\", \"judge-jev 0.5 still soft\", \"947 repos scored\", \"A 273 / B 302 / C 372\", \"LLM rubric ≠ benches\", \"No benchmark winner is claimed\", \"phishing: naive 62.6% vs regex 91.8%\", \"5-atomic + LR 95.0% *theirs*\", \"AITuber tension ±15\", \"README npm global\", \"repo is Rust\", \"git-confess code owns counting/blame/ratio\", \"httpx exhibit 11% (13/119) *theirs*\", \"90d trend +12.40% vs random +12.75% vs BH +41.71%\", \"5m win rate 25%\", \"Awesomejev 656 entries / 38,160 stars\", \"tracker likes 64 (+4) lastModified UNCHANGED\", \"Laya present\", \"Blackwood ABSENT\", \"Archer still promised_not_landed\", \"Blackwood tracker ABSENT; likes 2 gated manual\", \"r = c - p_a\", \"ECE 0.021; acc 0.807 vs warmup 0.746\", \"calibration beyond ~500 tokens unmeasured\", \"Independent primitive\", \"11.57s vs 54.10s · 4.67× · 120/128 *theirs*\", \"default path is pretrained Gemma probs not trained RLCD head\", \"GH Meanblock 404; lock leesk212/JEV-CPU\", \"softmax over letter slots ≠ Noul\", \"WANLI 0.741 vs openjev v2 0.77 *theirs*\", \"3-way NLI ≠ Noul\", \"priority 0.464 = majority floor\", \"banking77 contaminated\", \"raw margins not probabilities\", \"GH jev-haiku-benchmarking 404\", \"do not distill Jev as teacher of record (they distilled Haiku)\", \"“0.9 is not one number”\", \"ranking ≠ calibration\", \"banking77 0.8–0.9 stated 0.86 actual 0.73 over-confident *theirs*\", \"≠ Praveenrajus/jev-bench ≠ fstandhartinger/jevbench\", \"$0.0000153–$0.0000226 vs circulating $0.0004 (~20×)\", \"Score is 0..n-1 expectation not 0–1\", \"Noul has no confidence field\", \"TCP floor 198.8 ms\", \"type reliability is not a reason to choose Jev (json_schema 5/5)\", \"gateway tax not one number\", \"Function-only 5/8 vs hybrid 8/8\", \"4/8 without Jev\", \"8 designed cases not conversion lift\", \"≠ RadRebelSam/awesome-jev\", \"200-row pilot Jev 86.5% 173/200 vs Gemini Flash-Lite 86.0% 172/200 vs Pro 87.0% 174/200 *theirs*\", \"not a ranking\", \"NLI Tetris argmax P(entail)−P(contradict)\", \"情緒測謊器\", \"1q 396ms / 30q 567ms\", \"±0.03\", \"33q $0.000045 vs Gemini ~5× slower ~60× cost *theirs*\", \"≠ realZachi/jevtest\", \"8-example Jev vs GPT-5.6 Sol ~64× cost 5.4× latency *theirs*\", \"synthetic; no inference\", \"≠ JevBench v1.2 §78\", \"Judged 3317 / listed 2560\", \"Jev judges, code applies policy\", \"catalog ≠ endorsement\", \"APA “microsecond policy / zero hallucination” overclaim\", \"Client-side quiz; pointer from held docs; scanned-PDF warn\", \"CSP only api.typesafe.ai\", \"Jev judges / agent reasons / user decides\", \"selecting an option is not permission to implement\", \"degraded fallback\", \"pattern exact, judgement must clear floor\", \"no matching pattern → no model call\", \"not a correctness oracle\", \"$0.00022 vs chat $0.00306 *theirs*\", \"Spec vs artifact remainder\", \"treating 0.85 as 85% / minProbability hard-gate as Harbor\", \"VERIFY acquires discriminating evidence, never same-pool confidence-only rescoring\", \"fast/full/max are ceilings not sizes\", \"Solar writes, Jev chooses NEXT ACTION\", \"SemIf 2186★ (+20 vs §109 2166)\", \"jevlike 1038★ (+7 vs 1031)\", \"TypeAR 14★ flat\", \"AnotiaWang 96★ (+1 vs 95)\", \"yibie/awesome-jev 490★\", \"Laya likes 802 (was 783)\", \"tracker likes 64 flat, lastModified UNCHANGED\", \"do not reopen or amend PR #23/#24/#25/#26/#27.\", \"do not reopen or amend PR #23 or #24 or #25 or #26 or #27\", \"Calibration is not alpha\", \"NO CURRENT ALPHA CANDIDATE\", \"ΔR² approximately +0.00084\", \"Brier 0.2131387\", \"ECE 0.0421875\", \"Adding Jev probability to deterministic volatility improved Brier by only 1.4058e-05\", \"default 0.5 keeps zero non pinned\", \"keepResult median 0.14 to 0.17\", \"keepCall median 0.28 to 0.35\", \"usable range is about 0.10 to 0.25\", \"7.8% to 57.9%\", \"judges results it never sees\", \"task-finish eval not built yet\", \"$0.002 per compaction\", \"slavadubrov/sgr-judge-bench ≠ slavadubrov/jev-judge-bench\", \"Jev 108/120 $0.083 0.34 s\", \"Luna SGR 114/120\", \"paired Jev accuracy-difference intervals include zero\", \"not evidence of equivalence\", \"GLM SGR 26/120 93 format failures\", \"Terra-planned Jev hybrid 55/120\", \"rule-based by default, optionally Jev-backed\", \"empty README\", \"missing key cannot break the experience\", \"prefill plus exactly one decode\", \"softmax over A/B/C ≠ Noul\", \"BBQ 9,053/10,000 (90.53%)\", \"ECE 0.0890\", \"Mean confidence 0.9943\", \"overconfident\", \"score and noul not implemented\", \"DGUI 12 rows (was 6)\", \"INSTRUCT 119 rows likes 2\", \"encode the state once, decide everything in parallel\", \"0.740 accuracy against a 0.508 majority\", \"ECE 0.047\", \"fine-tune's advantage ends where its 384-token training data does\", \"jasonkneen/open-jev ≠ pngwn/open-jev\", \"same sha d41dc3cd\", \"Space does not call Jev\", \"recomputes routing from saved probabilities\", \"200-case Jev 97.0% / 100.0% / 95.0% / MAE 9.22\", \"synthetic repository benchmark\", \"Jev evaluations are advisory\", \"YehuiTang0316/jev-nlgrep ≠ Bentlybro/jevgrep ≠ can1357/jegrep ≠ uehaj/jev-semgrep\", \"default threshold 0.8 still soft\", \"40-line windows cannot prove whole function\", \"token-native sequential start/end Choice\", \"Gemini/Haiku stubs not configured yet\", \"handful of hand-written examples, not a benchmark\", \"Jev judged exactly what it was given\", \"laguagu/jev-skills ≠ laguagu/jev-evidence-lab ≠ Pleo2/awesome-jev-agent-skills\", \"contract_passed is not a claim of guaranteed factual truth\", \"Wilson lower bound 0.85 floor\", \"fixture mode no savings claim\", \"SemIf 2207★ (+21 vs §110 2186)\", \"jevlike 1043★ (+5 vs 1038)\", \"TypeAR 15★ (+1 vs 14)\", \"AnotiaWang 97★ (+1 vs 96)\", \"yibie/awesome-jev 506★ (+16 vs 490)\", \"Laya likes 822 (was 802)\", \"tracker likes 64 flat, lastModified UNCHANGED\", \"do not reopen or amend PR #23/#24/#25/#26/#27/#28\", \"people who compare Jev against GPT-5.6 has never fine-tuned BERTForXYZ for living and it shows\", \"zero shot classifiers\", \"scale them as much as decoder only models\", \"many problems solved with LLMs could have been solved with them, it was a skill issue\", \"opt for DeBERTa and ModernBERT ones\", \"BERTForXYZ → DeBERTa → ModernBERT\", \"Jev vs GPT-5.6 bakeoffs are a category error\", \"encoder / ZS classifiers\", \"institutional HF voice\", \"quote *theirs*\", \"do not invent accuracy numbers\", \"softmax/ZS scores still ≠ calibrated Noul\", \"soft scores ≠ hard gates\", \"@mervenoyann\", \"likes 421 / 189\", \"impressions 35498 / 9613\", \"multimodal image<>text ZS as perception front-end\", \"hf:MoritzLaurer/deberta-v3-large-zeroshot-v2.0 likes 139\", \"hf:MoritzLaurer/ModernBERT-large-zeroshot-v2.0 likes 72\", \"Bart, bert, deberta, modernbert, these are all LLMs\", \"Maziyar quoted\", \"Jev is exemplar not the mandate\", \"do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29\", \"Heman10x-NGU/Verdict-open-jev ≠ Heman10x-NGU/openJev-verdict-2.0\", \"TF-IDF + LogReg ECE 0.0207 vs Jev 0.1440\", \"Verdict-open-jev 48.07% vs Jev 90.80%\", \"abstention combined recall 10.00%\", \"p50 35.58 ms\", \"K=25 (maximum capacity) 72.00%\", \"0.85 coverage 84.60% selective risk 1.18%\", \"26.1× faster than standard Qwen JSON generation\", \"Jevify 90.0% / 167 ms CUDA graphs disabled\", \"Finding 1: Brier on stated confidence alone is a trap\", \"grpo_rlcr 0.78 / ECE 0.084\", \"reliability 0.007 but resolution 0.000\", \"27 900 schema-driven decisions\", \"13 600 / 13 600 questions\", \"candidate mass min 0.99999624\", \"22 configs · 166,054 rows · 4 calibration-gold\", \"sha a39eba3f\", \"Student B MAE 0.148 / Pearson 0.836 / 86.0%\", \"pngwn/open-jev-laya-bench README 404\", \"sha 9f69c742 likes 2\", \"HDFS 0.9933 (745/750) / retain 0.0084\", \"BGL ERROR/FATAL protection 1.0000\", \"2,479 / 2,500 HDFS uncertain\", \"cache hit 0.9648 (2412/2500)\", \"$0.153936 estimated\", \"E2 recomputes from saved probabilities\", \"Space sha eda59e0a\", \"MASSIVE English 0.783 / Khmer 0.033 / Hindi 0.133\", \"40–48 rows too small to ship T\", \"T never changes argmax\", \"siren2345/jev-single-decode-transformers ≠ siren2345/jev-single-decode\", \"Split Transformers experiment from llama.cpp runtime\", \"tanayvasishtha/jev-lab ≠ dairui1/jev-lab ≠ BrendanH18/jev-lab ≠ yibie/laya-jev-lab\", \"Four experiments stress-testing TypeSafe's Jev: calibration, bundle bias, label bias, and ensembling\", \"second pass must be $0.00 from cache\", \"The pages never call Jev\", \"Gemma 4 31B 77.0% / Jev 1.13.0 61.4% / Laya 322M 0.0%\", \"restriction state 95.0% against 84.4%\", \"None of the systems are particularly good at knowing when to stop and ask\", \"They skip the question and call a tool directly\", \"100% schema pass\", \"six-field joint 48.8% vs 72.8%\", \"ywchiu/jev_benchmark ≠ Running-Dolphins/jev-bench ≠ Praveenrajus/jev-bench\", \"ACT / REVIEW / FALLBACK\", \"A provider failure, timeout, malformed output, or missing answer is **not** a policy outcome\", \"confidence is descriptive provider output, not a substitute for probability\", \"Quality denominators include only valid scored answers\", \"an exact halfway tie chooses the lower level\", \"aiwithenoch/Jev-Skill ≠ simplosophy/jev-skill ≠ laguagu/jev-skills\", \"The local path does not claim to turn a smaller checkpoint into Jev\", \"Low support becomes decision: \"review\"\", \"MIT-0 SPDX NOASSERTION\", \"current-llm\", \"结构兼容,不是 Jev 模型能力\", \"altryne/jevify ≠ Mintzs/jevify ≠ gulagala001/jevify ≠ uspraveen/Jevify\", \"Find where Jev belongs. Design the questions. Measure the difference\", \"TypeAR-AI/TypeAR 301 → TypeLLM/TypeLLM\", \"TypeLLM/TypeLLM 16★\", \"SemIf 2241★ (+34 vs §111 2207)\", \"jevlike 1051★ (+8 vs 1043)\", \"AnotiaWang 98★ (+1 vs 97)\", \"yibie/awesome-jev 525★ (+19 vs 506)\", \"Laya likes 864 (was 822)\", \"tracker likes 67 (+3 vs 64)\", \"lastModified UNCHANGED `2026-09-20T04:29:16.000Z`\", \"do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#32\", \"soft Noul ≠ hard safety\" with a placement, not a stack replacement or a vendor how-to. Formal methods are one pillar., \"A hunch is a probability with a policy attached\", \"{ enter: 0.8, exit: 0.6 } is hysteresis\", \"replay a policy change without inference\", \"Decision models are providers, not the product\", \"huncho ≠ Kungie/gut ≠ carldaws/hunch ≠ tpellet/hunch\", \"pretrained Qwen2.5 base ECE 0.030 (0.5B) / 0.040 (7B)\", \"instruct 0.302 / 0.269\", \"70.9% → 70.0% mean conf 74.1% → 96.7%\", \"temperature scaling still matches it in-distribution\", \"No Jev API was called\", \"Qwen2.5 ≠ Archer\", \"Qwen/Qwen3.8-27B ≠ Archer\", \"calibration does not compose\", \"ECE has exactly zero statistical power to detect the failure mode that kills trajectories\", \"25–60× headline withdrawn\", \"P(all-correct): 0.0071 vs 0.0001\", \"TCE / AMS\", \"Deferred Crispification\", \"Qwen 3.8 sparring ≠ Archer\", \"pd.cut bins by equal width while jeval bins by quantile\", \"ECE 0.113 and ECE 0.076\", \"jeval drift is not implemented yet\", \"ranking ≠ calibration\", \"g0runmezadam/what-is-jev IS tunahansahin897/what-is-jev\", \"same GitHub id 1378007307\", \"947 repos scored\", \"A 273 · B 302 · C 372\", \"LLM rubric ≠ benches\", \"Probabilities are advisory, not calibrated guarantees\", \"light_cutoff_applied_to_combination 0\", \"AND: product (independence assumed and recorded in the trace)\", \"circuit-vl-4b ≠ Archer\", \"soundness theater / measurement theater\", \"hourly 0843 / notes.md §114\", , \"Jev Capability Resolver / NiazMorshed2007/jcr\", \"one tool to find documented deterministic commands in a nested capability tree\", \"returns context; **does not execute**\", \"skills = workflow+judgment; capabilities = individual operations\", \"format independent of Jev; proposed open standard exploration\", \"classify (Jev) → optional OpenAI decompose compound → beam search geometric mean of routing probs\", \"keep up to 3 paths ≥60% of best (JCR_BAND_RATIO 0.6)\", \"soft scores ≠ hard gates; 0.6 band is application policy\", \"routing ≠ permission; docs ≠ authority to run\", \"sol-vs-opus5-20 *theirs*; lookup+explain only; n=1 per cell; Not Harbor task-execution\", \"Claude Opus 5: agent input 108,585→15,819 (−85%)\", \"Codex GPT-5.6-Sol wall 25.3s→62.4s (Sol slower with JCR in 19/20)\", \"NiazMorshed2007/jcr ≠ skill-broker ≠ skillranker ≠ jev-sift ≠ jev-lens ≠ jevusher ≠ jev_select_capability\", \"do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33/#34\", \"notes.md §116\", \"soft Noul ≠ hard safety\" with a placement, not a stack replacement or a vendor how-to. Formal methods are one pillar. Not a substitute for the official typesafe-ai skill (live Jev API contracts)." +description: "Use when placing typed probabilistic judgment (Jev-class System One / decision models) with mathematical, logical, or algorithmic mental models — in AI, software, business, knowledge work, or life, not only SWE; deciding where a fast cheap categorization/classification/scoring model belongs versus generation, exact policy/code, or proof; applying expected utility, selective classification/abstention, calibration, cost-sensitive thresholds, value of information, MCDA, signal detection, search/control substitutions, or Leveson-style org/safety; using NATM/snap-fit/Norman as design intuition; designing mixed architecture (decision model + LLM writing); auditing an existing system, PR, workflow, or non-software practice for judgment-shaped holes and code smells; debugging a question that hovers near 0.5, clusters mid-scale, or hides two judgments; placing agent self-supervision gates (pre-action, output judge, done-check, stuck-detector, context sieve); coupling a typed judge as an optimizer metric (Ax, DSPy); choosing among TypeSafe Jev, open heads (Laya, kev, openjev-lm, Nimble, encoder DeBERTa, LoRA distill, open multimodal RLCD / blackwood), announced open decision-model (Watch — still not landed), constrained-AR (TypeAR, pcdServer, decision-token LoRA), diffusion structured reads, GLiNER/GLiClass/GLiGuard encoder family (locate vs categorize vs safety-schema classify vs local multi-head), listwise rankers, or vision scorers; placing judgment beside TLA+/Alloy/Apalache/Dafny/DST (Antithesis, Resonate, PufferLib) without laundering a Noul as a proof; answering \"it's just classification\", \"is Jev probabilistic programming\" (marginals vs joint, not a PPL), \"low/medium/high entropy\" (allocator, not a meter), \"perception specialist then judgment vs shared multimodal System One\", \"wait for Archer vs open multimodal RLCD\", \"screenshot/DOM candidates → typed Choice\", \"eval path\", \"jevals\", \"Harbor taskset\", \"shared bake-off ECE/NLL/Brier\", \"LLM-as-judge is not the System One score\", \"pipeline / measure / hill-climb perception into a decision\", \"Ax vs DSPy\", \"held-out\", \"correctness is not confidence\", \"is this only for software?\", Alloy vs Apalache, GLiNER vs Jev, \"is GLiGuard Jev?\", LLM-as-judge, paraphrase brittleness, allowlist then judge (allowlist *proves*; fail-open cannot block), \"missing other → confident wrong Choice\", \"lint the Jev request\", \"training confronts Choice other / none-of-the-above\", \"S1 reflex keeps control / optional S2 one-use advice\", \"soft AGENTS.md rules vs linter\" (Abide / jev-pref), \"edit-phase vs turn-phase observation window\", banded confidence fail-open preference lint, \"extractive selection / pointer-not-generator\", \"encoder GLiNER compaction vs Jev Score compaction (same job; pointer not summarizer)\", \"fail-closed keep_full under mutation envelope\", \"CI flaky-vs-real merge gate\", \"fail-open VOI wake/resume (Horvitz)\", \"claim/evidence Stop integrity\", \"S1 extract + escalate-S2 indexer\", \"Harbor on/off routing\", \"policy-as-judgment PR marshal\", \"shadow-mode compaction rollout\", \"Jev Ultrafast vs GLiNER Ultrafast (observe-score-act backend-agnostic)\", \"hybrid local decide + remote fill\", \"DONE ≠ verified success\", \"observed a11y/DOM candidates vs screenshot multimodal\", \"evidence-preserving stdout prune (not summarize)\", \"hard token/format envelope then soft Noul\", \"fail-safe keep original on prune failure\", \"stdout prune vs session compaction\", \"specialist S1 computer-use (Cua-S1 form-v0; not TypeSafe Jev)\", \"plan ≠ execute / dry-run default\", \"observed-element option head (fill/check/click/skip)\", \"local /v1/systemone drop-in (stub until hf scorer)\", \"dataframe-native semantic columns\", \"route≠memory\", \"advisory sidecar receipts\", \"structure induction over bags\", \"AST ∩ semantic lint\", \"extractable-from-state / retrieve first\", \"decision-model vs constrained-LLM bake-off\", \"dual-process S1 decide / S2 generate\", \"combinatorial grid ≠ extractive\", TOCTOU-of-Noul, vacuous specs, open weights vs constrained decoding vs encoder vs LoRA vs kev, whether a decision needs a model at all (meta-VOI), env-break vs policy-break, sqlite-jev / in-engine vs CLI store index, hard safety envelope (Jev proposes, code clamps), host-adapter routing (not MCP), distill-to-device memory gate, \"uncalibrated local likelihoods vs Noul / CUDA replica\", \"decision-native RAG retrieve wide then decide then evidence set\", \"classify-first MCP / read selectively\", \"living applied-mappings atlas / class patterns not a 342 hit list\", \"draft-gate silence as safer / heartbeat\", \"robotics text-state vs pixels\", \"verbatim session ledger / scored recall\", \"judgment as language primitive / English-as-config\", \"pre-registered AMBIGUOUS eval / cascade sign-flip\", \"healthcare Harbor-shaped S1+S2\", \"pre-exec tool gate allow/block/review\", \"productized public primitive / judgment wall\", \"meaning-search without embeddings\", \"attention≠correctness PR review\", \"skills→oxlint / AST prove ∩ remainder\", \"session-sticky first-prompt routing\", \"measured RAG rerank vs generative rerank\", \"Stagehand extract pick-and-copy / judge\", \"harness observe-score-act productization\", or \"formally verify with Jev\", \"capability kernel / secrets never in the agent\", \"Jev is SENSOR not policy\", \"type-safe ≠ correct\", \"typed control plane around DSPy\", \"native-probability calibration / Brier/ECE arena\", \"fan-out as measurement economics\", \"engine owns truth / Jev owns judgment\", \"human-confirmed kill gate\", \"train specialist when downstream reads p vs few-shot hosted when only argmax\", \"decide→policy→LLM leftover cascade\", \"Noul 0.5 cannot-tell never rounded\", \"calibration ≠ sortable / ORDER BY over Jev probs\", \"pairwise inversion / Score ordinality / two-decimal ties\", \"wire-compat self-hosted /v1/systemone GLiFormer\", \"class-backend economics\", \"loopback gateway hosted + local OpenJev\", \"do not distill Jev as teacher of record\", \"active-learning triage / training-data VOI\", \"index-once ask-many / citable evidence packets\", \"meaning-grep AND/OR/NOT line Nouls\", \"closed-vote-only computer-use / no planner LLM\", \"Jev vs local MLX PCD Harbor\", \"PCD O(1) speed ≠ calibrated Noul\", \"host-owned handlers × System One\", \"OMP/pi fail-open acceptance gate\", \"permission vs probability / operator owns the safety bar\", \"judgment ≠ permission / Jev never grants access\", \"eval integrity / instrument not score / dinostomp jev-as-if\", \"constrained optimizer + S1 features / never sole hot-path gate\", \"privilege ≠ verdict / effect contracts not tokens\", \"attention filter / VOI for human review / never blocks / never green unless sure\", \"measurement owns endorsement / evidence-gated question packs\", \"Jev supplies evidence / code owns authority\", \"ranking ≠ calibration / never hard-threshold raw p as frequency\", \"hot-click CU / indexed element table / S1 on click path\", \"Jev judges relevance / code decides structure / never rewrite\", \"local rules first then remainder / never auto-train on model's own hides\", \"combinators / System One as control plane / not chat turns\", \"receipts not leaderboard / type-safe ≠ correct jaggedness\", \"VOI over skill library / skillranker abstention\", \"OOD calibration / AUC ≠ ECE / sign of miscalibration by type\", \"Jev vs thinking-budget small models / frontier-100\", \"turnstile / replayable evidence≠authority\", \"MLX one-pass schema→JSON / Apple Silicon replica economics\", \"memory leases ended by new evidence\", \"never confidently wrong / TLA+ compose with judgment / escalate instead of hard-gate\", \"no seal no advance / coverage ledger / mint ≠ product brain\", \"skill-broker sibling turnstile/skillranker / judgment ≠ permission\", \"sureness / CERTAIN|CONFIDENT|LEANING|TORN|CLUELESS / max_prob is generous\", \"JevBench Harbor/jevals practice / calibration not in Main Score\", \"CI typed gate before expensive review / ci-gatekeeper\", \"Codex MCP host adapter / jev_select_capability\", \"judgment as attention redirect not merge blocker / jev-preflight\", \"compress-before-first-send / dizk jev-lens vs rashed attention filter\", \"tools≠use / SessionStart over hoping the model recalls\", \"observational memory / keep-kind verbatim / pi-om\", \"open-Jev class / openvons / JevPick menu decode\", \"physical-world System One / HA-Jev / not for locks\", \"judgment outside the store / jevql CLI\", \"landed-script trust / headless≠auto-approve\", \"digital-design combinators / extended Router Loop Retry Fallback Memory\", \"VOI cache admission / same-intent skip LLM\", \"BM25 vs Jev skill routing Harbor harness\", \"zeroshot vs BERT / contamination DiD / label-equivalence\", \"typed escalate continue abort baton / inverted loop\", \"worth-your-attention VOI / ThinkyMiner Winnow vs kevinpita winnow\", \"Jev WHETHER Python HOW LLM WHAT\", \"conflict vs ignorance / named Choice escape\", \"Playwright executes Jev chooses / sample-from-distribution\", \"OpenJev /v1/decide not TypeSafe drop-in\", \"SemIf wire-compat runoff\", \"decision-as-memory flywheel\", \"record/replay CI / jevassert\", \"failure-finding arena / jevarena ≠ jev-arena\", \"BBQ stereotype/uncertainty/cost\", \"decider≠executor / jeffrey\", \"sentence-as-rule lint / jevlint ≠ JevLint\", \"VOI hunk prune / prune-review\", \"whole-repo intent VERIFIED/VIOLATION/UNKNOWN\", \"GLiNER2 System One spec ≠ replica\", \"Rust/WebGPU grande / Clojure Laya byte parity / CPU SemIf\", \"ONNX ModernBERT local-jev measured not equivalent\", \"persist constraints across compaction / pi-heed\", \"calibration+cost as first-class gates\", \"Harbor-shaped Jev vs schema-guided LLM-as-judge / jev-judge-bench ≠ jevarena ≠ jevbench\", \"hand no-text steps to Jev / jev-use / Vercel drops confidence / margin fallback\", \"Pi System-One control plane / pi-jev-control\", \"generation as tree of Choices / never free-generates / jev-gpt\", \"OpenRouter recipe atlas / samples not benches / jev-cookbook\", \"personal history feed / no social graph / jevfeed\", \"competing NAR claims / dual-channel ECE / claim-verification / openJev-verdict ≠ OpenJev\", \"empty compaction-proxy skip / IPECTER\", \"throughput ≠ latency / like-for-like ECE\", \"1-token logprob endpoint ≠ Noul / coverage ≠ correctness / chakuho\", \"open replica engine / jevinf / argmax-parity ≠ ECE\", \"unofficial Elixir SDK ≠ OTP peer / dannote/jev\", \"jevex rename + n=16 SWE VOI / files-to-read\", \"commit pre-review attention≠verdict / middle band never rounded / commitjev\", \"Hermes plugin is Agnes not TypeSafe\", \"pi-jev-compact ≠ pi-jev-compaction / verbatim summarizer replacement\", \"empty Codex-proxy skip / IPECTER runway\", \"decision-native inbox / mailordinal / humans own ambiguity\", \"unofficial jev-cli not ready / ≠ jevql\", \"laya-multilingual / English checkpoint confident-wrong OOD / ships uncalibrated\", \"schema-conditioned DeBERTa scorer / peaked ranking ≠ calibration\", \"HF 401 access / GitHub 404 Hub-only\", \"productized System One HTTP / classifier.dev / label+confidence public contract\", \"escalate-under-threshold / smart tier 0.7 / multi-label ignores tier\", \"silent-fallback FALLBACK marker / granite 0.546 vs advertised 0.800\", \"vs_jev tracked JSON not transcription / read eval/README before quoting\", \"choxos/jev-reviewer ≠ egma-ai / systematic-review pointer-not-generator\", \"two-pass Choice+Noul / relative which-line + absolute does-this-line\", \"not-found is an answer / no paraphrase invent\", \"human check as productized judgment / checked never overwritten\", \"githubnext/localjev ≠ kunchenguid/local-jev / prompted JSON ≠ structured logit read\", \"wire-compat ≠ logit-equiv / self-reported probs / entropy confidence\", \"institutional open-replica / GitHub Next /v1/systemone\", \"Harbor-shaped bake-off AG News BoolQ SST-5 / 1200-request caveats\", \"LM Studio runner gap / structured-read primitives for OpenJev parity\", \"NandhaKishorM/laya packaging ≠ Hub-only / Router script-before-p\", \"post-T ECE ≠ raw ECE / Banking77 token-budget / 0.85 still soft / not TypeSafe drop-in / external census ≠ scored bake-off / GLiNER2+routers class-boundary / incomplete vs watch / Harbor honesty watch / JevBench v1.2 geometric-mean I/C/S/K / cal now ON rank / weight sensitivity / option-order 72→21 / instruction models class-boundary / ×2 latency assumption / est. costs / Laya absent gap / Qwen3.8 27B ≠ Archer\", \"hourly already-folded watch / apply-the-five / skip thin noise\", \"hard-gate Noul as PR gate is soundness theater / totally-tim/jev-gate ≠ jev-gateway\", \"S1 never stalls waiting / S2 one-use advisory\", \"purple telemetry = consumed not arrived\", \"Local controller ≠ githubnext/localjev\", \"seed = geometry not async replay\", \"20% starting gate still soft / schema-safe ≠ correct\", \"no pixels to either provider / confidence ≠ selected probability\", \"experimental viz not a flight controller / S2 never grants\", \"OCR+AX observe-score-act / typesafe-computer-use\", \"never send screenshot to frontier for the decision\", \"overlapping CU options = false low confidence\", \"split kind/item/site / offscreen\", \"writer/decider split + post-type Noul still soft\", \"155× one-screenshot Harbor-shaped ≠ taskset\", \"AX never sole / Spotify 0\", \"decision ≠ answer-reader capture\", \"typesafe-computer-use ≠ jev-ultrafast ≠ cua-s1 ≠ camoufox\", \"ASR observe-score-act / jev-voice-browser\", \"partial-speech VOI / complete Noul / free-text waits\", \"spoken confirm ≠ hard auth\", \"numbered overlay disambiguate without another model\", \"moritzkremb/jev-voice-browser ≠ jev-voice-control ≠ nikolas-j\", \"wrap-as-execution / AgentGhost ALLOW ASK DENY\", \"rules first then Jev remainder / ASK throws / fail-closed\", \"reddpy/AgentGhost ≠ jwen5419807/agentghost ≠ vventirozos\", \"JP genre atlas / studio_yebisu / stars ephemeral ≠ eval\", \"Jev Clearly Explained / akshay_pachaar / LLM hammer\", \"schema-safe ≠ correct / 200× 400× TypeSafe ceiling\", \"questions-as-code / shadow first / not a TypeSafe how-to\", \"proposition ≠ embedding / contrast-set refund\", \"boolean composition of soft Nouls / AND OR NOT after threshold\", \"uehaj/jev-semgrep ≠ semgrep.dev\", \"meaning-grep dedicated fold / not a gate\", \"decision-validated UI / Jev never authors text / gram-render\", \"decision-as-assert / jevtest ambiguous band\", \"typed decisions drive UI / jev2ui\", \"hybrid S1 closed verb menu / anima3 / jeff confidently flat\", \"pointer-not-generator search / JevFind\", \"jev-frontier-bench ≠ frontier-100 / ChaosNLI JS\", \"product bakeoff ≠ architecture duel / jev-gliclass-bench\", \"four engines same questions / majority floor / calibration ≠ discrimination\", \"authorship named escape / not courtroom evidence\", \"ha-switchboard HA remains execution / ≠ HA-Jev\", \"n8n classify/route/score / Low Confidence abstention\", \"fast-jev-compaction-pi ≠ pi-jev-compact ≠ pi-jev-compaction\", \"jevloop full-distribution optimizer / no LLM in the loop\", \"laya-vision SmolVLM / score untrained / ≠ blackwood ≠ Archer\", \"Cerebellum-2B /v1/decide ≠ TypeSafe / wire-compat vs agent-routing\", \"laya-grounded not drop-in / phishing regress / Platt not temperature\", \"GestaltLabs/Jeff-1 ≠ logan-markewich/jeff / acc vs ECE n=9730\", \"stanley-code empty findings ≠ approval / human promote\", \"findme ≠ JevFind / NL memory beam-search FS\", \"jevsubrouter price workers not conversation / counts ≠ dollars\", \"feelings .feels() default 0.5 is Noul-0.5-never-rounded / ≠ hunch ≠ Probably\", \"apa-agent-harness ≠ AntonioCoppe/jev-harness / unpublished npm\", \"grok-bot-jev skill cannot force a bot that ignores it / A/B proxies not tokens\", \"Essentiel-Jev never authority / human every action\", \"enzo-mcp independently falsifiable claims / ≠ jev-sift\", \"pigeonhole OTHER skip / decision-as-filing\", \"jev-reliability Nothing about accuracy\", \"clduab11/jev-test ≠ realZachi/jevtest / Nothing runs yet\", \"jev-rag-benchmark Jev wins is not an assumption\", \"dairui1/jev-lab ≠ BrendanH18/jev-lab\", \"jevmail gmail.readonly / mailjay archive/trash\", \"ZHUBoer/ego-jev reserved __none__\", \"runWorkflow completed ≠ success\", \"jsort scores are relative\", \"Noul not Choice for scale\", \"groundedness-judge-bench native vs schema-guided\", \"implicit_true included in yes\", \"jev_playground 0 promotions\", \"routing-backtest 0.0447%\", \"yuyang2230/jev-agent-skill jev-1.13-free\", \"jev-techstack-classifier stack_config.json\", \"s1_ruby collapse late\", \"undecided? abstain\", \"2389-research/judgement license null\", \"confidence ≠ winner p\", \"typesafeai-sdk-community not a new species\", \"tpellet/hunch exit 3\", \"never-execute list\", \"jev-file-search scores not calibrated accuracy\", \"jev-linkmap Jev never sees S2 prose\", \"muhammedilyasy/jev-mail metadata only\", \"tidy none-of-folders stay\", \"tab-bouncer pinned/audio/current never closed\", \"lkclean Show fail-open\", \"jev-yt-time-saver Show anyway\", \"ORIGIN pause-if-no-Jev\", \"validResponse sums-to-1\", \"jev-crawlers risk bands never raw boolean\", \"jevbrain AUTO_ACT is not a Noul\", \"judgekit YAML classify/score/route/verify\", \"typed-judge-kit verdict-in-code\", \"alsoleg89/decide packing VOI\", \"0.8 ≠ 80% accuracy\", \"Jev-Calibration Platt ECE 0.117→0.052\", \"jev-calibration-arena never acts\", \"ctmx/openrouter-jev-mcp Decision-as-Plugin\", \"FrancoisChastel/jev-code ≠ npm jev-code\", \"claudecode-jev-marketplace fail-open not hot path\", \"pedroknigge/mcp_jev packs not ask_jev\", \"cyrusasco/typesafe-mcp noul deadband 0.35–0.65\", \"codaaiteam/jev-skill jevtypesafeai.com ≠ TypeSafe\", \"hermes-switchyard ≠ hermes-jev-router ≠ hermes-plugin-jev\", \"nanoprune 2.8MB ECE 2.58%\", \"smartdio/jev-browser-agent ≠ ZHUBoer/ego-jev\", \"Dakai/omp-jev-web DONE ≠ proof\", \"hari007sh/jev ≠ dannote/jev\", \"0thernet/system-one-skills deterministic verify\", \"typed-gate band [0.40,0.60] is refusal\", \"pi-jev-gate fail-closed; choice is the verdict\", \"Foq ~25ms/2.2GB local\", \"rev prefill-only + HF jev-0.5b\", \"robfrase/jev planning memo\", \"typesafe_agent_gates 27/27 / 31/31\", \"EpicEric/safe-sh static remainder\", \"pastepilot Confirm before act\", \"Jev-Reranker live Jev not yet measured\", \"sessionwise opt-in relevance\", \"jev-search pointer sieve\", \"400ms Salesforce WebMCP\", \"typesafe-scheduler-diagnostics advisory\", \"droidjev screenshot-free\", \"Tewoto1 jevcu planner still writes\", \"ha-conversation-jev Jev→Grok\", \"dsh-jev can only gate\", \"jev-classification-benchmark specified not run\", \"jev-luna-pagerduty p≥0.50\", \"meldltd/meldecision laya-go ONNX\", \"laya-doom never pixels\", \"logixism/laya-api empty README\", \"akpsahan/laya ≠ Archer\", \"choxos/jevchess engine owns truth\", \"jev-drive sim not AV\", \"story-arc Jev never authors\", \"jev-hs-assistant HS6\", \"golergka/jev-plays-starcraft-2 UI-verified ≠ API Victory\", \"awesome-jev-use-cases catalog\", \"Nibir1/typesafe-go ≠ official\", \"fingerprint after redact\", \"recall vs decide\", \"publish fingerprints+answers\", \"CI replay as Harbor cousin\", \"Cache hit ≠ correctness\", \"hyperspaceai/jevcache ≠ kushals256/jevcache\", \"human labels only\", \"score never auto-accepts\", \"production capture flywheel\", \"sutro-sh/jev-align ≠ caiovicentino/jev-align\", \"guidance ≠ hook\", \"catalysts ≠ summaries\", \"compile-time System One\", \"unofficial ≠ TypeSafe\", \"format_version modernbert-jev/1\", \"Argos1111/jev_local ≠ us/jev-local ≠ kunchenguid/local-jev\", \"LFM default ≠ ModernBERT backend\", \"Nemotron ≠ TypeSafe Jev\", \"not a calibrated replacement\", \"djev-dev complements djev-spark\", \"images as Choice options\", \"Laya essay numbers *theirs*\", \"Router/OOD confidence\", \"hosted bootstrap ≠ silent TypeSafe\", \"difficulty + policy thresholds + JSONL trace\", \"jev-codex-pilot model + reasoning depth\", \"keep/shadow/hybrid/reject\", \"quarry evidence projection\", \"Frank-ZY-Dou/awesome-jev robotics/3D/control\", \"one-dollar-tahoe TypeSafe Jev defense eval\", \"jevguard calibrator/cache/escape\", \"jev-ci-selector CI shadow mode\", \"llama-jev llama.cpp replica\", \"petercr/jev-orchestrator ≠ FleeexCorp/jev-orchestrator\", \"seb4ez/jevguard ≠ AseemPrasad/JevGuard ≠ pablozr/JevGuard\", \"webNeat/llama-jev ≠ WiktorB2004/llama-index-jev\", \"OpenCode jev-pruner context sieve\", \"observe→score-candidates→prune\", \"jev-zen / jev-1.13-free\", \"zen-chat ≠ Noul\", \"fail-open original\", \"keepScore >0.1 floor\", \"host port of tamaratran/jev-pruner\", \"indiejoseph/opencode-jev-pruner ≠ nrdz-labs/fast-jev-opencode\", \"jev-webagent-bench empty stub\", \"Kiln-AI/jev_jsonschema noul_threshold 0.5\", \"NSStudent/JevSwiftSDK unofficial\", \"GLiNER2 native Apple path\", \"unofficial Swift/Core ML GLiNER 2.5-small\", \"entity spans + confidence\", \"not Choice/Score/Noul\", \"not TypeSafe\", \"label descriptions as schema\", \"on-device ANE economics\", \"honesty locks\", \"shershah1024/gliner-native-runtime ≠ Fastino\", \"≠ gliner25-compaction ≠ gliner2-ultrafast ≠ Eran-BA/Jev_from_GLiNER2 ≠ NSStudent/JevSwiftSDK ≠ jevmlx\", \"default threshold 0.1 still soft\", \"Decision Graph Protocol frame→assess→commit\", \"app retains permissions/effects\", \"Jev-first assessor-neutral\", \"guarded commit / receipt/next frame\", \"assessment batching\", \"hard-gating DGP as safety theater\", \"numerous-com/dgp ≠ TypeSafe official\", \"jegrep calibrated path+range Nouls\", \"no embeddings/index/daemon\", \"~$0.01–0.03 typical\", \"agent --json\", \"can1357/jegrep ≠ Bentlybro/jevgrep ≠ uehaj/jev-semgrep\", \"Archer-arch fidelity\", \"kev family OOD 0.76–0.77 vs Jev 0.86\", \"block-causal isolation\", \"pointer/readout CE-trained\", \"/v1/systemone drop-in\", \"replica honesty\", \"cost-sensitive decision theory × System One probabilities → control flow\", \"thresholds derived from costs not hard-coded\", \"YES / NO / UNSURE from cost_false_yes / cost_false_no / cost_human\", \"auto-batching same-object questions\", \"Kungie/gut ≠ tpellet/hunch ≠ carldaws/hunch\", \"judgment vs generation\", \"deterministic execution after probabilistic judgment\", \"exactly one app-owned callback\", \"explicit uncertain branch\", \"Illusion47586/judge ≠ lexingtonhibiki/judgekit ≠ Ascurse/typed-judge-kit\", \"variable-N option scoring as the trainable object\", \"dynamic candidate bags not fixed label sets\", \"zwliJay/jev-forge ≠ NanoJev\", \"open replica economics / latency vs closed Jev\", \"NAR local drop-in\", \"wfzyx/von late-catch HIGH\", \"competing NAR claims / replica honesty\", \"typed judgments vs chat judges on guardrailing\", \"ishaannk/llm-vs-jev cross-note only\", \"deeper integrity fold is rh-guard\", \"nothing wins outright\", \"can be argued out of guarding\", \"Jev IS the if-statement\", \"judgments/probabilities drive branches\", \"text model only writes prose\", \"interpreter owns variables/loops/budgets/replay\", \"otherwise maybe / confidence gate\", \"chaos samples after the gate\", \"southpolesteve/probably ≠ carldaws/hunch ≠ feelings ≠ Kungie/gut ≠ Illusion47586/judge ≠ tidymodels/probably\", \"133★ / forks 10 live\", \"build calibrated classifiers from human feedback\", \"retrieve by relevance not resemblance\", \"one calibrated yes/no per memory in one request\", \"pointer mode 17/18 19/20 *theirs*\", \"embedding resemblance misses the allergy\", \"samdotmak/jev-recall ≠ jev-search ≠ jev-sift ≠ carryforward ≠ chopratejas/invalidate\", \"memory leases ended by new evidence\", \"six Nouls then fixed rules in code\", \"0 of 157 false invalidations\", \"questions/plans/directives are not evidence\", \"unsure → review queue\", \"host keeps the store\", \"name↔body / comment truth / test-claims\", \"mizchi/jev-lint is mizchi/jevlint rename\", \"no shipped rule has severity error\", \"~1 in 5 findings wrong *theirs*\", \"mizchi/jev-lint ≠ huntedman/JevLint ≠ MichitoSugawara/jev-lint\", \"JSON Schema → typed JSON via Jev\", \"noul_threshold 0.5 decoder not a proof\", \"IncompatibleSchemaError lists every bad property\", \"on-device Laya CoreML ANE\", \"~5 ms P50 short decisions\", \"189/189 FP16 checkpoint parity\", \"10× not achieved\", \"mizorewww/laya-coreml ≠ gliner-native-runtime ≠ jevmlx ≠ NandhaKishorM/laya\", \"softmax over allowed tokens ≠ Noul\", \"question-first cache\", \"Micha0827/snapjudge ≠ githubnext/localjev ≠ jevmlx ≠ cendress/SnapJudge\", \"Jev-first Pi agent loop\", \"slow-LLM fallback\", \"explicit action menu / CandidateSource unimplemented\", \"62 tests wiring not quality\", \"direwolfiy/JevPi ≠ standardagents/jevpilot ≠ pi-jev-control\", \"resume-screening bias audit methodology\", \"name×resume factorial independent Nouls\", \"callback determined by resume quality\", \"mean-probability name gaps operationally negligible\", \"natemoo-re/bias-bench ≠ BBQ\", \"Plan/PRD panel → code-owned pass|review|block\", \"cheerleading out of scope\", \"austindixson/planalyzer ≠ single-goodness Noul\", \"cost-aware multi-model routing/escalation\", \"decide vs do\", \"successful-task cost\", \"cannacre8ive/switchboard-ai ≠ ha-switchboard ≠ hermes-switchyard\", \"frozen-protocol zero-shot bench\", \"TypeSafe Jev vs PrismNLI vs Laya\", \"contamination caveat\", \"elcronos/jev-vs-open-decision-models ≠ JevBench ≠ DMB\", \"context-window admission control\", \"VOI gate which tokens are worth the expensive model\", \"fail polarity per lens\", \"on small inputs lenses lose money\", \"cvsgireesh/jevusher ≠ jev-sift ≠ winnow\", \"typed decision control plane\", \"receipt ≠ authorization\", \"historical-v0 zero retained cases\", \"MokiMeow/jev-fabric ≠ jev-forge ≠ dgp\", \"live 15-dim typed rubric re-score per pause\", \"scoring economics exemplar\", \"OpenJev/Codiv ≠ TypeSafe hosted\", \"jose-troche/live-rubric ~$0.000004 desc / ~$0.000006 README\", \"adversarial pre-registered Jev eval\", \"28 predictions before data\", \"123,805 requests\", \"confidence does not track ignorance\", \"polite injection 65% / crude 0%\", \"willkelly/jev-evaluation ≠ jevals ≠ jev-baselines-eval\", \"provider-neutral Elixir/BEAM Noul/Choice/Score SDK\", \"class infrastructure\", \"nshkrdotcom/system_one_sdk ≠ typesafe_sdk ≠ dannote/jev\", \"question-linting of Jev questions themselves\", \"nine jaggedness rules, no API key, no labelled data\", \"static lint ≠ measured separation\", \"yodablocks/jevq ≠ tenbin ≠ JevLint ≠ commitjev\", \"open-weights Laya as class exemplar (binding)\", \"Nx/Bumblebee runtime\", \"host chooses backend\", \"ChristianAlexander/laya_ex ≠ system_one_sdk ≠ dannote/jev ≠ NandhaKishorM/laya\", \"on-chain/edge Laya deploy\", \"parity_verified stays false\", \"model output never grants Tx\", \"humandebri/IC-Laya ≠ laya_ex\", \"auditable weekend replica\", \"Jev outputs never used for training\", \"soft human-vote distributions\", \"unpaired 0.577 vs 0.727\", \"agilabs-ai/jev48 ≠ JevBench ≠ Mapika/decider\", \"adversarial dual-judge / framing attack surface\", \"comparative framing is the usable judgment\", \"prior injection crowds out evidence\", \"copyleftdev/ember ≠ ember.js\", \"Laya specialist fine-tune pipeline\", \"training still GPU-pending\", \"PIXELZX0/XERON ≠ convaiinnovations/laya\", \"Hub Laya replica drop\", \"daliborsb/laya ≠ convaiinnovations/laya ≠ NandhaKishorM/laya\", \"System One student distillation corpus\", \"gold is programmatic\", \"teacher is closed-API clone\", \"do not distill Jev as teacher of record\", \"MagaBitmex/jev-4b-distill-data ≠ missing student checkpoint\", \"non-LLM VIN System One\", \"planning depth not chat\", \"lewislululu/jevon ≠ douglance/jevon\", \"source-bound evidence checks\", \"local quote mismatch needs no API\", \"exit 0 ≠ claim truth\", \"WaynezProg/jev-kit ≠ jonathanavis96/jev-kit (Airlock) ≠ jev-use ≠ jev-mcp\", \"independent System One evidence catalog\", \"scores not one leaderboard\", \"no external record currently reproduced\", \"TokenTrim no-Jev matched hybrid 62.4%\", \"reachjalil/system-one-bench ≠ mallahyari/system-one-benchmark\", \"21 tasks · 134 items · 208 questions\", \"scenes from public GitHub contracts, not production logs\", \"SivletLabs/jev-eval ≠ willkelly/jev-evaluation ≠ 4esv/jev-eval ≠ xxkuboxx/jev-eval ≠ onlyoneaman/jev-eval ≠ dayhaysoos/jevals\", \"option isolation (sibling-blind)\", \"permutation-equivariant\", \"Hub OWNER not published\", \"nafisazizir/hev ≠ jaredpalmer/kev\", \"frozen local LLM logits, no trained decision head\", \"residual-head 9,222-param decreased 73/96→67/96\", \"confidence = 1−normalized entropy, not P(correct)\", \"yuki-oshio/mini-jev ≠ r-ms/mini-jev\", \"Jev classifier as autoregressive next-token predictor\", \"ChatJev-style soundness theater\", \"erik-dunteman/ChatJev ≠ dannote/jev ≠ jev-gpt\", \"calibrated decision head × AlphaProof value head\", \"implementation-layer isomorphism, semantic difference\", \"timeout = censoring\", \"do not launder Noul as proof\", \"parallel rank-prediction vs serial selection\", \"independent questions can conflict\", \"zzzzzec/jevsort ≠ keltokhy/jsort\", \"curated open System One ecosystem catalog\", \"rupeshpoojary9/awesome-open-system-one ≠ AnotiaWang/awesome-jev\", \"arXiv paper radar with Jev relevance scoring\", \"ranking ≠ calibration / 0.5 still soft\", \"fail-open failed evals not marked seen\", \"train calibrated ~27M from scratch\", \"typed Q→prob dist / one forward pass / no LLM decode\", \"hyusi2003/MiniSystemOne ≠ Colvin0315/MiniSystemOne\", \"description-only stub / size 5\", \"ESCI hard probe fails four of six\", \"jev_bool ECE 0.242 inversion 0.255\", \"do not re-fold §60 six-gates as new\", \"jobbyjev one-request-per-company from batch-size result\", \"find/design/evaluate TypeSafe Jev decision loops\", \"karanb192/jev-architect ≠ samtay32/jev-system-architect\", \"Jairik/jev-distiller size 1\", \"distill-Jev UI stub / do not distill Jev as teacher of record\", \"post-launch scored use-case map / Jev self-scores then human curation\", \"licensedsaucer9-web/jev-opportunities\", \"Jev-inize a use case into classifier/router\", \"gavinHuang/jevinize → simple-jev not TypeSafe\", \"featherless-ai/simple-jev\", \"compare saved decisions / same label can still change the branch\", \"VihaanAgarwal/jev-diff ≠ Saik0s/diffusiongemma-jev-macos\", \"not tested with a live Jev API key\", \"constrained logprob + temp/Platt ≠ Noul\", \"OpenJevPro pastes openjev-sglang JevBench as own\", \"zhangcy122/OpenJevPro ≠ IamBusy/OpenJev ≠ ekzhang/openjev-sglang\", \"PolyForm Noncommercial\", \"SmolLM-135M / sub-70ms / 0 output tokens\", \"demo P(True) 0.5052 / Choice conf 0.2872 / Score conf 0.0055\", \"README claims MIT / GitHub license null / no LICENSE file\", \"patelvishwa112/jev-system-one-rlcd ≠ arnabgho/rlcd-lite ≠ blackwood-rlcd\", \"source-backed Awesome Jev radar / 306+ commit-pinned\", \"logicrw/awesome-jev-projects ≠ AnotiaWang/awesome-jev ≠ yibie/awesome-jev ≠ cobanov/awesome-jev ≠ rupeshpoojary9/awesome-open-system-one\", \"auto GitHub sync / Issue-only submissions\", \"hashed n-gram encoder / rival-aware attention\", \"olanotolu/jevbetter vs jevlike starter\", \"synthetic hard menus top-1 0.916 vs 0.873 / ECE 0.0182 vs 0.0367 / 40 vs 4608 menus/sec\", \"shuffled-context control 0.335\", \"structured probability readouts\", \"distribution > argmax\", \"Noul 0.5 midpoint\", \"score is expectation not integer\", \"bare HTTP not SDK\", \"Arohtea/jev-readout\", \"Jev-style Choice/Score/Noul from ordinary models\", \"optional DSH plugin\", \"schema-valid ≠ calibrated\", \"gulagala001/jevify ≠ Mintzs/jevify\", \"Laya RLCD benchmark\", \"40.3% below constant-answer\", \"open-weight measurement\", \"mourad-ghafiri/laya-rlcd-benchmark ≠ yibie/laya-jev-lab\", \"cheap fail-open semantic edge\", \"second signal not sole\", \"FastLoopError catch\", \"SupremeDreamZ/jev-fastloop ≠ jev-ultrafast\", \"asking more questions in one call\", \"0.980 at every N\", \"nearly not fully deterministic\", \"TheWebDevel/jev-fanout\", \"Qwen3-VL perception + Jev decisions train RL\", \"0 model calls at deployment\", \"VLM alone 1.7 vs +Jev 4.4\", \"harneet2512/reflexrl ≠ khordoo/jev-reflex-autonomy-lab\", \"independent Jev API vs Laya\", \"cascade 0.60 matches 78% at 1.8×\", \"noul facts not judgements\", \"yibie/laya-jev-lab ≠ dairui1/jev-lab ≠ BrendanH18/jev-lab\", \"GLiNER vs GLiFormer vs Laya vs Jev\", \"extractors ≠ decision engines\", \"Laya dict-instructions collapse 58.3%\", \"umstek/zero-shot-ie-bench\", \"decisions-per-minute & cost\", \"204 moves vs 73\", \"throughput not intelligence\", \"angelgalvisc/snake-arena-jev-vs-llms ≠ vtrivedy/jev-plays-games\", \"behavioral contracts\", \"pin expectations eval upgrades\", \"raw 0.94 is not a release\", \"sathariels/jevcheck ≠ dayhaysoos/jevals ≠ SivletLabs/jev-eval\", \"evidence-linked dependency upgrade\", \"Jev never generates filenames\", \"no_direct_evidence ≠ safe to merge\", \"GaneshVG18/upgrade-radar ≠ LYchoon/paper-radar-jev\", \"discography theme/mood/complexity\", \"five atomic questions one call\", \"lirantal/discoprint\", \"Turn any open LLM into System-One Jev\", \"uspraveen/Jevify ≠ Mintzs/jevify ≠ gulagala001/jevify\", \"Jevify-any-LLM architecture probe\", \"description-only stub / size 0\", \"Train encoder-only calibrated decision models from a task sentence\", \"Exu is a toolkit, not a method\", \"strictly proper scoring rule\", \"Pre-alpha\", \"Ruivalim/exu-base\", \"scratch-trained calibrated decision model\", \"typed Q → probability dists\", \"Colvin0315/MiniSystemOne ≠ hyusi2003/MiniSystemOne\", \"no published weights download URL\", \"90.5 seconds / 29.2% pipeline evidence\", \"p_i/p_j independent of other candidates\", \"Recipe for calibrated decision models — small model out\", \"init → synth → train → eval → serve\", \"91.1 % / ECE 0.022 *theirs*\", \"Jev zero-shot 75.1\", \"scienthoon/luce\", \"Put Jev's three headline claims on trial\", \"0.5B local GPU\", \"46x speedup / accuracy identical\", \"ECE 0.624 sentiment catastrophe\", \"bigger model worse calibration\", \"RichardoMrMu/jev-mini ≠ yuki-oshio/mini-jev ≠ r-ms/mini-jev\", \"System-1 decision engine for local LLMs\", \"structured choices only\", \"JSON parse of generated text ≠ Noul\", \"TypefAI JEV / Journal Entry Voucher\", \"tapsin/jev-local ≠ us/jev-local ≠ Argos1111/jev_local\", \"Jev 1.13 reward-model eval across 8 benchmark tracks\", \"40,940 examples / 0 API errors\", \"RewardBench v1 92.58%\", \"Precise IF 50.63%\", \"goya4140/jev-reward-model-evaluation\", \"Scaffolding in progress\", \"Jev vs LLM support-ticket routing\", \"static + live decision bench\", \"TypeSafe's own published benchmark\", \"illustrative simulations, not live API calls\", \"JevBench v1 — smart/cheap/fast/reliable\", \"I/C/S/K 25% geometric mean\", \"classifier.dev fast tier 84.8 is Jev behind its own API\", \"do not re-fold §78 v1.2 board as new\", \"Laya (421M) 70.1 now on board\", \"Zero-shot/few-shot LLM routing\", \"hard budget filter before Jev\", \"Jev never asked to perform budget arithmetic\", \"Jev judges the next state, XState enforces transitions\", \"simulation uses synthetic keyword fixtures\", \"catalog gravity\", \"v-modal/awesome-jev-tools\", \"★339 live REST\", \"curation is not endorsement\", \"crawler-maintained directory\", \"Daily GitHub + npm sweep, human-merged\", \"RadRebelSam/awesome-jev ≠ AnotiaWang ≠ yibie ≠ cobanov ≠ logicrw ≠ v-modal\", \"HF peft SPLADE/BGE reranker\", \"rdxtremity/jev-reranking ≠ carlaiau/jev-reranking\", \"query-side encoders, not a Jev replica\", \"ONNX System One Qwen3.5-4B scorer\", \"source:pngwn/system-one-qwen3.5-4b-scorer\", \"CC-BY-NC-4.0\", \"temperature 1.75\", \"transformers.js AutoModel cannot load this graph\", \"Consistency benchmark Space\", \"This Space contains no benchmark result yet\", \"12-case plumbing fixture\", \"Benchmark-driven Jev router and judge\", \"cheap alone is not success\", \"Jev does not write, sum prices, or claim accuracy %\", \"Sol 94.2 / Luna 83.9 / Jev path 89.7\", \"19.2% Sol / 62.3% cost save / 4.5pp miss of 2pp non-inferiority\", \"p50 latency worse than Sol due to routing overhead\", \"erendikmenn/jev-llm-router-benchmark ≠ jev-rag-benchmark ≠ ryantsai/jev-llm-router\", \"Express + node:sqlite\", \"mock and Jev decision engines\", \"previous_ticket_count >= 3 is code\", \"MIN_CONFIDENCE 0.6 still soft\", \"substring false positives\", \"aesaganda/jev-ticket-router ≠ SarathChandraBellam/jev-vs-llm-ticket-router\", \"Universal Figure & Diagram Router\", \"confidence ≥ 0.85 hard-gate is theater\", \"generative AI banned from scientific plots\", \"six visual branches\", \"hoangngochuong24947-gif/jev-figure-router\", \"human-labeled (state, question, label)\", \"166,054 rows / 22 configs\", \"soft_label for human uncertainty\", \"Praveenrajus/jev-bench ≠ fstandhartinger/jevbench\", \"ternary bonsai System One GGUF\", \"openjev's mechanism, Bonsai's weights\", \"Hub does not ship weights\", \"100/100 easy T/F is not Harbor\", \"label_mass ≠ correctness\", \"stock llama.cpp Q2_0 silently gibberish\", \"NicolaiMTLassen/open-bonzi-jev ≠ NicolaiLassen\", \"transformers.js DeBERTa ONNX\", \"source:com-kotobalabs/open-jev-deberta-v3-large\", \"temperature 1.05\", \"AutoModel from_pretrained works\", \"onnx-community/open-jev-deberta-v3-large-ONNX ≠ system-one-qwen3.5-4b-scorer-ONNX\", \"107★ densify\", \"GH 151M vs README 149.6M\", \"PR #1 now closed unmerged\", \"do not re-fold §71 claim-audit as a beat\", \"typed decisions, RLCD, confidence-gated routing\", \"structured ≠ correct\", \"mock not live API\", \"26 tests\", \"wjdjdakf17/jev-study ≠ baekenough/jev-study\", \"bonzi-27b-v2 / ternary-8b / 27b-v1 GGUF family densify\", \"Hub still does not ship weights\", \"WANLI-256 74.6% / 65.2% / 71.1% *theirs*\", \"Bonsai 1 27B Q1_0 runs on stock llama.cpp\", \"ternary still needs PrismML fork\", \"hf:heman10x/openJev-verdict-2.0 twin tokenizer-only\", \"OpenJev Vision image classification + uncertainty\", \"CLEVR-4 held-out joint 0%\", \"hfdataset:IamBusy/OpenJev-Vision-Research-v0.1 12,832\", \"294,912 derived targets not independent samples\", \"Laya multilingual ONNX WebGPU typed-decisions port\", \"63/63 selected answers / 5.1e-4 CPU / 1.2e-2 WebGPU\", \"UpHash-Network/mini-jev is yuki-oshio transfer\", \"jev-injection-bench 11,900 labelled prompts\", \"Jev best ranking / Haiku better ECE 0.021 vs 0.058\", \"0.5–0.9 band is where Jev's numbers do not mean what they say\", \"Prompt wording moves panic 28%\", \"manojlds/jev-dspy-bench ≠ dspachos/jev-dspy ≠ jmanhype/jev-dspy-lab\", \"Jev agreement is similarity, never ground truth\", \"no aggregate quality grade or merge gate\", \"AbstentionBench-on-Jev rank 1 of 20 vs 2025 field\", \"question-asymmetry\", \"forward-looking 0.465 never extreme\", \"openkev calibration layer not a runtime\", \"ECE vs coverage independent\", \"select_threshold returns inf\", \"escalation catches uncertainty not ignorance\", \"misakaikato/openkev ≠ jaredpalmer/kev\", \"pdf-race Docling→Jev vs Gemini\", \"parser owns the wall clock\", \"12/12 tie is a tie\", \"titles selected not generated\", \"flopcheck 16 calibrated tweet judgments\", \"mechanical tells in code\", \"ZeroX-01/jev-atlas ≠ Zaious/jev-capability-atlas ≠ gorock007/jev-atlas\", \"catalog not endorsement\", \"Laya calibration lab Gradio MCP\", \"T never changes argmax\", \"confidence ≠ top-label p\", \"easy probe set refused\", \"40–48 rows too small to ship T\", \"Gemma-4 26B-A4B jevify classification+calibration\", \"LoRA adapter twin not independent eval\", \"Gemma-4 E4B jevify\", \"E4B LoRA stub card\", \"kushalpatil/jevify-gemma4 ≠ Mintzs/jevify ≠ gulagala001/jevify ≠ uspraveen/Jevify\", \"GH kushalpatil07/jevify 404\", \"PAWS 0.580/ece 0.288 is the weak cell\", \"smaller E4B slightly better OOD ECE than 26B-A4B\", \"Hub jevify merged LoRA ships weights\", \"bonzi Bonsai-8B v1 GGUF densify\", \"Bonsai-1.7B v1\", \"Bonsai-4B v1\", \"WANLI-256 64.5% / 60.2% / 52.0% *theirs*\", \"rank #4 / #5 / #6 of 6\", \"JulesHuisman/jev-eval scaffolding / README SHA c356a584 (was empty e69de29b)\", \"JulesHuisman/jev-eval ≠ SivletLabs/jev-eval ≠ willkelly/jev-evaluation ≠ 4esv ≠ xxkuboxx ≠ onlyoneaman ≠ dayhaysoos/jevals\", \"7 bands 6/10 vs 40 bands 0/10\", \"source receipts + confidence slider re-policy without re-inference\", \"32/32 synthetic is smoke not production\", \"classify HF datasets across typed semantic dimensions\", \"roadus2 watch misspelling\", \"lock roadius2/ultra_laya\", \"ultra_laya REVIEW defects\", \"default branch claude/laya-jev-review-gg5ppo\", \"XNLI EN 88.3% ECE 0.032 → RU 77.3% ECE 0.096\", \"Δ −11.0 pp [−14.2,−7.8]\", \"ECE +0.063\", \"MASSIVE no detectable difference at n=600\", \"confidence is function of p_max (r=1.000)\", \"pointer-not-generator 400 human-authored responses\", \"proposed ≠ authorized\", \"FewRel 160: Jev 85.0% vs lexical 13.125%\", \"gated 100% (95/95) coverage 59.375%\", \"J++ composable semantic computation language\", \"judge-jev 0.5 still soft\", \"947 repos scored\", \"A 273 / B 302 / C 372\", \"LLM rubric ≠ benches\", \"No benchmark winner is claimed\", \"phishing: naive 62.6% vs regex 91.8%\", \"5-atomic + LR 95.0% *theirs*\", \"AITuber tension ±15\", \"README npm global\", \"repo is Rust\", \"git-confess code owns counting/blame/ratio\", \"httpx exhibit 11% (13/119) *theirs*\", \"90d trend +12.40% vs random +12.75% vs BH +41.71%\", \"5m win rate 25%\", \"Awesomejev 656 entries / 38,160 stars\", \"tracker likes 64 (+4) lastModified UNCHANGED\", \"Laya present\", \"Blackwood ABSENT\", \"Archer still promised_not_landed\", \"Blackwood tracker ABSENT; likes 2 gated manual\", \"r = c - p_a\", \"ECE 0.021; acc 0.807 vs warmup 0.746\", \"calibration beyond ~500 tokens unmeasured\", \"Independent primitive\", \"11.57s vs 54.10s · 4.67× · 120/128 *theirs*\", \"default path is pretrained Gemma probs not trained RLCD head\", \"GH Meanblock 404; lock leesk212/JEV-CPU\", \"softmax over letter slots ≠ Noul\", \"WANLI 0.741 vs openjev v2 0.77 *theirs*\", \"3-way NLI ≠ Noul\", \"priority 0.464 = majority floor\", \"banking77 contaminated\", \"raw margins not probabilities\", \"GH jev-haiku-benchmarking 404\", \"do not distill Jev as teacher of record (they distilled Haiku)\", \"“0.9 is not one number”\", \"ranking ≠ calibration\", \"banking77 0.8–0.9 stated 0.86 actual 0.73 over-confident *theirs*\", \"≠ Praveenrajus/jev-bench ≠ fstandhartinger/jevbench\", \"$0.0000153–$0.0000226 vs circulating $0.0004 (~20×)\", \"Score is 0..n-1 expectation not 0–1\", \"Noul has no confidence field\", \"TCP floor 198.8 ms\", \"type reliability is not a reason to choose Jev (json_schema 5/5)\", \"gateway tax not one number\", \"Function-only 5/8 vs hybrid 8/8\", \"4/8 without Jev\", \"8 designed cases not conversion lift\", \"≠ RadRebelSam/awesome-jev\", \"200-row pilot Jev 86.5% 173/200 vs Gemini Flash-Lite 86.0% 172/200 vs Pro 87.0% 174/200 *theirs*\", \"not a ranking\", \"NLI Tetris argmax P(entail)−P(contradict)\", \"情緒測謊器\", \"1q 396ms / 30q 567ms\", \"±0.03\", \"33q $0.000045 vs Gemini ~5× slower ~60× cost *theirs*\", \"≠ realZachi/jevtest\", \"8-example Jev vs GPT-5.6 Sol ~64× cost 5.4× latency *theirs*\", \"synthetic; no inference\", \"≠ JevBench v1.2 §78\", \"Judged 3317 / listed 2560\", \"Jev judges, code applies policy\", \"catalog ≠ endorsement\", \"APA “microsecond policy / zero hallucination” overclaim\", \"Client-side quiz; pointer from held docs; scanned-PDF warn\", \"CSP only api.typesafe.ai\", \"Jev judges / agent reasons / user decides\", \"selecting an option is not permission to implement\", \"degraded fallback\", \"pattern exact, judgement must clear floor\", \"no matching pattern → no model call\", \"not a correctness oracle\", \"$0.00022 vs chat $0.00306 *theirs*\", \"Spec vs artifact remainder\", \"treating 0.85 as 85% / minProbability hard-gate as Harbor\", \"VERIFY acquires discriminating evidence, never same-pool confidence-only rescoring\", \"fast/full/max are ceilings not sizes\", \"Solar writes, Jev chooses NEXT ACTION\", \"SemIf 2186★ (+20 vs §109 2166)\", \"jevlike 1038★ (+7 vs 1031)\", \"TypeAR 14★ flat\", \"AnotiaWang 96★ (+1 vs 95)\", \"yibie/awesome-jev 490★\", \"Laya likes 802 (was 783)\", \"tracker likes 64 flat, lastModified UNCHANGED\", \"do not reopen or amend PR #23/#24/#25/#26/#27.\", \"do not reopen or amend PR #23 or #24 or #25 or #26 or #27\", \"Calibration is not alpha\", \"NO CURRENT ALPHA CANDIDATE\", \"ΔR² approximately +0.00084\", \"Brier 0.2131387\", \"ECE 0.0421875\", \"Adding Jev probability to deterministic volatility improved Brier by only 1.4058e-05\", \"default 0.5 keeps zero non pinned\", \"keepResult median 0.14 to 0.17\", \"keepCall median 0.28 to 0.35\", \"usable range is about 0.10 to 0.25\", \"7.8% to 57.9%\", \"judges results it never sees\", \"task-finish eval not built yet\", \"$0.002 per compaction\", \"slavadubrov/sgr-judge-bench ≠ slavadubrov/jev-judge-bench\", \"Jev 108/120 $0.083 0.34 s\", \"Luna SGR 114/120\", \"paired Jev accuracy-difference intervals include zero\", \"not evidence of equivalence\", \"GLM SGR 26/120 93 format failures\", \"Terra-planned Jev hybrid 55/120\", \"rule-based by default, optionally Jev-backed\", \"empty README\", \"missing key cannot break the experience\", \"prefill plus exactly one decode\", \"softmax over A/B/C ≠ Noul\", \"BBQ 9,053/10,000 (90.53%)\", \"ECE 0.0890\", \"Mean confidence 0.9943\", \"overconfident\", \"score and noul not implemented\", \"DGUI 12 rows (was 6)\", \"INSTRUCT 119 rows likes 2\", \"encode the state once, decide everything in parallel\", \"0.740 accuracy against a 0.508 majority\", \"ECE 0.047\", \"fine-tune's advantage ends where its 384-token training data does\", \"jasonkneen/open-jev ≠ pngwn/open-jev\", \"same sha d41dc3cd\", \"Space does not call Jev\", \"recomputes routing from saved probabilities\", \"200-case Jev 97.0% / 100.0% / 95.0% / MAE 9.22\", \"synthetic repository benchmark\", \"Jev evaluations are advisory\", \"YehuiTang0316/jev-nlgrep ≠ Bentlybro/jevgrep ≠ can1357/jegrep ≠ uehaj/jev-semgrep\", \"default threshold 0.8 still soft\", \"40-line windows cannot prove whole function\", \"token-native sequential start/end Choice\", \"Gemini/Haiku stubs not configured yet\", \"handful of hand-written examples, not a benchmark\", \"Jev judged exactly what it was given\", \"laguagu/jev-skills ≠ laguagu/jev-evidence-lab ≠ Pleo2/awesome-jev-agent-skills\", \"contract_passed is not a claim of guaranteed factual truth\", \"Wilson lower bound 0.85 floor\", \"fixture mode no savings claim\", \"SemIf 2207★ (+21 vs §110 2186)\", \"jevlike 1043★ (+5 vs 1038)\", \"TypeAR 15★ (+1 vs 14)\", \"AnotiaWang 97★ (+1 vs 96)\", \"yibie/awesome-jev 506★ (+16 vs 490)\", \"Laya likes 822 (was 802)\", \"tracker likes 64 flat, lastModified UNCHANGED\", \"do not reopen or amend PR #23/#24/#25/#26/#27/#28\", \"people who compare Jev against GPT-5.6 has never fine-tuned BERTForXYZ for living and it shows\", \"zero shot classifiers\", \"scale them as much as decoder only models\", \"many problems solved with LLMs could have been solved with them, it was a skill issue\", \"opt for DeBERTa and ModernBERT ones\", \"BERTForXYZ → DeBERTa → ModernBERT\", \"Jev vs GPT-5.6 bakeoffs are a category error\", \"encoder / ZS classifiers\", \"institutional HF voice\", \"quote *theirs*\", \"do not invent accuracy numbers\", \"softmax/ZS scores still ≠ calibrated Noul\", \"soft scores ≠ hard gates\", \"@mervenoyann\", \"likes 421 / 189\", \"impressions 35498 / 9613\", \"multimodal image<>text ZS as perception front-end\", \"hf:MoritzLaurer/deberta-v3-large-zeroshot-v2.0 likes 139\", \"hf:MoritzLaurer/ModernBERT-large-zeroshot-v2.0 likes 72\", \"Bart, bert, deberta, modernbert, these are all LLMs\", \"Maziyar quoted\", \"Jev is exemplar not the mandate\", \"do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29\", \"Heman10x-NGU/Verdict-open-jev ≠ Heman10x-NGU/openJev-verdict-2.0\", \"TF-IDF + LogReg ECE 0.0207 vs Jev 0.1440\", \"Verdict-open-jev 48.07% vs Jev 90.80%\", \"abstention combined recall 10.00%\", \"p50 35.58 ms\", \"K=25 (maximum capacity) 72.00%\", \"0.85 coverage 84.60% selective risk 1.18%\", \"26.1× faster than standard Qwen JSON generation\", \"Jevify 90.0% / 167 ms CUDA graphs disabled\", \"Finding 1: Brier on stated confidence alone is a trap\", \"grpo_rlcr 0.78 / ECE 0.084\", \"reliability 0.007 but resolution 0.000\", \"27 900 schema-driven decisions\", \"13 600 / 13 600 questions\", \"candidate mass min 0.99999624\", \"22 configs · 166,054 rows · 4 calibration-gold\", \"sha a39eba3f\", \"Student B MAE 0.148 / Pearson 0.836 / 86.0%\", \"pngwn/open-jev-laya-bench README 404\", \"sha 9f69c742 likes 2\", \"HDFS 0.9933 (745/750) / retain 0.0084\", \"BGL ERROR/FATAL protection 1.0000\", \"2,479 / 2,500 HDFS uncertain\", \"cache hit 0.9648 (2412/2500)\", \"$0.153936 estimated\", \"E2 recomputes from saved probabilities\", \"Space sha eda59e0a\", \"MASSIVE English 0.783 / Khmer 0.033 / Hindi 0.133\", \"40–48 rows too small to ship T\", \"T never changes argmax\", \"siren2345/jev-single-decode-transformers ≠ siren2345/jev-single-decode\", \"Split Transformers experiment from llama.cpp runtime\", \"tanayvasishtha/jev-lab ≠ dairui1/jev-lab ≠ BrendanH18/jev-lab ≠ yibie/laya-jev-lab\", \"Four experiments stress-testing TypeSafe's Jev: calibration, bundle bias, label bias, and ensembling\", \"second pass must be $0.00 from cache\", \"The pages never call Jev\", \"Gemma 4 31B 77.0% / Jev 1.13.0 61.4% / Laya 322M 0.0%\", \"restriction state 95.0% against 84.4%\", \"None of the systems are particularly good at knowing when to stop and ask\", \"They skip the question and call a tool directly\", \"100% schema pass\", \"six-field joint 48.8% vs 72.8%\", \"ywchiu/jev_benchmark ≠ Running-Dolphins/jev-bench ≠ Praveenrajus/jev-bench\", \"ACT / REVIEW / FALLBACK\", \"A provider failure, timeout, malformed output, or missing answer is **not** a policy outcome\", \"confidence is descriptive provider output, not a substitute for probability\", \"Quality denominators include only valid scored answers\", \"an exact halfway tie chooses the lower level\", \"aiwithenoch/Jev-Skill ≠ simplosophy/jev-skill ≠ laguagu/jev-skills\", \"The local path does not claim to turn a smaller checkpoint into Jev\", \"Low support becomes decision: \"review\"\", \"MIT-0 SPDX NOASSERTION\", \"current-llm\", \"结构兼容,不是 Jev 模型能力\", \"altryne/jevify ≠ Mintzs/jevify ≠ gulagala001/jevify ≠ uspraveen/Jevify\", \"Find where Jev belongs. Design the questions. Measure the difference\", \"TypeAR-AI/TypeAR 301 → TypeLLM/TypeLLM\", \"TypeLLM/TypeLLM 16★\", \"SemIf 2241★ (+34 vs §111 2207)\", \"jevlike 1051★ (+8 vs 1043)\", \"AnotiaWang 98★ (+1 vs 97)\", \"yibie/awesome-jev 525★ (+19 vs 506)\", \"Laya likes 864 (was 822)\", \"tracker likes 67 (+3 vs 64)\", \"lastModified UNCHANGED `2026-09-20T04:29:16.000Z`\", \"do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#32\", \"soft Noul ≠ hard safety\" with a placement, not a stack replacement or a vendor how-to. Formal methods are one pillar., \"A hunch is a probability with a policy attached\", \"{ enter: 0.8, exit: 0.6 } is hysteresis\", \"replay a policy change without inference\", \"Decision models are providers, not the product\", \"huncho ≠ Kungie/gut ≠ carldaws/hunch ≠ tpellet/hunch\", \"pretrained Qwen2.5 base ECE 0.030 (0.5B) / 0.040 (7B)\", \"instruct 0.302 / 0.269\", \"70.9% → 70.0% mean conf 74.1% → 96.7%\", \"temperature scaling still matches it in-distribution\", \"No Jev API was called\", \"Qwen2.5 ≠ Archer\", \"Qwen/Qwen3.8-27B ≠ Archer\", \"calibration does not compose\", \"ECE has exactly zero statistical power to detect the failure mode that kills trajectories\", \"25–60× headline withdrawn\", \"P(all-correct): 0.0071 vs 0.0001\", \"TCE / AMS\", \"Deferred Crispification\", \"Qwen 3.8 sparring ≠ Archer\", \"pd.cut bins by equal width while jeval bins by quantile\", \"ECE 0.113 and ECE 0.076\", \"jeval drift is not implemented yet\", \"ranking ≠ calibration\", \"g0runmezadam/what-is-jev IS tunahansahin897/what-is-jev\", \"same GitHub id 1378007307\", \"947 repos scored\", \"A 273 · B 302 · C 372\", \"LLM rubric ≠ benches\", \"Probabilities are advisory, not calibrated guarantees\", \"light_cutoff_applied_to_combination 0\", \"AND: product (independence assumed and recorded in the trace)\", \"circuit-vl-4b ≠ Archer\", \"soundness theater / measurement theater\", \"hourly 0843 / notes.md §114\", , \"Jev Capability Resolver / NiazMorshed2007/jcr\", \"one tool to find documented deterministic commands in a nested capability tree\", \"returns context; **does not execute**\", \"skills = workflow+judgment; capabilities = individual operations\", \"format independent of Jev; proposed open standard exploration\", \"classify (Jev) → optional OpenAI decompose compound → beam search geometric mean of routing probs\", \"keep up to 3 paths ≥60% of best (JCR_BAND_RATIO 0.6)\", \"soft scores ≠ hard gates; 0.6 band is application policy\", \"routing ≠ permission; docs ≠ authority to run\", \"sol-vs-opus5-20 *theirs*; lookup+explain only; n=1 per cell; Not Harbor task-execution\", \"Claude Opus 5: agent input 108,585→15,819 (−85%)\", \"Codex GPT-5.6-Sol wall 25.3s→62.4s (Sol slower with JCR in 19/20)\", \"NiazMorshed2007/jcr ≠ skill-broker ≠ skillranker ≠ jev-sift ≠ jev-lens ≠ jevusher ≠ jev_select_capability\", \"do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33/#34\", \"notes.md §116\", \"SemIf was formerly OpenJev\", \"independent; not affiliated with Jev or TypeSafe\", \"Direct option logits\", \"0 output tokens\", \"shared-state parallel\", \"MLX backend for Apple Silicon (`--backend mlx`)\", \"5.21×\", \"argmax agree 18/21\", \"systems comparison ≠ semantic equivalence\", \"Parallel suffixes 20.03 dec/s\", \"authored BA 0.813\", \"TypeSafe subset agreement 0.845 vs Published Jev 0.883\", \"Softmax over options ≠ calibrated Noul\", \"wire/agreement ≠ replica of TypeSafe\", \"SemIf ≠ kw2828/OpenJev playground\", \"≠ zhihz/openjev\", \"≠ apiplant/semif-rs port\", \"≠ dddanielliu/semif-serve\", \"live REST 2282★ / 140 forks\", \"HEAD ca3ba65f1429\", \"notes.md §117\", \"rename is densify not a second census\", \"JevBench 74.6 is §78 not this ladder\", \"typed output does not guarantee semantic correctness\", \"soft Noul ≠ hard safety\" with a placement, not a stack replacement or a vendor how-to. Formal methods are one pillar. Not a substitute for the official typesafe-ai skill (live Jev API contracts)." license: MIT metadata: version: 0.4.0 @@ -54,7 +54,7 @@ classical method you already trust, substitute it, classify the win "paraphrase brittleness", "allowlist then judge", "TOCTOU-of-Noul", "Jev inside the database / sqlite-jev", "Jev picks bitrate / join order / the model", "wait for Archer", "lint the request / missing - other", "training confronts Choice other / none-of-the-above", "soft AGENTS.md rules vs the linter", "screenshot Choice / omni System One", "extractive quotes / pointer not generator", "compaction summarize vs pointer", "encoder vs Jev compaction backend", "shadow-mode compaction rollout", "CI flaky-vs-real merge gate", "fail-open VOI wake/resume", "claim vs session evidence", "S1 indexer escalate-S2", "Harbor on/off routing", "fail-open vs fail-closed wake vs CI gate", "encoder vs Jev computer-use backend", "hybrid local decide + remote fill", "DONE vs verified success", "stdout prune vs session compaction", "OpenCode jev-pruner vs Claude jev-pruner", "zen-chat vs jev-zen Noul", "hard envelope then Noul prune", "Cua-S1 vs TypeSafe Jev", "plan vs execute dry-run", "specialist computer-use vs general agent", "local drop-in vs stub scorer", "route vs memory", "when does it hold / extractable from state", "decision model vs constrained LLM", "dual-process S1/S2", "combinatorial grid vs extractive", "uncalibrated local likelihoods", "decision-native RAG", "classify-first / read selectively", "living applied-mappings atlas / class patterns", "silence as safer / draft-gate heartbeat", "robotics text-state vs pixels", "verbatim ledger vs summary", "judgment as language primitive", "Stagehand extract pick-and-copy", "harness observe-score-act vs demo loop", "public judgment wall / six parallel questions", "meaning-search without embeddings", "attention ≠ correctness", "skills→oxlint / AST prove ∩ remainder", "session-sticky first-prompt routing", "measured RAG rerank vs generative rerank", "capability kernel / secrets never in the agent", "Jev is SENSOR not policy", "type-safe ≠ correct", "typed control plane around DSPy", "native vs verbalized confidence", "engine owns truth / Jev owns judgment", "human-confirmed kill gate", "train specialist vs few-shot hosted", "decide→policy→LLM leftover", "Noul 0.5 cannot-tell never rounded", "calibration ≠ sortable / ORDER BY", "pairwise inversion / Score ordinality / two-decimal ties", "wire-compat GLiFormer /v1/systemone", "class-backend economics", "loopback gateway hosted + local", "do not distill Jev as teacher", "active-learning triage", "evidence-packet explorer", "meaning-grep AND/OR/NOT", "closed-vote-only / no planner LLM", "Jev vs PCD Harbor", "PCD O(1) ≠ Noul", "host-owned handlers × System One", "OMP/pi fail-open gate", "permission vs probability / operator owns thresholds", "judgment ≠ permission / Jev never grants access", "eval integrity / instrument not score", "constrained optimizer + S1 features / never sole hot-path gate", "privilege ≠ verdict / effect contracts not tokens", "attention filter / VOI for human review / never blocks / never green unless sure", "measurement owns endorsement / evidence-gated question packs", "Jev supplies evidence / code owns authority", "ranking ≠ calibration / never hard-threshold raw p as frequency", "hot-click CU / indexed element table", "Jev judges relevance / code decides structure", "local rules first then remainder / never auto-train on own hides", "combinators / System One as control plane", "receipts not leaderboard / type-safe ≠ correct jaggedness", "VOI over skill library / skillranker abstention", "OOD calibration / AUC ≠ ECE", "Jev vs thinking-budget small models", "turnstile / replayable evidence≠authority", "MLX one-pass schema→JSON / Apple Silicon replica economics", "memory leases ended by new evidence", "never confidently wrong / TLA+ compose / escalate instead of hard-gate", "no seal no advance / coverage ledger / mint ≠ product brain", "skill-broker sibling / judgment ≠ permission", "sureness bands / max_prob is generous", "JevBench / calibration not in Main Score", "CI typed gate before expensive review", "Codex MCP host adapter", "judgment as attention redirect / jev-preflight", "compress-before-first-send / dizk jev-lens", "tools≠use / SessionStart over hoping", "observational memory / pi-om keep-kind", "open-Jev class / openvons / JevPick", "physical-world System One / HA-Jev / not for locks", "judgment outside the store / jevql", "landed-script trust / headless≠auto-approve", "digital-design combinators / extended five", "VOI cache admission / same-intent skip LLM", "BM25 vs Jev skill routing Harbor harness", "zeroshot vs BERT / contamination DiD", "typed escalate continue abort baton / inverted loop", "worth-your-attention VOI / ThinkyMiner Winnow", "Jev WHETHER Python HOW LLM WHAT", "conflict vs ignorance / named Choice escape", "Playwright executes Jev chooses", "OpenJev /v1/decide not drop-in", "SemIf wire-compat runoff", "decision-as-memory flywheel", "record/replay CI / jevassert", "failure-finding arena / jevarena ≠ jev-arena", "BBQ not a bias cert", "decider≠executor", "sentence-as-rule lint / jevlint", "sentence-as-rule lint / jev-lint is jevlint rename", "VOI hunk prune", "whole-repo intent VERIFIED/VIOLATION/UNKNOWN", "GLiNER2 spec ≠ replica", "open replica substrates / grande / laya-jolt / JEV-CPU", "ONNX local-jev not equivalent", "persist constraints across compaction / pi-heed", "calibration+cost first-class gates", "Harbor-shaped Jev vs SGR LLM-as-judge / jev-judge-bench ≠ jevarena ≠ jevbench", "hand no-text steps / jev-use / Vercel drops confidence", "Pi System-One control plane / pi-jev-control", "never free-generates / jev-gpt tree of Choices", "OpenRouter recipe atlas / samples not benches", "personal history feed / jevfeed / no social graph", "competing NAR claims / dual-channel ECE / openJev-verdict ≠ OpenJev", "empty compaction-proxy skip / IPECTER", "throughput ≠ latency / like-for-like ECE", "1-token logprob endpoint ≠ Noul / coverage ≠ correctness", "open replica engine / jevinf", "unofficial Elixir SDK ≠ OTP peer", "jevex n=16 files-to-read VOI", "commit pre-review attention≠verdict / middle band", "Hermes plugin is Agnes not TypeSafe", "pi-jev-compact ≠ pi-jev-compaction", "empty Codex-proxy skip / IPECTER runway", "decision-native inbox / mailordinal", "unofficial jev-cli not ready / ≠ jevql", "laya-multilingual / English checkpoint confident-wrong OOD", "schema-scorer peaked ranking ≠ calibration", "HF 401 / GitHub 404 Hub-only", "productized System One HTTP / classifier.dev", "escalate-under-threshold / smart tier / multi-label ignores", "silent FALLBACK / granite 0.546 vs advertised 0.800", "vs_jev tracked JSON / read eval/README", "choxos/jev-reviewer ≠ egma-ai / systematic-review pointer", "two-pass Choice+Noul evidence extraction", "not-found is an answer", "human check as productized judgment", "githubnext/localjev ≠ kunchenguid/local-jev", "wire-compat ≠ logit-equiv / prompted JSON ≠ structured read", "self-reported probs / entropy confidence", "GitHub Next local /v1/systemone", "LM Studio runner gap / structured-read primitives", "NandhaKishorM/laya packaging ≠ Hub-only / Router script-before-p", "post-T ECE ≠ raw ECE / Banking77 token-budget", "0.85 still soft / not TypeSafe drop-in", "external census ≠ scored bake-off", "GLiNER2+routers class-boundary", "incomplete openjev census vs watch", "Harbor honesty watch / silent fallback", "JevBench v1.2 geometric mean / cal ON rank / weight sensitivity", "option-order 72→21 / instruction models in the class table", "self-host latency ×2 assumption / est. costs", "Laya absent is a gap not a named exclusion", "Qwen3.8 27B ≠ Archer", "hourly already-folded watch / apply-the-five / skip thin noise", "hard-gate Noul as PR/quality gate is soundness theater", "S1 never stalls waiting / S2 one-use advisory", "Local controller ≠ githubnext/localjev", "purple telemetry = consumed not arrived", "seed = geometry not async replay", "20% starting gate still soft", "no pixels to either provider", "OCR+AX observe-score-act / typesafe-computer-use", "never send screenshot to frontier for the decision", "overlapping CU options = false low confidence", "split kind/item/site", "155× one-screenshot ≠ Harbor taskset", "decision ≠ answer-reader capture", "ASR observe-score-act / jev-voice-browser", "partial-speech VOI / free-text waits", "spoken confirm ≠ hard auth", "numbered overlay without another model", "wrap-as-execution / AgentGhost ALLOW ASK DENY", "rules first then Jev remainder / ASK throws / fail-closed", "reddpy/AgentGhost ≠ jwen5419807/agentghost ≠ vventirozos", "JP genre atlas / studio_yebisu / stars ephemeral ≠ eval", "Jev Clearly Explained / akshay_pachaar / LLM hammer", "schema-safe ≠ correct / 200× 400× TypeSafe ceiling", "questions-as-code / shadow first / not a TypeSafe how-to", "proposition ≠ embedding / contrast-set", "boolean composition of soft Nouls / AND OR NOT", "uehaj/jev-semgrep ≠ semgrep.dev", "meaning-grep dedicated fold / not a gate", "decision-validated UI / Jev never authors text", "decision-as-assert / jevtest ambiguous band", "typed decisions drive UI / jev2ui", "hybrid S1 closed verb menu / anima3", "pointer-not-generator search / JevFind", "jev-frontier-bench ≠ frontier-100", "product bakeoff ≠ architecture duel / GLiClass", "four engines same questions / majority floor", "authorship named escape / not evidence", "ha-switchboard HA remains execution", "n8n classify/route/score / Low Confidence", "fast-jev-compaction-pi ≠ pi-jev-compact ≠ pi-jev-compaction", "jevloop full-distribution optimizer / no LLM in the loop", "laya-vision SmolVLM / score untrained", "Cerebellum-2B /v1/decide ≠ TypeSafe / wire-compat vs agent-routing", "laya-grounded not drop-in / Platt not temperature", "GestaltLabs/Jeff-1 ≠ logan-markewich/jeff / acc vs ECE n=9730", "stanley-code empty findings ≠ approval / human promote", "findme ≠ JevFind / NL memory beam-search FS", "jevsubrouter price workers not conversation / counts ≠ dollars", "feelings .feels() default 0.5 is Noul-0.5-never-rounded / ≠ hunch ≠ Probably", "apa-agent-harness ≠ AntonioCoppe/jev-harness / unpublished npm", "grok-bot-jev skill cannot force a bot that ignores it / A/B proxies not tokens", "Essentiel-Jev never authority / human every action", "enzo-mcp independently falsifiable claims / ≠ jev-sift", "pigeonhole OTHER skip / decision-as-filing", "jev-reliability Nothing about accuracy", "clduab11/jev-test ≠ realZachi/jevtest / Nothing runs yet", "jev-rag-benchmark Jev wins is not an assumption", "dairui1/jev-lab ≠ BrendanH18/jev-lab", "jevmail gmail.readonly / mailjay archive/trash", "ZHUBoer/ego-jev reserved __none__", "runWorkflow completed ≠ success", "jsort scores are relative", "Noul not Choice for scale", "groundedness-judge-bench native vs schema-guided", "implicit_true included in yes", "jev_playground 0 promotions", "routing-backtest 0.0447%", "yuyang2230/jev-agent-skill jev-1.13-free", "jev-techstack-classifier stack_config.json", "s1_ruby collapse late", "undecided? abstain", "2389-research/judgement license null", "confidence ≠ winner p", "typesafeai-sdk-community not a new species", "tpellet/hunch exit 3", "never-execute list", "jev-file-search scores not calibrated accuracy", "jev-linkmap Jev never sees S2 prose", "muhammedilyasy/jev-mail metadata only", "tidy none-of-folders stay", "tab-bouncer pinned/audio/current never closed", "lkclean Show fail-open", "jev-yt-time-saver Show anyway", "ORIGIN pause-if-no-Jev", "validResponse sums-to-1", "jev-crawlers risk bands never raw boolean", "jevbrain AUTO_ACT is not a Noul", "judgekit YAML classify/score/route/verify", "typed-judge-kit verdict-in-code", "alsoleg89/decide packing VOI", "0.8 ≠ 80% accuracy", "Jev-Calibration Platt ECE 0.117→0.052", "jev-calibration-arena never acts", "ctmx/openrouter-jev-mcp Decision-as-Plugin", "FrancoisChastel/jev-code ≠ npm jev-code", "claudecode-jev-marketplace fail-open not hot path", "pedroknigge/mcp_jev packs not ask_jev", "cyrusasco/typesafe-mcp noul deadband 0.35–0.65", "codaaiteam/jev-skill jevtypesafeai.com ≠ TypeSafe", "hermes-switchyard ≠ hermes-jev-router ≠ hermes-plugin-jev", "nanoprune 2.8MB ECE 2.58%", "smartdio/jev-browser-agent ≠ ZHUBoer/ego-jev", "Dakai/omp-jev-web DONE ≠ proof", "hari007sh/jev ≠ dannote/jev", "0thernet/system-one-skills deterministic verify", "typed-gate band [0.40,0.60] is refusal", "pi-jev-gate fail-closed; choice is the verdict", "Foq ~25ms/2.2GB local", "rev prefill-only + HF jev-0.5b", "robfrase/jev planning memo", "typesafe_agent_gates 27/27 / 31/31", "EpicEric/safe-sh static remainder", "pastepilot Confirm before act", "Jev-Reranker live Jev not yet measured", "sessionwise opt-in relevance", "jev-search pointer sieve", "400ms Salesforce WebMCP", "typesafe-scheduler-diagnostics advisory", "droidjev screenshot-free", "Tewoto1 jevcu planner still writes", "ha-conversation-jev Jev→Grok", "dsh-jev can only gate", "jev-classification-benchmark specified not run", "jev-luna-pagerduty p≥0.50", "meldltd/meldecision laya-go ONNX", "laya-doom never pixels", "logixism/laya-api empty README", "akpsahan/laya ≠ Archer", "choxos/jevchess engine owns truth", "jev-drive sim not AV", "story-arc Jev never authors", "jev-hs-assistant HS6", "golergka/jev-plays-starcraft-2 UI-verified ≠ API Victory", "awesome-jev-use-cases catalog", "Nibir1/typesafe-go ≠ official", "fingerprint after redact", "recall vs decide", "publish fingerprints+answers", "CI replay as Harbor cousin", "Cache hit ≠ correctness", "hyperspaceai/jevcache ≠ kushals256/jevcache", "human labels only", "score never auto-accepts", "production capture flywheel", "sutro-sh/jev-align ≠ caiovicentino/jev-align", "guidance ≠ hook", "catalysts ≠ summaries", "compile-time System One", "unofficial ≠ TypeSafe", "format_version modernbert-jev/1", "Argos1111/jev_local ≠ us/jev-local ≠ kunchenguid/local-jev", "LFM default ≠ ModernBERT backend", "Nemotron ≠ TypeSafe Jev", "not a calibrated replacement", "djev-dev complements djev-spark", "images as Choice options", "Laya essay numbers *theirs*", "Router/OOD confidence", "hosted bootstrap ≠ silent TypeSafe", "difficulty + policy thresholds + JSONL trace", "jev-codex-pilot model + reasoning depth", "keep/shadow/hybrid/reject", "quarry evidence projection", "Frank-ZY-Dou/awesome-jev robotics/3D/control", "one-dollar-tahoe TypeSafe Jev defense eval", "jevguard calibrator/cache/escape", "jev-ci-selector CI shadow mode", "llama-jev llama.cpp replica", "petercr/jev-orchestrator ≠ FleeexCorp/jev-orchestrator", "seb4ez/jevguard ≠ AseemPrasad/JevGuard ≠ pablozr/JevGuard", "webNeat/llama-jev ≠ WiktorB2004/llama-index-jev", "OpenCode jev-pruner context sieve", "observe→score-candidates→prune", "jev-zen / jev-1.13-free", "zen-chat ≠ Noul", "fail-open original", "keepScore >0.1 floor", "host port of tamaratran/jev-pruner", "indiejoseph/opencode-jev-pruner ≠ nrdz-labs/fast-jev-opencode", "jev-webagent-bench empty stub", "Kiln-AI/jev_jsonschema noul_threshold 0.5", "NSStudent/JevSwiftSDK unofficial", "GLiNER2 native Apple path", "unofficial Swift/Core ML GLiNER 2.5-small", "entity spans + confidence", "not Choice/Score/Noul", "not TypeSafe", "label descriptions as schema", "on-device ANE economics", "honesty locks", "shershah1024/gliner-native-runtime ≠ Fastino", "≠ gliner25-compaction ≠ gliner2-ultrafast ≠ Eran-BA/Jev_from_GLiNER2 ≠ NSStudent/JevSwiftSDK ≠ jevmlx", "default threshold 0.1 still soft", "soft Noul ≠ hard safety", "Decision Graph Protocol frame→assess→commit", "app retains permissions/effects", "Jev-first assessor-neutral", "guarded commit / receipt/next frame", "assessment batching", "hard-gating DGP as safety theater", "numerous-com/dgp ≠ TypeSafe official", "jegrep calibrated path+range Nouls", "no embeddings/index/daemon", "~$0.01–0.03 typical", "agent --json", "can1357/jegrep ≠ Bentlybro/jevgrep ≠ uehaj/jev-semgrep", "Archer-arch fidelity", "kev family OOD 0.76–0.77 vs Jev 0.86", "block-causal isolation", "pointer/readout CE-trained", "/v1/systemone drop-in", "replica honesty", "cost-sensitive decision theory × System One probabilities → control flow", "thresholds derived from costs not hard-coded", "YES / NO / UNSURE from cost_false_yes / cost_false_no / cost_human", "auto-batching same-object questions", "Kungie/gut ≠ tpellet/hunch ≠ carldaws/hunch", "judgment vs generation", "deterministic execution after probabilistic judgment", "exactly one app-owned callback", "explicit uncertain branch", "Illusion47586/judge ≠ lexingtonhibiki/judgekit ≠ Ascurse/typed-judge-kit", "variable-N option scoring as the trainable object", "dynamic candidate bags not fixed label sets", "zwliJay/jev-forge ≠ NanoJev", "open replica economics / latency vs closed Jev", "NAR local drop-in", "wfzyx/von late-catch HIGH", "competing NAR claims / replica honesty", "typed judgments vs chat judges on guardrailing", "ishaannk/llm-vs-jev cross-note only", "deeper integrity fold is rh-guard", "nothing wins outright", "can be argued out of guarding"", "Jev IS the if-statement", "judgments/probabilities drive branches", "text model only writes prose", "interpreter owns variables/loops/budgets/replay", "otherwise maybe / confidence gate", "chaos samples after the gate", "southpolesteve/probably ≠ carldaws/hunch ≠ feelings ≠ Kungie/gut ≠ Illusion47586/judge ≠ tidymodels/probably", "133★ / forks 10 live", "build calibrated classifiers from human feedback", "retrieve by relevance not resemblance", "one calibrated yes/no per memory in one request", "pointer mode 17/18 19/20 *theirs*", "embedding resemblance misses the allergy", "samdotmak/jev-recall ≠ jev-search ≠ jev-sift ≠ carryforward ≠ chopratejas/invalidate", "memory leases ended by new evidence", "six Nouls then fixed rules in code", "0 of 157 false invalidations", "questions/plans/directives are not evidence", "unsure → review queue", "host keeps the store", "name↔body / comment truth / test-claims", "mizchi/jev-lint is mizchi/jevlint rename", "no shipped rule has severity error", "~1 in 5 findings wrong *theirs*", "mizchi/jev-lint ≠ huntedman/JevLint ≠ MichitoSugawara/jev-lint", "JSON Schema → typed JSON via Jev", "noul_threshold 0.5 decoder not a proof", "IncompatibleSchemaError lists every bad property", "on-device Laya CoreML ANE", "~5 ms P50 short decisions", "189/189 FP16 checkpoint parity", "10× not achieved", "mizorewww/laya-coreml ≠ gliner-native-runtime ≠ jevmlx ≠ NandhaKishorM/laya", "softmax over allowed tokens ≠ Noul", "question-first cache", "Micha0827/snapjudge ≠ githubnext/localjev ≠ jevmlx ≠ cendress/SnapJudge", "Jev-first Pi agent loop", "slow-LLM fallback", "explicit action menu / CandidateSource unimplemented", "62 tests wiring not quality", "direwolfiy/JevPi ≠ standardagents/jevpilot ≠ pi-jev-control", "resume-screening bias audit methodology", "name×resume factorial independent Nouls", "callback determined by resume quality", "mean-probability name gaps operationally negligible", "natemoo-re/bias-bench ≠ BBQ", "Plan/PRD panel → code-owned pass|review|block", "cheerleading out of scope", "austindixson/planalyzer ≠ single-goodness Noul", "cost-aware multi-model routing/escalation", "decide vs do", "successful-task cost", "cannacre8ive/switchboard-ai ≠ ha-switchboard ≠ hermes-switchyard", "frozen-protocol zero-shot bench", "TypeSafe Jev vs PrismNLI vs Laya", "contamination caveat", "elcronos/jev-vs-open-decision-models ≠ JevBench ≠ DMB", "context-window admission control", "VOI gate which tokens are worth the expensive model", "fail polarity per lens", "on small inputs lenses lose money", "cvsgireesh/jevusher ≠ jev-sift ≠ winnow", "typed decision control plane", "receipt ≠ authorization", "historical-v0 zero retained cases", "MokiMeow/jev-fabric ≠ jev-forge ≠ dgp", "live 15-dim typed rubric re-score per pause", "scoring economics exemplar", "OpenJev/Codiv ≠ TypeSafe hosted", "jose-troche/live-rubric ~$0.000004 desc / ~$0.000006 README", "adversarial pre-registered Jev eval", "28 predictions before data", "123,805 requests", "confidence does not track ignorance", "polite injection 65% / crude 0%", "willkelly/jev-evaluation ≠ jevals ≠ jev-baselines-eval", "provider-neutral Elixir/BEAM Noul/Choice/Score SDK", "class infrastructure", "nshkrdotcom/system_one_sdk ≠ typesafe_sdk ≠ dannote/jev", "question-linting of Jev questions themselves", "nine jaggedness rules, no API key, no labelled data", "static lint ≠ measured separation", "yodablocks/jevq ≠ tenbin ≠ JevLint ≠ commitjev", "open-weights Laya as class exemplar (binding)", "Nx/Bumblebee runtime", "host chooses backend", "ChristianAlexander/laya_ex ≠ system_one_sdk ≠ dannote/jev ≠ NandhaKishorM/laya", "on-chain/edge Laya deploy", "parity_verified stays false", "model output never grants Tx", "humandebri/IC-Laya ≠ laya_ex", "auditable weekend replica", "Jev outputs never used for training", "soft human-vote distributions", "unpaired 0.577 vs 0.727", "agilabs-ai/jev48 ≠ JevBench ≠ Mapika/decider", "adversarial dual-judge / framing attack surface", "comparative framing is the usable judgment", "prior injection crowds out evidence", "copyleftdev/ember ≠ ember.js", "Laya specialist fine-tune pipeline", "training still GPU-pending", "PIXELZX0/XERON ≠ convaiinnovations/laya", "Hub Laya replica drop", "daliborsb/laya ≠ convaiinnovations/laya ≠ NandhaKishorM/laya", "System One student distillation corpus", "gold is programmatic", "teacher is closed-API clone", "do not distill Jev as teacher of record", "MagaBitmex/jev-4b-distill-data ≠ missing student checkpoint", "non-LLM VIN System One", "planning depth not chat", "lewislululu/jevon ≠ douglance/jevon", "source-bound evidence checks", "local quote mismatch needs no API", "exit 0 ≠ claim truth", "WaynezProg/jev-kit ≠ jonathanavis96/jev-kit (Airlock) ≠ jev-use ≠ jev-mcp", "independent System One evidence catalog", "scores not one leaderboard", "no external record currently reproduced", "TokenTrim no-Jev matched hybrid 62.4%", "reachjalil/system-one-bench ≠ mallahyari/system-one-benchmark", "21 tasks · 134 items · 208 questions", "scenes from public GitHub contracts, not production logs", "SivletLabs/jev-eval ≠ willkelly/jev-evaluation ≠ 4esv/jev-eval ≠ xxkuboxx/jev-eval ≠ onlyoneaman/jev-eval ≠ dayhaysoos/jevals", "option isolation (sibling-blind)", "permutation-equivariant", "Hub OWNER not published", "nafisazizir/hev ≠ jaredpalmer/kev", "frozen local LLM logits, no trained decision head", "residual-head 9,222-param decreased 73/96→67/96", "confidence = 1−normalized entropy, not P(correct)", "yuki-oshio/mini-jev ≠ r-ms/mini-jev", "Jev classifier as autoregressive next-token predictor", "ChatJev-style soundness theater", "erik-dunteman/ChatJev ≠ dannote/jev ≠ jev-gpt", "calibrated decision head × AlphaProof value head", "implementation-layer isomorphism, semantic difference", "timeout = censoring", "do not launder Noul as proof", "parallel rank-prediction vs serial selection", "independent questions can conflict", "zzzzzec/jevsort ≠ keltokhy/jsort", "curated open System One ecosystem catalog", "rupeshpoojary9/awesome-open-system-one ≠ AnotiaWang/awesome-jev", "arXiv paper radar with Jev relevance scoring", "ranking ≠ calibration / 0.5 still soft", "fail-open failed evals not marked seen", "train calibrated ~27M from scratch", "typed Q→prob dist / one forward pass / no LLM decode", "hyusi2003/MiniSystemOne ≠ Colvin0315/MiniSystemOne", "description-only stub / size 5", "ESCI hard probe fails four of six", "jev_bool ECE 0.242 inversion 0.255", "do not re-fold §60 six-gates as new", "jobbyjev one-request-per-company from batch-size result", "find/design/evaluate TypeSafe Jev decision loops", "karanb192/jev-architect ≠ samtay32/jev-system-architect", "Jairik/jev-distiller size 1", "distill-Jev UI stub / do not distill Jev as teacher of record", "post-launch scored use-case map / Jev self-scores then human curation", "licensedsaucer9-web/jev-opportunities", "Jev-inize a use case into classifier/router", "gavinHuang/jevinize → simple-jev not TypeSafe", "featherless-ai/simple-jev", "compare saved decisions / same label can still change the branch", "VihaanAgarwal/jev-diff ≠ Saik0s/diffusiongemma-jev-macos", "not tested with a live Jev API key", "constrained logprob + temp/Platt ≠ Noul", "OpenJevPro pastes openjev-sglang JevBench as own", "zhangcy122/OpenJevPro ≠ IamBusy/OpenJev ≠ ekzhang/openjev-sglang", "PolyForm Noncommercial", "SmolLM-135M / sub-70ms / 0 output tokens", "demo P(True) 0.5052 / Choice conf 0.2872 / Score conf 0.0055", "README claims MIT / GitHub license null / no LICENSE file", "patelvishwa112/jev-system-one-rlcd ≠ arnabgho/rlcd-lite ≠ blackwood-rlcd", "source-backed Awesome Jev radar / 306+ commit-pinned", "logicrw/awesome-jev-projects ≠ AnotiaWang/awesome-jev ≠ yibie/awesome-jev ≠ cobanov/awesome-jev ≠ rupeshpoojary9/awesome-open-system-one", "auto GitHub sync / Issue-only submissions", "hashed n-gram encoder / rival-aware attention", "olanotolu/jevbetter vs jevlike starter", "synthetic hard menus top-1 0.916 vs 0.873 / ECE 0.0182 vs 0.0367 / 40 vs 4608 menus/sec", "shuffled-context control 0.335", "Turn any open LLM into System-One Jev", "uspraveen/Jevify ≠ Mintzs/jevify ≠ gulagala001/jevify", "Jevify-any-LLM architecture probe", "description-only stub / size 0", "Train encoder-only calibrated decision models from a task sentence", "Exu is a toolkit, not a method", "strictly proper scoring rule", "Pre-alpha", "Ruivalim/exu-base", "scratch-trained calibrated decision model", "typed Q → probability dists", "Colvin0315/MiniSystemOne ≠ hyusi2003/MiniSystemOne", "no published weights download URL", "90.5 seconds / 29.2% pipeline evidence", "p_i/p_j independent of other candidates", "Recipe for calibrated decision models — small model out", "init → synth → train → eval → serve", "91.1 % / ECE 0.022 *theirs*", "Jev zero-shot 75.1", "scienthoon/luce", "Put Jev's three headline claims on trial", "0.5B local GPU", "46x speedup / accuracy identical", "ECE 0.624 sentiment catastrophe", "bigger model worse calibration", "RichardoMrMu/jev-mini ≠ yuki-oshio/mini-jev ≠ r-ms/mini-jev", "System-1 decision engine for local LLMs", "structured choices only", "JSON parse of generated text ≠ Noul", "TypefAI JEV / Journal Entry Voucher", "tapsin/jev-local ≠ us/jev-local ≠ Argos1111/jev_local", "Jev 1.13 reward-model eval across 8 benchmark tracks", "40,940 examples / 0 API errors", "RewardBench v1 92.58%", "Precise IF 50.63%", "goya4140/jev-reward-model-evaluation", "Scaffolding in progress", "Jev vs LLM support-ticket routing", "static + live decision bench", "TypeSafe's own published benchmark", "illustrative simulations, not live API calls", "JevBench v1 — smart/cheap/fast/reliable", "I/C/S/K 25% geometric mean", "classifier.dev fast tier 84.8 is Jev behind its own API", "do not re-fold §78 v1.2 board as new", "Laya (421M) 70.1 now on board", "Zero-shot/few-shot LLM routing", "hard budget filter before Jev", "Jev never asked to perform budget arithmetic", "Jev judges the next state, XState enforces transitions", "simulation uses synthetic keyword fixtures", "catalog gravity", "v-modal/awesome-jev-tools", "★339 live REST", "curation is not endorsement", "crawler-maintained directory", "Daily GitHub + npm sweep, human-merged", "RadRebelSam/awesome-jev ≠ AnotiaWang ≠ yibie ≠ cobanov ≠ logicrw ≠ v-modal", "HF peft SPLADE/BGE reranker", "rdxtremity/jev-reranking ≠ carlaiau/jev-reranking", "query-side encoders, not a Jev replica", "ONNX System One Qwen3.5-4B scorer", "source:pngwn/system-one-qwen3.5-4b-scorer", "CC-BY-NC-4.0", "temperature 1.75", "transformers.js AutoModel cannot load this graph", "Consistency benchmark Space", "This Space contains no benchmark result yet", "12-case plumbing fixture", "Benchmark-driven Jev router and judge", "cheap alone is not success", "Jev does not write, sum prices, or claim accuracy %", "Sol 94.2 / Luna 83.9 / Jev path 89.7", "19.2% Sol / 62.3% cost save / 4.5pp miss of 2pp non-inferiority", "p50 latency worse than Sol due to routing overhead", "erendikmenn/jev-llm-router-benchmark ≠ jev-rag-benchmark ≠ ryantsai/jev-llm-router", "Express + node:sqlite", "mock and Jev decision engines", "previous_ticket_count >= 3 is code", "MIN_CONFIDENCE 0.6 still soft", "substring false positives", "aesaganda/jev-ticket-router ≠ SarathChandraBellam/jev-vs-llm-ticket-router", "Universal Figure & Diagram Router", "confidence ≥ 0.85 hard-gate is theater", "generative AI banned from scientific plots", "six visual branches", "hoangngochuong24947-gif/jev-figure-router", "human-labeled (state, question, label)", "166,054 rows / 22 configs", "soft_label for human uncertainty", "Praveenrajus/jev-bench ≠ fstandhartinger/jevbench", "ternary bonsai System One GGUF", "openjev's mechanism, Bonsai's weights", "Hub does not ship weights", "100/100 easy T/F is not Harbor", "label_mass ≠ correctness", "stock llama.cpp Q2_0 silently gibberish", "NicolaiMTLassen/open-bonzi-jev ≠ NicolaiLassen", "transformers.js DeBERTa ONNX", "source:com-kotobalabs/open-jev-deberta-v3-large", "temperature 1.05", "AutoModel from_pretrained works", "onnx-community/open-jev-deberta-v3-large-ONNX ≠ system-one-qwen3.5-4b-scorer-ONNX", "107★ densify", "GH 151M vs README 149.6M", "PR #1 now closed unmerged", "do not re-fold §71 claim-audit as a beat", "typed decisions, RLCD, confidence-gated routing", "structured ≠ correct", "mock not live API", "26 tests", "wjdjdakf17/jev-study ≠ baekenough/jev-study", "bonzi-27b-v2 / ternary-8b / 27b-v1 GGUF family densify", "WANLI-256 74.6% / 65.2% / 71.1% *theirs*", "Bonsai 1 27B Q1_0 runs on stock llama.cpp", "ternary still needs PrismML fork", "hf:heman10x/openJev-verdict-2.0 twin tokenizer-only", "OpenJev Vision image classification + uncertainty", "CLEVR-4 held-out joint 0%", "hfdataset:IamBusy/OpenJev-Vision-Research-v0.1 12,832", "294,912 derived targets not independent samples", "Laya multilingual ONNX WebGPU typed-decisions port", "63/63 selected answers / 5.1e-4 CPU / 1.2e-2 WebGPU", "UpHash-Network/mini-jev is yuki-oshio transfer", "jev-injection-bench 11,900 labelled prompts", "Jev best ranking / Haiku better ECE 0.021 vs 0.058", "0.5–0.9 band is where Jev's numbers do not mean what they say", "Prompt wording moves panic 28%", "manojlds/jev-dspy-bench ≠ dspachos/jev-dspy ≠ jmanhype/jev-dspy-lab", "Jev agreement is similarity, never ground truth", "no aggregate quality grade or merge gate", "AbstentionBench-on-Jev rank 1 of 20 vs 2025 field", "question-asymmetry", "forward-looking 0.465 never extreme", "openkev calibration layer not a runtime", "ECE vs coverage independent", "select_threshold returns inf", "escalation catches uncertainty not ignorance", "misakaikato/openkev ≠ jaredpalmer/kev", "pdf-race Docling→Jev vs Gemini", "parser owns the wall clock", "12/12 tie is a tie", "titles selected not generated", "flopcheck 16 calibrated tweet judgments", "mechanical tells in code", "ZeroX-01/jev-atlas ≠ Zaious/jev-capability-atlas ≠ gorock007/jev-atlas", "Laya calibration lab Gradio MCP", "T never changes argmax", "confidence ≠ top-label p", "easy probe set refused", "40–48 rows too small to ship T", "Gemma-4 26B-A4B jevify classification+calibration", "LoRA adapter twin not independent eval", "Gemma-4 E4B jevify", "E4B LoRA stub card", "kushalpatil/jevify-gemma4 ≠ Mintzs/jevify ≠ gulagala001/jevify ≠ uspraveen/Jevify", "GH kushalpatil07/jevify 404", "PAWS 0.580/ece 0.288 is the weak cell", "smaller E4B slightly better OOD ECE than 26B-A4B", "Hub jevify merged LoRA ships weights", "bonzi Bonsai-8B v1 GGUF densify", "Bonsai-1.7B v1", "Bonsai-4B v1", "WANLI-256 64.5% / 60.2% / 52.0% *theirs*", "rank #4 / #5 / #6 of 6", "JulesHuisman/jev-eval scaffolding / README SHA c356a584 (was empty e69de29b)", "JulesHuisman/jev-eval ≠ SivletLabs/jev-eval ≠ willkelly/jev-evaluation ≠ 4esv ≠ xxkuboxx ≠ onlyoneaman ≠ dayhaysoos/jevals", "7 bands 6/10 vs 40 bands 0/10", "source receipts + confidence slider re-policy without re-inference", "32/32 synthetic is smoke not production", "classify HF datasets across typed semantic dimensions", "roadus2 watch misspelling; lock roadius2/ultra_laya", "ultra_laya REVIEW defects", "default branch claude/laya-jev-review-gg5ppo", "XNLI EN 88.3% ECE 0.032 → RU 77.3% ECE 0.096", "Δ −11.0 pp [−14.2,−7.8]; ECE +0.063", "MASSIVE no detectable difference at n=600", "confidence is function of p_max (r=1.000)", "pointer-not-generator 400 human-authored responses", "proposed ≠ authorized", "FewRel 160: Jev 85.0% vs lexical 13.125%", "gated 100% (95/95) coverage 59.375%", "J++ composable semantic computation language", "judge-jev 0.5 still soft", "947 repos scored; A 273 / B 302 / C 372", "LLM rubric ≠ benches", "No benchmark winner is claimed", "phishing: naive 62.6% vs regex 91.8%; 5-atomic + LR 95.0% *theirs*", "AITuber tension ±15", "README npm global; repo is Rust", "git-confess code owns counting/blame/ratio", "httpx exhibit 11% (13/119) *theirs*", "90d trend +12.40% vs random +12.75% vs BH +41.71%", "5m win rate 25%", "Awesomejev 656 entries / 38,160 stars", "tracker likes 64 (+4) lastModified UNCHANGED", "Laya present; Blackwood ABSENT; Archer still promised_not_landed", "Blackwood tracker ABSENT; likes 2 gated manual", "r = c - p_a", "ECE 0.021; acc 0.807 vs warmup 0.746", "Independent primitive", "11.57s vs 54.10s · 4.67× · 120/128 *theirs*", "default path is pretrained Gemma probs not trained RLCD head", "GH Meanblock 404; lock leesk212/JEV-CPU", "softmax over letter slots ≠ Noul", "WANLI 0.741 vs openjev v2 0.77 *theirs*", "3-way NLI ≠ Noul", "priority 0.464 = majority floor", "banking77 contaminated", "raw margins not probabilities", "do not distill Jev as teacher of record (they distilled Haiku)", "“0.9 is not one number”", "ranking ≠ calibration", "banking77 0.8–0.9 stated 0.86 actual 0.73 over-confident *theirs*", "≠ Praveenrajus/jev-bench ≠ fstandhartinger/jevbench", "$0.0000153–$0.0000226 vs circulating $0.0004 (~20×)", "Score is 0..n-1 expectation not 0–1", "Noul has no confidence field", "TCP floor 198.8 ms", "type reliability is not a reason to choose Jev (json_schema 5/5)", "gateway tax not one number", "Function-only 5/8 vs hybrid 8/8", "4/8 without Jev", "8 designed cases not conversion lift", "200-row pilot Jev 86.5% 173/200 vs Gemini Flash-Lite 86.0% 172/200 vs Pro 87.0% 174/200 *theirs*", "not a ranking", "情緒測謊器", "8-example Jev vs GPT-5.6 Sol ~64× cost 5.4× latency *theirs*", "synthetic; no inference", "≠ JevBench v1.2 §78", "Judged 3317 / listed 2560", "Jev judges, code applies policy", "APA “microsecond policy / zero hallucination” overclaim", "Client-side quiz; pointer from held docs; scanned-PDF warn", "Jev judges / agent reasons / user decides", "selecting an option is not permission to implement", "pattern exact, judgement must clear floor", "no matching pattern → no model call", "not a correctness oracle", "Spec vs artifact remainder", "treating 0.85 as 85% / minProbability hard-gate as Harbor", "VERIFY acquires discriminating evidence, never same-pool confidence-only rescoring", "fast/full/max are ceilings not sizes", "Solar writes, Jev chooses NEXT ACTION", "do not reopen or amend PR #23 or #24 or #25 or #26 or #27", , "Calibration is not alpha", "NO CURRENT ALPHA CANDIDATE", "ΔR² approximately +0.00084", "Brier 0.2131387", "ECE 0.0421875", "Adding Jev probability to deterministic volatility improved Brier by only 1.4058e-05", "default 0.5 keeps zero non pinned", "keepResult median 0.14 to 0.17", "keepCall median 0.28 to 0.35", "usable range is about 0.10 to 0.25", "7.8% to 57.9%", "judges results it never sees", "task-finish eval not built yet", "$0.002 per compaction", "slavadubrov/sgr-judge-bench ≠ slavadubrov/jev-judge-bench", "Jev 108/120 $0.083 0.34 s", "Luna SGR 114/120", "paired Jev accuracy-difference intervals include zero", "not evidence of equivalence", "GLM SGR 26/120 93 format failures", "Terra-planned Jev hybrid 55/120", "rule-based by default, optionally Jev-backed", "empty README", "missing key cannot break the experience", "prefill plus exactly one decode", "softmax over A/B/C ≠ Noul", "BBQ 9,053/10,000 (90.53%)", "ECE 0.0890", "Mean confidence 0.9943", "overconfident", "score and noul not implemented", "DGUI 12 rows (was 6)", "INSTRUCT 119 rows likes 2", "encode the state once, decide everything in parallel", "0.740 accuracy against a 0.508 majority", "ECE 0.047", "fine-tune's advantage ends where its 384-token training data does", "jasonkneen/open-jev ≠ pngwn/open-jev", "same sha d41dc3cd", "Space does not call Jev", "recomputes routing from saved probabilities", "200-case Jev 97.0% / 100.0% / 95.0% / MAE 9.22", "synthetic repository benchmark", "Jev evaluations are advisory", "YehuiTang0316/jev-nlgrep ≠ Bentlybro/jevgrep ≠ can1357/jegrep ≠ uehaj/jev-semgrep", "default threshold 0.8 still soft", "40-line windows cannot prove whole function", "token-native sequential start/end Choice", "Gemini/Haiku stubs not configured yet", "handful of hand-written examples, not a benchmark", "Jev judged exactly what it was given", "laguagu/jev-skills ≠ laguagu/jev-evidence-lab ≠ Pleo2/awesome-jev-agent-skills", "contract_passed is not a claim of guaranteed factual truth", "Wilson lower bound 0.85 floor", "fixture mode no savings claim", "SemIf 2207★ (+21 vs §110 2186)", "jevlike 1043★ (+5 vs 1038)", "TypeAR 15★ (+1 vs 14)", "AnotiaWang 97★ (+1 vs 96)", "yibie/awesome-jev 506★ (+16 vs 490)", "Laya likes 822 (was 802)", "tracker likes 64 flat, lastModified UNCHANGED", "do not reopen or amend PR #23/#24/#25/#26/#27/#28", "Heman10x-NGU/Verdict-open-jev ≠ Heman10x-NGU/openJev-verdict-2.0", "TF-IDF + LogReg ECE 0.0207 vs Jev 0.1440", "Verdict-open-jev 48.07% vs Jev 90.80%", "abstention combined recall 10.00%", "p50 35.58 ms", "K=25 (maximum capacity) 72.00%", "0.85 coverage 84.60% selective risk 1.18%", "26.1× faster than standard Qwen JSON generation", "Jevify 90.0% / 167 ms CUDA graphs disabled", "Finding 1: Brier on stated confidence alone is a trap", "grpo_rlcr 0.78 / ECE 0.084", "reliability 0.007 but resolution 0.000", "27 900 schema-driven decisions", "13 600 / 13 600 questions", "candidate mass min 0.99999624", "22 configs · 166,054 rows · 4 calibration-gold", "sha a39eba3f", "Student B MAE 0.148 / Pearson 0.836 / 86.0%", "pngwn/open-jev-laya-bench README 404", "sha 9f69c742 likes 2", "HDFS 0.9933 (745/750) / retain 0.0084", "BGL ERROR/FATAL protection 1.0000", "2,479 / 2,500 HDFS uncertain", "cache hit 0.9648 (2412/2500)", "$0.153936 estimated", "E2 recomputes from saved probabilities", "Space sha eda59e0a", "MASSIVE English 0.783 / Khmer 0.033 / Hindi 0.133", "40–48 rows too small to ship T", "T never changes argmax", "siren2345/jev-single-decode-transformers ≠ siren2345/jev-single-decode", "Split Transformers experiment from llama.cpp runtime", "tanayvasishtha/jev-lab ≠ dairui1/jev-lab ≠ BrendanH18/jev-lab ≠ yibie/laya-jev-lab", "Four experiments stress-testing TypeSafe's Jev: calibration, bundle bias, label bias, and ensembling", "second pass must be $0.00 from cache", "The pages never call Jev", "Gemma 4 31B 77.0% / Jev 1.13.0 61.4% / Laya 322M 0.0%", "restriction state 95.0% against 84.4%", "None of the systems are particularly good at knowing when to stop and ask", "They skip the question and call a tool directly", "100% schema pass", "six-field joint 48.8% vs 72.8%", "ywchiu/jev_benchmark ≠ Running-Dolphins/jev-bench ≠ Praveenrajus/jev-bench", "ACT / REVIEW / FALLBACK", "A provider failure, timeout, malformed output, or missing answer is **not** a policy outcome", "confidence is descriptive provider output, not a substitute for probability", "Quality denominators include only valid scored answers", "an exact halfway tie chooses the lower level", "aiwithenoch/Jev-Skill ≠ simplosophy/jev-skill ≠ laguagu/jev-skills", "The local path does not claim to turn a smaller checkpoint into Jev", "Low support becomes decision: \"review\"", "MIT-0 SPDX NOASSERTION", "current-llm", "结构兼容,不是 Jev 模型能力", "altryne/jevify ≠ Mintzs/jevify ≠ gulagala001/jevify ≠ uspraveen/Jevify", "Find where Jev belongs. Design the questions. Measure the difference", "TypeAR-AI/TypeAR 301 → TypeLLM/TypeLLM", "TypeLLM/TypeLLM 16★", "SemIf 2241★ (+34 vs §111 2207)", "jevlike 1051★ (+8 vs 1043)", "AnotiaWang 98★ (+1 vs 97)", "yibie/awesome-jev 525★ (+19 vs 506)", "Laya likes 864 (was 822)", "tracker likes 67 (+3 vs 64)", "lastModified UNCHANGED `2026-09-20T04:29:16.000Z`", "do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#32", "hysteresis enter/exit / replay policy without inference", "calibration does not compose / hop-ECE permutation-invariant", "equal-width vs quantile ECE / ranking ≠ calibration", "Qwen2.5 ≠ Archer / Qwen 3.8 sparring ≠ Archer / Qwen/Qwen3.8-27B ≠ Archer", "Deferred Crispification / TCE / AMS", "g0runmezadam/what-is-jev IS tunahansahin897/what-is-jev", "pd.cut equal-width vs jeval quantile", "A hunch is a probability with a policy attached", "soundness theater / measurement theater / hourly 0843", , "Jev Capability Resolver / NiazMorshed2007/jcr", "one tool nested capability tree / returns context / does not execute", "skills vs capabilities / workflow+judgment vs operations", "format independent of Jev / proposed open standard", "JCR_BAND_RATIO 0.6 is application policy / soft scores ≠ hard gates", "routing ≠ permission / docs ≠ authority to run", "sol-vs-opus5-20 lookup+explain / n=1 / Not Harbor task-execution", "wall-time mixed / Sol slower with JCR in 19/20", "NiazMorshed2007/jcr ≠ skill-broker ≠ skillranker ≠ jev-sift ≠ jev-lens ≠ jevusher ≠ jev_select_capability", "do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33/#34", "notes.md §116", or "cascade sign-flip / calibration theater": read `references/faq.md`, + other", "training confronts Choice other / none-of-the-above", "soft AGENTS.md rules vs the linter", "screenshot Choice / omni System One", "extractive quotes / pointer not generator", "compaction summarize vs pointer", "encoder vs Jev compaction backend", "shadow-mode compaction rollout", "CI flaky-vs-real merge gate", "fail-open VOI wake/resume", "claim vs session evidence", "S1 indexer escalate-S2", "Harbor on/off routing", "fail-open vs fail-closed wake vs CI gate", "encoder vs Jev computer-use backend", "hybrid local decide + remote fill", "DONE vs verified success", "stdout prune vs session compaction", "OpenCode jev-pruner vs Claude jev-pruner", "zen-chat vs jev-zen Noul", "hard envelope then Noul prune", "Cua-S1 vs TypeSafe Jev", "plan vs execute dry-run", "specialist computer-use vs general agent", "local drop-in vs stub scorer", "route vs memory", "when does it hold / extractable from state", "decision model vs constrained LLM", "dual-process S1/S2", "combinatorial grid vs extractive", "uncalibrated local likelihoods", "decision-native RAG", "classify-first / read selectively", "living applied-mappings atlas / class patterns", "silence as safer / draft-gate heartbeat", "robotics text-state vs pixels", "verbatim ledger vs summary", "judgment as language primitive", "Stagehand extract pick-and-copy", "harness observe-score-act vs demo loop", "public judgment wall / six parallel questions", "meaning-search without embeddings", "attention ≠ correctness", "skills→oxlint / AST prove ∩ remainder", "session-sticky first-prompt routing", "measured RAG rerank vs generative rerank", "capability kernel / secrets never in the agent", "Jev is SENSOR not policy", "type-safe ≠ correct", "typed control plane around DSPy", "native vs verbalized confidence", "engine owns truth / Jev owns judgment", "human-confirmed kill gate", "train specialist vs few-shot hosted", "decide→policy→LLM leftover", "Noul 0.5 cannot-tell never rounded", "calibration ≠ sortable / ORDER BY", "pairwise inversion / Score ordinality / two-decimal ties", "wire-compat GLiFormer /v1/systemone", "class-backend economics", "loopback gateway hosted + local", "do not distill Jev as teacher", "active-learning triage", "evidence-packet explorer", "meaning-grep AND/OR/NOT", "closed-vote-only / no planner LLM", "Jev vs PCD Harbor", "PCD O(1) ≠ Noul", "host-owned handlers × System One", "OMP/pi fail-open gate", "permission vs probability / operator owns thresholds", "judgment ≠ permission / Jev never grants access", "eval integrity / instrument not score", "constrained optimizer + S1 features / never sole hot-path gate", "privilege ≠ verdict / effect contracts not tokens", "attention filter / VOI for human review / never blocks / never green unless sure", "measurement owns endorsement / evidence-gated question packs", "Jev supplies evidence / code owns authority", "ranking ≠ calibration / never hard-threshold raw p as frequency", "hot-click CU / indexed element table", "Jev judges relevance / code decides structure", "local rules first then remainder / never auto-train on own hides", "combinators / System One as control plane", "receipts not leaderboard / type-safe ≠ correct jaggedness", "VOI over skill library / skillranker abstention", "OOD calibration / AUC ≠ ECE", "Jev vs thinking-budget small models", "turnstile / replayable evidence≠authority", "MLX one-pass schema→JSON / Apple Silicon replica economics", "memory leases ended by new evidence", "never confidently wrong / TLA+ compose / escalate instead of hard-gate", "no seal no advance / coverage ledger / mint ≠ product brain", "skill-broker sibling / judgment ≠ permission", "sureness bands / max_prob is generous", "JevBench / calibration not in Main Score", "CI typed gate before expensive review", "Codex MCP host adapter", "judgment as attention redirect / jev-preflight", "compress-before-first-send / dizk jev-lens", "tools≠use / SessionStart over hoping", "observational memory / pi-om keep-kind", "open-Jev class / openvons / JevPick", "physical-world System One / HA-Jev / not for locks", "judgment outside the store / jevql", "landed-script trust / headless≠auto-approve", "digital-design combinators / extended five", "VOI cache admission / same-intent skip LLM", "BM25 vs Jev skill routing Harbor harness", "zeroshot vs BERT / contamination DiD", "typed escalate continue abort baton / inverted loop", "worth-your-attention VOI / ThinkyMiner Winnow", "Jev WHETHER Python HOW LLM WHAT", "conflict vs ignorance / named Choice escape", "Playwright executes Jev chooses", "OpenJev /v1/decide not drop-in", "SemIf wire-compat runoff; SemIf rename densify / MLX backend / 5.21× systems≠semantic / Softmax ≠ Noul (`notes.md` §117)", "decision-as-memory flywheel", "record/replay CI / jevassert", "failure-finding arena / jevarena ≠ jev-arena", "BBQ not a bias cert", "decider≠executor", "sentence-as-rule lint / jevlint", "sentence-as-rule lint / jev-lint is jevlint rename", "VOI hunk prune", "whole-repo intent VERIFIED/VIOLATION/UNKNOWN", "GLiNER2 spec ≠ replica", "open replica substrates / grande / laya-jolt / JEV-CPU", "ONNX local-jev not equivalent", "persist constraints across compaction / pi-heed", "calibration+cost first-class gates", "Harbor-shaped Jev vs SGR LLM-as-judge / jev-judge-bench ≠ jevarena ≠ jevbench", "hand no-text steps / jev-use / Vercel drops confidence", "Pi System-One control plane / pi-jev-control", "never free-generates / jev-gpt tree of Choices", "OpenRouter recipe atlas / samples not benches", "personal history feed / jevfeed / no social graph", "competing NAR claims / dual-channel ECE / openJev-verdict ≠ OpenJev", "empty compaction-proxy skip / IPECTER", "throughput ≠ latency / like-for-like ECE", "1-token logprob endpoint ≠ Noul / coverage ≠ correctness", "open replica engine / jevinf", "unofficial Elixir SDK ≠ OTP peer", "jevex n=16 files-to-read VOI", "commit pre-review attention≠verdict / middle band", "Hermes plugin is Agnes not TypeSafe", "pi-jev-compact ≠ pi-jev-compaction", "empty Codex-proxy skip / IPECTER runway", "decision-native inbox / mailordinal", "unofficial jev-cli not ready / ≠ jevql", "laya-multilingual / English checkpoint confident-wrong OOD", "schema-scorer peaked ranking ≠ calibration", "HF 401 / GitHub 404 Hub-only", "productized System One HTTP / classifier.dev", "escalate-under-threshold / smart tier / multi-label ignores", "silent FALLBACK / granite 0.546 vs advertised 0.800", "vs_jev tracked JSON / read eval/README", "choxos/jev-reviewer ≠ egma-ai / systematic-review pointer", "two-pass Choice+Noul evidence extraction", "not-found is an answer", "human check as productized judgment", "githubnext/localjev ≠ kunchenguid/local-jev", "wire-compat ≠ logit-equiv / prompted JSON ≠ structured read", "self-reported probs / entropy confidence", "GitHub Next local /v1/systemone", "LM Studio runner gap / structured-read primitives", "NandhaKishorM/laya packaging ≠ Hub-only / Router script-before-p", "post-T ECE ≠ raw ECE / Banking77 token-budget", "0.85 still soft / not TypeSafe drop-in", "external census ≠ scored bake-off", "GLiNER2+routers class-boundary", "incomplete openjev census vs watch", "Harbor honesty watch / silent fallback", "JevBench v1.2 geometric mean / cal ON rank / weight sensitivity", "option-order 72→21 / instruction models in the class table", "self-host latency ×2 assumption / est. costs", "Laya absent is a gap not a named exclusion", "Qwen3.8 27B ≠ Archer", "hourly already-folded watch / apply-the-five / skip thin noise", "hard-gate Noul as PR/quality gate is soundness theater", "S1 never stalls waiting / S2 one-use advisory", "Local controller ≠ githubnext/localjev", "purple telemetry = consumed not arrived", "seed = geometry not async replay", "20% starting gate still soft", "no pixels to either provider", "OCR+AX observe-score-act / typesafe-computer-use", "never send screenshot to frontier for the decision", "overlapping CU options = false low confidence", "split kind/item/site", "155× one-screenshot ≠ Harbor taskset", "decision ≠ answer-reader capture", "ASR observe-score-act / jev-voice-browser", "partial-speech VOI / free-text waits", "spoken confirm ≠ hard auth", "numbered overlay without another model", "wrap-as-execution / AgentGhost ALLOW ASK DENY", "rules first then Jev remainder / ASK throws / fail-closed", "reddpy/AgentGhost ≠ jwen5419807/agentghost ≠ vventirozos", "JP genre atlas / studio_yebisu / stars ephemeral ≠ eval", "Jev Clearly Explained / akshay_pachaar / LLM hammer", "schema-safe ≠ correct / 200× 400× TypeSafe ceiling", "questions-as-code / shadow first / not a TypeSafe how-to", "proposition ≠ embedding / contrast-set", "boolean composition of soft Nouls / AND OR NOT", "uehaj/jev-semgrep ≠ semgrep.dev", "meaning-grep dedicated fold / not a gate", "decision-validated UI / Jev never authors text", "decision-as-assert / jevtest ambiguous band", "typed decisions drive UI / jev2ui", "hybrid S1 closed verb menu / anima3", "pointer-not-generator search / JevFind", "jev-frontier-bench ≠ frontier-100", "product bakeoff ≠ architecture duel / GLiClass", "four engines same questions / majority floor", "authorship named escape / not evidence", "ha-switchboard HA remains execution", "n8n classify/route/score / Low Confidence", "fast-jev-compaction-pi ≠ pi-jev-compact ≠ pi-jev-compaction", "jevloop full-distribution optimizer / no LLM in the loop", "laya-vision SmolVLM / score untrained", "Cerebellum-2B /v1/decide ≠ TypeSafe / wire-compat vs agent-routing", "laya-grounded not drop-in / Platt not temperature", "GestaltLabs/Jeff-1 ≠ logan-markewich/jeff / acc vs ECE n=9730", "stanley-code empty findings ≠ approval / human promote", "findme ≠ JevFind / NL memory beam-search FS", "jevsubrouter price workers not conversation / counts ≠ dollars", "feelings .feels() default 0.5 is Noul-0.5-never-rounded / ≠ hunch ≠ Probably", "apa-agent-harness ≠ AntonioCoppe/jev-harness / unpublished npm", "grok-bot-jev skill cannot force a bot that ignores it / A/B proxies not tokens", "Essentiel-Jev never authority / human every action", "enzo-mcp independently falsifiable claims / ≠ jev-sift", "pigeonhole OTHER skip / decision-as-filing", "jev-reliability Nothing about accuracy", "clduab11/jev-test ≠ realZachi/jevtest / Nothing runs yet", "jev-rag-benchmark Jev wins is not an assumption", "dairui1/jev-lab ≠ BrendanH18/jev-lab", "jevmail gmail.readonly / mailjay archive/trash", "ZHUBoer/ego-jev reserved __none__", "runWorkflow completed ≠ success", "jsort scores are relative", "Noul not Choice for scale", "groundedness-judge-bench native vs schema-guided", "implicit_true included in yes", "jev_playground 0 promotions", "routing-backtest 0.0447%", "yuyang2230/jev-agent-skill jev-1.13-free", "jev-techstack-classifier stack_config.json", "s1_ruby collapse late", "undecided? abstain", "2389-research/judgement license null", "confidence ≠ winner p", "typesafeai-sdk-community not a new species", "tpellet/hunch exit 3", "never-execute list", "jev-file-search scores not calibrated accuracy", "jev-linkmap Jev never sees S2 prose", "muhammedilyasy/jev-mail metadata only", "tidy none-of-folders stay", "tab-bouncer pinned/audio/current never closed", "lkclean Show fail-open", "jev-yt-time-saver Show anyway", "ORIGIN pause-if-no-Jev", "validResponse sums-to-1", "jev-crawlers risk bands never raw boolean", "jevbrain AUTO_ACT is not a Noul", "judgekit YAML classify/score/route/verify", "typed-judge-kit verdict-in-code", "alsoleg89/decide packing VOI", "0.8 ≠ 80% accuracy", "Jev-Calibration Platt ECE 0.117→0.052", "jev-calibration-arena never acts", "ctmx/openrouter-jev-mcp Decision-as-Plugin", "FrancoisChastel/jev-code ≠ npm jev-code", "claudecode-jev-marketplace fail-open not hot path", "pedroknigge/mcp_jev packs not ask_jev", "cyrusasco/typesafe-mcp noul deadband 0.35–0.65", "codaaiteam/jev-skill jevtypesafeai.com ≠ TypeSafe", "hermes-switchyard ≠ hermes-jev-router ≠ hermes-plugin-jev", "nanoprune 2.8MB ECE 2.58%", "smartdio/jev-browser-agent ≠ ZHUBoer/ego-jev", "Dakai/omp-jev-web DONE ≠ proof", "hari007sh/jev ≠ dannote/jev", "0thernet/system-one-skills deterministic verify", "typed-gate band [0.40,0.60] is refusal", "pi-jev-gate fail-closed; choice is the verdict", "Foq ~25ms/2.2GB local", "rev prefill-only + HF jev-0.5b", "robfrase/jev planning memo", "typesafe_agent_gates 27/27 / 31/31", "EpicEric/safe-sh static remainder", "pastepilot Confirm before act", "Jev-Reranker live Jev not yet measured", "sessionwise opt-in relevance", "jev-search pointer sieve", "400ms Salesforce WebMCP", "typesafe-scheduler-diagnostics advisory", "droidjev screenshot-free", "Tewoto1 jevcu planner still writes", "ha-conversation-jev Jev→Grok", "dsh-jev can only gate", "jev-classification-benchmark specified not run", "jev-luna-pagerduty p≥0.50", "meldltd/meldecision laya-go ONNX", "laya-doom never pixels", "logixism/laya-api empty README", "akpsahan/laya ≠ Archer", "choxos/jevchess engine owns truth", "jev-drive sim not AV", "story-arc Jev never authors", "jev-hs-assistant HS6", "golergka/jev-plays-starcraft-2 UI-verified ≠ API Victory", "awesome-jev-use-cases catalog", "Nibir1/typesafe-go ≠ official", "fingerprint after redact", "recall vs decide", "publish fingerprints+answers", "CI replay as Harbor cousin", "Cache hit ≠ correctness", "hyperspaceai/jevcache ≠ kushals256/jevcache", "human labels only", "score never auto-accepts", "production capture flywheel", "sutro-sh/jev-align ≠ caiovicentino/jev-align", "guidance ≠ hook", "catalysts ≠ summaries", "compile-time System One", "unofficial ≠ TypeSafe", "format_version modernbert-jev/1", "Argos1111/jev_local ≠ us/jev-local ≠ kunchenguid/local-jev", "LFM default ≠ ModernBERT backend", "Nemotron ≠ TypeSafe Jev", "not a calibrated replacement", "djev-dev complements djev-spark", "images as Choice options", "Laya essay numbers *theirs*", "Router/OOD confidence", "hosted bootstrap ≠ silent TypeSafe", "difficulty + policy thresholds + JSONL trace", "jev-codex-pilot model + reasoning depth", "keep/shadow/hybrid/reject", "quarry evidence projection", "Frank-ZY-Dou/awesome-jev robotics/3D/control", "one-dollar-tahoe TypeSafe Jev defense eval", "jevguard calibrator/cache/escape", "jev-ci-selector CI shadow mode", "llama-jev llama.cpp replica", "petercr/jev-orchestrator ≠ FleeexCorp/jev-orchestrator", "seb4ez/jevguard ≠ AseemPrasad/JevGuard ≠ pablozr/JevGuard", "webNeat/llama-jev ≠ WiktorB2004/llama-index-jev", "OpenCode jev-pruner context sieve", "observe→score-candidates→prune", "jev-zen / jev-1.13-free", "zen-chat ≠ Noul", "fail-open original", "keepScore >0.1 floor", "host port of tamaratran/jev-pruner", "indiejoseph/opencode-jev-pruner ≠ nrdz-labs/fast-jev-opencode", "jev-webagent-bench empty stub", "Kiln-AI/jev_jsonschema noul_threshold 0.5", "NSStudent/JevSwiftSDK unofficial", "GLiNER2 native Apple path", "unofficial Swift/Core ML GLiNER 2.5-small", "entity spans + confidence", "not Choice/Score/Noul", "not TypeSafe", "label descriptions as schema", "on-device ANE economics", "honesty locks", "shershah1024/gliner-native-runtime ≠ Fastino", "≠ gliner25-compaction ≠ gliner2-ultrafast ≠ Eran-BA/Jev_from_GLiNER2 ≠ NSStudent/JevSwiftSDK ≠ jevmlx", "default threshold 0.1 still soft", "soft Noul ≠ hard safety", "Decision Graph Protocol frame→assess→commit", "app retains permissions/effects", "Jev-first assessor-neutral", "guarded commit / receipt/next frame", "assessment batching", "hard-gating DGP as safety theater", "numerous-com/dgp ≠ TypeSafe official", "jegrep calibrated path+range Nouls", "no embeddings/index/daemon", "~$0.01–0.03 typical", "agent --json", "can1357/jegrep ≠ Bentlybro/jevgrep ≠ uehaj/jev-semgrep", "Archer-arch fidelity", "kev family OOD 0.76–0.77 vs Jev 0.86", "block-causal isolation", "pointer/readout CE-trained", "/v1/systemone drop-in", "replica honesty", "cost-sensitive decision theory × System One probabilities → control flow", "thresholds derived from costs not hard-coded", "YES / NO / UNSURE from cost_false_yes / cost_false_no / cost_human", "auto-batching same-object questions", "Kungie/gut ≠ tpellet/hunch ≠ carldaws/hunch", "judgment vs generation", "deterministic execution after probabilistic judgment", "exactly one app-owned callback", "explicit uncertain branch", "Illusion47586/judge ≠ lexingtonhibiki/judgekit ≠ Ascurse/typed-judge-kit", "variable-N option scoring as the trainable object", "dynamic candidate bags not fixed label sets", "zwliJay/jev-forge ≠ NanoJev", "open replica economics / latency vs closed Jev", "NAR local drop-in", "wfzyx/von late-catch HIGH", "competing NAR claims / replica honesty", "typed judgments vs chat judges on guardrailing", "ishaannk/llm-vs-jev cross-note only", "deeper integrity fold is rh-guard", "nothing wins outright", "can be argued out of guarding"", "Jev IS the if-statement", "judgments/probabilities drive branches", "text model only writes prose", "interpreter owns variables/loops/budgets/replay", "otherwise maybe / confidence gate", "chaos samples after the gate", "southpolesteve/probably ≠ carldaws/hunch ≠ feelings ≠ Kungie/gut ≠ Illusion47586/judge ≠ tidymodels/probably", "133★ / forks 10 live", "build calibrated classifiers from human feedback", "retrieve by relevance not resemblance", "one calibrated yes/no per memory in one request", "pointer mode 17/18 19/20 *theirs*", "embedding resemblance misses the allergy", "samdotmak/jev-recall ≠ jev-search ≠ jev-sift ≠ carryforward ≠ chopratejas/invalidate", "memory leases ended by new evidence", "six Nouls then fixed rules in code", "0 of 157 false invalidations", "questions/plans/directives are not evidence", "unsure → review queue", "host keeps the store", "name↔body / comment truth / test-claims", "mizchi/jev-lint is mizchi/jevlint rename", "no shipped rule has severity error", "~1 in 5 findings wrong *theirs*", "mizchi/jev-lint ≠ huntedman/JevLint ≠ MichitoSugawara/jev-lint", "JSON Schema → typed JSON via Jev", "noul_threshold 0.5 decoder not a proof", "IncompatibleSchemaError lists every bad property", "on-device Laya CoreML ANE", "~5 ms P50 short decisions", "189/189 FP16 checkpoint parity", "10× not achieved", "mizorewww/laya-coreml ≠ gliner-native-runtime ≠ jevmlx ≠ NandhaKishorM/laya", "softmax over allowed tokens ≠ Noul", "question-first cache", "Micha0827/snapjudge ≠ githubnext/localjev ≠ jevmlx ≠ cendress/SnapJudge", "Jev-first Pi agent loop", "slow-LLM fallback", "explicit action menu / CandidateSource unimplemented", "62 tests wiring not quality", "direwolfiy/JevPi ≠ standardagents/jevpilot ≠ pi-jev-control", "resume-screening bias audit methodology", "name×resume factorial independent Nouls", "callback determined by resume quality", "mean-probability name gaps operationally negligible", "natemoo-re/bias-bench ≠ BBQ", "Plan/PRD panel → code-owned pass|review|block", "cheerleading out of scope", "austindixson/planalyzer ≠ single-goodness Noul", "cost-aware multi-model routing/escalation", "decide vs do", "successful-task cost", "cannacre8ive/switchboard-ai ≠ ha-switchboard ≠ hermes-switchyard", "frozen-protocol zero-shot bench", "TypeSafe Jev vs PrismNLI vs Laya", "contamination caveat", "elcronos/jev-vs-open-decision-models ≠ JevBench ≠ DMB", "context-window admission control", "VOI gate which tokens are worth the expensive model", "fail polarity per lens", "on small inputs lenses lose money", "cvsgireesh/jevusher ≠ jev-sift ≠ winnow", "typed decision control plane", "receipt ≠ authorization", "historical-v0 zero retained cases", "MokiMeow/jev-fabric ≠ jev-forge ≠ dgp", "live 15-dim typed rubric re-score per pause", "scoring economics exemplar", "OpenJev/Codiv ≠ TypeSafe hosted", "jose-troche/live-rubric ~$0.000004 desc / ~$0.000006 README", "adversarial pre-registered Jev eval", "28 predictions before data", "123,805 requests", "confidence does not track ignorance", "polite injection 65% / crude 0%", "willkelly/jev-evaluation ≠ jevals ≠ jev-baselines-eval", "provider-neutral Elixir/BEAM Noul/Choice/Score SDK", "class infrastructure", "nshkrdotcom/system_one_sdk ≠ typesafe_sdk ≠ dannote/jev", "question-linting of Jev questions themselves", "nine jaggedness rules, no API key, no labelled data", "static lint ≠ measured separation", "yodablocks/jevq ≠ tenbin ≠ JevLint ≠ commitjev", "open-weights Laya as class exemplar (binding)", "Nx/Bumblebee runtime", "host chooses backend", "ChristianAlexander/laya_ex ≠ system_one_sdk ≠ dannote/jev ≠ NandhaKishorM/laya", "on-chain/edge Laya deploy", "parity_verified stays false", "model output never grants Tx", "humandebri/IC-Laya ≠ laya_ex", "auditable weekend replica", "Jev outputs never used for training", "soft human-vote distributions", "unpaired 0.577 vs 0.727", "agilabs-ai/jev48 ≠ JevBench ≠ Mapika/decider", "adversarial dual-judge / framing attack surface", "comparative framing is the usable judgment", "prior injection crowds out evidence", "copyleftdev/ember ≠ ember.js", "Laya specialist fine-tune pipeline", "training still GPU-pending", "PIXELZX0/XERON ≠ convaiinnovations/laya", "Hub Laya replica drop", "daliborsb/laya ≠ convaiinnovations/laya ≠ NandhaKishorM/laya", "System One student distillation corpus", "gold is programmatic", "teacher is closed-API clone", "do not distill Jev as teacher of record", "MagaBitmex/jev-4b-distill-data ≠ missing student checkpoint", "non-LLM VIN System One", "planning depth not chat", "lewislululu/jevon ≠ douglance/jevon", "source-bound evidence checks", "local quote mismatch needs no API", "exit 0 ≠ claim truth", "WaynezProg/jev-kit ≠ jonathanavis96/jev-kit (Airlock) ≠ jev-use ≠ jev-mcp", "independent System One evidence catalog", "scores not one leaderboard", "no external record currently reproduced", "TokenTrim no-Jev matched hybrid 62.4%", "reachjalil/system-one-bench ≠ mallahyari/system-one-benchmark", "21 tasks · 134 items · 208 questions", "scenes from public GitHub contracts, not production logs", "SivletLabs/jev-eval ≠ willkelly/jev-evaluation ≠ 4esv/jev-eval ≠ xxkuboxx/jev-eval ≠ onlyoneaman/jev-eval ≠ dayhaysoos/jevals", "option isolation (sibling-blind)", "permutation-equivariant", "Hub OWNER not published", "nafisazizir/hev ≠ jaredpalmer/kev", "frozen local LLM logits, no trained decision head", "residual-head 9,222-param decreased 73/96→67/96", "confidence = 1−normalized entropy, not P(correct)", "yuki-oshio/mini-jev ≠ r-ms/mini-jev", "Jev classifier as autoregressive next-token predictor", "ChatJev-style soundness theater", "erik-dunteman/ChatJev ≠ dannote/jev ≠ jev-gpt", "calibrated decision head × AlphaProof value head", "implementation-layer isomorphism, semantic difference", "timeout = censoring", "do not launder Noul as proof", "parallel rank-prediction vs serial selection", "independent questions can conflict", "zzzzzec/jevsort ≠ keltokhy/jsort", "curated open System One ecosystem catalog", "rupeshpoojary9/awesome-open-system-one ≠ AnotiaWang/awesome-jev", "arXiv paper radar with Jev relevance scoring", "ranking ≠ calibration / 0.5 still soft", "fail-open failed evals not marked seen", "train calibrated ~27M from scratch", "typed Q→prob dist / one forward pass / no LLM decode", "hyusi2003/MiniSystemOne ≠ Colvin0315/MiniSystemOne", "description-only stub / size 5", "ESCI hard probe fails four of six", "jev_bool ECE 0.242 inversion 0.255", "do not re-fold §60 six-gates as new", "jobbyjev one-request-per-company from batch-size result", "find/design/evaluate TypeSafe Jev decision loops", "karanb192/jev-architect ≠ samtay32/jev-system-architect", "Jairik/jev-distiller size 1", "distill-Jev UI stub / do not distill Jev as teacher of record", "post-launch scored use-case map / Jev self-scores then human curation", "licensedsaucer9-web/jev-opportunities", "Jev-inize a use case into classifier/router", "gavinHuang/jevinize → simple-jev not TypeSafe", "featherless-ai/simple-jev", "compare saved decisions / same label can still change the branch", "VihaanAgarwal/jev-diff ≠ Saik0s/diffusiongemma-jev-macos", "not tested with a live Jev API key", "constrained logprob + temp/Platt ≠ Noul", "OpenJevPro pastes openjev-sglang JevBench as own", "zhangcy122/OpenJevPro ≠ IamBusy/OpenJev ≠ ekzhang/openjev-sglang", "PolyForm Noncommercial", "SmolLM-135M / sub-70ms / 0 output tokens", "demo P(True) 0.5052 / Choice conf 0.2872 / Score conf 0.0055", "README claims MIT / GitHub license null / no LICENSE file", "patelvishwa112/jev-system-one-rlcd ≠ arnabgho/rlcd-lite ≠ blackwood-rlcd", "source-backed Awesome Jev radar / 306+ commit-pinned", "logicrw/awesome-jev-projects ≠ AnotiaWang/awesome-jev ≠ yibie/awesome-jev ≠ cobanov/awesome-jev ≠ rupeshpoojary9/awesome-open-system-one", "auto GitHub sync / Issue-only submissions", "hashed n-gram encoder / rival-aware attention", "olanotolu/jevbetter vs jevlike starter", "synthetic hard menus top-1 0.916 vs 0.873 / ECE 0.0182 vs 0.0367 / 40 vs 4608 menus/sec", "shuffled-context control 0.335", "Turn any open LLM into System-One Jev", "uspraveen/Jevify ≠ Mintzs/jevify ≠ gulagala001/jevify", "Jevify-any-LLM architecture probe", "description-only stub / size 0", "Train encoder-only calibrated decision models from a task sentence", "Exu is a toolkit, not a method", "strictly proper scoring rule", "Pre-alpha", "Ruivalim/exu-base", "scratch-trained calibrated decision model", "typed Q → probability dists", "Colvin0315/MiniSystemOne ≠ hyusi2003/MiniSystemOne", "no published weights download URL", "90.5 seconds / 29.2% pipeline evidence", "p_i/p_j independent of other candidates", "Recipe for calibrated decision models — small model out", "init → synth → train → eval → serve", "91.1 % / ECE 0.022 *theirs*", "Jev zero-shot 75.1", "scienthoon/luce", "Put Jev's three headline claims on trial", "0.5B local GPU", "46x speedup / accuracy identical", "ECE 0.624 sentiment catastrophe", "bigger model worse calibration", "RichardoMrMu/jev-mini ≠ yuki-oshio/mini-jev ≠ r-ms/mini-jev", "System-1 decision engine for local LLMs", "structured choices only", "JSON parse of generated text ≠ Noul", "TypefAI JEV / Journal Entry Voucher", "tapsin/jev-local ≠ us/jev-local ≠ Argos1111/jev_local", "Jev 1.13 reward-model eval across 8 benchmark tracks", "40,940 examples / 0 API errors", "RewardBench v1 92.58%", "Precise IF 50.63%", "goya4140/jev-reward-model-evaluation", "Scaffolding in progress", "Jev vs LLM support-ticket routing", "static + live decision bench", "TypeSafe's own published benchmark", "illustrative simulations, not live API calls", "JevBench v1 — smart/cheap/fast/reliable", "I/C/S/K 25% geometric mean", "classifier.dev fast tier 84.8 is Jev behind its own API", "do not re-fold §78 v1.2 board as new", "Laya (421M) 70.1 now on board", "Zero-shot/few-shot LLM routing", "hard budget filter before Jev", "Jev never asked to perform budget arithmetic", "Jev judges the next state, XState enforces transitions", "simulation uses synthetic keyword fixtures", "catalog gravity", "v-modal/awesome-jev-tools", "★339 live REST", "curation is not endorsement", "crawler-maintained directory", "Daily GitHub + npm sweep, human-merged", "RadRebelSam/awesome-jev ≠ AnotiaWang ≠ yibie ≠ cobanov ≠ logicrw ≠ v-modal", "HF peft SPLADE/BGE reranker", "rdxtremity/jev-reranking ≠ carlaiau/jev-reranking", "query-side encoders, not a Jev replica", "ONNX System One Qwen3.5-4B scorer", "source:pngwn/system-one-qwen3.5-4b-scorer", "CC-BY-NC-4.0", "temperature 1.75", "transformers.js AutoModel cannot load this graph", "Consistency benchmark Space", "This Space contains no benchmark result yet", "12-case plumbing fixture", "Benchmark-driven Jev router and judge", "cheap alone is not success", "Jev does not write, sum prices, or claim accuracy %", "Sol 94.2 / Luna 83.9 / Jev path 89.7", "19.2% Sol / 62.3% cost save / 4.5pp miss of 2pp non-inferiority", "p50 latency worse than Sol due to routing overhead", "erendikmenn/jev-llm-router-benchmark ≠ jev-rag-benchmark ≠ ryantsai/jev-llm-router", "Express + node:sqlite", "mock and Jev decision engines", "previous_ticket_count >= 3 is code", "MIN_CONFIDENCE 0.6 still soft", "substring false positives", "aesaganda/jev-ticket-router ≠ SarathChandraBellam/jev-vs-llm-ticket-router", "Universal Figure & Diagram Router", "confidence ≥ 0.85 hard-gate is theater", "generative AI banned from scientific plots", "six visual branches", "hoangngochuong24947-gif/jev-figure-router", "human-labeled (state, question, label)", "166,054 rows / 22 configs", "soft_label for human uncertainty", "Praveenrajus/jev-bench ≠ fstandhartinger/jevbench", "ternary bonsai System One GGUF", "openjev's mechanism, Bonsai's weights", "Hub does not ship weights", "100/100 easy T/F is not Harbor", "label_mass ≠ correctness", "stock llama.cpp Q2_0 silently gibberish", "NicolaiMTLassen/open-bonzi-jev ≠ NicolaiLassen", "transformers.js DeBERTa ONNX", "source:com-kotobalabs/open-jev-deberta-v3-large", "temperature 1.05", "AutoModel from_pretrained works", "onnx-community/open-jev-deberta-v3-large-ONNX ≠ system-one-qwen3.5-4b-scorer-ONNX", "107★ densify", "GH 151M vs README 149.6M", "PR #1 now closed unmerged", "do not re-fold §71 claim-audit as a beat", "typed decisions, RLCD, confidence-gated routing", "structured ≠ correct", "mock not live API", "26 tests", "wjdjdakf17/jev-study ≠ baekenough/jev-study", "bonzi-27b-v2 / ternary-8b / 27b-v1 GGUF family densify", "WANLI-256 74.6% / 65.2% / 71.1% *theirs*", "Bonsai 1 27B Q1_0 runs on stock llama.cpp", "ternary still needs PrismML fork", "hf:heman10x/openJev-verdict-2.0 twin tokenizer-only", "OpenJev Vision image classification + uncertainty", "CLEVR-4 held-out joint 0%", "hfdataset:IamBusy/OpenJev-Vision-Research-v0.1 12,832", "294,912 derived targets not independent samples", "Laya multilingual ONNX WebGPU typed-decisions port", "63/63 selected answers / 5.1e-4 CPU / 1.2e-2 WebGPU", "UpHash-Network/mini-jev is yuki-oshio transfer", "jev-injection-bench 11,900 labelled prompts", "Jev best ranking / Haiku better ECE 0.021 vs 0.058", "0.5–0.9 band is where Jev's numbers do not mean what they say", "Prompt wording moves panic 28%", "manojlds/jev-dspy-bench ≠ dspachos/jev-dspy ≠ jmanhype/jev-dspy-lab", "Jev agreement is similarity, never ground truth", "no aggregate quality grade or merge gate", "AbstentionBench-on-Jev rank 1 of 20 vs 2025 field", "question-asymmetry", "forward-looking 0.465 never extreme", "openkev calibration layer not a runtime", "ECE vs coverage independent", "select_threshold returns inf", "escalation catches uncertainty not ignorance", "misakaikato/openkev ≠ jaredpalmer/kev", "pdf-race Docling→Jev vs Gemini", "parser owns the wall clock", "12/12 tie is a tie", "titles selected not generated", "flopcheck 16 calibrated tweet judgments", "mechanical tells in code", "ZeroX-01/jev-atlas ≠ Zaious/jev-capability-atlas ≠ gorock007/jev-atlas", "Laya calibration lab Gradio MCP", "T never changes argmax", "confidence ≠ top-label p", "easy probe set refused", "40–48 rows too small to ship T", "Gemma-4 26B-A4B jevify classification+calibration", "LoRA adapter twin not independent eval", "Gemma-4 E4B jevify", "E4B LoRA stub card", "kushalpatil/jevify-gemma4 ≠ Mintzs/jevify ≠ gulagala001/jevify ≠ uspraveen/Jevify", "GH kushalpatil07/jevify 404", "PAWS 0.580/ece 0.288 is the weak cell", "smaller E4B slightly better OOD ECE than 26B-A4B", "Hub jevify merged LoRA ships weights", "bonzi Bonsai-8B v1 GGUF densify", "Bonsai-1.7B v1", "Bonsai-4B v1", "WANLI-256 64.5% / 60.2% / 52.0% *theirs*", "rank #4 / #5 / #6 of 6", "JulesHuisman/jev-eval scaffolding / README SHA c356a584 (was empty e69de29b)", "JulesHuisman/jev-eval ≠ SivletLabs/jev-eval ≠ willkelly/jev-evaluation ≠ 4esv ≠ xxkuboxx ≠ onlyoneaman ≠ dayhaysoos/jevals", "7 bands 6/10 vs 40 bands 0/10", "source receipts + confidence slider re-policy without re-inference", "32/32 synthetic is smoke not production", "classify HF datasets across typed semantic dimensions", "roadus2 watch misspelling; lock roadius2/ultra_laya", "ultra_laya REVIEW defects", "default branch claude/laya-jev-review-gg5ppo", "XNLI EN 88.3% ECE 0.032 → RU 77.3% ECE 0.096", "Δ −11.0 pp [−14.2,−7.8]; ECE +0.063", "MASSIVE no detectable difference at n=600", "confidence is function of p_max (r=1.000)", "pointer-not-generator 400 human-authored responses", "proposed ≠ authorized", "FewRel 160: Jev 85.0% vs lexical 13.125%", "gated 100% (95/95) coverage 59.375%", "J++ composable semantic computation language", "judge-jev 0.5 still soft", "947 repos scored; A 273 / B 302 / C 372", "LLM rubric ≠ benches", "No benchmark winner is claimed", "phishing: naive 62.6% vs regex 91.8%; 5-atomic + LR 95.0% *theirs*", "AITuber tension ±15", "README npm global; repo is Rust", "git-confess code owns counting/blame/ratio", "httpx exhibit 11% (13/119) *theirs*", "90d trend +12.40% vs random +12.75% vs BH +41.71%", "5m win rate 25%", "Awesomejev 656 entries / 38,160 stars", "tracker likes 64 (+4) lastModified UNCHANGED", "Laya present; Blackwood ABSENT; Archer still promised_not_landed", "Blackwood tracker ABSENT; likes 2 gated manual", "r = c - p_a", "ECE 0.021; acc 0.807 vs warmup 0.746", "Independent primitive", "11.57s vs 54.10s · 4.67× · 120/128 *theirs*", "default path is pretrained Gemma probs not trained RLCD head", "GH Meanblock 404; lock leesk212/JEV-CPU", "softmax over letter slots ≠ Noul", "WANLI 0.741 vs openjev v2 0.77 *theirs*", "3-way NLI ≠ Noul", "priority 0.464 = majority floor", "banking77 contaminated", "raw margins not probabilities", "do not distill Jev as teacher of record (they distilled Haiku)", "“0.9 is not one number”", "ranking ≠ calibration", "banking77 0.8–0.9 stated 0.86 actual 0.73 over-confident *theirs*", "≠ Praveenrajus/jev-bench ≠ fstandhartinger/jevbench", "$0.0000153–$0.0000226 vs circulating $0.0004 (~20×)", "Score is 0..n-1 expectation not 0–1", "Noul has no confidence field", "TCP floor 198.8 ms", "type reliability is not a reason to choose Jev (json_schema 5/5)", "gateway tax not one number", "Function-only 5/8 vs hybrid 8/8", "4/8 without Jev", "8 designed cases not conversion lift", "200-row pilot Jev 86.5% 173/200 vs Gemini Flash-Lite 86.0% 172/200 vs Pro 87.0% 174/200 *theirs*", "not a ranking", "情緒測謊器", "8-example Jev vs GPT-5.6 Sol ~64× cost 5.4× latency *theirs*", "synthetic; no inference", "≠ JevBench v1.2 §78", "Judged 3317 / listed 2560", "Jev judges, code applies policy", "APA “microsecond policy / zero hallucination” overclaim", "Client-side quiz; pointer from held docs; scanned-PDF warn", "Jev judges / agent reasons / user decides", "selecting an option is not permission to implement", "pattern exact, judgement must clear floor", "no matching pattern → no model call", "not a correctness oracle", "Spec vs artifact remainder", "treating 0.85 as 85% / minProbability hard-gate as Harbor", "VERIFY acquires discriminating evidence, never same-pool confidence-only rescoring", "fast/full/max are ceilings not sizes", "Solar writes, Jev chooses NEXT ACTION", "do not reopen or amend PR #23 or #24 or #25 or #26 or #27", , "Calibration is not alpha", "NO CURRENT ALPHA CANDIDATE", "ΔR² approximately +0.00084", "Brier 0.2131387", "ECE 0.0421875", "Adding Jev probability to deterministic volatility improved Brier by only 1.4058e-05", "default 0.5 keeps zero non pinned", "keepResult median 0.14 to 0.17", "keepCall median 0.28 to 0.35", "usable range is about 0.10 to 0.25", "7.8% to 57.9%", "judges results it never sees", "task-finish eval not built yet", "$0.002 per compaction", "slavadubrov/sgr-judge-bench ≠ slavadubrov/jev-judge-bench", "Jev 108/120 $0.083 0.34 s", "Luna SGR 114/120", "paired Jev accuracy-difference intervals include zero", "not evidence of equivalence", "GLM SGR 26/120 93 format failures", "Terra-planned Jev hybrid 55/120", "rule-based by default, optionally Jev-backed", "empty README", "missing key cannot break the experience", "prefill plus exactly one decode", "softmax over A/B/C ≠ Noul", "BBQ 9,053/10,000 (90.53%)", "ECE 0.0890", "Mean confidence 0.9943", "overconfident", "score and noul not implemented", "DGUI 12 rows (was 6)", "INSTRUCT 119 rows likes 2", "encode the state once, decide everything in parallel", "0.740 accuracy against a 0.508 majority", "ECE 0.047", "fine-tune's advantage ends where its 384-token training data does", "jasonkneen/open-jev ≠ pngwn/open-jev", "same sha d41dc3cd", "Space does not call Jev", "recomputes routing from saved probabilities", "200-case Jev 97.0% / 100.0% / 95.0% / MAE 9.22", "synthetic repository benchmark", "Jev evaluations are advisory", "YehuiTang0316/jev-nlgrep ≠ Bentlybro/jevgrep ≠ can1357/jegrep ≠ uehaj/jev-semgrep", "default threshold 0.8 still soft", "40-line windows cannot prove whole function", "token-native sequential start/end Choice", "Gemini/Haiku stubs not configured yet", "handful of hand-written examples, not a benchmark", "Jev judged exactly what it was given", "laguagu/jev-skills ≠ laguagu/jev-evidence-lab ≠ Pleo2/awesome-jev-agent-skills", "contract_passed is not a claim of guaranteed factual truth", "Wilson lower bound 0.85 floor", "fixture mode no savings claim", "SemIf 2207★ (+21 vs §110 2186)", "jevlike 1043★ (+5 vs 1038)", "TypeAR 15★ (+1 vs 14)", "AnotiaWang 97★ (+1 vs 96)", "yibie/awesome-jev 506★ (+16 vs 490)", "Laya likes 822 (was 802)", "tracker likes 64 flat, lastModified UNCHANGED", "do not reopen or amend PR #23/#24/#25/#26/#27/#28", "Heman10x-NGU/Verdict-open-jev ≠ Heman10x-NGU/openJev-verdict-2.0", "TF-IDF + LogReg ECE 0.0207 vs Jev 0.1440", "Verdict-open-jev 48.07% vs Jev 90.80%", "abstention combined recall 10.00%", "p50 35.58 ms", "K=25 (maximum capacity) 72.00%", "0.85 coverage 84.60% selective risk 1.18%", "26.1× faster than standard Qwen JSON generation", "Jevify 90.0% / 167 ms CUDA graphs disabled", "Finding 1: Brier on stated confidence alone is a trap", "grpo_rlcr 0.78 / ECE 0.084", "reliability 0.007 but resolution 0.000", "27 900 schema-driven decisions", "13 600 / 13 600 questions", "candidate mass min 0.99999624", "22 configs · 166,054 rows · 4 calibration-gold", "sha a39eba3f", "Student B MAE 0.148 / Pearson 0.836 / 86.0%", "pngwn/open-jev-laya-bench README 404", "sha 9f69c742 likes 2", "HDFS 0.9933 (745/750) / retain 0.0084", "BGL ERROR/FATAL protection 1.0000", "2,479 / 2,500 HDFS uncertain", "cache hit 0.9648 (2412/2500)", "$0.153936 estimated", "E2 recomputes from saved probabilities", "Space sha eda59e0a", "MASSIVE English 0.783 / Khmer 0.033 / Hindi 0.133", "40–48 rows too small to ship T", "T never changes argmax", "siren2345/jev-single-decode-transformers ≠ siren2345/jev-single-decode", "Split Transformers experiment from llama.cpp runtime", "tanayvasishtha/jev-lab ≠ dairui1/jev-lab ≠ BrendanH18/jev-lab ≠ yibie/laya-jev-lab", "Four experiments stress-testing TypeSafe's Jev: calibration, bundle bias, label bias, and ensembling", "second pass must be $0.00 from cache", "The pages never call Jev", "Gemma 4 31B 77.0% / Jev 1.13.0 61.4% / Laya 322M 0.0%", "restriction state 95.0% against 84.4%", "None of the systems are particularly good at knowing when to stop and ask", "They skip the question and call a tool directly", "100% schema pass", "six-field joint 48.8% vs 72.8%", "ywchiu/jev_benchmark ≠ Running-Dolphins/jev-bench ≠ Praveenrajus/jev-bench", "ACT / REVIEW / FALLBACK", "A provider failure, timeout, malformed output, or missing answer is **not** a policy outcome", "confidence is descriptive provider output, not a substitute for probability", "Quality denominators include only valid scored answers", "an exact halfway tie chooses the lower level", "aiwithenoch/Jev-Skill ≠ simplosophy/jev-skill ≠ laguagu/jev-skills", "The local path does not claim to turn a smaller checkpoint into Jev", "Low support becomes decision: \"review\"", "MIT-0 SPDX NOASSERTION", "current-llm", "结构兼容,不是 Jev 模型能力", "altryne/jevify ≠ Mintzs/jevify ≠ gulagala001/jevify ≠ uspraveen/Jevify", "Find where Jev belongs. Design the questions. Measure the difference", "TypeAR-AI/TypeAR 301 → TypeLLM/TypeLLM", "TypeLLM/TypeLLM 16★", "SemIf 2241★ (+34 vs §111 2207)", "jevlike 1051★ (+8 vs 1043)", "AnotiaWang 98★ (+1 vs 97)", "yibie/awesome-jev 525★ (+19 vs 506)", "Laya likes 864 (was 822)", "tracker likes 67 (+3 vs 64)", "lastModified UNCHANGED `2026-09-20T04:29:16.000Z`", "do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#32", "hysteresis enter/exit / replay policy without inference", "calibration does not compose / hop-ECE permutation-invariant", "equal-width vs quantile ECE / ranking ≠ calibration", "Qwen2.5 ≠ Archer / Qwen 3.8 sparring ≠ Archer / Qwen/Qwen3.8-27B ≠ Archer", "Deferred Crispification / TCE / AMS", "g0runmezadam/what-is-jev IS tunahansahin897/what-is-jev", "pd.cut equal-width vs jeval quantile", "A hunch is a probability with a policy attached", "soundness theater / measurement theater / hourly 0843", , "Jev Capability Resolver / NiazMorshed2007/jcr", "one tool nested capability tree / returns context / does not execute", "skills vs capabilities / workflow+judgment vs operations", "format independent of Jev / proposed open standard", "JCR_BAND_RATIO 0.6 is application policy / soft scores ≠ hard gates", "routing ≠ permission / docs ≠ authority to run", "sol-vs-opus5-20 lookup+explain / n=1 / Not Harbor task-execution", "wall-time mixed / Sol slower with JCR in 19/20", "NiazMorshed2007/jcr ≠ skill-broker ≠ skillranker ≠ jev-sift ≠ jev-lens ≠ jevusher ≠ jev_select_capability", "do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33/#34", "notes.md §116", "copy the SemIf/MLX installer?", "quote 5.21× as beating Jev?", "treat 0.845 as a TypeSafe replica?", "collapse SemIf into kw2828/zhihz/semif-rs/semif-serve", "softmax over options as a Noul", or "cascade sign-flip / calibration theater": read `references/faq.md`, then `references/mental-models.md`, then `references/mixed-architecture.md`, then `references/judgment-class.md` before any mapping. Proof, @@ -263,3 +263,4 @@ User-provided 0915 uniqueness lock: TianyuCodings/NanoJev unified-games-v1 densi User-provided 0920 jcr uniqueness lock: NiazMorshed2007/jcr MIT; site https://jcr.niazmorshed.dev; topics ai-agents,jev,mcp; **4★**; HEAD `138b3832`; README SHA `2a49dbc1`; LICENSE SHA `46231303`; size **14850**; Jev Capability Resolver; one tool to find documented deterministic commands in a nested capability tree; returns context; **does not execute**; skills = workflow+judgment; capabilities = individual operations; format independent of Jev; proposed open standard exploration; classify (Jev) → optional OpenAI decompose compound → beam search geometric mean of routing probs; keep up to 3 paths ≥60% of best (JCR_BAND_RATIO 0.6); ambiguity / no-match / depth-limit explicit; soft scores ≠ hard gates; 0.6 band is application policy; routing ≠ permission; docs ≠ authority to run; sol-vs-opus5-20 *theirs*: 20 scenarios × 4 variants = 80 runs; lookup+explain only, no execution; Claude Opus 5: agent input 108,585→15,819 (−85%), cost $0.3700→$0.1222 (−67%), wall 105.5s→77.7s; Codex GPT-5.6-Sol: 61,952→47,669 (−23%), $0.1377→$0.1151 (−16%), wall 25.3s→62.4s (Sol slower with JCR in 19/20); One Sol outlier 372.6s / 193 Jev calls; n=1 per cell; Not Harbor task-execution; Claude/Codex harnesses; compare mode; 50 scenarios bundled; 11 groups, 960 nodes, 11,360 items; 16 routing rounds per step; NiazMorshed2007/jcr ≠ skill-broker ≠ skillranker ≠ jev-sift ≠ jev-lens ≠ jevusher ≠ jev_select_capability; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33/#34; notes.md §116 +User-provided 0922 uniqueness lock: SemIf was formerly OpenJev; independent; not affiliated with Jev or TypeSafe; homepage openjev.com; default master; MIT; HEAD ca3ba65f1429; Tolerate float roundoff in MLX evidence verification; pushed 2026-09-19; live REST 2282★ / 140 forks; size 9177; README SHA 74ab7f7f; LICENSE SHA ca562883; interface pattern reproduction with open models; does not reproduce Jev undisclosed model/training; Direct option logits; 0 output tokens; shared-state parallel; MLX backend for Apple Silicon (`--backend mlx`); Qwen3.5-4B 3090 direct 1.023s vs AR JSON 5.332s (**5.21×**); argmax agree 18/21; systems comparison ≠ semantic equivalence; Parallel suffixes 20.03 dec/s on 777 decisions; Browser ladder Qwen3.5-4B authored BA 0.813, pert 0.766, TypeSafe subset agreement 0.845 vs Published Jev 0.883 (102 across 20 cases); Softmax over options ≠ calibrated Noul; typed output does not guarantee semantic correctness; wire/agreement ≠ replica of TypeSafe; SemIf ≠ kw2828/OpenJev playground ≠ zhihz/openjev ≠ apiplant/semif-rs port ≠ dddanielliu/semif-serve; rename is densify not a second census; JevBench 74.6 is §78 not this ladder; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33/#34/#35/#36/#38; do not push onto open #39/#40; notes.md §117 diff --git a/.agents/skills/augustus/references/agent-self-assessment.md b/.agents/skills/augustus/references/agent-self-assessment.md index 87ce5c5..8b2ef20 100644 --- a/.agents/skills/augustus/references/agent-self-assessment.md +++ b/.agents/skills/augustus/references/agent-self-assessment.md @@ -936,3 +936,6 @@ User-provided 0915 uniqueness lock: TianyuCodings/NanoJev unified-games-v1 densi **0920 jcr:** pre-action lookup is not a done-check and not a permission gate. Ambiguity / no-match / depth-limit are explicit abstention paths. NiazMorshed2007/jcr. `notes.md` §116. User-provided 0920 jcr uniqueness lock: NiazMorshed2007/jcr MIT; site https://jcr.niazmorshed.dev; topics ai-agents,jev,mcp; **4★**; HEAD `138b3832`; README SHA `2a49dbc1`; LICENSE SHA `46231303`; size **14850**; Jev Capability Resolver; one tool to find documented deterministic commands in a nested capability tree; returns context; **does not execute**; skills = workflow+judgment; capabilities = individual operations; format independent of Jev; proposed open standard exploration; classify (Jev) → optional OpenAI decompose compound → beam search geometric mean of routing probs; keep up to 3 paths ≥60% of best (JCR_BAND_RATIO 0.6); ambiguity / no-match / depth-limit explicit; soft scores ≠ hard gates; 0.6 band is application policy; routing ≠ permission; docs ≠ authority to run; sol-vs-opus5-20 *theirs*: 20 scenarios × 4 variants = 80 runs; lookup+explain only, no execution; Claude Opus 5: agent input 108,585→15,819 (−85%), cost $0.3700→$0.1222 (−67%), wall 105.5s→77.7s; Codex GPT-5.6-Sol: 61,952→47,669 (−23%), $0.1377→$0.1151 (−16%), wall 25.3s→62.4s (Sol slower with JCR in 19/20); One Sol outlier 372.6s / 193 Jev calls; n=1 per cell; Not Harbor task-execution; Claude/Codex harnesses; compare mode; 50 scenarios bundled; 11 groups, 960 nodes, 11,360 items; 16 routing rounds per step; NiazMorshed2007/jcr ≠ skill-broker ≠ skillranker ≠ jev-sift ≠ jev-lens ≠ jevusher ≠ jev_select_capability; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33/#34; notes.md §116 + +User-provided 0922 (`notes.md` §117 / items 330–336 / batch #100). SemIf was formerly OpenJev; MLX backend; 5.21× systems≠semantic; Softmax over options ≠ calibrated Noul; live REST 2282★. Do not reopen #23–#36/#38; do not push onto open #39/#40. +User-provided 0922 uniqueness lock: SemIf was formerly OpenJev; independent; not affiliated with Jev or TypeSafe; homepage openjev.com; default master; MIT; HEAD ca3ba65f1429; Tolerate float roundoff in MLX evidence verification; pushed 2026-09-19; live REST 2282★ / 140 forks; size 9177; README SHA 74ab7f7f; LICENSE SHA ca562883; interface pattern reproduction with open models; does not reproduce Jev undisclosed model/training; Direct option logits; 0 output tokens; shared-state parallel; MLX backend for Apple Silicon (`--backend mlx`); Qwen3.5-4B 3090 direct 1.023s vs AR JSON 5.332s (**5.21×**); argmax agree 18/21; systems comparison ≠ semantic equivalence; Parallel suffixes 20.03 dec/s on 777 decisions; Browser ladder Qwen3.5-4B authored BA 0.813, pert 0.766, TypeSafe subset agreement 0.845 vs Published Jev 0.883 (102 across 20 cases); Softmax over options ≠ calibrated Noul; typed output does not guarantee semantic correctness; wire/agreement ≠ replica of TypeSafe; SemIf ≠ kw2828/OpenJev playground ≠ zhihz/openjev ≠ apiplant/semif-rs port ≠ dddanielliu/semif-serve; rename is densify not a second census; JevBench 74.6 is §78 not this ladder; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33/#34/#35/#36/#38; do not push onto open #39/#40; notes.md §117 diff --git a/.agents/skills/augustus/references/applied-mappings.md b/.agents/skills/augustus/references/applied-mappings.md index 70688ca..e042d9d 100644 --- a/.agents/skills/augustus/references/applied-mappings.md +++ b/.agents/skills/augustus/references/applied-mappings.md @@ -2505,3 +2505,6 @@ User-provided 0915 uniqueness lock: TianyuCodings/NanoJev unified-games-v1 densi User-provided 0920 jcr uniqueness lock: NiazMorshed2007/jcr MIT; site https://jcr.niazmorshed.dev; topics ai-agents,jev,mcp; **4★**; HEAD `138b3832`; README SHA `2a49dbc1`; LICENSE SHA `46231303`; size **14850**; Jev Capability Resolver; one tool to find documented deterministic commands in a nested capability tree; returns context; **does not execute**; skills = workflow+judgment; capabilities = individual operations; format independent of Jev; proposed open standard exploration; classify (Jev) → optional OpenAI decompose compound → beam search geometric mean of routing probs; keep up to 3 paths ≥60% of best (JCR_BAND_RATIO 0.6); ambiguity / no-match / depth-limit explicit; soft scores ≠ hard gates; 0.6 band is application policy; routing ≠ permission; docs ≠ authority to run; sol-vs-opus5-20 *theirs*: 20 scenarios × 4 variants = 80 runs; lookup+explain only, no execution; Claude Opus 5: agent input 108,585→15,819 (−85%), cost $0.3700→$0.1222 (−67%), wall 105.5s→77.7s; Codex GPT-5.6-Sol: 61,952→47,669 (−23%), $0.1377→$0.1151 (−16%), wall 25.3s→62.4s (Sol slower with JCR in 19/20); One Sol outlier 372.6s / 193 Jev calls; n=1 per cell; Not Harbor task-execution; Claude/Codex harnesses; compare mode; 50 scenarios bundled; 11 groups, 960 nodes, 11,360 items; 16 routing rounds per step; NiazMorshed2007/jcr ≠ skill-broker ≠ skillranker ≠ jev-sift ≠ jev-lens ≠ jevusher ≠ jev_select_capability; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33/#34; notes.md §116 + +User-provided 0922 (`notes.md` §117 / items 330–336 / batch #100). SemIf was formerly OpenJev; MLX backend; 5.21× systems≠semantic; Softmax over options ≠ calibrated Noul; live REST 2282★. Do not reopen #23–#36/#38; do not push onto open #39/#40. +User-provided 0922 uniqueness lock: SemIf was formerly OpenJev; independent; not affiliated with Jev or TypeSafe; homepage openjev.com; default master; MIT; HEAD ca3ba65f1429; Tolerate float roundoff in MLX evidence verification; pushed 2026-09-19; live REST 2282★ / 140 forks; size 9177; README SHA 74ab7f7f; LICENSE SHA ca562883; interface pattern reproduction with open models; does not reproduce Jev undisclosed model/training; Direct option logits; 0 output tokens; shared-state parallel; MLX backend for Apple Silicon (`--backend mlx`); Qwen3.5-4B 3090 direct 1.023s vs AR JSON 5.332s (**5.21×**); argmax agree 18/21; systems comparison ≠ semantic equivalence; Parallel suffixes 20.03 dec/s on 777 decisions; Browser ladder Qwen3.5-4B authored BA 0.813, pert 0.766, TypeSafe subset agreement 0.845 vs Published Jev 0.883 (102 across 20 cases); Softmax over options ≠ calibrated Noul; typed output does not guarantee semantic correctness; wire/agreement ≠ replica of TypeSafe; SemIf ≠ kw2828/OpenJev playground ≠ zhihz/openjev ≠ apiplant/semif-rs port ≠ dddanielliu/semif-serve; rename is densify not a second census; JevBench 74.6 is §78 not this ladder; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33/#34/#35/#36/#38; do not push onto open #39/#40; notes.md §117 diff --git a/.agents/skills/augustus/references/composition-algebra.md b/.agents/skills/augustus/references/composition-algebra.md index e12327b..2c58ea4 100644 --- a/.agents/skills/augustus/references/composition-algebra.md +++ b/.agents/skills/augustus/references/composition-algebra.md @@ -2088,6 +2088,48 @@ User-provided 0915 uniqueness lock: TianyuCodings/NanoJev unified-games-v1 densi positions 8 (Metric) × 11 (Bounds). Claude/Codex harnesses. compare mode. 50 scenarios bundled. Full cards: `validation.md`. + +330. **SemIf rename densify** (TheoLeeCJ/SemIf): + position 1 (Operand). SemIf was formerly OpenJev. rename is densify not a second census. + independent; not affiliated with Jev or TypeSafe. homepage openjev.com. default master. MIT. + live REST 2282★ / 140 forks. HEAD ca3ba65f1429. + Full cards: `judgment-class.md`, `faq.md`, `mental-models.md`. +331. **Interface pattern ≠ replica** (TheoLeeCJ/SemIf): + position 1 (Operand). interface pattern reproduction with open models; + does not reproduce Jev undisclosed model/training. + wire/agreement ≠ replica of TypeSafe. + Full cards: `judgment-class.md`, `faq.md`. +332. **Direct logits / 0 tokens / shared-state** (TheoLeeCJ/SemIf): + positions 1 (Operand) × 8 (Metric). Direct option logits; 0 output tokens; shared-state parallel. + Full cards: `mixed-architecture.md`, `judgment-class.md`. +333. **MLX Apple Silicon densify** (TheoLeeCJ/SemIf): + position 1 (Operand). MLX backend for Apple Silicon (`--backend mlx`). + Tolerate float roundoff in MLX evidence verification. + typed output does not guarantee semantic correctness. + Full cards: `judgment-class.md`, `faq.md`. +334. **Speed 5.21× systems comparison** (TheoLeeCJ/SemIf): + position 8 (Metric). Qwen3.5-4B 3090 direct 1.023s vs AR JSON 5.332s (**5.21×**); + argmax agree 18/21; systems comparison ≠ semantic equivalence. + Full cards: `validation.md`, `faq.md`. +335. **Parallel suffixes 20.03 densify** (TheoLeeCJ/SemIf): + position 8 (Metric). Parallel suffixes 20.03 dec/s on 777 decisions. + Full cards: `validation.md`. +336. **Browser ladder + cousins** (TheoLeeCJ/SemIf): + positions 8 (Metric) × 11 (Bounds). authored BA 0.813, pert 0.766, + TypeSafe subset agreement 0.845 vs Published Jev 0.883 (102 across 20 cases). + Softmax over options ≠ calibrated Noul. JevBench 74.6 is §78 not this ladder. + SemIf ≠ kw2828/OpenJev playground ≠ zhihz/openjev ≠ apiplant/semif-rs port ≠ dddanielliu/semif-serve. + Full cards: `validation.md`, `faq.md`, `mental-models.md`. + +User-provided 0922 items 330–336 (`notes.md` §117). Do **not** +re-fold merged #36 §115 items 303–308 / §114 items 289–302 / §113 items 273–288 +/ merged #38 §116 items 309–316. Skip open #40 items 322–329. +Soft Noul ≠ hard safety. Ranking ≠ calibration theater. +SemIf 2282★ (+41 vs §113 2241; +45 vs §114 2237); +do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33/#34/#35/#36/#38. +do not push onto open #39/#40. + + User-provided 0920 jcr items 309–316 (`notes.md` §116). Do **not** re-fold 0743 items 273–288 / merged #30 items 268–272 / 0646 items 248–267. Merged #35 owns §114 / items 289–302 / batch #97. Merged #36 owns §115 / 303–308 / #98. Open #37 owns §117 / 315–321 / #100 (item overlap 315–316 is #37's remap). @@ -2192,3 +2234,4 @@ Soft Noul ≠ hard safety. Hourly 0843 uniqueness lock: A hunch is a probability with a policy attached; { enter: 0.8, exit: 0.6 } is hysteresis; replay a policy change without inference; Decision models are providers, not the product; huncho ≠ Kungie/gut ≠ carldaws/hunch ≠ tpellet/hunch; pretrained Qwen2.5 base ECE 0.030 (0.5B) / 0.040 (7B); instruct 0.302 / 0.269; 70.9% → 70.0% mean conf 74.1% → 96.7%; temperature scaling still matches it in-distribution; No Jev API was called; Qwen2.5 ≠ Archer; Qwen/Qwen3.8-27B ≠ Archer; 学習済みモデル v0.1 は準備中です; bool AUROC 0.523; 先頭だと0件、末尾だと250件; 温度を渡さない場合、確率は較正されていません; このリポジトリには Jev を呼ぶコードが存在しません; g0runmezadam/what-is-jev IS tunahansahin897/what-is-jev (same GitHub id 1378007307); 947 repos scored; A 273 · B 302 · C 372; LLM rubric ≠ benches; Data as of 2026-09-20; HEAD 895b9498; README SHA 3ae98c56; 13 focused checks and one mutually exclusive outcome; Probabilities are advisory, not calibrated guarantees; omni-/ask-jev ≠ pedroknigge/mcp_jev; pd.cut bins by equal width while jeval bins by quantile; ECE 0.113 and ECE 0.076; jeval drift is not implemented yet; rlaope/jeval ≠ dayhaysoos/jevals; calibration does not compose; ECE has exactly zero statistical power to detect the failure mode that kills trajectories; 25–60× headline withdrawn; P(all-correct): 0.0071 vs 0.0001; TCE / AMS; Qwen 3.8 sparring ≠ Archer; Deferred Crispification; light_cutoff_applied_to_combination 0; recorded run, kinematic animation; BANKING77 Accuracy BERT-Base 93.02 Jev 79.90; Analyse jev calibration (NLL, ECE) backlog; BERT figures are published supervised references, not zero-shot; 档位措辞效应 分数极差中位 0.50、最大 1.32; 修好后对照组是 0.01; 不是 benchmark; 概率没做 calibration; ~1,430 API calls, about $0.15; xiaohuaxi/jev-study ≠ wjdjdakf17/jev-study ≠ baekenough/jev-study; AND: product (independence assumed and recorded in the trace); chat model's stated confidence is not calibrated; circuit-vl-4b ≠ Archer; Bring your own API key; vamsikrishna2421/jev-usecases ≠ whyashthakker/awesome-jev-use-cases; catalog ≠ endorsement; SemIf 2237★ (+30 vs §111 2207); jevlike 1049★ (+6 vs 1043); TypeLLM/TypeLLM 16★; AnotiaWang 98★; yibie/awesome-jev 520★ (+14 vs 506); Laya likes 861 (was 822); tracker likes 66 (+2 vs 64) lastModified UNCHANGED; Blackwood likes 2 gated manual; Archer still promised_not_landed; Hub archerhume/4rcherhume HTTP 401; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33; notes.md §114 User-provided 0920 jcr uniqueness lock: NiazMorshed2007/jcr MIT; site https://jcr.niazmorshed.dev; topics ai-agents,jev,mcp; **4★**; HEAD `138b3832`; README SHA `2a49dbc1`; LICENSE SHA `46231303`; size **14850**; Jev Capability Resolver; one tool to find documented deterministic commands in a nested capability tree; returns context; **does not execute**; skills = workflow+judgment; capabilities = individual operations; format independent of Jev; proposed open standard exploration; classify (Jev) → optional OpenAI decompose compound → beam search geometric mean of routing probs; keep up to 3 paths ≥60% of best (JCR_BAND_RATIO 0.6); ambiguity / no-match / depth-limit explicit; soft scores ≠ hard gates; 0.6 band is application policy; routing ≠ permission; docs ≠ authority to run; sol-vs-opus5-20 *theirs*: 20 scenarios × 4 variants = 80 runs; lookup+explain only, no execution; Claude Opus 5: agent input 108,585→15,819 (−85%), cost $0.3700→$0.1222 (−67%), wall 105.5s→77.7s; Codex GPT-5.6-Sol: 61,952→47,669 (−23%), $0.1377→$0.1151 (−16%), wall 25.3s→62.4s (Sol slower with JCR in 19/20); One Sol outlier 372.6s / 193 Jev calls; n=1 per cell; Not Harbor task-execution; Claude/Codex harnesses; compare mode; 50 scenarios bundled; 11 groups, 960 nodes, 11,360 items; 16 routing rounds per step; NiazMorshed2007/jcr ≠ skill-broker ≠ skillranker ≠ jev-sift ≠ jev-lens ≠ jevusher ≠ jev_select_capability; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33/#34; notes.md §116 +User-provided 0922 uniqueness lock: SemIf was formerly OpenJev; independent; not affiliated with Jev or TypeSafe; homepage openjev.com; default master; MIT; HEAD ca3ba65f1429; Tolerate float roundoff in MLX evidence verification; pushed 2026-09-19; live REST 2282★ / 140 forks; size 9177; README SHA 74ab7f7f; LICENSE SHA ca562883; interface pattern reproduction with open models; does not reproduce Jev undisclosed model/training; Direct option logits; 0 output tokens; shared-state parallel; MLX backend for Apple Silicon (`--backend mlx`); Qwen3.5-4B 3090 direct 1.023s vs AR JSON 5.332s (**5.21×**); argmax agree 18/21; systems comparison ≠ semantic equivalence; Parallel suffixes 20.03 dec/s on 777 decisions; Browser ladder Qwen3.5-4B authored BA 0.813, pert 0.766, TypeSafe subset agreement 0.845 vs Published Jev 0.883 (102 across 20 cases); Softmax over options ≠ calibrated Noul; typed output does not guarantee semantic correctness; wire/agreement ≠ replica of TypeSafe; SemIf ≠ kw2828/OpenJev playground ≠ zhihz/openjev ≠ apiplant/semif-rs port ≠ dddanielliu/semif-serve; rename is densify not a second census; JevBench 74.6 is §78 not this ladder; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33/#34/#35/#36/#38; do not push onto open #39/#40; notes.md §117 diff --git a/.agents/skills/augustus/references/faq.md b/.agents/skills/augustus/references/faq.md index 7e2f4b8..c7fd7d0 100644 --- a/.agents/skills/augustus/references/faq.md +++ b/.agents/skills/augustus/references/faq.md @@ -3622,4 +3622,27 @@ Hourly 0646 uniqueness lock: Calibration is not alpha; NO CURRENT ALPHA CANDIDAT Hourly 0843 uniqueness lock: A hunch is a probability with a policy attached; { enter: 0.8, exit: 0.6 } is hysteresis; replay a policy change without inference; Decision models are providers, not the product; huncho ≠ Kungie/gut ≠ carldaws/hunch ≠ tpellet/hunch; pretrained Qwen2.5 base ECE 0.030 (0.5B) / 0.040 (7B); instruct 0.302 / 0.269; 70.9% → 70.0% mean conf 74.1% → 96.7%; temperature scaling still matches it in-distribution; No Jev API was called; Qwen2.5 ≠ Archer; Qwen/Qwen3.8-27B ≠ Archer; 学習済みモデル v0.1 は準備中です; bool AUROC 0.523; 先頭だと0件、末尾だと250件; 温度を渡さない場合、確率は較正されていません; このリポジトリには Jev を呼ぶコードが存在しません; g0runmezadam/what-is-jev IS tunahansahin897/what-is-jev (same GitHub id 1378007307); 947 repos scored; A 273 · B 302 · C 372; LLM rubric ≠ benches; Data as of 2026-09-20; HEAD 895b9498; README SHA 3ae98c56; 13 focused checks and one mutually exclusive outcome; Probabilities are advisory, not calibrated guarantees; omni-/ask-jev ≠ pedroknigge/mcp_jev; pd.cut bins by equal width while jeval bins by quantile; ECE 0.113 and ECE 0.076; jeval drift is not implemented yet; rlaope/jeval ≠ dayhaysoos/jevals; calibration does not compose; ECE has exactly zero statistical power to detect the failure mode that kills trajectories; 25–60× headline withdrawn; P(all-correct): 0.0071 vs 0.0001; TCE / AMS; Qwen 3.8 sparring ≠ Archer; Deferred Crispification; light_cutoff_applied_to_combination 0; recorded run, kinematic animation; BANKING77 Accuracy BERT-Base 93.02 Jev 79.90; Analyse jev calibration (NLL, ECE) backlog; BERT figures are published supervised references, not zero-shot; 档位措辞效应 分数极差中位 0.50、最大 1.32; 修好后对照组是 0.01; 不是 benchmark; 概率没做 calibration; ~1,430 API calls, about $0.15; xiaohuaxi/jev-study ≠ wjdjdakf17/jev-study ≠ baekenough/jev-study; AND: product (independence assumed and recorded in the trace); chat model's stated confidence is not calibrated; circuit-vl-4b ≠ Archer; Bring your own API key; vamsikrishna2421/jev-usecases ≠ whyashthakker/awesome-jev-use-cases; catalog ≠ endorsement; SemIf 2237★ (+30 vs §111 2207); jevlike 1049★ (+6 vs 1043); TypeLLM/TypeLLM 16★; AnotiaWang 98★; yibie/awesome-jev 520★ (+14 vs 506); Laya likes 861 (was 822); tracker likes 66 (+2 vs 64) lastModified UNCHANGED; Blackwood likes 2 gated manual; Archer still promised_not_landed; Hub archerhume/4rcherhume HTTP 401; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33; notes.md §114 + +## Collapse SemIf into a new OpenJev census? Is the rename a second species? + +No. SemIf was formerly OpenJev. rename is densify not a second census. independent; not affiliated with Jev or TypeSafe. homepage openjev.com. default master. MIT. live REST 2282★ / 140 forks. HEAD ca3ba65f1429. `notes.md` §117. + +## Copy the SemIf / MLX installer? Treat `--backend mlx` as a new class? + +No. Docs/skill only. MLX backend for Apple Silicon (`--backend mlx`) is a backend, not a species. Tolerate float roundoff in MLX evidence verification. typed output does not guarantee semantic correctness. Do not copy `pip` / `venv`. `notes.md` §117. + +## Quote 5.21× as beating Jev? Treat 18/21 as semantic equivalence? + +No. Qwen3.5-4B 3090 direct 1.023s vs AR JSON 5.332s (**5.21×**). argmax agree 18/21. systems comparison ≠ semantic equivalence. Parallel suffixes 20.03 dec/s on 777 decisions. Direct option logits; 0 output tokens; shared-state parallel. `notes.md` §117. + +## Treat 0.813 / 0.845 as Harbor? Merge the ladder with JevBench 74.6? + +No. authored BA 0.813, pert 0.766, TypeSafe subset agreement 0.845 vs Published Jev 0.883 (102 across 20 cases). Softmax over options ≠ calibrated Noul. wire/agreement ≠ replica of TypeSafe. JevBench 74.6 is §78 not this ladder. `notes.md` §117. + +## Collapse SemIf into kw2828 / zhihz / semif-rs / semif-serve? Reopen #23–#38? Push onto #39/#40? + +No. SemIf ≠ kw2828/OpenJev playground ≠ zhihz/openjev ≠ apiplant/semif-rs port ≠ dddanielliu/semif-serve. interface pattern reproduction with open models; does not reproduce Jev undisclosed model/training. do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33/#34/#35/#36/#38. do not push onto open #39/#40. `invented_signal: false`. `notes.md` §117. + + User-provided 0920 jcr uniqueness lock: NiazMorshed2007/jcr MIT; site https://jcr.niazmorshed.dev; topics ai-agents,jev,mcp; **4★**; HEAD `138b3832`; README SHA `2a49dbc1`; LICENSE SHA `46231303`; size **14850**; Jev Capability Resolver; one tool to find documented deterministic commands in a nested capability tree; returns context; **does not execute**; skills = workflow+judgment; capabilities = individual operations; format independent of Jev; proposed open standard exploration; classify (Jev) → optional OpenAI decompose compound → beam search geometric mean of routing probs; keep up to 3 paths ≥60% of best (JCR_BAND_RATIO 0.6); ambiguity / no-match / depth-limit explicit; soft scores ≠ hard gates; 0.6 band is application policy; routing ≠ permission; docs ≠ authority to run; sol-vs-opus5-20 *theirs*: 20 scenarios × 4 variants = 80 runs; lookup+explain only, no execution; Claude Opus 5: agent input 108,585→15,819 (−85%), cost $0.3700→$0.1222 (−67%), wall 105.5s→77.7s; Codex GPT-5.6-Sol: 61,952→47,669 (−23%), $0.1377→$0.1151 (−16%), wall 25.3s→62.4s (Sol slower with JCR in 19/20); One Sol outlier 372.6s / 193 Jev calls; n=1 per cell; Not Harbor task-execution; Claude/Codex harnesses; compare mode; 50 scenarios bundled; 11 groups, 960 nodes, 11,360 items; 16 routing rounds per step; NiazMorshed2007/jcr ≠ skill-broker ≠ skillranker ≠ jev-sift ≠ jev-lens ≠ jevusher ≠ jev_select_capability; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33/#34; notes.md §116 +User-provided 0922 uniqueness lock: SemIf was formerly OpenJev; independent; not affiliated with Jev or TypeSafe; homepage openjev.com; default master; MIT; HEAD ca3ba65f1429; Tolerate float roundoff in MLX evidence verification; pushed 2026-09-19; live REST 2282★ / 140 forks; size 9177; README SHA 74ab7f7f; LICENSE SHA ca562883; interface pattern reproduction with open models; does not reproduce Jev undisclosed model/training; Direct option logits; 0 output tokens; shared-state parallel; MLX backend for Apple Silicon (`--backend mlx`); Qwen3.5-4B 3090 direct 1.023s vs AR JSON 5.332s (**5.21×**); argmax agree 18/21; systems comparison ≠ semantic equivalence; Parallel suffixes 20.03 dec/s on 777 decisions; Browser ladder Qwen3.5-4B authored BA 0.813, pert 0.766, TypeSafe subset agreement 0.845 vs Published Jev 0.883 (102 across 20 cases); Softmax over options ≠ calibrated Noul; typed output does not guarantee semantic correctness; wire/agreement ≠ replica of TypeSafe; SemIf ≠ kw2828/OpenJev playground ≠ zhihz/openjev ≠ apiplant/semif-rs port ≠ dddanielliu/semif-serve; rename is densify not a second census; JevBench 74.6 is §78 not this ladder; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33/#34/#35/#36/#38; do not push onto open #39/#40; notes.md §117 diff --git a/.agents/skills/augustus/references/formal-methods.md b/.agents/skills/augustus/references/formal-methods.md index 75e795f..ecafb22 100644 --- a/.agents/skills/augustus/references/formal-methods.md +++ b/.agents/skills/augustus/references/formal-methods.md @@ -1365,3 +1365,6 @@ preference lint; `faq.md`; `boundary-audit.md`. Hourly 0843 uniqueness lock: A hunch is a probability with a policy attached; { enter: 0.8, exit: 0.6 } is hysteresis; replay a policy change without inference; Decision models are providers, not the product; huncho ≠ Kungie/gut ≠ carldaws/hunch ≠ tpellet/hunch; pretrained Qwen2.5 base ECE 0.030 (0.5B) / 0.040 (7B); instruct 0.302 / 0.269; 70.9% → 70.0% mean conf 74.1% → 96.7%; temperature scaling still matches it in-distribution; No Jev API was called; Qwen2.5 ≠ Archer; Qwen/Qwen3.8-27B ≠ Archer; 学習済みモデル v0.1 は準備中です; bool AUROC 0.523; 先頭だと0件、末尾だと250件; 温度を渡さない場合、確率は較正されていません; このリポジトリには Jev を呼ぶコードが存在しません; g0runmezadam/what-is-jev IS tunahansahin897/what-is-jev (same GitHub id 1378007307); 947 repos scored; A 273 · B 302 · C 372; LLM rubric ≠ benches; Data as of 2026-09-20; HEAD 895b9498; README SHA 3ae98c56; 13 focused checks and one mutually exclusive outcome; Probabilities are advisory, not calibrated guarantees; omni-/ask-jev ≠ pedroknigge/mcp_jev; pd.cut bins by equal width while jeval bins by quantile; ECE 0.113 and ECE 0.076; jeval drift is not implemented yet; rlaope/jeval ≠ dayhaysoos/jevals; calibration does not compose; ECE has exactly zero statistical power to detect the failure mode that kills trajectories; 25–60× headline withdrawn; P(all-correct): 0.0071 vs 0.0001; TCE / AMS; Qwen 3.8 sparring ≠ Archer; Deferred Crispification; light_cutoff_applied_to_combination 0; recorded run, kinematic animation; BANKING77 Accuracy BERT-Base 93.02 Jev 79.90; Analyse jev calibration (NLL, ECE) backlog; BERT figures are published supervised references, not zero-shot; 档位措辞效应 分数极差中位 0.50、最大 1.32; 修好后对照组是 0.01; 不是 benchmark; 概率没做 calibration; ~1,430 API calls, about $0.15; xiaohuaxi/jev-study ≠ wjdjdakf17/jev-study ≠ baekenough/jev-study; AND: product (independence assumed and recorded in the trace); chat model's stated confidence is not calibrated; circuit-vl-4b ≠ Archer; Bring your own API key; vamsikrishna2421/jev-usecases ≠ whyashthakker/awesome-jev-use-cases; catalog ≠ endorsement; SemIf 2237★ (+30 vs §111 2207); jevlike 1049★ (+6 vs 1043); TypeLLM/TypeLLM 16★; AnotiaWang 98★; yibie/awesome-jev 520★ (+14 vs 506); Laya likes 861 (was 822); tracker likes 66 (+2 vs 64) lastModified UNCHANGED; Blackwood likes 2 gated manual; Archer still promised_not_landed; Hub archerhume/4rcherhume HTTP 401; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33; notes.md §114 User-provided 0920 jcr uniqueness lock: NiazMorshed2007/jcr MIT; site https://jcr.niazmorshed.dev; topics ai-agents,jev,mcp; **4★**; HEAD `138b3832`; README SHA `2a49dbc1`; LICENSE SHA `46231303`; size **14850**; Jev Capability Resolver; one tool to find documented deterministic commands in a nested capability tree; returns context; **does not execute**; skills = workflow+judgment; capabilities = individual operations; format independent of Jev; proposed open standard exploration; classify (Jev) → optional OpenAI decompose compound → beam search geometric mean of routing probs; keep up to 3 paths ≥60% of best (JCR_BAND_RATIO 0.6); ambiguity / no-match / depth-limit explicit; soft scores ≠ hard gates; 0.6 band is application policy; routing ≠ permission; docs ≠ authority to run; sol-vs-opus5-20 *theirs*: 20 scenarios × 4 variants = 80 runs; lookup+explain only, no execution; Claude Opus 5: agent input 108,585→15,819 (−85%), cost $0.3700→$0.1222 (−67%), wall 105.5s→77.7s; Codex GPT-5.6-Sol: 61,952→47,669 (−23%), $0.1377→$0.1151 (−16%), wall 25.3s→62.4s (Sol slower with JCR in 19/20); One Sol outlier 372.6s / 193 Jev calls; n=1 per cell; Not Harbor task-execution; Claude/Codex harnesses; compare mode; 50 scenarios bundled; 11 groups, 960 nodes, 11,360 items; 16 routing rounds per step; NiazMorshed2007/jcr ≠ skill-broker ≠ skillranker ≠ jev-sift ≠ jev-lens ≠ jevusher ≠ jev_select_capability; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33/#34; notes.md §116 + +User-provided 0922 (`notes.md` §117 / items 330–336 / batch #100). SemIf was formerly OpenJev; MLX backend; 5.21× systems≠semantic; Softmax over options ≠ calibrated Noul; live REST 2282★. Do not reopen #23–#36/#38; do not push onto open #39/#40. +User-provided 0922 uniqueness lock: SemIf was formerly OpenJev; independent; not affiliated with Jev or TypeSafe; homepage openjev.com; default master; MIT; HEAD ca3ba65f1429; Tolerate float roundoff in MLX evidence verification; pushed 2026-09-19; live REST 2282★ / 140 forks; size 9177; README SHA 74ab7f7f; LICENSE SHA ca562883; interface pattern reproduction with open models; does not reproduce Jev undisclosed model/training; Direct option logits; 0 output tokens; shared-state parallel; MLX backend for Apple Silicon (`--backend mlx`); Qwen3.5-4B 3090 direct 1.023s vs AR JSON 5.332s (**5.21×**); argmax agree 18/21; systems comparison ≠ semantic equivalence; Parallel suffixes 20.03 dec/s on 777 decisions; Browser ladder Qwen3.5-4B authored BA 0.813, pert 0.766, TypeSafe subset agreement 0.845 vs Published Jev 0.883 (102 across 20 cases); Softmax over options ≠ calibrated Noul; typed output does not guarantee semantic correctness; wire/agreement ≠ replica of TypeSafe; SemIf ≠ kw2828/OpenJev playground ≠ zhihz/openjev ≠ apiplant/semif-rs port ≠ dddanielliu/semif-serve; rename is densify not a second census; JevBench 74.6 is §78 not this ladder; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33/#34/#35/#36/#38; do not push onto open #39/#40; notes.md §117 diff --git a/.agents/skills/augustus/references/formal-semi-formal.md b/.agents/skills/augustus/references/formal-semi-formal.md index 413bf64..f67331d 100644 --- a/.agents/skills/augustus/references/formal-semi-formal.md +++ b/.agents/skills/augustus/references/formal-semi-formal.md @@ -78,3 +78,6 @@ User-provided 0915 uniqueness lock: TianyuCodings/NanoJev unified-games-v1 densi JCR (`notes.md` §116): docs ≠ authority to run. routing ≠ permission. A Noul / routing probability is a SENSOR, not a grant. User-provided 0920 jcr uniqueness lock: NiazMorshed2007/jcr MIT; site https://jcr.niazmorshed.dev; topics ai-agents,jev,mcp; **4★**; HEAD `138b3832`; README SHA `2a49dbc1`; LICENSE SHA `46231303`; size **14850**; Jev Capability Resolver; one tool to find documented deterministic commands in a nested capability tree; returns context; **does not execute**; skills = workflow+judgment; capabilities = individual operations; format independent of Jev; proposed open standard exploration; classify (Jev) → optional OpenAI decompose compound → beam search geometric mean of routing probs; keep up to 3 paths ≥60% of best (JCR_BAND_RATIO 0.6); ambiguity / no-match / depth-limit explicit; soft scores ≠ hard gates; 0.6 band is application policy; routing ≠ permission; docs ≠ authority to run; sol-vs-opus5-20 *theirs*: 20 scenarios × 4 variants = 80 runs; lookup+explain only, no execution; Claude Opus 5: agent input 108,585→15,819 (−85%), cost $0.3700→$0.1222 (−67%), wall 105.5s→77.7s; Codex GPT-5.6-Sol: 61,952→47,669 (−23%), $0.1377→$0.1151 (−16%), wall 25.3s→62.4s (Sol slower with JCR in 19/20); One Sol outlier 372.6s / 193 Jev calls; n=1 per cell; Not Harbor task-execution; Claude/Codex harnesses; compare mode; 50 scenarios bundled; 11 groups, 960 nodes, 11,360 items; 16 routing rounds per step; NiazMorshed2007/jcr ≠ skill-broker ≠ skillranker ≠ jev-sift ≠ jev-lens ≠ jevusher ≠ jev_select_capability; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33/#34; notes.md §116 + +User-provided 0922 (`notes.md` §117 / items 330–336 / batch #100). SemIf was formerly OpenJev; MLX backend; 5.21× systems≠semantic; Softmax over options ≠ calibrated Noul; live REST 2282★. Do not reopen #23–#36/#38; do not push onto open #39/#40. +User-provided 0922 uniqueness lock: SemIf was formerly OpenJev; independent; not affiliated with Jev or TypeSafe; homepage openjev.com; default master; MIT; HEAD ca3ba65f1429; Tolerate float roundoff in MLX evidence verification; pushed 2026-09-19; live REST 2282★ / 140 forks; size 9177; README SHA 74ab7f7f; LICENSE SHA ca562883; interface pattern reproduction with open models; does not reproduce Jev undisclosed model/training; Direct option logits; 0 output tokens; shared-state parallel; MLX backend for Apple Silicon (`--backend mlx`); Qwen3.5-4B 3090 direct 1.023s vs AR JSON 5.332s (**5.21×**); argmax agree 18/21; systems comparison ≠ semantic equivalence; Parallel suffixes 20.03 dec/s on 777 decisions; Browser ladder Qwen3.5-4B authored BA 0.813, pert 0.766, TypeSafe subset agreement 0.845 vs Published Jev 0.883 (102 across 20 cases); Softmax over options ≠ calibrated Noul; typed output does not guarantee semantic correctness; wire/agreement ≠ replica of TypeSafe; SemIf ≠ kw2828/OpenJev playground ≠ zhihz/openjev ≠ apiplant/semif-rs port ≠ dddanielliu/semif-serve; rename is densify not a second census; JevBench 74.6 is §78 not this ladder; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33/#34/#35/#36/#38; do not push onto open #39/#40; notes.md §117 diff --git a/.agents/skills/augustus/references/judgment-class.md b/.agents/skills/augustus/references/judgment-class.md index 5f716e1..9e9d0b0 100644 --- a/.agents/skills/augustus/references/judgment-class.md +++ b/.agents/skills/augustus/references/judgment-class.md @@ -1458,3 +1458,6 @@ User-provided 0915 uniqueness lock: TianyuCodings/NanoJev unified-games-v1 densi User-provided 0920 jcr uniqueness lock: NiazMorshed2007/jcr MIT; site https://jcr.niazmorshed.dev; topics ai-agents,jev,mcp; **4★**; HEAD `138b3832`; README SHA `2a49dbc1`; LICENSE SHA `46231303`; size **14850**; Jev Capability Resolver; one tool to find documented deterministic commands in a nested capability tree; returns context; **does not execute**; skills = workflow+judgment; capabilities = individual operations; format independent of Jev; proposed open standard exploration; classify (Jev) → optional OpenAI decompose compound → beam search geometric mean of routing probs; keep up to 3 paths ≥60% of best (JCR_BAND_RATIO 0.6); ambiguity / no-match / depth-limit explicit; soft scores ≠ hard gates; 0.6 band is application policy; routing ≠ permission; docs ≠ authority to run; sol-vs-opus5-20 *theirs*: 20 scenarios × 4 variants = 80 runs; lookup+explain only, no execution; Claude Opus 5: agent input 108,585→15,819 (−85%), cost $0.3700→$0.1222 (−67%), wall 105.5s→77.7s; Codex GPT-5.6-Sol: 61,952→47,669 (−23%), $0.1377→$0.1151 (−16%), wall 25.3s→62.4s (Sol slower with JCR in 19/20); One Sol outlier 372.6s / 193 Jev calls; n=1 per cell; Not Harbor task-execution; Claude/Codex harnesses; compare mode; 50 scenarios bundled; 11 groups, 960 nodes, 11,360 items; 16 routing rounds per step; NiazMorshed2007/jcr ≠ skill-broker ≠ skillranker ≠ jev-sift ≠ jev-lens ≠ jevusher ≠ jev_select_capability; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33/#34; notes.md §116 + +User-provided 0922 (`notes.md` §117 / items 330–336 / batch #100). SemIf was formerly OpenJev; MLX backend; 5.21× systems≠semantic; Softmax over options ≠ calibrated Noul; live REST 2282★. Do not reopen #23–#36/#38; do not push onto open #39/#40. +User-provided 0922 uniqueness lock: SemIf was formerly OpenJev; independent; not affiliated with Jev or TypeSafe; homepage openjev.com; default master; MIT; HEAD ca3ba65f1429; Tolerate float roundoff in MLX evidence verification; pushed 2026-09-19; live REST 2282★ / 140 forks; size 9177; README SHA 74ab7f7f; LICENSE SHA ca562883; interface pattern reproduction with open models; does not reproduce Jev undisclosed model/training; Direct option logits; 0 output tokens; shared-state parallel; MLX backend for Apple Silicon (`--backend mlx`); Qwen3.5-4B 3090 direct 1.023s vs AR JSON 5.332s (**5.21×**); argmax agree 18/21; systems comparison ≠ semantic equivalence; Parallel suffixes 20.03 dec/s on 777 decisions; Browser ladder Qwen3.5-4B authored BA 0.813, pert 0.766, TypeSafe subset agreement 0.845 vs Published Jev 0.883 (102 across 20 cases); Softmax over options ≠ calibrated Noul; typed output does not guarantee semantic correctness; wire/agreement ≠ replica of TypeSafe; SemIf ≠ kw2828/OpenJev playground ≠ zhihz/openjev ≠ apiplant/semif-rs port ≠ dddanielliu/semif-serve; rename is densify not a second census; JevBench 74.6 is §78 not this ladder; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33/#34/#35/#36/#38; do not push onto open #39/#40; notes.md §117 diff --git a/.agents/skills/augustus/references/mappings.md b/.agents/skills/augustus/references/mappings.md index ad7db70..55688a9 100644 --- a/.agents/skills/augustus/references/mappings.md +++ b/.agents/skills/augustus/references/mappings.md @@ -2470,3 +2470,6 @@ User-provided 0915 uniqueness lock: TianyuCodings/NanoJev unified-games-v1 densi User-provided 0920 jcr uniqueness lock: NiazMorshed2007/jcr MIT; site https://jcr.niazmorshed.dev; topics ai-agents,jev,mcp; **4★**; HEAD `138b3832`; README SHA `2a49dbc1`; LICENSE SHA `46231303`; size **14850**; Jev Capability Resolver; one tool to find documented deterministic commands in a nested capability tree; returns context; **does not execute**; skills = workflow+judgment; capabilities = individual operations; format independent of Jev; proposed open standard exploration; classify (Jev) → optional OpenAI decompose compound → beam search geometric mean of routing probs; keep up to 3 paths ≥60% of best (JCR_BAND_RATIO 0.6); ambiguity / no-match / depth-limit explicit; soft scores ≠ hard gates; 0.6 band is application policy; routing ≠ permission; docs ≠ authority to run; sol-vs-opus5-20 *theirs*: 20 scenarios × 4 variants = 80 runs; lookup+explain only, no execution; Claude Opus 5: agent input 108,585→15,819 (−85%), cost $0.3700→$0.1222 (−67%), wall 105.5s→77.7s; Codex GPT-5.6-Sol: 61,952→47,669 (−23%), $0.1377→$0.1151 (−16%), wall 25.3s→62.4s (Sol slower with JCR in 19/20); One Sol outlier 372.6s / 193 Jev calls; n=1 per cell; Not Harbor task-execution; Claude/Codex harnesses; compare mode; 50 scenarios bundled; 11 groups, 960 nodes, 11,360 items; 16 routing rounds per step; NiazMorshed2007/jcr ≠ skill-broker ≠ skillranker ≠ jev-sift ≠ jev-lens ≠ jevusher ≠ jev_select_capability; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33/#34; notes.md §116 + +User-provided 0922 (`notes.md` §117 / items 330–336 / batch #100). SemIf was formerly OpenJev; MLX backend; 5.21× systems≠semantic; Softmax over options ≠ calibrated Noul; live REST 2282★. Do not reopen #23–#36/#38; do not push onto open #39/#40. +User-provided 0922 uniqueness lock: SemIf was formerly OpenJev; independent; not affiliated with Jev or TypeSafe; homepage openjev.com; default master; MIT; HEAD ca3ba65f1429; Tolerate float roundoff in MLX evidence verification; pushed 2026-09-19; live REST 2282★ / 140 forks; size 9177; README SHA 74ab7f7f; LICENSE SHA ca562883; interface pattern reproduction with open models; does not reproduce Jev undisclosed model/training; Direct option logits; 0 output tokens; shared-state parallel; MLX backend for Apple Silicon (`--backend mlx`); Qwen3.5-4B 3090 direct 1.023s vs AR JSON 5.332s (**5.21×**); argmax agree 18/21; systems comparison ≠ semantic equivalence; Parallel suffixes 20.03 dec/s on 777 decisions; Browser ladder Qwen3.5-4B authored BA 0.813, pert 0.766, TypeSafe subset agreement 0.845 vs Published Jev 0.883 (102 across 20 cases); Softmax over options ≠ calibrated Noul; typed output does not guarantee semantic correctness; wire/agreement ≠ replica of TypeSafe; SemIf ≠ kw2828/OpenJev playground ≠ zhihz/openjev ≠ apiplant/semif-rs port ≠ dddanielliu/semif-serve; rename is densify not a second census; JevBench 74.6 is §78 not this ladder; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33/#34/#35/#36/#38; do not push onto open #39/#40; notes.md §117 diff --git a/.agents/skills/augustus/references/mental-models.md b/.agents/skills/augustus/references/mental-models.md index f4ec716..03c252a 100644 --- a/.agents/skills/augustus/references/mental-models.md +++ b/.agents/skills/augustus/references/mental-models.md @@ -2485,6 +2485,54 @@ do not reopen or amend PR #23/#24/#25/#26/#27. + +## Apply 0922 (`notes.md` §117) + +Same pillars, user-provided HIGH densify (TheoLeeCJ/SemIf rename + +MLX + accuracy ladder — not a first sighting; AI / SWE / business / +knowledge work, not SWE-only). Do **not** re-fold §78 JevBench 74.6 / +§69 semif-serve / §113 census 2241★ / §114 hourly 2237★ / merged #36 +NanoJev §115 / merged #38 jcr §116 as if they were this ladder. +Reconstruct onto main after merged **#34/#35/#36/#38**; never reopen +merged #7–**#36** or **#38**. Do not push onto open **#39** (hygiene) +or **#40** (llm-to-jev §118 / 322–329 / #101). This fold is `notes.md` +§117 / composition **330–336** / batch **#100**. Skip Archer rewrite +(still **promised_not_landed**). Quote *theirs*. Soft Noul ≠ hard +safety. Formal methods **compose** with scoring; a Noul is a SENSOR. +Jev is the hot exemplar, not the whole mandate. Treating 0.813 / +0.845 / 5.21× as Harbor, 18/21 as semantic equivalence, or +softmax-over-options as a Noul is the anti-pattern. + +1. **Rename is densify** — SemIf was formerly OpenJev. independent; + not affiliated with Jev or TypeSafe. homepage openjev.com. + live REST 2282★ / 140 forks. HEAD ca3ba65f1429. +2. **Interface pattern ≠ replica** — interface pattern reproduction + with open models; does not reproduce Jev undisclosed model/training. + wire/agreement ≠ replica of TypeSafe. +3. **Direct logits / MLX** — Direct option logits; 0 output tokens; + shared-state parallel; MLX backend for Apple Silicon (`--backend mlx`). + typed output does not guarantee semantic correctness. +4. **Systems comparison ≠ semantic equivalence** — direct 1.023s vs + AR JSON 5.332s (**5.21×**); argmax agree 18/21. Parallel suffixes + 20.03 dec/s on 777 decisions. +5. **Browser ladder ≠ Harbor; cousins ≠ host** — authored BA 0.813, + pert 0.766, TypeSafe subset agreement 0.845 vs Published Jev 0.883 + (102 across 20 cases). Softmax over options ≠ calibrated Noul. + JevBench 74.6 is §78 not this ladder. SemIf ≠ kw2828/OpenJev + playground ≠ zhihz/openjev ≠ apiplant/semif-rs port ≠ + dddanielliu/semif-serve. + +Soft Noul ≠ hard safety: 0.813 / 0.845 / 0.883 / 5.21× / 18/21 are +**sensors**. Treating subset agreement as a replica or 5.21× as +TypeSafe latency is the same theater as jev-gate §79. + +Formal methods **compose** with scoring. A Noul is a SENSOR. +Option-logit readout / prefix reuse / MLX float roundoff are exact +work. Ranking ≠ calibration theater. + +Unique consecutive fragments (this fold): +User-provided 0922 uniqueness lock: SemIf was formerly OpenJev; independent; not affiliated with Jev or TypeSafe; homepage openjev.com; default master; MIT; HEAD ca3ba65f1429; Tolerate float roundoff in MLX evidence verification; pushed 2026-09-19; live REST 2282★ / 140 forks; size 9177; README SHA 74ab7f7f; LICENSE SHA ca562883; interface pattern reproduction with open models; does not reproduce Jev undisclosed model/training; Direct option logits; 0 output tokens; shared-state parallel; MLX backend for Apple Silicon (`--backend mlx`); Qwen3.5-4B 3090 direct 1.023s vs AR JSON 5.332s (**5.21×**); argmax agree 18/21; systems comparison ≠ semantic equivalence; Parallel suffixes 20.03 dec/s on 777 decisions; Browser ladder Qwen3.5-4B authored BA 0.813, pert 0.766, TypeSafe subset agreement 0.845 vs Published Jev 0.883 (102 across 20 cases); Softmax over options ≠ calibrated Noul; typed output does not guarantee semantic correctness; wire/agreement ≠ replica of TypeSafe; SemIf ≠ kw2828/OpenJev playground ≠ zhihz/openjev ≠ apiplant/semif-rs port ≠ dddanielliu/semif-serve; rename is densify not a second census; JevBench 74.6 is §78 not this ladder; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33/#34/#35/#36/#38; do not push onto open #39/#40; notes.md §117 + ## Apply 0920 jcr (`notes.md` §116) Same pillars, one productized HIGH (capability-tree @@ -2879,3 +2927,4 @@ Related: `mappings.md` §1–§18, `methods-catalog.md`, `toolbox-mapping.md`, Hourly 0843 uniqueness lock: A hunch is a probability with a policy attached; { enter: 0.8, exit: 0.6 } is hysteresis; replay a policy change without inference; Decision models are providers, not the product; huncho ≠ Kungie/gut ≠ carldaws/hunch ≠ tpellet/hunch; pretrained Qwen2.5 base ECE 0.030 (0.5B) / 0.040 (7B); instruct 0.302 / 0.269; 70.9% → 70.0% mean conf 74.1% → 96.7%; temperature scaling still matches it in-distribution; No Jev API was called; Qwen2.5 ≠ Archer; Qwen/Qwen3.8-27B ≠ Archer; 学習済みモデル v0.1 は準備中です; bool AUROC 0.523; 先頭だと0件、末尾だと250件; 温度を渡さない場合、確率は較正されていません; このリポジトリには Jev を呼ぶコードが存在しません; g0runmezadam/what-is-jev IS tunahansahin897/what-is-jev (same GitHub id 1378007307); 947 repos scored; A 273 · B 302 · C 372; LLM rubric ≠ benches; Data as of 2026-09-20; HEAD 895b9498; README SHA 3ae98c56; 13 focused checks and one mutually exclusive outcome; Probabilities are advisory, not calibrated guarantees; omni-/ask-jev ≠ pedroknigge/mcp_jev; pd.cut bins by equal width while jeval bins by quantile; ECE 0.113 and ECE 0.076; jeval drift is not implemented yet; rlaope/jeval ≠ dayhaysoos/jevals; calibration does not compose; ECE has exactly zero statistical power to detect the failure mode that kills trajectories; 25–60× headline withdrawn; P(all-correct): 0.0071 vs 0.0001; TCE / AMS; Qwen 3.8 sparring ≠ Archer; Deferred Crispification; light_cutoff_applied_to_combination 0; recorded run, kinematic animation; BANKING77 Accuracy BERT-Base 93.02 Jev 79.90; Analyse jev calibration (NLL, ECE) backlog; BERT figures are published supervised references, not zero-shot; 档位措辞效应 分数极差中位 0.50、最大 1.32; 修好后对照组是 0.01; 不是 benchmark; 概率没做 calibration; ~1,430 API calls, about $0.15; xiaohuaxi/jev-study ≠ wjdjdakf17/jev-study ≠ baekenough/jev-study; AND: product (independence assumed and recorded in the trace); chat model's stated confidence is not calibrated; circuit-vl-4b ≠ Archer; Bring your own API key; vamsikrishna2421/jev-usecases ≠ whyashthakker/awesome-jev-use-cases; catalog ≠ endorsement; SemIf 2237★ (+30 vs §111 2207); jevlike 1049★ (+6 vs 1043); TypeLLM/TypeLLM 16★; AnotiaWang 98★; yibie/awesome-jev 520★ (+14 vs 506); Laya likes 861 (was 822); tracker likes 66 (+2 vs 64) lastModified UNCHANGED; Blackwood likes 2 gated manual; Archer still promised_not_landed; Hub archerhume/4rcherhume HTTP 401; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33; notes.md §114 User-provided 0920 jcr uniqueness lock: NiazMorshed2007/jcr MIT; site https://jcr.niazmorshed.dev; topics ai-agents,jev,mcp; **4★**; HEAD `138b3832`; README SHA `2a49dbc1`; LICENSE SHA `46231303`; size **14850**; Jev Capability Resolver; one tool to find documented deterministic commands in a nested capability tree; returns context; **does not execute**; skills = workflow+judgment; capabilities = individual operations; format independent of Jev; proposed open standard exploration; classify (Jev) → optional OpenAI decompose compound → beam search geometric mean of routing probs; keep up to 3 paths ≥60% of best (JCR_BAND_RATIO 0.6); ambiguity / no-match / depth-limit explicit; soft scores ≠ hard gates; 0.6 band is application policy; routing ≠ permission; docs ≠ authority to run; sol-vs-opus5-20 *theirs*: 20 scenarios × 4 variants = 80 runs; lookup+explain only, no execution; Claude Opus 5: agent input 108,585→15,819 (−85%), cost $0.3700→$0.1222 (−67%), wall 105.5s→77.7s; Codex GPT-5.6-Sol: 61,952→47,669 (−23%), $0.1377→$0.1151 (−16%), wall 25.3s→62.4s (Sol slower with JCR in 19/20); One Sol outlier 372.6s / 193 Jev calls; n=1 per cell; Not Harbor task-execution; Claude/Codex harnesses; compare mode; 50 scenarios bundled; 11 groups, 960 nodes, 11,360 items; 16 routing rounds per step; NiazMorshed2007/jcr ≠ skill-broker ≠ skillranker ≠ jev-sift ≠ jev-lens ≠ jevusher ≠ jev_select_capability; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33/#34; notes.md §116 +User-provided 0922 uniqueness lock: SemIf was formerly OpenJev; independent; not affiliated with Jev or TypeSafe; homepage openjev.com; default master; MIT; HEAD ca3ba65f1429; Tolerate float roundoff in MLX evidence verification; pushed 2026-09-19; live REST 2282★ / 140 forks; size 9177; README SHA 74ab7f7f; LICENSE SHA ca562883; interface pattern reproduction with open models; does not reproduce Jev undisclosed model/training; Direct option logits; 0 output tokens; shared-state parallel; MLX backend for Apple Silicon (`--backend mlx`); Qwen3.5-4B 3090 direct 1.023s vs AR JSON 5.332s (**5.21×**); argmax agree 18/21; systems comparison ≠ semantic equivalence; Parallel suffixes 20.03 dec/s on 777 decisions; Browser ladder Qwen3.5-4B authored BA 0.813, pert 0.766, TypeSafe subset agreement 0.845 vs Published Jev 0.883 (102 across 20 cases); Softmax over options ≠ calibrated Noul; typed output does not guarantee semantic correctness; wire/agreement ≠ replica of TypeSafe; SemIf ≠ kw2828/OpenJev playground ≠ zhihz/openjev ≠ apiplant/semif-rs port ≠ dddanielliu/semif-serve; rename is densify not a second census; JevBench 74.6 is §78 not this ladder; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33/#34/#35/#36/#38; do not push onto open #39/#40; notes.md §117 diff --git a/.agents/skills/augustus/references/methods-catalog.md b/.agents/skills/augustus/references/methods-catalog.md index bf79223..2201466 100644 --- a/.agents/skills/augustus/references/methods-catalog.md +++ b/.agents/skills/augustus/references/methods-catalog.md @@ -272,3 +272,6 @@ Hourly 0843 uniqueness lock: A hunch is a probability with a policy attached; { **0920 jcr:** geometric-mean path score is a control statistic (cookbook beam cousin), not a calibrated Noul and not a permission bit. `notes.md` §116. User-provided 0920 jcr uniqueness lock: NiazMorshed2007/jcr MIT; site https://jcr.niazmorshed.dev; topics ai-agents,jev,mcp; **4★**; HEAD `138b3832`; README SHA `2a49dbc1`; LICENSE SHA `46231303`; size **14850**; Jev Capability Resolver; one tool to find documented deterministic commands in a nested capability tree; returns context; **does not execute**; skills = workflow+judgment; capabilities = individual operations; format independent of Jev; proposed open standard exploration; classify (Jev) → optional OpenAI decompose compound → beam search geometric mean of routing probs; keep up to 3 paths ≥60% of best (JCR_BAND_RATIO 0.6); ambiguity / no-match / depth-limit explicit; soft scores ≠ hard gates; 0.6 band is application policy; routing ≠ permission; docs ≠ authority to run; sol-vs-opus5-20 *theirs*: 20 scenarios × 4 variants = 80 runs; lookup+explain only, no execution; Claude Opus 5: agent input 108,585→15,819 (−85%), cost $0.3700→$0.1222 (−67%), wall 105.5s→77.7s; Codex GPT-5.6-Sol: 61,952→47,669 (−23%), $0.1377→$0.1151 (−16%), wall 25.3s→62.4s (Sol slower with JCR in 19/20); One Sol outlier 372.6s / 193 Jev calls; n=1 per cell; Not Harbor task-execution; Claude/Codex harnesses; compare mode; 50 scenarios bundled; 11 groups, 960 nodes, 11,360 items; 16 routing rounds per step; NiazMorshed2007/jcr ≠ skill-broker ≠ skillranker ≠ jev-sift ≠ jev-lens ≠ jevusher ≠ jev_select_capability; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33/#34; notes.md §116 + +User-provided 0922 (`notes.md` §117 / items 330–336 / batch #100). SemIf was formerly OpenJev; MLX backend; 5.21× systems≠semantic; Softmax over options ≠ calibrated Noul; live REST 2282★. Do not reopen #23–#36/#38; do not push onto open #39/#40. +User-provided 0922 uniqueness lock: SemIf was formerly OpenJev; independent; not affiliated with Jev or TypeSafe; homepage openjev.com; default master; MIT; HEAD ca3ba65f1429; Tolerate float roundoff in MLX evidence verification; pushed 2026-09-19; live REST 2282★ / 140 forks; size 9177; README SHA 74ab7f7f; LICENSE SHA ca562883; interface pattern reproduction with open models; does not reproduce Jev undisclosed model/training; Direct option logits; 0 output tokens; shared-state parallel; MLX backend for Apple Silicon (`--backend mlx`); Qwen3.5-4B 3090 direct 1.023s vs AR JSON 5.332s (**5.21×**); argmax agree 18/21; systems comparison ≠ semantic equivalence; Parallel suffixes 20.03 dec/s on 777 decisions; Browser ladder Qwen3.5-4B authored BA 0.813, pert 0.766, TypeSafe subset agreement 0.845 vs Published Jev 0.883 (102 across 20 cases); Softmax over options ≠ calibrated Noul; typed output does not guarantee semantic correctness; wire/agreement ≠ replica of TypeSafe; SemIf ≠ kw2828/OpenJev playground ≠ zhihz/openjev ≠ apiplant/semif-rs port ≠ dddanielliu/semif-serve; rename is densify not a second census; JevBench 74.6 is §78 not this ladder; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33/#34/#35/#36/#38; do not push onto open #39/#40; notes.md §117 diff --git a/.agents/skills/augustus/references/mixed-architecture.md b/.agents/skills/augustus/references/mixed-architecture.md index 44cb020..2f41886 100644 --- a/.agents/skills/augustus/references/mixed-architecture.md +++ b/.agents/skills/augustus/references/mixed-architecture.md @@ -1435,3 +1435,6 @@ classifier-step), recommend one, and name the experiment that kills it. Hourly 0843 uniqueness lock: A hunch is a probability with a policy attached; { enter: 0.8, exit: 0.6 } is hysteresis; replay a policy change without inference; Decision models are providers, not the product; huncho ≠ Kungie/gut ≠ carldaws/hunch ≠ tpellet/hunch; pretrained Qwen2.5 base ECE 0.030 (0.5B) / 0.040 (7B); instruct 0.302 / 0.269; 70.9% → 70.0% mean conf 74.1% → 96.7%; temperature scaling still matches it in-distribution; No Jev API was called; Qwen2.5 ≠ Archer; Qwen/Qwen3.8-27B ≠ Archer; 学習済みモデル v0.1 は準備中です; bool AUROC 0.523; 先頭だと0件、末尾だと250件; 温度を渡さない場合、確率は較正されていません; このリポジトリには Jev を呼ぶコードが存在しません; g0runmezadam/what-is-jev IS tunahansahin897/what-is-jev (same GitHub id 1378007307); 947 repos scored; A 273 · B 302 · C 372; LLM rubric ≠ benches; Data as of 2026-09-20; HEAD 895b9498; README SHA 3ae98c56; 13 focused checks and one mutually exclusive outcome; Probabilities are advisory, not calibrated guarantees; omni-/ask-jev ≠ pedroknigge/mcp_jev; pd.cut bins by equal width while jeval bins by quantile; ECE 0.113 and ECE 0.076; jeval drift is not implemented yet; rlaope/jeval ≠ dayhaysoos/jevals; calibration does not compose; ECE has exactly zero statistical power to detect the failure mode that kills trajectories; 25–60× headline withdrawn; P(all-correct): 0.0071 vs 0.0001; TCE / AMS; Qwen 3.8 sparring ≠ Archer; Deferred Crispification; light_cutoff_applied_to_combination 0; recorded run, kinematic animation; BANKING77 Accuracy BERT-Base 93.02 Jev 79.90; Analyse jev calibration (NLL, ECE) backlog; BERT figures are published supervised references, not zero-shot; 档位措辞效应 分数极差中位 0.50、最大 1.32; 修好后对照组是 0.01; 不是 benchmark; 概率没做 calibration; ~1,430 API calls, about $0.15; xiaohuaxi/jev-study ≠ wjdjdakf17/jev-study ≠ baekenough/jev-study; AND: product (independence assumed and recorded in the trace); chat model's stated confidence is not calibrated; circuit-vl-4b ≠ Archer; Bring your own API key; vamsikrishna2421/jev-usecases ≠ whyashthakker/awesome-jev-use-cases; catalog ≠ endorsement; SemIf 2237★ (+30 vs §111 2207); jevlike 1049★ (+6 vs 1043); TypeLLM/TypeLLM 16★; AnotiaWang 98★; yibie/awesome-jev 520★ (+14 vs 506); Laya likes 861 (was 822); tracker likes 66 (+2 vs 64) lastModified UNCHANGED; Blackwood likes 2 gated manual; Archer still promised_not_landed; Hub archerhume/4rcherhume HTTP 401; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33; notes.md §114 User-provided 0920 jcr uniqueness lock: NiazMorshed2007/jcr MIT; site https://jcr.niazmorshed.dev; topics ai-agents,jev,mcp; **4★**; HEAD `138b3832`; README SHA `2a49dbc1`; LICENSE SHA `46231303`; size **14850**; Jev Capability Resolver; one tool to find documented deterministic commands in a nested capability tree; returns context; **does not execute**; skills = workflow+judgment; capabilities = individual operations; format independent of Jev; proposed open standard exploration; classify (Jev) → optional OpenAI decompose compound → beam search geometric mean of routing probs; keep up to 3 paths ≥60% of best (JCR_BAND_RATIO 0.6); ambiguity / no-match / depth-limit explicit; soft scores ≠ hard gates; 0.6 band is application policy; routing ≠ permission; docs ≠ authority to run; sol-vs-opus5-20 *theirs*: 20 scenarios × 4 variants = 80 runs; lookup+explain only, no execution; Claude Opus 5: agent input 108,585→15,819 (−85%), cost $0.3700→$0.1222 (−67%), wall 105.5s→77.7s; Codex GPT-5.6-Sol: 61,952→47,669 (−23%), $0.1377→$0.1151 (−16%), wall 25.3s→62.4s (Sol slower with JCR in 19/20); One Sol outlier 372.6s / 193 Jev calls; n=1 per cell; Not Harbor task-execution; Claude/Codex harnesses; compare mode; 50 scenarios bundled; 11 groups, 960 nodes, 11,360 items; 16 routing rounds per step; NiazMorshed2007/jcr ≠ skill-broker ≠ skillranker ≠ jev-sift ≠ jev-lens ≠ jevusher ≠ jev_select_capability; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33/#34; notes.md §116 + +User-provided 0922 (`notes.md` §117 / items 330–336 / batch #100). SemIf was formerly OpenJev; MLX backend; 5.21× systems≠semantic; Softmax over options ≠ calibrated Noul; live REST 2282★. Do not reopen #23–#36/#38; do not push onto open #39/#40. +User-provided 0922 uniqueness lock: SemIf was formerly OpenJev; independent; not affiliated with Jev or TypeSafe; homepage openjev.com; default master; MIT; HEAD ca3ba65f1429; Tolerate float roundoff in MLX evidence verification; pushed 2026-09-19; live REST 2282★ / 140 forks; size 9177; README SHA 74ab7f7f; LICENSE SHA ca562883; interface pattern reproduction with open models; does not reproduce Jev undisclosed model/training; Direct option logits; 0 output tokens; shared-state parallel; MLX backend for Apple Silicon (`--backend mlx`); Qwen3.5-4B 3090 direct 1.023s vs AR JSON 5.332s (**5.21×**); argmax agree 18/21; systems comparison ≠ semantic equivalence; Parallel suffixes 20.03 dec/s on 777 decisions; Browser ladder Qwen3.5-4B authored BA 0.813, pert 0.766, TypeSafe subset agreement 0.845 vs Published Jev 0.883 (102 across 20 cases); Softmax over options ≠ calibrated Noul; typed output does not guarantee semantic correctness; wire/agreement ≠ replica of TypeSafe; SemIf ≠ kw2828/OpenJev playground ≠ zhihz/openjev ≠ apiplant/semif-rs port ≠ dddanielliu/semif-serve; rename is densify not a second census; JevBench 74.6 is §78 not this ladder; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33/#34/#35/#36/#38; do not push onto open #39/#40; notes.md §117 diff --git a/.agents/skills/augustus/references/question-design.md b/.agents/skills/augustus/references/question-design.md index c9e734b..6df6ee2 100644 --- a/.agents/skills/augustus/references/question-design.md +++ b/.agents/skills/augustus/references/question-design.md @@ -414,3 +414,6 @@ Confidence-gated routing (doc defaults): act / confirm-or-flag / hand off, floor Hourly 0843 uniqueness lock: A hunch is a probability with a policy attached; { enter: 0.8, exit: 0.6 } is hysteresis; replay a policy change without inference; Decision models are providers, not the product; huncho ≠ Kungie/gut ≠ carldaws/hunch ≠ tpellet/hunch; pretrained Qwen2.5 base ECE 0.030 (0.5B) / 0.040 (7B); instruct 0.302 / 0.269; 70.9% → 70.0% mean conf 74.1% → 96.7%; temperature scaling still matches it in-distribution; No Jev API was called; Qwen2.5 ≠ Archer; Qwen/Qwen3.8-27B ≠ Archer; 学習済みモデル v0.1 は準備中です; bool AUROC 0.523; 先頭だと0件、末尾だと250件; 温度を渡さない場合、確率は較正されていません; このリポジトリには Jev を呼ぶコードが存在しません; g0runmezadam/what-is-jev IS tunahansahin897/what-is-jev (same GitHub id 1378007307); 947 repos scored; A 273 · B 302 · C 372; LLM rubric ≠ benches; Data as of 2026-09-20; HEAD 895b9498; README SHA 3ae98c56; 13 focused checks and one mutually exclusive outcome; Probabilities are advisory, not calibrated guarantees; omni-/ask-jev ≠ pedroknigge/mcp_jev; pd.cut bins by equal width while jeval bins by quantile; ECE 0.113 and ECE 0.076; jeval drift is not implemented yet; rlaope/jeval ≠ dayhaysoos/jevals; calibration does not compose; ECE has exactly zero statistical power to detect the failure mode that kills trajectories; 25–60× headline withdrawn; P(all-correct): 0.0071 vs 0.0001; TCE / AMS; Qwen 3.8 sparring ≠ Archer; Deferred Crispification; light_cutoff_applied_to_combination 0; recorded run, kinematic animation; BANKING77 Accuracy BERT-Base 93.02 Jev 79.90; Analyse jev calibration (NLL, ECE) backlog; BERT figures are published supervised references, not zero-shot; 档位措辞效应 分数极差中位 0.50、最大 1.32; 修好后对照组是 0.01; 不是 benchmark; 概率没做 calibration; ~1,430 API calls, about $0.15; xiaohuaxi/jev-study ≠ wjdjdakf17/jev-study ≠ baekenough/jev-study; AND: product (independence assumed and recorded in the trace); chat model's stated confidence is not calibrated; circuit-vl-4b ≠ Archer; Bring your own API key; vamsikrishna2421/jev-usecases ≠ whyashthakker/awesome-jev-use-cases; catalog ≠ endorsement; SemIf 2237★ (+30 vs §111 2207); jevlike 1049★ (+6 vs 1043); TypeLLM/TypeLLM 16★; AnotiaWang 98★; yibie/awesome-jev 520★ (+14 vs 506); Laya likes 861 (was 822); tracker likes 66 (+2 vs 64) lastModified UNCHANGED; Blackwood likes 2 gated manual; Archer still promised_not_landed; Hub archerhume/4rcherhume HTTP 401; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33; notes.md §114 User-provided 0920 jcr uniqueness lock: NiazMorshed2007/jcr MIT; site https://jcr.niazmorshed.dev; topics ai-agents,jev,mcp; **4★**; HEAD `138b3832`; README SHA `2a49dbc1`; LICENSE SHA `46231303`; size **14850**; Jev Capability Resolver; one tool to find documented deterministic commands in a nested capability tree; returns context; **does not execute**; skills = workflow+judgment; capabilities = individual operations; format independent of Jev; proposed open standard exploration; classify (Jev) → optional OpenAI decompose compound → beam search geometric mean of routing probs; keep up to 3 paths ≥60% of best (JCR_BAND_RATIO 0.6); ambiguity / no-match / depth-limit explicit; soft scores ≠ hard gates; 0.6 band is application policy; routing ≠ permission; docs ≠ authority to run; sol-vs-opus5-20 *theirs*: 20 scenarios × 4 variants = 80 runs; lookup+explain only, no execution; Claude Opus 5: agent input 108,585→15,819 (−85%), cost $0.3700→$0.1222 (−67%), wall 105.5s→77.7s; Codex GPT-5.6-Sol: 61,952→47,669 (−23%), $0.1377→$0.1151 (−16%), wall 25.3s→62.4s (Sol slower with JCR in 19/20); One Sol outlier 372.6s / 193 Jev calls; n=1 per cell; Not Harbor task-execution; Claude/Codex harnesses; compare mode; 50 scenarios bundled; 11 groups, 960 nodes, 11,360 items; 16 routing rounds per step; NiazMorshed2007/jcr ≠ skill-broker ≠ skillranker ≠ jev-sift ≠ jev-lens ≠ jevusher ≠ jev_select_capability; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33/#34; notes.md §116 + +User-provided 0922 (`notes.md` §117 / items 330–336 / batch #100). SemIf was formerly OpenJev; MLX backend; 5.21× systems≠semantic; Softmax over options ≠ calibrated Noul; live REST 2282★. Do not reopen #23–#36/#38; do not push onto open #39/#40. +User-provided 0922 uniqueness lock: SemIf was formerly OpenJev; independent; not affiliated with Jev or TypeSafe; homepage openjev.com; default master; MIT; HEAD ca3ba65f1429; Tolerate float roundoff in MLX evidence verification; pushed 2026-09-19; live REST 2282★ / 140 forks; size 9177; README SHA 74ab7f7f; LICENSE SHA ca562883; interface pattern reproduction with open models; does not reproduce Jev undisclosed model/training; Direct option logits; 0 output tokens; shared-state parallel; MLX backend for Apple Silicon (`--backend mlx`); Qwen3.5-4B 3090 direct 1.023s vs AR JSON 5.332s (**5.21×**); argmax agree 18/21; systems comparison ≠ semantic equivalence; Parallel suffixes 20.03 dec/s on 777 decisions; Browser ladder Qwen3.5-4B authored BA 0.813, pert 0.766, TypeSafe subset agreement 0.845 vs Published Jev 0.883 (102 across 20 cases); Softmax over options ≠ calibrated Noul; typed output does not guarantee semantic correctness; wire/agreement ≠ replica of TypeSafe; SemIf ≠ kw2828/OpenJev playground ≠ zhihz/openjev ≠ apiplant/semif-rs port ≠ dddanielliu/semif-serve; rename is densify not a second census; JevBench 74.6 is §78 not this ladder; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33/#34/#35/#36/#38; do not push onto open #39/#40; notes.md §117 diff --git a/.agents/skills/augustus/references/toolbox-mapping.md b/.agents/skills/augustus/references/toolbox-mapping.md index 223f42e..176e8e2 100644 --- a/.agents/skills/augustus/references/toolbox-mapping.md +++ b/.agents/skills/augustus/references/toolbox-mapping.md @@ -348,3 +348,6 @@ User-provided 0915 uniqueness lock: TianyuCodings/NanoJev unified-games-v1 densi **0920 jcr sweep:** judgment-shaped hole is *which documented operation fits this request* on a tree the code already holds. Substitute hierarchical Choice + geometric-mean beam; keep execution exact. NiazMorshed2007/jcr. `notes.md` §116. User-provided 0920 jcr uniqueness lock: NiazMorshed2007/jcr MIT; site https://jcr.niazmorshed.dev; topics ai-agents,jev,mcp; **4★**; HEAD `138b3832`; README SHA `2a49dbc1`; LICENSE SHA `46231303`; size **14850**; Jev Capability Resolver; one tool to find documented deterministic commands in a nested capability tree; returns context; **does not execute**; skills = workflow+judgment; capabilities = individual operations; format independent of Jev; proposed open standard exploration; classify (Jev) → optional OpenAI decompose compound → beam search geometric mean of routing probs; keep up to 3 paths ≥60% of best (JCR_BAND_RATIO 0.6); ambiguity / no-match / depth-limit explicit; soft scores ≠ hard gates; 0.6 band is application policy; routing ≠ permission; docs ≠ authority to run; sol-vs-opus5-20 *theirs*: 20 scenarios × 4 variants = 80 runs; lookup+explain only, no execution; Claude Opus 5: agent input 108,585→15,819 (−85%), cost $0.3700→$0.1222 (−67%), wall 105.5s→77.7s; Codex GPT-5.6-Sol: 61,952→47,669 (−23%), $0.1377→$0.1151 (−16%), wall 25.3s→62.4s (Sol slower with JCR in 19/20); One Sol outlier 372.6s / 193 Jev calls; n=1 per cell; Not Harbor task-execution; Claude/Codex harnesses; compare mode; 50 scenarios bundled; 11 groups, 960 nodes, 11,360 items; 16 routing rounds per step; NiazMorshed2007/jcr ≠ skill-broker ≠ skillranker ≠ jev-sift ≠ jev-lens ≠ jevusher ≠ jev_select_capability; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33/#34; notes.md §116 + +User-provided 0922 (`notes.md` §117 / items 330–336 / batch #100). SemIf was formerly OpenJev; MLX backend; 5.21× systems≠semantic; Softmax over options ≠ calibrated Noul; live REST 2282★. Do not reopen #23–#36/#38; do not push onto open #39/#40. +User-provided 0922 uniqueness lock: SemIf was formerly OpenJev; independent; not affiliated with Jev or TypeSafe; homepage openjev.com; default master; MIT; HEAD ca3ba65f1429; Tolerate float roundoff in MLX evidence verification; pushed 2026-09-19; live REST 2282★ / 140 forks; size 9177; README SHA 74ab7f7f; LICENSE SHA ca562883; interface pattern reproduction with open models; does not reproduce Jev undisclosed model/training; Direct option logits; 0 output tokens; shared-state parallel; MLX backend for Apple Silicon (`--backend mlx`); Qwen3.5-4B 3090 direct 1.023s vs AR JSON 5.332s (**5.21×**); argmax agree 18/21; systems comparison ≠ semantic equivalence; Parallel suffixes 20.03 dec/s on 777 decisions; Browser ladder Qwen3.5-4B authored BA 0.813, pert 0.766, TypeSafe subset agreement 0.845 vs Published Jev 0.883 (102 across 20 cases); Softmax over options ≠ calibrated Noul; typed output does not guarantee semantic correctness; wire/agreement ≠ replica of TypeSafe; SemIf ≠ kw2828/OpenJev playground ≠ zhihz/openjev ≠ apiplant/semif-rs port ≠ dddanielliu/semif-serve; rename is densify not a second census; JevBench 74.6 is §78 not this ladder; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33/#34/#35/#36/#38; do not push onto open #39/#40; notes.md §117 diff --git a/.agents/skills/augustus/references/validation.md b/.agents/skills/augustus/references/validation.md index 6fa6554..67971dc 100644 --- a/.agents/skills/augustus/references/validation.md +++ b/.agents/skills/augustus/references/validation.md @@ -1149,3 +1149,6 @@ User-provided 0806 uniqueness lock: people who compare Jev against GPT-5.6 has n Hourly 0843 uniqueness lock: A hunch is a probability with a policy attached; { enter: 0.8, exit: 0.6 } is hysteresis; replay a policy change without inference; Decision models are providers, not the product; huncho ≠ Kungie/gut ≠ carldaws/hunch ≠ tpellet/hunch; pretrained Qwen2.5 base ECE 0.030 (0.5B) / 0.040 (7B); instruct 0.302 / 0.269; 70.9% → 70.0% mean conf 74.1% → 96.7%; temperature scaling still matches it in-distribution; No Jev API was called; Qwen2.5 ≠ Archer; Qwen/Qwen3.8-27B ≠ Archer; 学習済みモデル v0.1 は準備中です; bool AUROC 0.523; 先頭だと0件、末尾だと250件; 温度を渡さない場合、確率は較正されていません; このリポジトリには Jev を呼ぶコードが存在しません; g0runmezadam/what-is-jev IS tunahansahin897/what-is-jev (same GitHub id 1378007307); 947 repos scored; A 273 · B 302 · C 372; LLM rubric ≠ benches; Data as of 2026-09-20; HEAD 895b9498; README SHA 3ae98c56; 13 focused checks and one mutually exclusive outcome; Probabilities are advisory, not calibrated guarantees; omni-/ask-jev ≠ pedroknigge/mcp_jev; pd.cut bins by equal width while jeval bins by quantile; ECE 0.113 and ECE 0.076; jeval drift is not implemented yet; rlaope/jeval ≠ dayhaysoos/jevals; calibration does not compose; ECE has exactly zero statistical power to detect the failure mode that kills trajectories; 25–60× headline withdrawn; P(all-correct): 0.0071 vs 0.0001; TCE / AMS; Qwen 3.8 sparring ≠ Archer; Deferred Crispification; light_cutoff_applied_to_combination 0; recorded run, kinematic animation; BANKING77 Accuracy BERT-Base 93.02 Jev 79.90; Analyse jev calibration (NLL, ECE) backlog; BERT figures are published supervised references, not zero-shot; 档位措辞效应 分数极差中位 0.50、最大 1.32; 修好后对照组是 0.01; 不是 benchmark; 概率没做 calibration; ~1,430 API calls, about $0.15; xiaohuaxi/jev-study ≠ wjdjdakf17/jev-study ≠ baekenough/jev-study; AND: product (independence assumed and recorded in the trace); chat model's stated confidence is not calibrated; circuit-vl-4b ≠ Archer; Bring your own API key; vamsikrishna2421/jev-usecases ≠ whyashthakker/awesome-jev-use-cases; catalog ≠ endorsement; SemIf 2237★ (+30 vs §111 2207); jevlike 1049★ (+6 vs 1043); TypeLLM/TypeLLM 16★; AnotiaWang 98★; yibie/awesome-jev 520★ (+14 vs 506); Laya likes 861 (was 822); tracker likes 66 (+2 vs 64) lastModified UNCHANGED; Blackwood likes 2 gated manual; Archer still promised_not_landed; Hub archerhume/4rcherhume HTTP 401; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33; notes.md §114 User-provided 0920 jcr uniqueness lock: NiazMorshed2007/jcr MIT; site https://jcr.niazmorshed.dev; topics ai-agents,jev,mcp; **4★**; HEAD `138b3832`; README SHA `2a49dbc1`; LICENSE SHA `46231303`; size **14850**; Jev Capability Resolver; one tool to find documented deterministic commands in a nested capability tree; returns context; **does not execute**; skills = workflow+judgment; capabilities = individual operations; format independent of Jev; proposed open standard exploration; classify (Jev) → optional OpenAI decompose compound → beam search geometric mean of routing probs; keep up to 3 paths ≥60% of best (JCR_BAND_RATIO 0.6); ambiguity / no-match / depth-limit explicit; soft scores ≠ hard gates; 0.6 band is application policy; routing ≠ permission; docs ≠ authority to run; sol-vs-opus5-20 *theirs*: 20 scenarios × 4 variants = 80 runs; lookup+explain only, no execution; Claude Opus 5: agent input 108,585→15,819 (−85%), cost $0.3700→$0.1222 (−67%), wall 105.5s→77.7s; Codex GPT-5.6-Sol: 61,952→47,669 (−23%), $0.1377→$0.1151 (−16%), wall 25.3s→62.4s (Sol slower with JCR in 19/20); One Sol outlier 372.6s / 193 Jev calls; n=1 per cell; Not Harbor task-execution; Claude/Codex harnesses; compare mode; 50 scenarios bundled; 11 groups, 960 nodes, 11,360 items; 16 routing rounds per step; NiazMorshed2007/jcr ≠ skill-broker ≠ skillranker ≠ jev-sift ≠ jev-lens ≠ jevusher ≠ jev_select_capability; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33/#34; notes.md §116 + +User-provided 0922 (`notes.md` §117 / items 330–336 / batch #100). SemIf was formerly OpenJev; MLX backend; 5.21× systems≠semantic; Softmax over options ≠ calibrated Noul; live REST 2282★. Do not reopen #23–#36/#38; do not push onto open #39/#40. +User-provided 0922 uniqueness lock: SemIf was formerly OpenJev; independent; not affiliated with Jev or TypeSafe; homepage openjev.com; default master; MIT; HEAD ca3ba65f1429; Tolerate float roundoff in MLX evidence verification; pushed 2026-09-19; live REST 2282★ / 140 forks; size 9177; README SHA 74ab7f7f; LICENSE SHA ca562883; interface pattern reproduction with open models; does not reproduce Jev undisclosed model/training; Direct option logits; 0 output tokens; shared-state parallel; MLX backend for Apple Silicon (`--backend mlx`); Qwen3.5-4B 3090 direct 1.023s vs AR JSON 5.332s (**5.21×**); argmax agree 18/21; systems comparison ≠ semantic equivalence; Parallel suffixes 20.03 dec/s on 777 decisions; Browser ladder Qwen3.5-4B authored BA 0.813, pert 0.766, TypeSafe subset agreement 0.845 vs Published Jev 0.883 (102 across 20 cases); Softmax over options ≠ calibrated Noul; typed output does not guarantee semantic correctness; wire/agreement ≠ replica of TypeSafe; SemIf ≠ kw2828/OpenJev playground ≠ zhihz/openjev ≠ apiplant/semif-rs port ≠ dddanielliu/semif-serve; rename is densify not a second census; JevBench 74.6 is §78 not this ladder; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33/#34/#35/#36/#38; do not push onto open #39/#40; notes.md §117 diff --git a/.agents/skills/augustus/scripts/uniqueness_gate.py b/.agents/skills/augustus/scripts/uniqueness_gate.py index 6a15897..41c48da 100644 --- a/.agents/skills/augustus/scripts/uniqueness_gate.py +++ b/.agents/skills/augustus/scripts/uniqueness_gate.py @@ -1,13 +1,13 @@ #!/usr/bin/env python3 """Uniqueness gate for merged 0843 (§114), merged 0915 NanoJev (§115), -and user-provided 0920 jcr (§116). +user-provided 0920 jcr (§116), and user-provided 0922 SemIf (§117). Each lock must appear as one consecutive substring in every listed overlay. Fragments scattered across files do not count. -Also: YAML-parse SKILL.md frontmatter; notes.md owns §114, §115, and §116; -composition items 289–302, 303–308, and 309–316 exist; findings batches -#97, #98, and #99 exist. Does not fetch the network. Does not treat a lock +Also: YAML-parse SKILL.md frontmatter; notes.md owns §114, §115, §116, and §117; +composition items 289–316 and 330–336 exist; findings batches +#97, #98, #99, and #100 exist. §118 and items 317–329 are reserved. Does not fetch the network. Does not treat a lock as a Harbor score. """ @@ -30,6 +30,11 @@ 'User-provided 0920 jcr uniqueness lock: NiazMorshed2007/jcr MIT; site https://jcr.niazmorshed.dev; topics ai-agents,jev,mcp; **4★**; HEAD `138b3832`; README SHA `2a49dbc1`; LICENSE SHA `46231303`; size **14850**; Jev Capability Resolver; one tool to find documented deterministic commands in a nested capability tree; returns context; **does not execute**; skills = workflow+judgment; capabilities = individual operations; format independent of Jev; proposed open standard exploration; classify (Jev) → optional OpenAI decompose compound → beam search geometric mean of routing probs; keep up to 3 paths ≥60% of best (JCR_BAND_RATIO 0.6); ambiguity / no-match / depth-limit explicit; soft scores ≠ hard gates; 0.6 band is application policy; routing ≠ permission; docs ≠ authority to run; sol-vs-opus5-20 *theirs*: 20 scenarios × 4 variants = 80 runs; lookup+explain only, no execution; Claude Opus 5: agent input 108,585→15,819 (−85%), cost $0.3700→$0.1222 (−67%), wall 105.5s→77.7s; Codex GPT-5.6-Sol: 61,952→47,669 (−23%), $0.1377→$0.1151 (−16%), wall 25.3s→62.4s (Sol slower with JCR in 19/20); One Sol outlier 372.6s / 193 Jev calls; n=1 per cell; Not Harbor task-execution; Claude/Codex harnesses; compare mode; 50 scenarios bundled; 11 groups, 960 nodes, 11,360 items; 16 routing rounds per step; NiazMorshed2007/jcr ≠ skill-broker ≠ skillranker ≠ jev-sift ≠ jev-lens ≠ jevusher ≠ jev_select_capability; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33/#34; notes.md §116' ) + +UNIQ_0922 = ( +'User-provided 0922 uniqueness lock: SemIf was formerly OpenJev; independent; not affiliated with Jev or TypeSafe; homepage openjev.com; default master; MIT; HEAD ca3ba65f1429; Tolerate float roundoff in MLX evidence verification; pushed 2026-09-19; live REST 2282★ / 140 forks; size 9177; README SHA 74ab7f7f; LICENSE SHA ca562883; interface pattern reproduction with open models; does not reproduce Jev undisclosed model/training; Direct option logits; 0 output tokens; shared-state parallel; MLX backend for Apple Silicon (`--backend mlx`); Qwen3.5-4B 3090 direct 1.023s vs AR JSON 5.332s (**5.21×**); argmax agree 18/21; systems comparison ≠ semantic equivalence; Parallel suffixes 20.03 dec/s on 777 decisions; Browser ladder Qwen3.5-4B authored BA 0.813, pert 0.766, TypeSafe subset agreement 0.845 vs Published Jev 0.883 (102 across 20 cases); Softmax over options ≠ calibrated Noul; typed output does not guarantee semantic correctness; wire/agreement ≠ replica of TypeSafe; SemIf ≠ kw2828/OpenJev playground ≠ zhihz/openjev ≠ apiplant/semif-rs port ≠ dddanielliu/semif-serve; rename is densify not a second census; JevBench 74.6 is §78 not this ladder; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33/#34/#35/#36/#38; do not push onto open #39/#40; notes.md §117' +) + OVERLAYS = [ "research/notes.md", "research/changelog-hourly.md", @@ -81,6 +86,8 @@ def main() -> int: failed.append(f"0915 lock missing as one substring: {rel}") if UNIQ_JCR not in body: failed.append(f"jcr lock missing as one substring: {rel}") + if UNIQ_0922 not in body: + failed.append(f"0922 lock missing as one substring: {rel}") notes = (ROOT / "research/notes.md").read_text(encoding="utf-8") if "## 114. Hourly 0843 HIGH" not in notes: failed.append("notes.md missing §114 heading") @@ -88,15 +95,23 @@ def main() -> int: failed.append("notes.md missing §115 heading") if "## 116. User-provided HIGH — NiazMorshed2007/jcr" not in notes: failed.append("notes.md missing §116 heading") + if "## 117. User-provided HIGH" not in notes: + failed.append("notes.md missing §117 heading") + if "## 118." in notes: + failed.append("notes.md stole reserved §118") algebra = (ROOT / ".agents/skills/augustus/references/composition-algebra.md").read_text( encoding="utf-8" ) - for n in list(range(289, 303)) + list(range(303, 309)) + list(range(309, 317)): + for n in list(range(289, 317)) + list(range(330, 337)): needle = f"{n}. **" if needle not in algebra: failed.append(f"composition-algebra missing item {n}") + for n in range(317, 330): + needle = f"{n}. **" + if needle in algebra: + failed.append(f"composition-algebra stole reserved item {n}") findings = (ROOT / "research/archive/findings.md").read_text(encoding="utf-8") - for batch in ("## Batch #97", "## Batch #98", "## Batch #99"): + for batch in ("## Batch #97", "## Batch #98", "## Batch #99", "## Batch #100"): if batch not in findings: failed.append(f"findings.md missing {batch}") skill = (ROOT / ".agents/skills/augustus/SKILL.md").read_text(encoding="utf-8") @@ -120,6 +135,11 @@ def main() -> int: "docs ≠ authority to run", "Not Harbor task-execution", "TianyuCodings/NanoJev", + "SemIf was formerly OpenJev", + "systems comparison ≠ semantic equivalence", + "Softmax over options ≠ calibrated Noul", + "live REST 2282★", + "JevBench 74.6 is §78 not this ladder", ): if frag not in combined: failed.append(f"SKILL.md missing fragment {frag!r}") @@ -131,7 +151,7 @@ def main() -> int: print("uniqueness-gate ok") print( f"0843 chars={len(UNIQ_0843)} 0915 chars={len(UNIQ_0915)} " - f"jcr chars={len(UNIQ_JCR)} overlays={len(OVERLAYS)}" + f"jcr chars={len(UNIQ_JCR)} lock0922 chars={len(UNIQ_0922)} overlays={len(OVERLAYS)}" ) return 0 diff --git a/CHANGELOG.md b/CHANGELOG.md index 588622d..ea082c7 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -16,6 +16,39 @@ folds: `research/notes.md`. ## [Unreleased] +User-provided 0922 HIGH (`research/notes.md` §117 / composition +items 330–336 / findings batch #100). SemIf rename + MLX + +accuracy-ladder densify onto latest main after merged **#38** +(jcr / §116). Does **not** bump the 0.4.0 pin. +Do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33/#34/#35/#36/#38. +Do not push onto open #39/#40. Merged #38 owns §116 / 309–316 / #99. +Merged #35 owns §114 / 289–302 / #97. Merged #36 owns §115 / +303–308 / #98 — leave them alone. Open #40 claims §118 / 322–329 / #101. + +### Added + +- **SemIf densify (PRIMARY, `notes.md` §117).** [TheoLeeCJ/SemIf](https://github.com/TheoLeeCJ/SemIf) + MIT; homepage openjev.com; default **master**; live REST + **2282★** / **140** forks; HEAD `ca3ba65f1429` (Tolerate float + roundoff in MLX evidence verification, 2026-09-19). SemIf was + formerly OpenJev; independent; not affiliated with Jev or + TypeSafe. Interface pattern reproduction with open models; + does not reproduce Jev undisclosed model/training. Direct option + logits; 0 output tokens; shared-state parallel; MLX backend + (`--backend mlx`). Speed *theirs* Qwen3.5-4B 3090: direct + 1.023s vs AR JSON 5.332s (**5.21×**); argmax agree 18/21 + (systems comparison ≠ semantic equivalence). Parallel suffixes + 20.03 dec/s on 777. Browser ladder *theirs*: authored BA 0.813, + pert 0.766, TypeSafe subset 0.845 vs Published Jev 0.883 + (102 across 20 cases). Softmax over options ≠ calibrated Noul; + wire/agreement ≠ replica. SemIf ≠ kw2828/OpenJev playground ≠ + zhihz/openjev ≠ apiplant/semif-rs ≠ dddanielliu/semif-serve. + Rename is densify not a second census. JevBench 74.6 is §78 + not this ladder. Skip Archer rewrite. Quote *theirs*. + `invented_signal: false`. Composition items 330–336 / batch #100. + Uniqueness dump in [`research/changelog-hourly.md`](research/changelog-hourly.md). + + Hourly 0843 HIGH (`research/notes.md` §114 / composition items 289–302 / findings batch #97) plus merged #36 NanoJev (`research/notes.md` §115 / items 303–308 / batch #98) plus diff --git a/README.md b/README.md index 0479a65..ca130db 100644 --- a/README.md +++ b/README.md @@ -125,3 +125,4 @@ MIT. See [LICENSE](LICENSE). Security reports: [SECURITY.md](SECURITY.md). Contributions: [CONTRIBUTING.md](CONTRIBUTING.md). User-provided 0920 jcr uniqueness lock: NiazMorshed2007/jcr MIT; site https://jcr.niazmorshed.dev; topics ai-agents,jev,mcp; **4★**; HEAD `138b3832`; README SHA `2a49dbc1`; LICENSE SHA `46231303`; size **14850**; Jev Capability Resolver; one tool to find documented deterministic commands in a nested capability tree; returns context; **does not execute**; skills = workflow+judgment; capabilities = individual operations; format independent of Jev; proposed open standard exploration; classify (Jev) → optional OpenAI decompose compound → beam search geometric mean of routing probs; keep up to 3 paths ≥60% of best (JCR_BAND_RATIO 0.6); ambiguity / no-match / depth-limit explicit; soft scores ≠ hard gates; 0.6 band is application policy; routing ≠ permission; docs ≠ authority to run; sol-vs-opus5-20 *theirs*: 20 scenarios × 4 variants = 80 runs; lookup+explain only, no execution; Claude Opus 5: agent input 108,585→15,819 (−85%), cost $0.3700→$0.1222 (−67%), wall 105.5s→77.7s; Codex GPT-5.6-Sol: 61,952→47,669 (−23%), $0.1377→$0.1151 (−16%), wall 25.3s→62.4s (Sol slower with JCR in 19/20); One Sol outlier 372.6s / 193 Jev calls; n=1 per cell; Not Harbor task-execution; Claude/Codex harnesses; compare mode; 50 scenarios bundled; 11 groups, 960 nodes, 11,360 items; 16 routing rounds per step; NiazMorshed2007/jcr ≠ skill-broker ≠ skillranker ≠ jev-sift ≠ jev-lens ≠ jevusher ≠ jev_select_capability; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33/#34; notes.md §116 +User-provided 0922 uniqueness lock: SemIf was formerly OpenJev; independent; not affiliated with Jev or TypeSafe; homepage openjev.com; default master; MIT; HEAD ca3ba65f1429; Tolerate float roundoff in MLX evidence verification; pushed 2026-09-19; live REST 2282★ / 140 forks; size 9177; README SHA 74ab7f7f; LICENSE SHA ca562883; interface pattern reproduction with open models; does not reproduce Jev undisclosed model/training; Direct option logits; 0 output tokens; shared-state parallel; MLX backend for Apple Silicon (`--backend mlx`); Qwen3.5-4B 3090 direct 1.023s vs AR JSON 5.332s (**5.21×**); argmax agree 18/21; systems comparison ≠ semantic equivalence; Parallel suffixes 20.03 dec/s on 777 decisions; Browser ladder Qwen3.5-4B authored BA 0.813, pert 0.766, TypeSafe subset agreement 0.845 vs Published Jev 0.883 (102 across 20 cases); Softmax over options ≠ calibrated Noul; typed output does not guarantee semantic correctness; wire/agreement ≠ replica of TypeSafe; SemIf ≠ kw2828/OpenJev playground ≠ zhihz/openjev ≠ apiplant/semif-rs port ≠ dddanielliu/semif-serve; rename is densify not a second census; JevBench 74.6 is §78 not this ladder; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33/#34/#35/#36/#38; do not push onto open #39/#40; notes.md §117 diff --git a/docs/ecosystem.md b/docs/ecosystem.md index c6d48f9..d8f5dc3 100644 --- a/docs/ecosystem.md +++ b/docs/ecosystem.md @@ -1116,4 +1116,16 @@ question-design, validation, and optimizer coupling. Hourly 0843 uniqueness lock: A hunch is a probability with a policy attached; { enter: 0.8, exit: 0.6 } is hysteresis; replay a policy change without inference; Decision models are providers, not the product; huncho ≠ Kungie/gut ≠ carldaws/hunch ≠ tpellet/hunch; pretrained Qwen2.5 base ECE 0.030 (0.5B) / 0.040 (7B); instruct 0.302 / 0.269; 70.9% → 70.0% mean conf 74.1% → 96.7%; temperature scaling still matches it in-distribution; No Jev API was called; Qwen2.5 ≠ Archer; Qwen/Qwen3.8-27B ≠ Archer; 学習済みモデル v0.1 は準備中です; bool AUROC 0.523; 先頭だと0件、末尾だと250件; 温度を渡さない場合、確率は較正されていません; このリポジトリには Jev を呼ぶコードが存在しません; g0runmezadam/what-is-jev IS tunahansahin897/what-is-jev (same GitHub id 1378007307); 947 repos scored; A 273 · B 302 · C 372; LLM rubric ≠ benches; Data as of 2026-09-20; HEAD 895b9498; README SHA 3ae98c56; 13 focused checks and one mutually exclusive outcome; Probabilities are advisory, not calibrated guarantees; omni-/ask-jev ≠ pedroknigge/mcp_jev; pd.cut bins by equal width while jeval bins by quantile; ECE 0.113 and ECE 0.076; jeval drift is not implemented yet; rlaope/jeval ≠ dayhaysoos/jevals; calibration does not compose; ECE has exactly zero statistical power to detect the failure mode that kills trajectories; 25–60× headline withdrawn; P(all-correct): 0.0071 vs 0.0001; TCE / AMS; Qwen 3.8 sparring ≠ Archer; Deferred Crispification; light_cutoff_applied_to_combination 0; recorded run, kinematic animation; BANKING77 Accuracy BERT-Base 93.02 Jev 79.90; Analyse jev calibration (NLL, ECE) backlog; BERT figures are published supervised references, not zero-shot; 档位措辞效应 分数极差中位 0.50、最大 1.32; 修好后对照组是 0.01; 不是 benchmark; 概率没做 calibration; ~1,430 API calls, about $0.15; xiaohuaxi/jev-study ≠ wjdjdakf17/jev-study ≠ baekenough/jev-study; AND: product (independence assumed and recorded in the trace); chat model's stated confidence is not calibrated; circuit-vl-4b ≠ Archer; Bring your own API key; vamsikrishna2421/jev-usecases ≠ whyashthakker/awesome-jev-use-cases; catalog ≠ endorsement; SemIf 2237★ (+30 vs §111 2207); jevlike 1049★ (+6 vs 1043); TypeLLM/TypeLLM 16★; AnotiaWang 98★; yibie/awesome-jev 520★ (+14 vs 506); Laya likes 861 (was 822); tracker likes 66 (+2 vs 64) lastModified UNCHANGED; Blackwood likes 2 gated manual; Archer still promised_not_landed; Hub archerhume/4rcherhume HTTP 401; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33; notes.md §114 + +### SemIf rename + MLX + ladder (`notes.md` §117) + +[TheoLeeCJ/SemIf](https://github.com/TheoLeeCJ/SemIf) densify (not a first +sighting). SemIf was formerly OpenJev; independent; not affiliated with Jev +or TypeSafe. Direct option logits; 0 output tokens; MLX backend +(`--backend mlx`). Speed *theirs* 5.21× (1.023s vs 5.332s); argmax 18/21 — +systems comparison ≠ semantic equivalence. Browser ladder authored BA 0.813 / +subset 0.845 vs Jev 0.883 *theirs*. Softmax over options ≠ calibrated Noul. +Live REST **2282★**. JevBench 74.6 is §78 not this ladder. + User-provided 0920 jcr uniqueness lock: NiazMorshed2007/jcr MIT; site https://jcr.niazmorshed.dev; topics ai-agents,jev,mcp; **4★**; HEAD `138b3832`; README SHA `2a49dbc1`; LICENSE SHA `46231303`; size **14850**; Jev Capability Resolver; one tool to find documented deterministic commands in a nested capability tree; returns context; **does not execute**; skills = workflow+judgment; capabilities = individual operations; format independent of Jev; proposed open standard exploration; classify (Jev) → optional OpenAI decompose compound → beam search geometric mean of routing probs; keep up to 3 paths ≥60% of best (JCR_BAND_RATIO 0.6); ambiguity / no-match / depth-limit explicit; soft scores ≠ hard gates; 0.6 band is application policy; routing ≠ permission; docs ≠ authority to run; sol-vs-opus5-20 *theirs*: 20 scenarios × 4 variants = 80 runs; lookup+explain only, no execution; Claude Opus 5: agent input 108,585→15,819 (−85%), cost $0.3700→$0.1222 (−67%), wall 105.5s→77.7s; Codex GPT-5.6-Sol: 61,952→47,669 (−23%), $0.1377→$0.1151 (−16%), wall 25.3s→62.4s (Sol slower with JCR in 19/20); One Sol outlier 372.6s / 193 Jev calls; n=1 per cell; Not Harbor task-execution; Claude/Codex harnesses; compare mode; 50 scenarios bundled; 11 groups, 960 nodes, 11,360 items; 16 routing rounds per step; NiazMorshed2007/jcr ≠ skill-broker ≠ skillranker ≠ jev-sift ≠ jev-lens ≠ jevusher ≠ jev_select_capability; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33/#34; notes.md §116 +User-provided 0922 uniqueness lock: SemIf was formerly OpenJev; independent; not affiliated with Jev or TypeSafe; homepage openjev.com; default master; MIT; HEAD ca3ba65f1429; Tolerate float roundoff in MLX evidence verification; pushed 2026-09-19; live REST 2282★ / 140 forks; size 9177; README SHA 74ab7f7f; LICENSE SHA ca562883; interface pattern reproduction with open models; does not reproduce Jev undisclosed model/training; Direct option logits; 0 output tokens; shared-state parallel; MLX backend for Apple Silicon (`--backend mlx`); Qwen3.5-4B 3090 direct 1.023s vs AR JSON 5.332s (**5.21×**); argmax agree 18/21; systems comparison ≠ semantic equivalence; Parallel suffixes 20.03 dec/s on 777 decisions; Browser ladder Qwen3.5-4B authored BA 0.813, pert 0.766, TypeSafe subset agreement 0.845 vs Published Jev 0.883 (102 across 20 cases); Softmax over options ≠ calibrated Noul; typed output does not guarantee semantic correctness; wire/agreement ≠ replica of TypeSafe; SemIf ≠ kw2828/OpenJev playground ≠ zhihz/openjev ≠ apiplant/semif-rs port ≠ dddanielliu/semif-serve; rename is densify not a second census; JevBench 74.6 is §78 not this ladder; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33/#34/#35/#36/#38; do not push onto open #39/#40; notes.md §117 diff --git a/research/archive/findings.md b/research/archive/findings.md index f9eccfd..2a0271a 100644 --- a/research/archive/findings.md +++ b/research/archive/findings.md @@ -1,9 +1,44 @@ # Deep-read findings (evidence for research/notes.md) + +## Batch #100 (2026-09-20 ~15:22 UTC / ~09:22 Boise) — user-provided SemIf densify + +User-provided 0922 uniqueness lock: SemIf was formerly OpenJev; independent; not affiliated with Jev or TypeSafe; homepage openjev.com; default master; MIT; HEAD ca3ba65f1429; Tolerate float roundoff in MLX evidence verification; pushed 2026-09-19; live REST 2282★ / 140 forks; size 9177; README SHA 74ab7f7f; LICENSE SHA ca562883; interface pattern reproduction with open models; does not reproduce Jev undisclosed model/training; Direct option logits; 0 output tokens; shared-state parallel; MLX backend for Apple Silicon (`--backend mlx`); Qwen3.5-4B 3090 direct 1.023s vs AR JSON 5.332s (**5.21×**); argmax agree 18/21; systems comparison ≠ semantic equivalence; Parallel suffixes 20.03 dec/s on 777 decisions; Browser ladder Qwen3.5-4B authored BA 0.813, pert 0.766, TypeSafe subset agreement 0.845 vs Published Jev 0.883 (102 across 20 cases); Softmax over options ≠ calibrated Noul; typed output does not guarantee semantic correctness; wire/agreement ≠ replica of TypeSafe; SemIf ≠ kw2828/OpenJev playground ≠ zhihz/openjev ≠ apiplant/semif-rs port ≠ dddanielliu/semif-serve; rename is densify not a second census; JevBench 74.6 is §78 not this ladder; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33/#34/#35/#36/#38; do not push onto open #39/#40; notes.md §117 + +Note: `research/notes.md` §117. Docs-only reconstructed onto latest main +after merged #34 (Pages), merged #35 (hourly 0843), merged #36 (NanoJev), +and merged #38 (jcr). Merged #35 owns `notes.md` §114 / items 289–302 / +batch #97. Merged #36 owns `notes.md` §115 / items 303–308 / batch #98. +Merged #38 owns `notes.md` §116 / items 309–316 / batch #99. This fold is +§117 / items 330–336 / batch #100. ID skip §118 / items 322–329 / batch #101 +for open #40 llm-to-jev. +**HARD RULE:** do not reopen or amend PR #23–#36 or #38. +Do not push onto open #39/#40. Never reopen merged #7–**#36** / **#38**. +Do **not** re-fold §78 JevBench 74.6 / §69 semif-serve / §113 census / +§114 hourly / §115 NanoJev / §116 jcr. Skip Archer rewrite. No invented +metrics. Hunches labeled. Quote READMEs. Soft Noul ≠ hard safety. Augustus +owns placement. `invented_signal: false`. + +- **TheoLeeCJ/SemIf (PRIMARY densify).** Python MIT; **2282★** / **140** forks; + HEAD `ca3ba65f1429`; size **9177**; homepage openjev.com; default master. + SemIf was formerly OpenJev. independent; not affiliated with Jev or TypeSafe. + interface pattern reproduction with open models; does not reproduce Jev undisclosed model/training. + Direct option logits; 0 output tokens; shared-state parallel; + MLX backend for Apple Silicon (`--backend mlx`). + Qwen3.5-4B 3090 direct 1.023s vs AR JSON 5.332s (**5.21×**); argmax agree 18/21; + systems comparison ≠ semantic equivalence. Parallel suffixes 20.03 dec/s on 777 decisions. + authored BA 0.813, pert 0.766, TypeSafe subset agreement 0.845 vs Published Jev 0.883 (102 across 20 cases). + Softmax over options ≠ calibrated Noul; typed output does not guarantee semantic correctness; + wire/agreement ≠ replica of TypeSafe. + SemIf ≠ kw2828/OpenJev playground ≠ zhihz/openjev ≠ apiplant/semif-rs port ≠ dddanielliu/semif-serve. + JevBench 74.6 is §78 not this ladder. rename is densify not a second census. + + ## Batch #99 (2026-09-20 ~15:20 UTC / ~09:20 Boise) — user-provided HIGH NiazMorshed2007/jcr User-provided 0920 jcr uniqueness lock: NiazMorshed2007/jcr MIT; site https://jcr.niazmorshed.dev; topics ai-agents,jev,mcp; **4★**; HEAD `138b3832`; README SHA `2a49dbc1`; LICENSE SHA `46231303`; size **14850**; Jev Capability Resolver; one tool to find documented deterministic commands in a nested capability tree; returns context; **does not execute**; skills = workflow+judgment; capabilities = individual operations; format independent of Jev; proposed open standard exploration; classify (Jev) → optional OpenAI decompose compound → beam search geometric mean of routing probs; keep up to 3 paths ≥60% of best (JCR_BAND_RATIO 0.6); ambiguity / no-match / depth-limit explicit; soft scores ≠ hard gates; 0.6 band is application policy; routing ≠ permission; docs ≠ authority to run; sol-vs-opus5-20 *theirs*: 20 scenarios × 4 variants = 80 runs; lookup+explain only, no execution; Claude Opus 5: agent input 108,585→15,819 (−85%), cost $0.3700→$0.1222 (−67%), wall 105.5s→77.7s; Codex GPT-5.6-Sol: 61,952→47,669 (−23%), $0.1377→$0.1151 (−16%), wall 25.3s→62.4s (Sol slower with JCR in 19/20); One Sol outlier 372.6s / 193 Jev calls; n=1 per cell; Not Harbor task-execution; Claude/Codex harnesses; compare mode; 50 scenarios bundled; 11 groups, 960 nodes, 11,360 items; 16 routing rounds per step; NiazMorshed2007/jcr ≠ skill-broker ≠ skillranker ≠ jev-sift ≠ jev-lens ≠ jevusher ≠ jev_select_capability; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33/#34; notes.md §116 +User-provided 0922 uniqueness lock: SemIf was formerly OpenJev; independent; not affiliated with Jev or TypeSafe; homepage openjev.com; default master; MIT; HEAD ca3ba65f1429; Tolerate float roundoff in MLX evidence verification; pushed 2026-09-19; live REST 2282★ / 140 forks; size 9177; README SHA 74ab7f7f; LICENSE SHA ca562883; interface pattern reproduction with open models; does not reproduce Jev undisclosed model/training; Direct option logits; 0 output tokens; shared-state parallel; MLX backend for Apple Silicon (`--backend mlx`); Qwen3.5-4B 3090 direct 1.023s vs AR JSON 5.332s (**5.21×**); argmax agree 18/21; systems comparison ≠ semantic equivalence; Parallel suffixes 20.03 dec/s on 777 decisions; Browser ladder Qwen3.5-4B authored BA 0.813, pert 0.766, TypeSafe subset agreement 0.845 vs Published Jev 0.883 (102 across 20 cases); Softmax over options ≠ calibrated Noul; typed output does not guarantee semantic correctness; wire/agreement ≠ replica of TypeSafe; SemIf ≠ kw2828/OpenJev playground ≠ zhihz/openjev ≠ apiplant/semif-rs port ≠ dddanielliu/semif-serve; rename is densify not a second census; JevBench 74.6 is §78 not this ladder; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33/#34/#35/#36/#38; do not push onto open #39/#40; notes.md §117 Note: `research/notes.md` §116. Docs-only, rebased onto merged #35 (`0189825` / §114 / 289–302 / #97) after diff --git a/research/changelog-hourly.md b/research/changelog-hourly.md index 36c54a0..ce5770a 100644 --- a/research/changelog-hourly.md +++ b/research/changelog-hourly.md @@ -1,4 +1,4 @@ -# Hourly uniqueness dump (pre-0.4.0 + 0743 + 0843 + 0915 + 0920 jcr) +# Hourly uniqueness dump (pre-0.4.0 + 0743 + 0843 + 0915 + 0920 jcr + 0922 SemIf) This is the pre-0.4.0 `CHANGELOG.md` after hourly folds (#2–#30 / notes §44–§112) stuffed uniqueness locks into Keep-a-Changelog sections, plus @@ -15,6 +15,18 @@ the merged **0743 HIGH** dump (PR #31 / notes.md §113 / items 273–288 --- +## User-provided 0922 HIGH (SemIf densify, notes.md §117) + +- Docs-only reconstructed onto latest main after merged #34/#35/#36/#38. + **HARD RULE:** do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33/#34/#35/#36/#38. + Do not push onto open #39/#40. Keep merged #38 §116 / items 309–316 / batch #99. + Skip §118 / 322–329 / #101. Not a first sighting. Rename + MLX + accuracy ladder densify. + Quote *theirs*. No wrappers. `invented_signal: false`. +- User-provided 0922 uniqueness lock: SemIf was formerly OpenJev; independent; not affiliated with Jev or TypeSafe; homepage openjev.com; default master; MIT; HEAD ca3ba65f1429; Tolerate float roundoff in MLX evidence verification; pushed 2026-09-19; live REST 2282★ / 140 forks; size 9177; README SHA 74ab7f7f; LICENSE SHA ca562883; interface pattern reproduction with open models; does not reproduce Jev undisclosed model/training; Direct option logits; 0 output tokens; shared-state parallel; MLX backend for Apple Silicon (`--backend mlx`); Qwen3.5-4B 3090 direct 1.023s vs AR JSON 5.332s (**5.21×**); argmax agree 18/21; systems comparison ≠ semantic equivalence; Parallel suffixes 20.03 dec/s on 777 decisions; Browser ladder Qwen3.5-4B authored BA 0.813, pert 0.766, TypeSafe subset agreement 0.845 vs Published Jev 0.883 (102 across 20 cases); Softmax over options ≠ calibrated Noul; typed output does not guarantee semantic correctness; wire/agreement ≠ replica of TypeSafe; SemIf ≠ kw2828/OpenJev playground ≠ zhihz/openjev ≠ apiplant/semif-rs port ≠ dddanielliu/semif-serve; rename is densify not a second census; JevBench 74.6 is §78 not this ladder; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33/#34/#35/#36/#38; do not push onto open #39/#40; notes.md §117 + +--- + + ## User-provided 0920 jcr HIGH (notes.md §116) - User-provided HIGH NiazMorshed2007/jcr (`research/notes.md` §116): @@ -3734,3 +3746,4 @@ The dated passes below are how 0.1.0 was assembled. Hourly 0843 uniqueness lock: A hunch is a probability with a policy attached; { enter: 0.8, exit: 0.6 } is hysteresis; replay a policy change without inference; Decision models are providers, not the product; huncho ≠ Kungie/gut ≠ carldaws/hunch ≠ tpellet/hunch; pretrained Qwen2.5 base ECE 0.030 (0.5B) / 0.040 (7B); instruct 0.302 / 0.269; 70.9% → 70.0% mean conf 74.1% → 96.7%; temperature scaling still matches it in-distribution; No Jev API was called; Qwen2.5 ≠ Archer; Qwen/Qwen3.8-27B ≠ Archer; 学習済みモデル v0.1 は準備中です; bool AUROC 0.523; 先頭だと0件、末尾だと250件; 温度を渡さない場合、確率は較正されていません; このリポジトリには Jev を呼ぶコードが存在しません; g0runmezadam/what-is-jev IS tunahansahin897/what-is-jev (same GitHub id 1378007307); 947 repos scored; A 273 · B 302 · C 372; LLM rubric ≠ benches; Data as of 2026-09-20; HEAD 895b9498; README SHA 3ae98c56; 13 focused checks and one mutually exclusive outcome; Probabilities are advisory, not calibrated guarantees; omni-/ask-jev ≠ pedroknigge/mcp_jev; pd.cut bins by equal width while jeval bins by quantile; ECE 0.113 and ECE 0.076; jeval drift is not implemented yet; rlaope/jeval ≠ dayhaysoos/jevals; calibration does not compose; ECE has exactly zero statistical power to detect the failure mode that kills trajectories; 25–60× headline withdrawn; P(all-correct): 0.0071 vs 0.0001; TCE / AMS; Qwen 3.8 sparring ≠ Archer; Deferred Crispification; light_cutoff_applied_to_combination 0; recorded run, kinematic animation; BANKING77 Accuracy BERT-Base 93.02 Jev 79.90; Analyse jev calibration (NLL, ECE) backlog; BERT figures are published supervised references, not zero-shot; 档位措辞效应 分数极差中位 0.50、最大 1.32; 修好后对照组是 0.01; 不是 benchmark; 概率没做 calibration; ~1,430 API calls, about $0.15; xiaohuaxi/jev-study ≠ wjdjdakf17/jev-study ≠ baekenough/jev-study; AND: product (independence assumed and recorded in the trace); chat model's stated confidence is not calibrated; circuit-vl-4b ≠ Archer; Bring your own API key; vamsikrishna2421/jev-usecases ≠ whyashthakker/awesome-jev-use-cases; catalog ≠ endorsement; SemIf 2237★ (+30 vs §111 2207); jevlike 1049★ (+6 vs 1043); TypeLLM/TypeLLM 16★; AnotiaWang 98★; yibie/awesome-jev 520★ (+14 vs 506); Laya likes 861 (was 822); tracker likes 66 (+2 vs 64) lastModified UNCHANGED; Blackwood likes 2 gated manual; Archer still promised_not_landed; Hub archerhume/4rcherhume HTTP 401; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33; notes.md §114 User-provided 0920 jcr uniqueness lock: NiazMorshed2007/jcr MIT; site https://jcr.niazmorshed.dev; topics ai-agents,jev,mcp; **4★**; HEAD `138b3832`; README SHA `2a49dbc1`; LICENSE SHA `46231303`; size **14850**; Jev Capability Resolver; one tool to find documented deterministic commands in a nested capability tree; returns context; **does not execute**; skills = workflow+judgment; capabilities = individual operations; format independent of Jev; proposed open standard exploration; classify (Jev) → optional OpenAI decompose compound → beam search geometric mean of routing probs; keep up to 3 paths ≥60% of best (JCR_BAND_RATIO 0.6); ambiguity / no-match / depth-limit explicit; soft scores ≠ hard gates; 0.6 band is application policy; routing ≠ permission; docs ≠ authority to run; sol-vs-opus5-20 *theirs*: 20 scenarios × 4 variants = 80 runs; lookup+explain only, no execution; Claude Opus 5: agent input 108,585→15,819 (−85%), cost $0.3700→$0.1222 (−67%), wall 105.5s→77.7s; Codex GPT-5.6-Sol: 61,952→47,669 (−23%), $0.1377→$0.1151 (−16%), wall 25.3s→62.4s (Sol slower with JCR in 19/20); One Sol outlier 372.6s / 193 Jev calls; n=1 per cell; Not Harbor task-execution; Claude/Codex harnesses; compare mode; 50 scenarios bundled; 11 groups, 960 nodes, 11,360 items; 16 routing rounds per step; NiazMorshed2007/jcr ≠ skill-broker ≠ skillranker ≠ jev-sift ≠ jev-lens ≠ jevusher ≠ jev_select_capability; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33/#34; notes.md §116 +User-provided 0922 uniqueness lock: SemIf was formerly OpenJev; independent; not affiliated with Jev or TypeSafe; homepage openjev.com; default master; MIT; HEAD ca3ba65f1429; Tolerate float roundoff in MLX evidence verification; pushed 2026-09-19; live REST 2282★ / 140 forks; size 9177; README SHA 74ab7f7f; LICENSE SHA ca562883; interface pattern reproduction with open models; does not reproduce Jev undisclosed model/training; Direct option logits; 0 output tokens; shared-state parallel; MLX backend for Apple Silicon (`--backend mlx`); Qwen3.5-4B 3090 direct 1.023s vs AR JSON 5.332s (**5.21×**); argmax agree 18/21; systems comparison ≠ semantic equivalence; Parallel suffixes 20.03 dec/s on 777 decisions; Browser ladder Qwen3.5-4B authored BA 0.813, pert 0.766, TypeSafe subset agreement 0.845 vs Published Jev 0.883 (102 across 20 cases); Softmax over options ≠ calibrated Noul; typed output does not guarantee semantic correctness; wire/agreement ≠ replica of TypeSafe; SemIf ≠ kw2828/OpenJev playground ≠ zhihz/openjev ≠ apiplant/semif-rs port ≠ dddanielliu/semif-serve; rename is densify not a second census; JevBench 74.6 is §78 not this ladder; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33/#34/#35/#36/#38; do not push onto open #39/#40; notes.md §117 diff --git a/research/notes.md b/research/notes.md index 98780a7..e3ee008 100644 --- a/research/notes.md +++ b/research/notes.md @@ -28906,3 +28906,234 @@ Cards: `mixed-architecture.md` (fail table); User-provided 0920 jcr uniqueness lock: NiazMorshed2007/jcr MIT; site https://jcr.niazmorshed.dev; topics ai-agents,jev,mcp; **4★**; HEAD `138b3832`; README SHA `2a49dbc1`; LICENSE SHA `46231303`; size **14850**; Jev Capability Resolver; one tool to find documented deterministic commands in a nested capability tree; returns context; **does not execute**; skills = workflow+judgment; capabilities = individual operations; format independent of Jev; proposed open standard exploration; classify (Jev) → optional OpenAI decompose compound → beam search geometric mean of routing probs; keep up to 3 paths ≥60% of best (JCR_BAND_RATIO 0.6); ambiguity / no-match / depth-limit explicit; soft scores ≠ hard gates; 0.6 band is application policy; routing ≠ permission; docs ≠ authority to run; sol-vs-opus5-20 *theirs*: 20 scenarios × 4 variants = 80 runs; lookup+explain only, no execution; Claude Opus 5: agent input 108,585→15,819 (−85%), cost $0.3700→$0.1222 (−67%), wall 105.5s→77.7s; Codex GPT-5.6-Sol: 61,952→47,669 (−23%), $0.1377→$0.1151 (−16%), wall 25.3s→62.4s (Sol slower with JCR in 19/20); One Sol outlier 372.6s / 193 Jev calls; n=1 per cell; Not Harbor task-execution; Claude/Codex harnesses; compare mode; 50 scenarios bundled; 11 groups, 960 nodes, 11,360 items; 16 routing rounds per step; NiazMorshed2007/jcr ≠ skill-broker ≠ skillranker ≠ jev-sift ≠ jev-lens ≠ jevusher ≠ jev_select_capability; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33/#34; notes.md §116 +User-provided 0922 uniqueness lock: SemIf was formerly OpenJev; independent; not affiliated with Jev or TypeSafe; homepage openjev.com; default master; MIT; HEAD ca3ba65f1429; Tolerate float roundoff in MLX evidence verification; pushed 2026-09-19; live REST 2282★ / 140 forks; size 9177; README SHA 74ab7f7f; LICENSE SHA ca562883; interface pattern reproduction with open models; does not reproduce Jev undisclosed model/training; Direct option logits; 0 output tokens; shared-state parallel; MLX backend for Apple Silicon (`--backend mlx`); Qwen3.5-4B 3090 direct 1.023s vs AR JSON 5.332s (**5.21×**); argmax agree 18/21; systems comparison ≠ semantic equivalence; Parallel suffixes 20.03 dec/s on 777 decisions; Browser ladder Qwen3.5-4B authored BA 0.813, pert 0.766, TypeSafe subset agreement 0.845 vs Published Jev 0.883 (102 across 20 cases); Softmax over options ≠ calibrated Noul; typed output does not guarantee semantic correctness; wire/agreement ≠ replica of TypeSafe; SemIf ≠ kw2828/OpenJev playground ≠ zhihz/openjev ≠ apiplant/semif-rs port ≠ dddanielliu/semif-serve; rename is densify not a second census; JevBench 74.6 is §78 not this ladder; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33/#34/#35/#36/#38; do not push onto open #39/#40; notes.md §117 + +## 117. User-provided HIGH — SemIf rename + MLX + ladder densify (2026-09-20 ~09:22 Boise / 2026-09-20T15:22Z) + +Docs-only reconstructed onto latest `main` after merged **#34** (Pages), +merged **#35** (hourly 0843 / §114), merged **#36** (NanoJev / §115), +and merged **#38** (jcr / §116). +Branch `cursor/densify-semif-rename-mlx-ladder-1d85` (PR #37). +**Never reopen merged** Augustus PR #7 / **#8** / +**#9** / **#10** / **#12** / **#13** / **#14** / +**#15** / **#16** / **#17** / **#18** / **#19** / +**#20** / **#21** / **#22** / **#23** / **#24** / +**#25** / **#26** / **#27** / **#28** / **#29** / +**#30** / **#31** / **#32** / **#33** / **#34** / +**#35** / **#36** / **#38**. **HARD RULE:** do not reopen or amend +PR #23–#36 or **#38** (merged #35 owns `notes.md` §114 / items 289–302 / +batch #97; merged **#36** owns `notes.md` §115 / items 303–308 / batch #98; +merged **#38** owns `notes.md` §116 / items 309–316 / batch #99 — +leave them alone). Do **not** push onto open **#39** (post-#35 hygiene; +IDs unchanged §114) or open **#40** (llm-to-jev; claims §118 / +items 322–329 / batch #101). This fold **keeps** `notes.md` §116 +from merged #38 and **skips** §118 / composition items 317–329 / +findings batch #101. Remap: merged #38 took 309–316 (overrunning the +original 315–321 claim) and open #40 took 322–329, so this fold's +composition items stay **330–336** (still `notes.md` §117 / batch **#100**). + +**Not a first sighting.** SemIf is already in the +census (stars since §48 spotcheck 1551★; §113 live +REST **2241★**; §114 hourly pulse **2237★**) and on +the JevBench v1.2 board (§78: SemIf Qwen3.5-4B **74.6**, +geometric mean — **do not merge** that protocol with +this ladder). This pass densifies **rename + MLX + +accuracy ladder**. Do **not** re-fold §69 semif-serve +runoff (`1164 vs 178 ms`, wire-compat ≠ replica) as if +it were this card. Do **not** re-fold §78 JevBench +74.6 as this authored BA. Do **not** re-fold +IamBusy/OpenJev `/v1/decide` 45/60. Quote **their** +README / `docs/MLX.md`. Mark *theirs*. No invented +metrics. Hunches labeled. No wrappers, `pip` / +`venv` / `semif-score` install recipes / +`TYPESAFE_API_KEY` / `.env`. `invented_signal: +false`. Skip Archer rewrite. Qwen3.8-27B ≠ Archer. +X MCP **not** used this pass; no invented tweets. + +Lane is Augustus: **mental models / architecture / +class lineage / measurement honesty**. Backend-agnostic +categorization/scoring/decision class. SemIf is an +**open-model interface-pattern reproduction**, not a +TypeSafe replica and not a new species. Softmax over +options ≠ calibrated Noul. Wire/agreement ≠ replica. +Systems comparison ≠ semantic equivalence. Formal +methods **compose** with scoring; a Noul is a SENSOR. +Jev is the hot exemplar, not the whole mandate. +Mathematical / logical / algorithmic mental models +across AI, SWE, business, knowledge work — not +SWE-only. + +### Live REST relock (this pass; quote over watch) + +[`TheoLeeCJ/SemIf`](https://github.com/TheoLeeCJ/SemIf) +MIT; homepage [openjev.com](https://openjev.com); +default **master**; **2282★** / **140** forks; GitHub +size **9177**; HEAD +`ca3ba65f142967030ecb453346e94d6f476a69df` +(`ca3ba65f1429`); commit *theirs* **Tolerate float +roundoff in MLX evidence verification** +(2026-09-19T04:46:33Z); pushed 2026-09-19T04:46:36Z. +README SHA `74ab7f7ffa3d492c2dc905c1e75410dca3e0677e`; +LICENSE SHA `ca562883550941229de6555a8374fe2c83a18e08`; +`docs/MLX.md` SHA +`e1d81a75a04aa457deeea1d601686c5529fc5854`. +Pulse vs §113: **2282★** (+41 vs **2241**). Pulse vs +§114 hourly: **2282★** (+45 vs **2237**). Prior PR #37 +lock cited 2282★ — **quote live 2282**. + +Cousins (live REST this pass; distinction, not +census-as-endorsement): +[`kw2828/OpenJev`](https://github.com/kw2828/OpenJev) +MIT; **1★**; playground; size **656095**; +[`zhihz/openjev`](https://github.com/zhihz/openjev) +**19★**; +[`apiplant/semif-rs`](https://github.com/apiplant/semif-rs) +**0★** Rust+candle port; +[`dddanielliu/semif-serve`](https://github.com/dddanielliu/semif-serve) +**0★** (already §69). Live `gh api +repos/IamBusy/OpenJev` **resolved to** +[`IamBusy/OpenJev-Vision`](https://github.com/IamBusy/OpenJev-Vision) +this pass — do **not** collapse SemIf into Vision, +and do **not** treat that resolve as a deletion of +the §69 `/v1/decide` card. + +### HIGH (densify, one host) + +1. **[`TheoLeeCJ/SemIf`](https://github.com/TheoLeeCJ/SemIf) + (PRIMARY densify)** — Python MIT; formerly OpenJev + (**rename is IS, not ≠**). Independent; not + affiliated with Jev or TypeSafe. README *theirs*: + SemIf was formerly called OpenJev. It is not + affiliated with or endorsed by TypeSafe. Interface + pattern reproduction with open models; **does not + reproduce Jev's undisclosed model or training**. + Direct option logits; **0 output tokens**; + shared-state parallel; MLX backend for Apple + Silicon (`--backend mlx`). Speed *theirs* + Qwen3.5-4B 3090: direct **1.023 s** vs AR JSON + **5.332 s** (**5.21×**); argmax agree **18/21** + with the generative path — *theirs*: systems + comparison rather than a claim that the two + readouts are semantically equivalent. Parallel + suffixes **20.03** dec/s on **777** decisions + *theirs*. Browser ladder *theirs*: Qwen3.5-4B + authored BA **0.813**, pert **0.766**, TypeSafe + subset agreement **0.845** vs Published Jev + **0.883** (102 across 20 cases). Softmax over + options ≠ calibrated Noul. Wire/agreement ≠ replica + of TypeSafe. SemIf ≠ kw2828/OpenJev playground ≠ + zhihz/openjev ≠ apiplant/semif-rs port ≠ + dddanielliu/semif-serve. Do not copy `pip` / + `semif-score` / MLX install. Do not dump + `docs/MLX.md` recipes. MLX *theirs*: typed output + does not guarantee semantic correctness, and + softmax scores are not calibrated confidence. + HEAD this pass is the float-roundoff tolerate + commit on MLX evidence verification. + + **Hunch:** the product is **0 output tokens** plus + a named serving surface, not a closed-model clone. + Agreement 0.845 on 102 aligned public rows is a + **subset agreement**, not Harbor and not ECE. + 5.21× is a **systems** ratio on their owned 21 + binary criteria, not TypeSafe's latency envelope. + +### How-to-apply (five placements / one host) + +When someone pastes SemIf / OpenJev.com / "the open +Jev", extract the **rename + interface-pattern + +measurement-honesty** theses. Do not steal a Harbor +number. Do not copy the installer. + +1. **Rename is densify, not a second census** — + SemIf **IS** formerly OpenJev (same GitHub + `TheoLeeCJ/SemIf`; homepage still openjev.com). + Quote *theirs*. Do not open a new species row. + Do not treat the rename as a retraction of prior + star pulses (§48 1551★ → §113 2241★ → §114 2237★ + → this pass 2282★). +2. **Interface pattern ≠ replica** — reproduces the + **typed-decision interface** (state + runtime + criteria + option logits) with **open models**. + Explicitly does **not** reproduce Jev's + undisclosed model or training. Softmax over + supplied options is conditional on those options + *theirs* — calibrate on the workload. ≠ TypeSafe + hosted Noul. +3. **Direct logits / 0 tokens / shared-state / MLX** + — one forward pass reads declared option logits; + no answer token is sampled. Shared-state prefill + once, then branch criteria. Apple Silicon path is + `--backend mlx` (PR #8 / float-roundoff HEAD). + MLX is a **backend**, not a new class. Do not + copy install. Default backend remains Torch/CUDA + *theirs*. MLX reranker mode is explicitly + unsupported *theirs*. +4. **Systems comparison ≠ semantic equivalence** — + 1.023 s vs 5.332 s (**5.21×**), 0 vs 111 output + tokens, argmax **18/21**. Quote the disagreement + as the honesty. Parallel suffixes **20.03** dec/s + on 777; BF16 reuse changed 5–6 of 777 argmaxes + *theirs* (experimental). Do not paste 5.21× as + "beats Jev" or as a replica latency. +5. **Browser ladder ≠ Harbor; cousins ≠ host** — + Qwen3.5-4B authored BA **0.813** / pert **0.766** + / TypeSafe subset **0.845** vs Published Jev + **0.883** on **102 across 20 cases**. Jev number + *theirs* is read from TypeSafe's published + records; they did not run a live Jev endpoint. + **Do not merge** with §78 JevBench **74.6** + (different protocol, 534-row geometric mean). + WANLI-256 BA **0.637** *theirs* is not bonzi + WANLI-256 64.5%. SemIf ≠ kw2828 playground ≠ + zhihz/openjev ≠ semif-rs port ≠ semif-serve + `/v1/systemone` wire. + +### Theater (do not) + +Treat 0.813 / 0.845 / 0.883 as Harbor or as +calibrated Noul; treat 5.21× as TypeSafe latency or +as semantic equivalence; treat 18/21 as "the same +model"; softmax-over-options as a Noul; copy `pip` +/ `venv` / `--backend mlx` install as a recipe; +collapse SemIf into kw2828 / zhihz / semif-rs / +semif-serve / IamBusy `/v1/decide` / OpenJev-Vision; +re-fold §78 74.6 as this ladder; re-fold §69 +1164 vs 178 ms as this speed table; treat MLX as a +new species; invent tweets / Archer drop; reopen +merged #23–#36 or #38; push onto open #39/#40; steal +§114 from merged #35, §115 from merged #36, §116 +from merged #38, or §118 / 322–329 from #40. + +### Census + +Live REST quoted above. SemIf **2282★** (+34 vs +§113 **2241**; +38 vs §114 **2237**). Cousin stars +quoted for distinction only. Awesomejev / tracker / +Laya / jevlike **not re-derived** this pass +(user-linked SemIf densify, not an hourly watch). +Archer still **NOT landed**. `invented_signal: false`. + +### Not + +Not a TypeSafe how-to. Not a SemIf/MLX install +guide. Not a second OpenJev census. Not a JevBench +v1.2 rerun. Not wrappers. Do not copy keys / +commands. Do not dump source / weights / eval JSON. +Skip Archer rewrite. + +### Overlay set + +SKILL.md YAML+protocol+mapping-index, mental-models +Apply 0922, composition-algebra items 330–336, faq, +mixed-architecture, validation, toolbox-mapping, +methods-catalog, formal-methods, formal-semi-formal, +applied-mappings, judgment-class, question-design, +agent-self-assessment, mappings, CHANGELOG, README, +docs/ecosystem, findings batch #100, refresh-log, +sources.json, changelog-hourly.md, uniqueness_gate.py. + +Offline check: uniqueness_gate.py (0843 + 0915 + jcr + 0922 +consecutive locks) and `evaluate_decisions.py +--self-test`. No live Jev key. No wrappers. + diff --git a/research/refresh-log.md b/research/refresh-log.md index d776938..3f35fd1 100644 --- a/research/refresh-log.md +++ b/research/refresh-log.md @@ -3091,6 +3091,7 @@ User-provided 0920 jcr uniqueness lock: NiazMorshed2007/jcr MIT; site https://jcr.niazmorshed.dev; topics ai-agents,jev,mcp; **4★**; HEAD `138b3832`; README SHA `2a49dbc1`; LICENSE SHA `46231303`; size **14850**; Jev Capability Resolver; one tool to find documented deterministic commands in a nested capability tree; returns context; **does not execute**; skills = workflow+judgment; capabilities = individual operations; format independent of Jev; proposed open standard exploration; classify (Jev) → optional OpenAI decompose compound → beam search geometric mean of routing probs; keep up to 3 paths ≥60% of best (JCR_BAND_RATIO 0.6); ambiguity / no-match / depth-limit explicit; soft scores ≠ hard gates; 0.6 band is application policy; routing ≠ permission; docs ≠ authority to run; sol-vs-opus5-20 *theirs*: 20 scenarios × 4 variants = 80 runs; lookup+explain only, no execution; Claude Opus 5: agent input 108,585→15,819 (−85%), cost $0.3700→$0.1222 (−67%), wall 105.5s→77.7s; Codex GPT-5.6-Sol: 61,952→47,669 (−23%), $0.1377→$0.1151 (−16%), wall 25.3s→62.4s (Sol slower with JCR in 19/20); One Sol outlier 372.6s / 193 Jev calls; n=1 per cell; Not Harbor task-execution; Claude/Codex harnesses; compare mode; 50 scenarios bundled; 11 groups, 960 nodes, 11,360 items; 16 routing rounds per step; NiazMorshed2007/jcr ≠ skill-broker ≠ skillranker ≠ jev-sift ≠ jev-lens ≠ jevusher ≠ jev_select_capability; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33/#34; notes.md §116 +User-provided 0922 uniqueness lock: SemIf was formerly OpenJev; independent; not affiliated with Jev or TypeSafe; homepage openjev.com; default master; MIT; HEAD ca3ba65f1429; Tolerate float roundoff in MLX evidence verification; pushed 2026-09-19; live REST 2282★ / 140 forks; size 9177; README SHA 74ab7f7f; LICENSE SHA ca562883; interface pattern reproduction with open models; does not reproduce Jev undisclosed model/training; Direct option logits; 0 output tokens; shared-state parallel; MLX backend for Apple Silicon (`--backend mlx`); Qwen3.5-4B 3090 direct 1.023s vs AR JSON 5.332s (**5.21×**); argmax agree 18/21; systems comparison ≠ semantic equivalence; Parallel suffixes 20.03 dec/s on 777 decisions; Browser ladder Qwen3.5-4B authored BA 0.813, pert 0.766, TypeSafe subset agreement 0.845 vs Published Jev 0.883 (102 across 20 cases); Softmax over options ≠ calibrated Noul; typed output does not guarantee semantic correctness; wire/agreement ≠ replica of TypeSafe; SemIf ≠ kw2828/OpenJev playground ≠ zhihz/openjev ≠ apiplant/semif-rs port ≠ dddanielliu/semif-serve; rename is densify not a second census; JevBench 74.6 is §78 not this ladder; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33/#34/#35/#36/#38; do not push onto open #39/#40; notes.md §117 ## 2026-09-20 ~14:43 UTC — hourly 0843 HIGH measurement / judgment - Docs + evaluator on a **fresh PR off main** (`cursor/hourly-0843-augustus-fold-220d`). @@ -3148,3 +3149,24 @@ Hourly 0843 uniqueness lock: A hunch is a probability with a policy attached; { - Quote *theirs*. Do not invent accuracy numbers. `invented_signal: false`. Do **not** merge from this review. - User-provided 0915 uniqueness lock: TianyuCodings/NanoJev unified-games-v1 densify; A 0.6B parallel decision model: states and questions in, complete probability distributions out. Zero output-token decoding.; One model, four games; ViZDoom Basic 128/128 vs Jev 56/128; Predict Position 27/128 vs Jev 11/128; Maze 225 attempts vs Jev 2738; Snake 30 food / 256 steps; held-out Maze 4/10 Snake 8/8 Basic 128/128 Predict 27/128; Untuned Qwen3-0.6B baseline; 18,760 questions per variant; 16,333 ViZDoom; 896 Predict Position expert episodes; hard_lr1e5; mix weights 1/3, 1/3, 1/6, 1/6; Hub C-Tianyu/NanoJev revision unified-games-v1 likes 58; dataset C-Tianyu/NanoJev-Data likes 5; HEAD 618cea6d906d54e128360786d12f703fff2b1245; 1289★ / 158 forks / size 64035; README SHA 4190093c64ee75b26e9726daa3b00cbcf6d3157a; MIT; caijinchun/nanojev-arena ≠ liao96312/jev-arena-nanojev ≠ zwliJay/jev-forge ≠ NanoJev; not TypeSafe Jev; open replica / specialist gameplay S1; soft scores ≠ hard gates; Game success ≠ calibrated Noul; local type boolean ≠ TypeSafe noul; invented_signal false; do not reopen or amend PR #31/#32/#33/#35; notes.md §115 + +## 2026-09-20 ~16:10 UTC / ~10:10 Boise — User-provided 0922 SemIf densify (review relock onto #38) +- Reconstruct onto latest `main` after merged #34 (Pages), merged #35 + (hourly 0843 / `notes.md` §114 / items 289–302 / batch #97), merged + **#36** (NanoJev / `notes.md` §115 / items 303–308 / batch #98), and + merged **#38** (jcr / `notes.md` §116 / items 309–316 / batch #99). + **HARD RULE:** do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33/#34/#35/#36/#38. + Do not push onto open #39/#40. This fold remains `notes.md` §117 / items 330–336 / batch #100. +- Live REST relock: TheoLeeCJ/SemIf MIT; homepage openjev.com; default master; + **2282★** / **140** forks; size **9177**; HEAD `ca3ba65f142967030ecb453346e94d6f476a69df` + (Tolerate float roundoff in MLX evidence verification, 2026-09-19T04:46:33Z); + README SHA `74ab7f7ffa3d492c2dc905c1e75410dca3e0677e`; LICENSE SHA `ca562883`. +- Cousins: kw2828/OpenJev **1★** playground; zhihz/openjev **19★**; + apiplant/semif-rs **0★**; dddanielliu/semif-serve **0★**. IamBusy/OpenJev + API resolved to OpenJev-Vision this pass — not a §69 deletion. +- Folded into `notes.md` §117, SKILL.md, mental-models Apply 0922, faq, + composition-algebra items 330–336, findings batch #100, uniqueness lock + across 21 overlays, uniqueness_gate.py (0843 + 0915 + jcr + 0922). Not a first sighting. + No wrapper. Quote *theirs*. `invented_signal: false`. X MCP not used; no invented tweets. +- User-provided 0922 uniqueness lock: SemIf was formerly OpenJev; independent; not affiliated with Jev or TypeSafe; homepage openjev.com; default master; MIT; HEAD ca3ba65f1429; Tolerate float roundoff in MLX evidence verification; pushed 2026-09-19; live REST 2282★ / 140 forks; size 9177; README SHA 74ab7f7f; LICENSE SHA ca562883; interface pattern reproduction with open models; does not reproduce Jev undisclosed model/training; Direct option logits; 0 output tokens; shared-state parallel; MLX backend for Apple Silicon (`--backend mlx`); Qwen3.5-4B 3090 direct 1.023s vs AR JSON 5.332s (**5.21×**); argmax agree 18/21; systems comparison ≠ semantic equivalence; Parallel suffixes 20.03 dec/s on 777 decisions; Browser ladder Qwen3.5-4B authored BA 0.813, pert 0.766, TypeSafe subset agreement 0.845 vs Published Jev 0.883 (102 across 20 cases); Softmax over options ≠ calibrated Noul; typed output does not guarantee semantic correctness; wire/agreement ≠ replica of TypeSafe; SemIf ≠ kw2828/OpenJev playground ≠ zhihz/openjev ≠ apiplant/semif-rs port ≠ dddanielliu/semif-serve; rename is densify not a second census; JevBench 74.6 is §78 not this ladder; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33/#34/#35/#36/#38; do not push onto open #39/#40; notes.md §117 + diff --git a/research/sources.json b/research/sources.json index 2f974eb..d1fba4d 100644 --- a/research/sources.json +++ b/research/sources.json @@ -1,6 +1,6 @@ { "refresh_cadence": "hourly", - "retrieved": "2026-09-20T14:43Z", + "retrieved": "2026-09-20T16:10Z", "sources": [ { "kind": "docs", @@ -1428,7 +1428,7 @@ "kind": "github", "title": "TheoLeeCJ/SemIf (star spotcheck)", "url": "https://github.com/TheoLeeCJ/SemIf", - "note": "1551 stars this pass (2026-09-18T20:51Z); +60 vs awesome claim 1491. Independent. notes.md §48." + "note": "PRIMARY densify 0922. formerly OpenJev; independent not TypeSafe; MIT; homepage openjev.com; default master; live REST 2282★ / 140 forks; size 9177; HEAD ca3ba65f1429; README SHA 74ab7f7f; LICENSE SHA ca562883. 5.21× systems≠semantic; Softmax ≠ Noul. notes.md §117 / items 330–336 / batch #100." }, { "kind": "github", @@ -4945,6 +4945,30 @@ "title": "How I’m using Jev in an agent harness — JCR", "url": "https://jcr.niazmorshed.dev", "note": "Author article 2026-09-20. Quote *theirs*. Complement to README. notes.md §116." + }, + { + "kind": "web", + "title": "openjev.com", + "url": "https://openjev.com", + "note": "SemIf homepage field. rename densify not a second census. notes.md §117." + }, + { + "kind": "github", + "title": "kw2828/OpenJev", + "url": "https://github.com/kw2828/OpenJev", + "note": "cousin playground ≠ SemIf. 1★. notes.md §117." + }, + { + "kind": "github", + "title": "zhihz/openjev", + "url": "https://github.com/zhihz/openjev", + "note": "cousin ≠ SemIf. 19★. notes.md §117." + }, + { + "kind": "github", + "title": "apiplant/semif-rs", + "url": "https://github.com/apiplant/semif-rs", + "note": "Rust+candle port ≠ SemIf. 0★. notes.md §117." } ] }