diff --git a/.agents/skills/augustus/SKILL.md b/.agents/skills/augustus/SKILL.md index 96e35a02..4c526651 100644 --- a/.agents/skills/augustus/SKILL.md +++ b/.agents/skills/augustus/SKILL.md @@ -1,6 +1,6 @@ --- name: augustus -description: "Use when placing typed probabilistic judgment (Jev-class System One / decision models) with mathematical, logical, or algorithmic mental models — in AI, software, business, knowledge work, or life, not only SWE; deciding where a fast cheap categorization/classification/scoring model belongs versus generation, exact policy/code, or proof; applying expected utility, selective classification/abstention, calibration, cost-sensitive thresholds, value of information, MCDA, signal detection, search/control substitutions, or Leveson-style org/safety; using NATM/snap-fit/Norman as design intuition; designing mixed architecture (decision model + LLM writing); auditing an existing system, PR, workflow, or non-software practice for judgment-shaped holes and code smells; debugging a question that hovers near 0.5, clusters mid-scale, or hides two judgments; placing agent self-supervision gates (pre-action, output judge, done-check, stuck-detector, context sieve); coupling a typed judge as an optimizer metric (Ax, DSPy); choosing among TypeSafe Jev, open heads (Laya, kev, openjev-lm, Nimble, encoder DeBERTa, LoRA distill, open multimodal RLCD / blackwood), announced open decision-model (Watch — still not landed), constrained-AR (TypeAR, pcdServer, decision-token LoRA), diffusion structured reads, GLiNER/GLiClass/GLiGuard encoder family (locate vs categorize vs safety-schema classify vs local multi-head), listwise rankers, or vision scorers; placing judgment beside TLA+/Alloy/Apalache/Dafny/DST (Antithesis, Resonate, PufferLib) without laundering a Noul as a proof; answering \"it's just classification\", \"is Jev probabilistic programming\" (marginals vs joint, not a PPL), \"low/medium/high entropy\" (allocator, not a meter), \"perception specialist then judgment vs shared multimodal System One\", \"wait for Archer vs open multimodal RLCD\", \"screenshot/DOM candidates → typed Choice\", \"eval path\", \"jevals\", \"Harbor taskset\", \"shared bake-off ECE/NLL/Brier\", \"LLM-as-judge is not the System One score\", \"pipeline / measure / hill-climb perception into a decision\", \"Ax vs DSPy\", \"held-out\", \"correctness is not confidence\", \"is this only for software?\", Alloy vs Apalache, GLiNER vs Jev, \"is GLiGuard Jev?\", LLM-as-judge, paraphrase brittleness, allowlist then judge (allowlist *proves*; fail-open cannot block), \"missing other → confident wrong Choice\", \"lint the Jev request\", \"training confronts Choice other / none-of-the-above\", \"S1 reflex keeps control / optional S2 one-use advice\", \"soft AGENTS.md rules vs linter\" (Abide / jev-pref), \"edit-phase vs turn-phase observation window\", banded confidence fail-open preference lint, \"extractive selection / pointer-not-generator\", \"encoder GLiNER compaction vs Jev Score compaction (same job; pointer not summarizer)\", \"fail-closed keep_full under mutation envelope\", \"CI flaky-vs-real merge gate\", \"fail-open VOI wake/resume (Horvitz)\", \"claim/evidence Stop integrity\", \"S1 extract + escalate-S2 indexer\", \"Harbor on/off routing\", \"policy-as-judgment PR marshal\", \"shadow-mode compaction rollout\", \"Jev Ultrafast vs GLiNER Ultrafast (observe-score-act backend-agnostic)\", \"hybrid local decide + remote fill\", \"DONE ≠ verified success\", \"observed a11y/DOM candidates vs screenshot multimodal\", \"evidence-preserving stdout prune (not summarize)\", \"hard token/format envelope then soft Noul\", \"fail-safe keep original on prune failure\", \"stdout prune vs session compaction\", \"specialist S1 computer-use (Cua-S1 form-v0; not TypeSafe Jev)\", \"plan ≠ execute / dry-run default\", \"observed-element option head (fill/check/click/skip)\", \"local /v1/systemone drop-in (stub until hf scorer)\", \"dataframe-native semantic columns\", \"route≠memory\", \"advisory sidecar receipts\", \"structure induction over bags\", \"AST ∩ semantic lint\", \"extractable-from-state / retrieve first\", \"decision-model vs constrained-LLM bake-off\", \"dual-process S1 decide / S2 generate\", \"combinatorial grid ≠ extractive\", TOCTOU-of-Noul, vacuous specs, open weights vs constrained decoding vs encoder vs LoRA vs kev, whether a decision needs a model at all (meta-VOI), env-break vs policy-break, sqlite-jev / in-engine vs CLI store index, hard safety envelope (Jev proposes, code clamps), host-adapter routing (not MCP), distill-to-device memory gate, \"uncalibrated local likelihoods vs Noul / CUDA replica\", \"decision-native RAG retrieve wide then decide then evidence set\", \"classify-first MCP / read selectively\", \"living applied-mappings atlas / class patterns not a 342 hit list\", \"draft-gate silence as safer / heartbeat\", \"robotics text-state vs pixels\", \"verbatim session ledger / scored recall\", \"judgment as language primitive / English-as-config\", \"pre-registered AMBIGUOUS eval / cascade sign-flip\", \"healthcare Harbor-shaped S1+S2\", \"pre-exec tool gate allow/block/review\", \"productized public primitive / judgment wall\", \"meaning-search without embeddings\", \"attention≠correctness PR review\", \"skills→oxlint / AST prove ∩ remainder\", \"session-sticky first-prompt routing\", \"measured RAG rerank vs generative rerank\", \"Stagehand extract pick-and-copy / judge\", \"harness observe-score-act productization\", or \"formally verify with Jev\", \"capability kernel / secrets never in the agent\", \"Jev is SENSOR not policy\", \"type-safe ≠ correct\", \"typed control plane around DSPy\", \"native-probability calibration / Brier/ECE arena\", \"fan-out as measurement economics\", \"engine owns truth / Jev owns judgment\", \"human-confirmed kill gate\", \"train specialist when downstream reads p vs few-shot hosted when only argmax\", \"decide→policy→LLM leftover cascade\", \"Noul 0.5 cannot-tell never rounded\", \"calibration ≠ sortable / ORDER BY over Jev probs\", \"pairwise inversion / Score ordinality / two-decimal ties\", \"wire-compat self-hosted /v1/systemone GLiFormer\", \"class-backend economics\", \"loopback gateway hosted + local OpenJev\", \"do not distill Jev as teacher of record\", \"active-learning triage / training-data VOI\", \"index-once ask-many / citable evidence packets\", \"meaning-grep AND/OR/NOT line Nouls\", \"closed-vote-only computer-use / no planner LLM\", \"Jev vs local MLX PCD Harbor\", \"PCD O(1) speed ≠ calibrated Noul\", \"host-owned handlers × System One\", \"OMP/pi fail-open acceptance gate\", \"permission vs probability / operator owns the safety bar\", \"judgment ≠ permission / Jev never grants access\", \"eval integrity / instrument not score / dinostomp jev-as-if\", \"constrained optimizer + S1 features / never sole hot-path gate\", \"privilege ≠ verdict / effect contracts not tokens\", \"attention filter / VOI for human review / never blocks / never green unless sure\", \"measurement owns endorsement / evidence-gated question packs\", \"Jev supplies evidence / code owns authority\", \"ranking ≠ calibration / never hard-threshold raw p as frequency\", \"hot-click CU / indexed element table / S1 on click path\", \"Jev judges relevance / code decides structure / never rewrite\", \"local rules first then remainder / never auto-train on model's own hides\", \"combinators / System One as control plane / not chat turns\", \"receipts not leaderboard / type-safe ≠ correct jaggedness\", \"VOI over skill library / skillranker abstention\", \"OOD calibration / AUC ≠ ECE / sign of miscalibration by type\", \"Jev vs thinking-budget small models / frontier-100\", \"turnstile / replayable evidence≠authority\", \"MLX one-pass schema→JSON / Apple Silicon replica economics\", \"memory leases ended by new evidence\", \"never confidently wrong / TLA+ compose with judgment / escalate instead of hard-gate\", \"no seal no advance / coverage ledger / mint ≠ product brain\", \"skill-broker sibling turnstile/skillranker / judgment ≠ permission\", \"sureness / CERTAIN|CONFIDENT|LEANING|TORN|CLUELESS / max_prob is generous\", \"JevBench Harbor/jevals practice / calibration not in Main Score\", \"CI typed gate before expensive review / ci-gatekeeper\", \"Codex MCP host adapter / jev_select_capability\", \"judgment as attention redirect not merge blocker / jev-preflight\", \"compress-before-first-send / dizk jev-lens vs rashed attention filter\", \"tools≠use / SessionStart over hoping the model recalls\", \"observational memory / keep-kind verbatim / pi-om\", \"open-Jev class / openvons / JevPick menu decode\", \"physical-world System One / HA-Jev / not for locks\", \"judgment outside the store / jevql CLI\", \"landed-script trust / headless≠auto-approve\", \"digital-design combinators / extended Router Loop Retry Fallback Memory\", \"VOI cache admission / same-intent skip LLM\", \"BM25 vs Jev skill routing Harbor harness\", \"zeroshot vs BERT / contamination DiD / label-equivalence\", \"typed escalate continue abort baton / inverted loop\", \"worth-your-attention VOI / ThinkyMiner Winnow vs kevinpita winnow\", \"Jev WHETHER Python HOW LLM WHAT\", \"conflict vs ignorance / named Choice escape\", \"Playwright executes Jev chooses / sample-from-distribution\", \"OpenJev /v1/decide not TypeSafe drop-in\", \"SemIf wire-compat runoff\", \"decision-as-memory flywheel\", \"record/replay CI / jevassert\", \"failure-finding arena / jevarena ≠ jev-arena\", \"BBQ stereotype/uncertainty/cost\", \"decider≠executor / jeffrey\", \"sentence-as-rule lint / jevlint ≠ JevLint\", \"VOI hunk prune / prune-review\", \"whole-repo intent VERIFIED/VIOLATION/UNKNOWN\", \"GLiNER2 System One spec ≠ replica\", \"Rust/WebGPU grande / Clojure Laya byte parity / CPU SemIf\", \"ONNX ModernBERT local-jev measured not equivalent\", \"persist constraints across compaction / pi-heed\", \"calibration+cost as first-class gates\", \"Harbor-shaped Jev vs schema-guided LLM-as-judge / jev-judge-bench ≠ jevarena ≠ jevbench\", \"hand no-text steps to Jev / jev-use / Vercel drops confidence / margin fallback\", \"Pi System-One control plane / pi-jev-control\", \"generation as tree of Choices / never free-generates / jev-gpt\", \"OpenRouter recipe atlas / samples not benches / jev-cookbook\", \"personal history feed / no social graph / jevfeed\", \"competing NAR claims / dual-channel ECE / claim-verification / openJev-verdict ≠ OpenJev\", \"empty compaction-proxy skip / IPECTER\", \"throughput ≠ latency / like-for-like ECE\", \"1-token logprob endpoint ≠ Noul / coverage ≠ correctness / chakuho\", \"open replica engine / jevinf / argmax-parity ≠ ECE\", \"unofficial Elixir SDK ≠ OTP peer / dannote/jev\", \"jevex rename + n=16 SWE VOI / files-to-read\", \"commit pre-review attention≠verdict / middle band never rounded / commitjev\", \"Hermes plugin is Agnes not TypeSafe\", \"pi-jev-compact ≠ pi-jev-compaction / verbatim summarizer replacement\", \"empty Codex-proxy skip / IPECTER runway\", \"decision-native inbox / mailordinal / humans own ambiguity\", \"unofficial jev-cli not ready / ≠ jevql\", \"laya-multilingual / English checkpoint confident-wrong OOD / ships uncalibrated\", \"schema-conditioned DeBERTa scorer / peaked ranking ≠ calibration\", \"HF 401 access / GitHub 404 Hub-only\", \"productized System One HTTP / classifier.dev / label+confidence public contract\", \"escalate-under-threshold / smart tier 0.7 / multi-label ignores tier\", \"silent-fallback FALLBACK marker / granite 0.546 vs advertised 0.800\", \"vs_jev tracked JSON not transcription / read eval/README before quoting\", \"choxos/jev-reviewer ≠ egma-ai / systematic-review pointer-not-generator\", \"two-pass Choice+Noul / relative which-line + absolute does-this-line\", \"not-found is an answer / no paraphrase invent\", \"human check as productized judgment / checked never overwritten\", \"githubnext/localjev ≠ kunchenguid/local-jev / prompted JSON ≠ structured logit read\", \"wire-compat ≠ logit-equiv / self-reported probs / entropy confidence\", \"institutional open-replica / GitHub Next /v1/systemone\", \"Harbor-shaped bake-off AG News BoolQ SST-5 / 1200-request caveats\", \"LM Studio runner gap / structured-read primitives for OpenJev parity\", \"NandhaKishorM/laya packaging ≠ Hub-only / Router script-before-p\", \"post-T ECE ≠ raw ECE / Banking77 token-budget / 0.85 still soft / not TypeSafe drop-in / external census ≠ scored bake-off / GLiNER2+routers class-boundary / incomplete vs watch / Harbor honesty watch / JevBench v1.2 geometric-mean I/C/S/K / cal now ON rank / weight sensitivity / option-order 72→21 / instruction models class-boundary / ×2 latency assumption / est. costs / Laya absent gap / Qwen3.8 27B ≠ Archer\", \"hourly already-folded watch / apply-the-five / skip thin noise\", \"hard-gate Noul as PR gate is soundness theater / totally-tim/jev-gate ≠ jev-gateway\", \"S1 never stalls waiting / S2 one-use advisory\", \"purple telemetry = consumed not arrived\", \"Local controller ≠ githubnext/localjev\", \"seed = geometry not async replay\", \"20% starting gate still soft / schema-safe ≠ correct\", \"no pixels to either provider / confidence ≠ selected probability\", \"experimental viz not a flight controller / S2 never grants\", \"OCR+AX observe-score-act / typesafe-computer-use\", \"never send screenshot to frontier for the decision\", \"overlapping CU options = false low confidence\", \"split kind/item/site / offscreen\", \"writer/decider split + post-type Noul still soft\", \"155× one-screenshot Harbor-shaped ≠ taskset\", \"AX never sole / Spotify 0\", \"decision ≠ answer-reader capture\", \"typesafe-computer-use ≠ jev-ultrafast ≠ cua-s1 ≠ camoufox\", \"ASR observe-score-act / jev-voice-browser\", \"partial-speech VOI / complete Noul / free-text waits\", \"spoken confirm ≠ hard auth\", \"numbered overlay disambiguate without another model\", \"moritzkremb/jev-voice-browser ≠ jev-voice-control ≠ nikolas-j\", \"wrap-as-execution / AgentGhost ALLOW ASK DENY\", \"rules first then Jev remainder / ASK throws / fail-closed\", \"reddpy/AgentGhost ≠ jwen5419807/agentghost ≠ vventirozos\", \"JP genre atlas / studio_yebisu / stars ephemeral ≠ eval\", \"Jev Clearly Explained / akshay_pachaar / LLM hammer\", \"schema-safe ≠ correct / 200× 400× TypeSafe ceiling\", \"questions-as-code / shadow first / not a TypeSafe how-to\", \"proposition ≠ embedding / contrast-set refund\", \"boolean composition of soft Nouls / AND OR NOT after threshold\", \"uehaj/jev-semgrep ≠ semgrep.dev\", \"meaning-grep dedicated fold / not a gate\", \"decision-validated UI / Jev never authors text / gram-render\", \"decision-as-assert / jevtest ambiguous band\", \"typed decisions drive UI / jev2ui\", \"hybrid S1 closed verb menu / anima3 / jeff confidently flat\", \"pointer-not-generator search / JevFind\", \"jev-frontier-bench ≠ frontier-100 / ChaosNLI JS\", \"product bakeoff ≠ architecture duel / jev-gliclass-bench\", \"four engines same questions / majority floor / calibration ≠ discrimination\", \"authorship named escape / not courtroom evidence\", \"ha-switchboard HA remains execution / ≠ HA-Jev\", \"n8n classify/route/score / Low Confidence abstention\", \"fast-jev-compaction-pi ≠ pi-jev-compact ≠ pi-jev-compaction\", \"jevloop full-distribution optimizer / no LLM in the loop\", \"laya-vision SmolVLM / score untrained / ≠ blackwood ≠ Archer\", \"Cerebellum-2B /v1/decide ≠ TypeSafe / wire-compat vs agent-routing\", \"laya-grounded not drop-in / phishing regress / Platt not temperature\", \"GestaltLabs/Jeff-1 ≠ logan-markewich/jeff / acc vs ECE n=9730\", \"stanley-code empty findings ≠ approval / human promote\", \"findme ≠ JevFind / NL memory beam-search FS\", \"jevsubrouter price workers not conversation / counts ≠ dollars\", \"feelings .feels() default 0.5 is Noul-0.5-never-rounded / ≠ hunch ≠ Probably\", \"apa-agent-harness ≠ AntonioCoppe/jev-harness / unpublished npm\", \"grok-bot-jev skill cannot force a bot that ignores it / A/B proxies not tokens\", \"Essentiel-Jev never authority / human every action\", \"enzo-mcp independently falsifiable claims / ≠ jev-sift\", \"pigeonhole OTHER skip / decision-as-filing\", \"jev-reliability Nothing about accuracy\", \"clduab11/jev-test ≠ realZachi/jevtest / Nothing runs yet\", \"jev-rag-benchmark Jev wins is not an assumption\", \"dairui1/jev-lab ≠ BrendanH18/jev-lab\", \"jevmail gmail.readonly / mailjay archive/trash\", \"ZHUBoer/ego-jev reserved __none__\", \"runWorkflow completed ≠ success\", \"jsort scores are relative\", \"Noul not Choice for scale\", \"groundedness-judge-bench native vs schema-guided\", \"implicit_true included in yes\", \"jev_playground 0 promotions\", \"routing-backtest 0.0447%\", \"yuyang2230/jev-agent-skill jev-1.13-free\", \"jev-techstack-classifier stack_config.json\", \"s1_ruby collapse late\", \"undecided? abstain\", \"2389-research/judgement license null\", \"confidence ≠ winner p\", \"typesafeai-sdk-community not a new species\", \"tpellet/hunch exit 3\", \"never-execute list\", \"jev-file-search scores not calibrated accuracy\", \"jev-linkmap Jev never sees S2 prose\", \"muhammedilyasy/jev-mail metadata only\", \"tidy none-of-folders stay\", \"tab-bouncer pinned/audio/current never closed\", \"lkclean Show fail-open\", \"jev-yt-time-saver Show anyway\", \"ORIGIN pause-if-no-Jev\", \"validResponse sums-to-1\", \"jev-crawlers risk bands never raw boolean\", \"jevbrain AUTO_ACT is not a Noul\", \"judgekit YAML classify/score/route/verify\", \"typed-judge-kit verdict-in-code\", \"alsoleg89/decide packing VOI\", \"0.8 ≠ 80% accuracy\", \"Jev-Calibration Platt ECE 0.117→0.052\", \"jev-calibration-arena never acts\", \"ctmx/openrouter-jev-mcp Decision-as-Plugin\", \"FrancoisChastel/jev-code ≠ npm jev-code\", \"claudecode-jev-marketplace fail-open not hot path\", \"pedroknigge/mcp_jev packs not ask_jev\", \"cyrusasco/typesafe-mcp noul deadband 0.35–0.65\", \"codaaiteam/jev-skill jevtypesafeai.com ≠ TypeSafe\", \"hermes-switchyard ≠ hermes-jev-router ≠ hermes-plugin-jev\", \"nanoprune 2.8MB ECE 2.58%\", \"smartdio/jev-browser-agent ≠ ZHUBoer/ego-jev\", \"Dakai/omp-jev-web DONE ≠ proof\", \"hari007sh/jev ≠ dannote/jev\", \"0thernet/system-one-skills deterministic verify\", \"typed-gate band [0.40,0.60] is refusal\", \"pi-jev-gate fail-closed; choice is the verdict\", \"Foq ~25ms/2.2GB local\", \"rev prefill-only + HF jev-0.5b\", \"robfrase/jev planning memo\", \"typesafe_agent_gates 27/27 / 31/31\", \"EpicEric/safe-sh static remainder\", \"pastepilot Confirm before act\", \"Jev-Reranker live Jev not yet measured\", \"sessionwise opt-in relevance\", \"jev-search pointer sieve\", \"400ms Salesforce WebMCP\", \"typesafe-scheduler-diagnostics advisory\", \"droidjev screenshot-free\", \"Tewoto1 jevcu planner still writes\", \"ha-conversation-jev Jev→Grok\", \"dsh-jev can only gate\", \"jev-classification-benchmark specified not run\", \"jev-luna-pagerduty p≥0.50\", \"meldltd/meldecision laya-go ONNX\", \"laya-doom never pixels\", \"logixism/laya-api empty README\", \"akpsahan/laya ≠ Archer\", \"choxos/jevchess engine owns truth\", \"jev-drive sim not AV\", \"story-arc Jev never authors\", \"jev-hs-assistant HS6\", \"golergka/jev-plays-starcraft-2 UI-verified ≠ API Victory\", \"awesome-jev-use-cases catalog\", \"Nibir1/typesafe-go ≠ official\", \"fingerprint after redact\", \"recall vs decide\", \"publish fingerprints+answers\", \"CI replay as Harbor cousin\", \"Cache hit ≠ correctness\", \"hyperspaceai/jevcache ≠ kushals256/jevcache\", \"human labels only\", \"score never auto-accepts\", \"production capture flywheel\", \"sutro-sh/jev-align ≠ caiovicentino/jev-align\", \"guidance ≠ hook\", \"catalysts ≠ summaries\", \"compile-time System One\", \"unofficial ≠ TypeSafe\", \"format_version modernbert-jev/1\", \"Argos1111/jev_local ≠ us/jev-local ≠ kunchenguid/local-jev\", \"LFM default ≠ ModernBERT backend\", \"Nemotron ≠ TypeSafe Jev\", \"not a calibrated replacement\", \"djev-dev complements djev-spark\", \"images as Choice options\", \"Laya essay numbers *theirs*\", \"Router/OOD confidence\", \"hosted bootstrap ≠ silent TypeSafe\", \"difficulty + policy thresholds + JSONL trace\", \"jev-codex-pilot model + reasoning depth\", \"keep/shadow/hybrid/reject\", \"quarry evidence projection\", \"Frank-ZY-Dou/awesome-jev robotics/3D/control\", \"one-dollar-tahoe TypeSafe Jev defense eval\", \"jevguard calibrator/cache/escape\", \"jev-ci-selector CI shadow mode\", \"llama-jev llama.cpp replica\", \"petercr/jev-orchestrator ≠ FleeexCorp/jev-orchestrator\", \"seb4ez/jevguard ≠ AseemPrasad/JevGuard ≠ pablozr/JevGuard\", \"webNeat/llama-jev ≠ WiktorB2004/llama-index-jev\", \"OpenCode jev-pruner context sieve\", \"observe→score-candidates→prune\", \"jev-zen / jev-1.13-free\", \"zen-chat ≠ Noul\", \"fail-open original\", \"keepScore >0.1 floor\", \"host port of tamaratran/jev-pruner\", \"indiejoseph/opencode-jev-pruner ≠ nrdz-labs/fast-jev-opencode\", \"jev-webagent-bench empty stub\", \"Kiln-AI/jev_jsonschema noul_threshold 0.5\", \"NSStudent/JevSwiftSDK unofficial\", \"GLiNER2 native Apple path\", \"unofficial Swift/Core ML GLiNER 2.5-small\", \"entity spans + confidence\", \"not Choice/Score/Noul\", \"not TypeSafe\", \"label descriptions as schema\", \"on-device ANE economics\", \"honesty locks\", \"shershah1024/gliner-native-runtime ≠ Fastino\", \"≠ gliner25-compaction ≠ gliner2-ultrafast ≠ Eran-BA/Jev_from_GLiNER2 ≠ NSStudent/JevSwiftSDK ≠ jevmlx\", \"default threshold 0.1 still soft\", \"Decision Graph Protocol frame→assess→commit\", \"app retains permissions/effects\", \"Jev-first assessor-neutral\", \"guarded commit / receipt/next frame\", \"assessment batching\", \"hard-gating DGP as safety theater\", \"numerous-com/dgp ≠ TypeSafe official\", \"jegrep calibrated path+range Nouls\", \"no embeddings/index/daemon\", \"~$0.01–0.03 typical\", \"agent --json\", \"can1357/jegrep ≠ Bentlybro/jevgrep ≠ uehaj/jev-semgrep\", \"Archer-arch fidelity\", \"kev family OOD 0.76–0.77 vs Jev 0.86\", \"block-causal isolation\", \"pointer/readout CE-trained\", \"/v1/systemone drop-in\", \"replica honesty\", \"cost-sensitive decision theory × System One probabilities → control flow\", \"thresholds derived from costs not hard-coded\", \"YES / NO / UNSURE from cost_false_yes / cost_false_no / cost_human\", \"auto-batching same-object questions\", \"Kungie/gut ≠ tpellet/hunch ≠ carldaws/hunch\", \"judgment vs generation\", \"deterministic execution after probabilistic judgment\", \"exactly one app-owned callback\", \"explicit uncertain branch\", \"Illusion47586/judge ≠ lexingtonhibiki/judgekit ≠ Ascurse/typed-judge-kit\", \"variable-N option scoring as the trainable object\", \"dynamic candidate bags not fixed label sets\", \"zwliJay/jev-forge ≠ NanoJev\", \"open replica economics / latency vs closed Jev\", \"NAR local drop-in\", \"wfzyx/von late-catch HIGH\", \"competing NAR claims / replica honesty\", \"typed judgments vs chat judges on guardrailing\", \"ishaannk/llm-vs-jev cross-note only\", \"deeper integrity fold is rh-guard\", \"nothing wins outright\", \"can be argued out of guarding\", \"Jev IS the if-statement\", \"judgments/probabilities drive branches\", \"text model only writes prose\", \"interpreter owns variables/loops/budgets/replay\", \"otherwise maybe / confidence gate\", \"chaos samples after the gate\", \"southpolesteve/probably ≠ carldaws/hunch ≠ feelings ≠ Kungie/gut ≠ Illusion47586/judge ≠ tidymodels/probably\", \"133★ / forks 10 live\", \"build calibrated classifiers from human feedback\", \"retrieve by relevance not resemblance\", \"one calibrated yes/no per memory in one request\", \"pointer mode 17/18 19/20 *theirs*\", \"embedding resemblance misses the allergy\", \"samdotmak/jev-recall ≠ jev-search ≠ jev-sift ≠ carryforward ≠ chopratejas/invalidate\", \"memory leases ended by new evidence\", \"six Nouls then fixed rules in code\", \"0 of 157 false invalidations\", \"questions/plans/directives are not evidence\", \"unsure → review queue\", \"host keeps the store\", \"name↔body / comment truth / test-claims\", \"mizchi/jev-lint is mizchi/jevlint rename\", \"no shipped rule has severity error\", \"~1 in 5 findings wrong *theirs*\", \"mizchi/jev-lint ≠ huntedman/JevLint ≠ MichitoSugawara/jev-lint\", \"JSON Schema → typed JSON via Jev\", \"noul_threshold 0.5 decoder not a proof\", \"IncompatibleSchemaError lists every bad property\", \"on-device Laya CoreML ANE\", \"~5 ms P50 short decisions\", \"189/189 FP16 checkpoint parity\", \"10× not achieved\", \"mizorewww/laya-coreml ≠ gliner-native-runtime ≠ jevmlx ≠ NandhaKishorM/laya\", \"softmax over allowed tokens ≠ Noul\", \"question-first cache\", \"Micha0827/snapjudge ≠ githubnext/localjev ≠ jevmlx ≠ cendress/SnapJudge\", \"Jev-first Pi agent loop\", \"slow-LLM fallback\", \"explicit action menu / CandidateSource unimplemented\", \"62 tests wiring not quality\", \"direwolfiy/JevPi ≠ standardagents/jevpilot ≠ pi-jev-control\", \"resume-screening bias audit methodology\", \"name×resume factorial independent Nouls\", \"callback determined by resume quality\", \"mean-probability name gaps operationally negligible\", \"natemoo-re/bias-bench ≠ BBQ\", \"Plan/PRD panel → code-owned pass|review|block\", \"cheerleading out of scope\", \"austindixson/planalyzer ≠ single-goodness Noul\", \"cost-aware multi-model routing/escalation\", \"decide vs do\", \"successful-task cost\", \"cannacre8ive/switchboard-ai ≠ ha-switchboard ≠ hermes-switchyard\", \"frozen-protocol zero-shot bench\", \"TypeSafe Jev vs PrismNLI vs Laya\", \"contamination caveat\", \"elcronos/jev-vs-open-decision-models ≠ JevBench ≠ DMB\", \"context-window admission control\", \"VOI gate which tokens are worth the expensive model\", \"fail polarity per lens\", \"on small inputs lenses lose money\", \"cvsgireesh/jevusher ≠ jev-sift ≠ winnow\", \"typed decision control plane\", \"receipt ≠ authorization\", \"historical-v0 zero retained cases\", \"MokiMeow/jev-fabric ≠ jev-forge ≠ dgp\", \"live 15-dim typed rubric re-score per pause\", \"scoring economics exemplar\", \"OpenJev/Codiv ≠ TypeSafe hosted\", \"jose-troche/live-rubric ~$0.000004 desc / ~$0.000006 README\", \"adversarial pre-registered Jev eval\", \"28 predictions before data\", \"123,805 requests\", \"confidence does not track ignorance\", \"polite injection 65% / crude 0%\", \"willkelly/jev-evaluation ≠ jevals ≠ jev-baselines-eval\", \"provider-neutral Elixir/BEAM Noul/Choice/Score SDK\", \"class infrastructure\", \"nshkrdotcom/system_one_sdk ≠ typesafe_sdk ≠ dannote/jev\", \"question-linting of Jev questions themselves\", \"nine jaggedness rules, no API key, no labelled data\", \"static lint ≠ measured separation\", \"yodablocks/jevq ≠ tenbin ≠ JevLint ≠ commitjev\", \"open-weights Laya as class exemplar (binding)\", \"Nx/Bumblebee runtime\", \"host chooses backend\", \"ChristianAlexander/laya_ex ≠ system_one_sdk ≠ dannote/jev ≠ NandhaKishorM/laya\", \"on-chain/edge Laya deploy\", \"parity_verified stays false\", \"model output never grants Tx\", \"humandebri/IC-Laya ≠ laya_ex\", \"auditable weekend replica\", \"Jev outputs never used for training\", \"soft human-vote distributions\", \"unpaired 0.577 vs 0.727\", \"agilabs-ai/jev48 ≠ JevBench ≠ Mapika/decider\", \"adversarial dual-judge / framing attack surface\", \"comparative framing is the usable judgment\", \"prior injection crowds out evidence\", \"copyleftdev/ember ≠ ember.js\", \"Laya specialist fine-tune pipeline\", \"training still GPU-pending\", \"PIXELZX0/XERON ≠ convaiinnovations/laya\", \"Hub Laya replica drop\", \"daliborsb/laya ≠ convaiinnovations/laya ≠ NandhaKishorM/laya\", \"System One student distillation corpus\", \"gold is programmatic\", \"teacher is closed-API clone\", \"do not distill Jev as teacher of record\", \"MagaBitmex/jev-4b-distill-data ≠ missing student checkpoint\", \"non-LLM VIN System One\", \"planning depth not chat\", \"lewislululu/jevon ≠ douglance/jevon\", \"source-bound evidence checks\", \"local quote mismatch needs no API\", \"exit 0 ≠ claim truth\", \"WaynezProg/jev-kit ≠ jonathanavis96/jev-kit (Airlock) ≠ jev-use ≠ jev-mcp\", \"independent System One evidence catalog\", \"scores not one leaderboard\", \"no external record currently reproduced\", \"TokenTrim no-Jev matched hybrid 62.4%\", \"reachjalil/system-one-bench ≠ mallahyari/system-one-benchmark\", \"21 tasks · 134 items · 208 questions\", \"scenes from public GitHub contracts, not production logs\", \"SivletLabs/jev-eval ≠ willkelly/jev-evaluation ≠ 4esv/jev-eval ≠ xxkuboxx/jev-eval ≠ onlyoneaman/jev-eval ≠ dayhaysoos/jevals\", \"option isolation (sibling-blind)\", \"permutation-equivariant\", \"Hub OWNER not published\", \"nafisazizir/hev ≠ jaredpalmer/kev\", \"frozen local LLM logits, no trained decision head\", \"residual-head 9,222-param decreased 73/96→67/96\", \"confidence = 1−normalized entropy, not P(correct)\", \"yuki-oshio/mini-jev ≠ r-ms/mini-jev\", \"Jev classifier as autoregressive next-token predictor\", \"ChatJev-style soundness theater\", \"erik-dunteman/ChatJev ≠ dannote/jev ≠ jev-gpt\", \"calibrated decision head × AlphaProof value head\", \"implementation-layer isomorphism, semantic difference\", \"timeout = censoring\", \"do not launder Noul as proof\", \"parallel rank-prediction vs serial selection\", \"independent questions can conflict\", \"zzzzzec/jevsort ≠ keltokhy/jsort\", \"curated open System One ecosystem catalog\", \"rupeshpoojary9/awesome-open-system-one ≠ AnotiaWang/awesome-jev\", \"arXiv paper radar with Jev relevance scoring\", \"ranking ≠ calibration / 0.5 still soft\", \"fail-open failed evals not marked seen\", \"train calibrated ~27M from scratch\", \"typed Q→prob dist / one forward pass / no LLM decode\", \"hyusi2003/MiniSystemOne ≠ Colvin0315/MiniSystemOne\", \"description-only stub / size 5\", \"ESCI hard probe fails four of six\", \"jev_bool ECE 0.242 inversion 0.255\", \"do not re-fold §60 six-gates as new\", \"jobbyjev one-request-per-company from batch-size result\", \"find/design/evaluate TypeSafe Jev decision loops\", \"karanb192/jev-architect ≠ samtay32/jev-system-architect\", \"Jairik/jev-distiller size 1\", \"distill-Jev UI stub / do not distill Jev as teacher of record\", \"post-launch scored use-case map / Jev self-scores then human curation\", \"licensedsaucer9-web/jev-opportunities\", \"Jev-inize a use case into classifier/router\", \"gavinHuang/jevinize → simple-jev not TypeSafe\", \"featherless-ai/simple-jev\", \"compare saved decisions / same label can still change the branch\", \"VihaanAgarwal/jev-diff ≠ Saik0s/diffusiongemma-jev-macos\", \"not tested with a live Jev API key\", \"constrained logprob + temp/Platt ≠ Noul\", \"OpenJevPro pastes openjev-sglang JevBench as own\", \"zhangcy122/OpenJevPro ≠ IamBusy/OpenJev ≠ ekzhang/openjev-sglang\", \"PolyForm Noncommercial\", \"SmolLM-135M / sub-70ms / 0 output tokens\", \"demo P(True) 0.5052 / Choice conf 0.2872 / Score conf 0.0055\", \"README claims MIT / GitHub license null / no LICENSE file\", \"patelvishwa112/jev-system-one-rlcd ≠ arnabgho/rlcd-lite ≠ blackwood-rlcd\", \"source-backed Awesome Jev radar / 306+ commit-pinned\", \"logicrw/awesome-jev-projects ≠ AnotiaWang/awesome-jev ≠ yibie/awesome-jev ≠ cobanov/awesome-jev ≠ rupeshpoojary9/awesome-open-system-one\", \"auto GitHub sync / Issue-only submissions\", \"hashed n-gram encoder / rival-aware attention\", \"olanotolu/jevbetter vs jevlike starter\", \"synthetic hard menus top-1 0.916 vs 0.873 / ECE 0.0182 vs 0.0367 / 40 vs 4608 menus/sec\", \"shuffled-context control 0.335\", \"structured probability readouts\", \"distribution > argmax\", \"Noul 0.5 midpoint\", \"score is expectation not integer\", \"bare HTTP not SDK\", \"Arohtea/jev-readout\", \"Jev-style Choice/Score/Noul from ordinary models\", \"optional DSH plugin\", \"schema-valid ≠ calibrated\", \"gulagala001/jevify ≠ Mintzs/jevify\", \"Laya RLCD benchmark\", \"40.3% below constant-answer\", \"open-weight measurement\", \"mourad-ghafiri/laya-rlcd-benchmark ≠ yibie/laya-jev-lab\", \"cheap fail-open semantic edge\", \"second signal not sole\", \"FastLoopError catch\", \"SupremeDreamZ/jev-fastloop ≠ jev-ultrafast\", \"asking more questions in one call\", \"0.980 at every N\", \"nearly not fully deterministic\", \"TheWebDevel/jev-fanout\", \"Qwen3-VL perception + Jev decisions train RL\", \"0 model calls at deployment\", \"VLM alone 1.7 vs +Jev 4.4\", \"harneet2512/reflexrl ≠ khordoo/jev-reflex-autonomy-lab\", \"independent Jev API vs Laya\", \"cascade 0.60 matches 78% at 1.8×\", \"noul facts not judgements\", \"yibie/laya-jev-lab ≠ dairui1/jev-lab ≠ BrendanH18/jev-lab\", \"GLiNER vs GLiFormer vs Laya vs Jev\", \"extractors ≠ decision engines\", \"Laya dict-instructions collapse 58.3%\", \"umstek/zero-shot-ie-bench\", \"decisions-per-minute & cost\", \"204 moves vs 73\", \"throughput not intelligence\", \"angelgalvisc/snake-arena-jev-vs-llms ≠ vtrivedy/jev-plays-games\", \"behavioral contracts\", \"pin expectations eval upgrades\", \"raw 0.94 is not a release\", \"sathariels/jevcheck ≠ dayhaysoos/jevals ≠ SivletLabs/jev-eval\", \"evidence-linked dependency upgrade\", \"Jev never generates filenames\", \"no_direct_evidence ≠ safe to merge\", \"GaneshVG18/upgrade-radar ≠ LYchoon/paper-radar-jev\", \"discography theme/mood/complexity\", \"five atomic questions one call\", \"lirantal/discoprint\", \"Turn any open LLM into System-One Jev\", \"uspraveen/Jevify ≠ Mintzs/jevify ≠ gulagala001/jevify\", \"Jevify-any-LLM architecture probe\", \"description-only stub / size 0\", \"Train encoder-only calibrated decision models from a task sentence\", \"Exu is a toolkit, not a method\", \"strictly proper scoring rule\", \"Pre-alpha\", \"Ruivalim/exu-base\", \"scratch-trained calibrated decision model\", \"typed Q → probability dists\", \"Colvin0315/MiniSystemOne ≠ hyusi2003/MiniSystemOne\", \"no published weights download URL\", \"90.5 seconds / 29.2% pipeline evidence\", \"p_i/p_j independent of other candidates\", \"Recipe for calibrated decision models — small model out\", \"init → synth → train → eval → serve\", \"91.1 % / ECE 0.022 *theirs*\", \"Jev zero-shot 75.1\", \"scienthoon/luce\", \"Put Jev's three headline claims on trial\", \"0.5B local GPU\", \"46x speedup / accuracy identical\", \"ECE 0.624 sentiment catastrophe\", \"bigger model worse calibration\", \"RichardoMrMu/jev-mini ≠ yuki-oshio/mini-jev ≠ r-ms/mini-jev\", \"System-1 decision engine for local LLMs\", \"structured choices only\", \"JSON parse of generated text ≠ Noul\", \"TypefAI JEV / Journal Entry Voucher\", \"tapsin/jev-local ≠ us/jev-local ≠ Argos1111/jev_local\", \"Jev 1.13 reward-model eval across 8 benchmark tracks\", \"40,940 examples / 0 API errors\", \"RewardBench v1 92.58%\", \"Precise IF 50.63%\", \"goya4140/jev-reward-model-evaluation\", \"Scaffolding in progress\", \"Jev vs LLM support-ticket routing\", \"static + live decision bench\", \"TypeSafe's own published benchmark\", \"illustrative simulations, not live API calls\", \"JevBench v1 — smart/cheap/fast/reliable\", \"I/C/S/K 25% geometric mean\", \"classifier.dev fast tier 84.8 is Jev behind its own API\", \"do not re-fold §78 v1.2 board as new\", \"Laya (421M) 70.1 now on board\", \"Zero-shot/few-shot LLM routing\", \"hard budget filter before Jev\", \"Jev never asked to perform budget arithmetic\", \"Jev judges the next state, XState enforces transitions\", \"simulation uses synthetic keyword fixtures\", \"catalog gravity\", \"v-modal/awesome-jev-tools\", \"★339 live REST\", \"curation is not endorsement\", \"crawler-maintained directory\", \"Daily GitHub + npm sweep, human-merged\", \"RadRebelSam/awesome-jev ≠ AnotiaWang ≠ yibie ≠ cobanov ≠ logicrw ≠ v-modal\", \"HF peft SPLADE/BGE reranker\", \"rdxtremity/jev-reranking ≠ carlaiau/jev-reranking\", \"query-side encoders, not a Jev replica\", \"ONNX System One Qwen3.5-4B scorer\", \"source:pngwn/system-one-qwen3.5-4b-scorer\", \"CC-BY-NC-4.0\", \"temperature 1.75\", \"transformers.js AutoModel cannot load this graph\", \"Consistency benchmark Space\", \"This Space contains no benchmark result yet\", \"12-case plumbing fixture\", \"Benchmark-driven Jev router and judge\", \"cheap alone is not success\", \"Jev does not write, sum prices, or claim accuracy %\", \"Sol 94.2 / Luna 83.9 / Jev path 89.7\", \"19.2% Sol / 62.3% cost save / 4.5pp miss of 2pp non-inferiority\", \"p50 latency worse than Sol due to routing overhead\", \"erendikmenn/jev-llm-router-benchmark ≠ jev-rag-benchmark ≠ ryantsai/jev-llm-router\", \"Express + node:sqlite\", \"mock and Jev decision engines\", \"previous_ticket_count >= 3 is code\", \"MIN_CONFIDENCE 0.6 still soft\", \"substring false positives\", \"aesaganda/jev-ticket-router ≠ SarathChandraBellam/jev-vs-llm-ticket-router\", \"Universal Figure & Diagram Router\", \"confidence ≥ 0.85 hard-gate is theater\", \"generative AI banned from scientific plots\", \"six visual branches\", \"hoangngochuong24947-gif/jev-figure-router\", \"human-labeled (state, question, label)\", \"166,054 rows / 22 configs\", \"soft_label for human uncertainty\", \"Praveenrajus/jev-bench ≠ fstandhartinger/jevbench\", \"ternary bonsai System One GGUF\", \"openjev's mechanism, Bonsai's weights\", \"Hub does not ship weights\", \"100/100 easy T/F is not Harbor\", \"label_mass ≠ correctness\", \"stock llama.cpp Q2_0 silently gibberish\", \"NicolaiMTLassen/open-bonzi-jev ≠ NicolaiLassen\", \"transformers.js DeBERTa ONNX\", \"source:com-kotobalabs/open-jev-deberta-v3-large\", \"temperature 1.05\", \"AutoModel from_pretrained works\", \"onnx-community/open-jev-deberta-v3-large-ONNX ≠ system-one-qwen3.5-4b-scorer-ONNX\", \"107★ densify\", \"GH 151M vs README 149.6M\", \"PR #1 now closed unmerged\", \"do not re-fold §71 claim-audit as a beat\", \"typed decisions, RLCD, confidence-gated routing\", \"structured ≠ correct\", \"mock not live API\", \"26 tests\", \"wjdjdakf17/jev-study ≠ baekenough/jev-study\", \"bonzi-27b-v2 / ternary-8b / 27b-v1 GGUF family densify\", \"Hub still does not ship weights\", \"WANLI-256 74.6% / 65.2% / 71.1% *theirs*\", \"Bonsai 1 27B Q1_0 runs on stock llama.cpp\", \"ternary still needs PrismML fork\", \"hf:heman10x/openJev-verdict-2.0 twin tokenizer-only\", \"OpenJev Vision image classification + uncertainty\", \"CLEVR-4 held-out joint 0%\", \"hfdataset:IamBusy/OpenJev-Vision-Research-v0.1 12,832\", \"294,912 derived targets not independent samples\", \"Laya multilingual ONNX WebGPU typed-decisions port\", \"63/63 selected answers / 5.1e-4 CPU / 1.2e-2 WebGPU\", \"UpHash-Network/mini-jev is yuki-oshio transfer\", \"jev-injection-bench 11,900 labelled prompts\", \"Jev best ranking / Haiku better ECE 0.021 vs 0.058\", \"0.5–0.9 band is where Jev's numbers do not mean what they say\", \"Prompt wording moves panic 28%\", \"manojlds/jev-dspy-bench ≠ dspachos/jev-dspy ≠ jmanhype/jev-dspy-lab\", \"Jev agreement is similarity, never ground truth\", \"no aggregate quality grade or merge gate\", \"AbstentionBench-on-Jev rank 1 of 20 vs 2025 field\", \"question-asymmetry\", \"forward-looking 0.465 never extreme\", \"openkev calibration layer not a runtime\", \"ECE vs coverage independent\", \"select_threshold returns inf\", \"escalation catches uncertainty not ignorance\", \"misakaikato/openkev ≠ jaredpalmer/kev\", \"pdf-race Docling→Jev vs Gemini\", \"parser owns the wall clock\", \"12/12 tie is a tie\", \"titles selected not generated\", \"flopcheck 16 calibrated tweet judgments\", \"mechanical tells in code\", \"ZeroX-01/jev-atlas ≠ Zaious/jev-capability-atlas ≠ gorock007/jev-atlas\", \"catalog not endorsement\", \"Laya calibration lab Gradio MCP\", \"T never changes argmax\", \"confidence ≠ top-label p\", \"easy probe set refused\", \"40–48 rows too small to ship T\", \"Gemma-4 26B-A4B jevify classification+calibration\", \"LoRA adapter twin not independent eval\", \"Gemma-4 E4B jevify\", \"E4B LoRA stub card\", \"kushalpatil/jevify-gemma4 ≠ Mintzs/jevify ≠ gulagala001/jevify ≠ uspraveen/Jevify\", \"GH kushalpatil07/jevify 404\", \"PAWS 0.580/ece 0.288 is the weak cell\", \"smaller E4B slightly better OOD ECE than 26B-A4B\", \"Hub jevify merged LoRA ships weights\", \"bonzi Bonsai-8B v1 GGUF densify\", \"Bonsai-1.7B v1\", \"Bonsai-4B v1\", \"WANLI-256 64.5% / 60.2% / 52.0% *theirs*\", \"rank #4 / #5 / #6 of 6\", \"JulesHuisman/jev-eval scaffolding / README SHA c356a584 (was empty e69de29b)\", \"JulesHuisman/jev-eval ≠ SivletLabs/jev-eval ≠ willkelly/jev-evaluation ≠ 4esv ≠ xxkuboxx ≠ onlyoneaman ≠ dayhaysoos/jevals\", \"7 bands 6/10 vs 40 bands 0/10\", \"source receipts + confidence slider re-policy without re-inference\", \"32/32 synthetic is smoke not production\", \"classify HF datasets across typed semantic dimensions\", \"roadus2 watch misspelling\", \"lock roadius2/ultra_laya\", \"ultra_laya REVIEW defects\", \"default branch claude/laya-jev-review-gg5ppo\", \"XNLI EN 88.3% ECE 0.032 → RU 77.3% ECE 0.096\", \"Δ −11.0 pp [−14.2,−7.8]\", \"ECE +0.063\", \"MASSIVE no detectable difference at n=600\", \"confidence is function of p_max (r=1.000)\", \"pointer-not-generator 400 human-authored responses\", \"proposed ≠ authorized\", \"FewRel 160: Jev 85.0% vs lexical 13.125%\", \"gated 100% (95/95) coverage 59.375%\", \"J++ composable semantic computation language\", \"judge-jev 0.5 still soft\", \"947 repos scored\", \"A 273 / B 302 / C 372\", \"LLM rubric ≠ benches\", \"No benchmark winner is claimed\", \"phishing: naive 62.6% vs regex 91.8%\", \"5-atomic + LR 95.0% *theirs*\", \"AITuber tension ±15\", \"README npm global\", \"repo is Rust\", \"git-confess code owns counting/blame/ratio\", \"httpx exhibit 11% (13/119) *theirs*\", \"90d trend +12.40% vs random +12.75% vs BH +41.71%\", \"5m win rate 25%\", \"Awesomejev 656 entries / 38,160 stars\", \"tracker likes 64 (+4) lastModified UNCHANGED\", \"Laya present\", \"Blackwood ABSENT\", \"Archer still promised_not_landed\", \"Blackwood tracker ABSENT; likes 2 gated manual\", \"r = c - p_a\", \"ECE 0.021; acc 0.807 vs warmup 0.746\", \"calibration beyond ~500 tokens unmeasured\", \"Independent primitive\", \"11.57s vs 54.10s · 4.67× · 120/128 *theirs*\", \"default path is pretrained Gemma probs not trained RLCD head\", \"GH Meanblock 404; lock leesk212/JEV-CPU\", \"softmax over letter slots ≠ Noul\", \"WANLI 0.741 vs openjev v2 0.77 *theirs*\", \"3-way NLI ≠ Noul\", \"priority 0.464 = majority floor\", \"banking77 contaminated\", \"raw margins not probabilities\", \"GH jev-haiku-benchmarking 404\", \"do not distill Jev as teacher of record (they distilled Haiku)\", \"“0.9 is not one number”\", \"ranking ≠ calibration\", \"banking77 0.8–0.9 stated 0.86 actual 0.73 over-confident *theirs*\", \"≠ Praveenrajus/jev-bench ≠ fstandhartinger/jevbench\", \"$0.0000153–$0.0000226 vs circulating $0.0004 (~20×)\", \"Score is 0..n-1 expectation not 0–1\", \"Noul has no confidence field\", \"TCP floor 198.8 ms\", \"type reliability is not a reason to choose Jev (json_schema 5/5)\", \"gateway tax not one number\", \"Function-only 5/8 vs hybrid 8/8\", \"4/8 without Jev\", \"8 designed cases not conversion lift\", \"≠ RadRebelSam/awesome-jev\", \"200-row pilot Jev 86.5% 173/200 vs Gemini Flash-Lite 86.0% 172/200 vs Pro 87.0% 174/200 *theirs*\", \"not a ranking\", \"NLI Tetris argmax P(entail)−P(contradict)\", \"情緒測謊器\", \"1q 396ms / 30q 567ms\", \"±0.03\", \"33q $0.000045 vs Gemini ~5× slower ~60× cost *theirs*\", \"≠ realZachi/jevtest\", \"8-example Jev vs GPT-5.6 Sol ~64× cost 5.4× latency *theirs*\", \"synthetic; no inference\", \"≠ JevBench v1.2 §78\", \"Judged 3317 / listed 2560\", \"Jev judges, code applies policy\", \"catalog ≠ endorsement\", \"APA “microsecond policy / zero hallucination” overclaim\", \"Client-side quiz; pointer from held docs; scanned-PDF warn\", \"CSP only api.typesafe.ai\", \"Jev judges / agent reasons / user decides\", \"selecting an option is not permission to implement\", \"degraded fallback\", \"pattern exact, judgement must clear floor\", \"no matching pattern → no model call\", \"not a correctness oracle\", \"$0.00022 vs chat $0.00306 *theirs*\", \"Spec vs artifact remainder\", \"treating 0.85 as 85% / minProbability hard-gate as Harbor\", \"VERIFY acquires discriminating evidence, never same-pool confidence-only rescoring\", \"fast/full/max are ceilings not sizes\", \"Solar writes, Jev chooses NEXT ACTION\", \"SemIf 2186★ (+20 vs §109 2166)\", \"jevlike 1038★ (+7 vs 1031)\", \"TypeAR 14★ flat\", \"AnotiaWang 96★ (+1 vs 95)\", \"yibie/awesome-jev 490★\", \"Laya likes 802 (was 783)\", \"tracker likes 64 flat, lastModified UNCHANGED\", \"do not reopen or amend PR #23/#24/#25/#26/#27.\", \"do not reopen or amend PR #23 or #24 or #25 or #26 or #27\", \"Calibration is not alpha\", \"NO CURRENT ALPHA CANDIDATE\", \"ΔR² approximately +0.00084\", \"Brier 0.2131387\", \"ECE 0.0421875\", \"Adding Jev probability to deterministic volatility improved Brier by only 1.4058e-05\", \"default 0.5 keeps zero non pinned\", \"keepResult median 0.14 to 0.17\", \"keepCall median 0.28 to 0.35\", \"usable range is about 0.10 to 0.25\", \"7.8% to 57.9%\", \"judges results it never sees\", \"task-finish eval not built yet\", \"$0.002 per compaction\", \"slavadubrov/sgr-judge-bench ≠ slavadubrov/jev-judge-bench\", \"Jev 108/120 $0.083 0.34 s\", \"Luna SGR 114/120\", \"paired Jev accuracy-difference intervals include zero\", \"not evidence of equivalence\", \"GLM SGR 26/120 93 format failures\", \"Terra-planned Jev hybrid 55/120\", \"rule-based by default, optionally Jev-backed\", \"empty README\", \"missing key cannot break the experience\", \"prefill plus exactly one decode\", \"softmax over A/B/C ≠ Noul\", \"BBQ 9,053/10,000 (90.53%)\", \"ECE 0.0890\", \"Mean confidence 0.9943\", \"overconfident\", \"score and noul not implemented\", \"DGUI 12 rows (was 6)\", \"INSTRUCT 119 rows likes 2\", \"encode the state once, decide everything in parallel\", \"0.740 accuracy against a 0.508 majority\", \"ECE 0.047\", \"fine-tune's advantage ends where its 384-token training data does\", \"jasonkneen/open-jev ≠ pngwn/open-jev\", \"same sha d41dc3cd\", \"Space does not call Jev\", \"recomputes routing from saved probabilities\", \"200-case Jev 97.0% / 100.0% / 95.0% / MAE 9.22\", \"synthetic repository benchmark\", \"Jev evaluations are advisory\", \"YehuiTang0316/jev-nlgrep ≠ Bentlybro/jevgrep ≠ can1357/jegrep ≠ uehaj/jev-semgrep\", \"default threshold 0.8 still soft\", \"40-line windows cannot prove whole function\", \"token-native sequential start/end Choice\", \"Gemini/Haiku stubs not configured yet\", \"handful of hand-written examples, not a benchmark\", \"Jev judged exactly what it was given\", \"laguagu/jev-skills ≠ laguagu/jev-evidence-lab ≠ Pleo2/awesome-jev-agent-skills\", \"contract_passed is not a claim of guaranteed factual truth\", \"Wilson lower bound 0.85 floor\", \"fixture mode no savings claim\", \"SemIf 2207★ (+21 vs §110 2186)\", \"jevlike 1043★ (+5 vs 1038)\", \"TypeAR 15★ (+1 vs 14)\", \"AnotiaWang 97★ (+1 vs 96)\", \"yibie/awesome-jev 506★ (+16 vs 490)\", \"Laya likes 822 (was 802)\", \"tracker likes 64 flat, lastModified UNCHANGED\", \"do not reopen or amend PR #23/#24/#25/#26/#27/#28\", \"people who compare Jev against GPT-5.6 has never fine-tuned BERTForXYZ for living and it shows\", \"zero shot classifiers\", \"scale them as much as decoder only models\", \"many problems solved with LLMs could have been solved with them, it was a skill issue\", \"opt for DeBERTa and ModernBERT ones\", \"BERTForXYZ → DeBERTa → ModernBERT\", \"Jev vs GPT-5.6 bakeoffs are a category error\", \"encoder / ZS classifiers\", \"institutional HF voice\", \"quote *theirs*\", \"do not invent accuracy numbers\", \"softmax/ZS scores still ≠ calibrated Noul\", \"soft scores ≠ hard gates\", \"@mervenoyann\", \"likes 421 / 189\", \"impressions 35498 / 9613\", \"multimodal image<>text ZS as perception front-end\", \"hf:MoritzLaurer/deberta-v3-large-zeroshot-v2.0 likes 139\", \"hf:MoritzLaurer/ModernBERT-large-zeroshot-v2.0 likes 72\", \"Bart, bert, deberta, modernbert, these are all LLMs\", \"Maziyar quoted\", \"Jev is exemplar not the mandate\", \"do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29\", \"Heman10x-NGU/Verdict-open-jev ≠ Heman10x-NGU/openJev-verdict-2.0\", \"TF-IDF + LogReg ECE 0.0207 vs Jev 0.1440\", \"Verdict-open-jev 48.07% vs Jev 90.80%\", \"abstention combined recall 10.00%\", \"p50 35.58 ms\", \"K=25 (maximum capacity) 72.00%\", \"0.85 coverage 84.60% selective risk 1.18%\", \"26.1× faster than standard Qwen JSON generation\", \"Jevify 90.0% / 167 ms CUDA graphs disabled\", \"Finding 1: Brier on stated confidence alone is a trap\", \"grpo_rlcr 0.78 / ECE 0.084\", \"reliability 0.007 but resolution 0.000\", \"27 900 schema-driven decisions\", \"13 600 / 13 600 questions\", \"candidate mass min 0.99999624\", \"22 configs · 166,054 rows · 4 calibration-gold\", \"sha a39eba3f\", \"Student B MAE 0.148 / Pearson 0.836 / 86.0%\", \"pngwn/open-jev-laya-bench README 404\", \"sha 9f69c742 likes 2\", \"HDFS 0.9933 (745/750) / retain 0.0084\", \"BGL ERROR/FATAL protection 1.0000\", \"2,479 / 2,500 HDFS uncertain\", \"cache hit 0.9648 (2412/2500)\", \"$0.153936 estimated\", \"E2 recomputes from saved probabilities\", \"Space sha eda59e0a\", \"MASSIVE English 0.783 / Khmer 0.033 / Hindi 0.133\", \"40–48 rows too small to ship T\", \"T never changes argmax\", \"siren2345/jev-single-decode-transformers ≠ siren2345/jev-single-decode\", \"Split Transformers experiment from llama.cpp runtime\", \"tanayvasishtha/jev-lab ≠ dairui1/jev-lab ≠ BrendanH18/jev-lab ≠ yibie/laya-jev-lab\", \"Four experiments stress-testing TypeSafe's Jev: calibration, bundle bias, label bias, and ensembling\", \"second pass must be $0.00 from cache\", \"The pages never call Jev\", \"Gemma 4 31B 77.0% / Jev 1.13.0 61.4% / Laya 322M 0.0%\", \"restriction state 95.0% against 84.4%\", \"None of the systems are particularly good at knowing when to stop and ask\", \"They skip the question and call a tool directly\", \"100% schema pass\", \"six-field joint 48.8% vs 72.8%\", \"ywchiu/jev_benchmark ≠ Running-Dolphins/jev-bench ≠ Praveenrajus/jev-bench\", \"ACT / REVIEW / FALLBACK\", \"A provider failure, timeout, malformed output, or missing answer is **not** a policy outcome\", \"confidence is descriptive provider output, not a substitute for probability\", \"Quality denominators include only valid scored answers\", \"an exact halfway tie chooses the lower level\", \"aiwithenoch/Jev-Skill ≠ simplosophy/jev-skill ≠ laguagu/jev-skills\", \"The local path does not claim to turn a smaller checkpoint into Jev\", \"Low support becomes decision: \"review\"\", \"MIT-0 SPDX NOASSERTION\", \"current-llm\", \"结构兼容,不是 Jev 模型能力\", \"altryne/jevify ≠ Mintzs/jevify ≠ gulagala001/jevify ≠ uspraveen/Jevify\", \"Find where Jev belongs. Design the questions. Measure the difference\", \"TypeAR-AI/TypeAR 301 → TypeLLM/TypeLLM\", \"TypeLLM/TypeLLM 16★\", \"SemIf 2241★ (+34 vs §111 2207)\", \"jevlike 1051★ (+8 vs 1043)\", \"AnotiaWang 98★ (+1 vs 97)\", \"yibie/awesome-jev 525★ (+19 vs 506)\", \"Laya likes 864 (was 822)\", \"tracker likes 67 (+3 vs 64)\", \"lastModified UNCHANGED `2026-09-20T04:29:16.000Z`\", \"do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#32\", \"soft Noul ≠ hard safety\" with a placement, not a stack replacement or a vendor how-to. Formal methods are one pillar., \"A hunch is a probability with a policy attached\", \"{ enter: 0.8, exit: 0.6 } is hysteresis\", \"replay a policy change without inference\", \"Decision models are providers, not the product\", \"huncho ≠ Kungie/gut ≠ carldaws/hunch ≠ tpellet/hunch\", \"pretrained Qwen2.5 base ECE 0.030 (0.5B) / 0.040 (7B)\", \"instruct 0.302 / 0.269\", \"70.9% → 70.0% mean conf 74.1% → 96.7%\", \"temperature scaling still matches it in-distribution\", \"No Jev API was called\", \"Qwen2.5 ≠ Archer\", \"Qwen/Qwen3.8-27B ≠ Archer\", \"calibration does not compose\", \"ECE has exactly zero statistical power to detect the failure mode that kills trajectories\", \"25–60× headline withdrawn\", \"P(all-correct): 0.0071 vs 0.0001\", \"TCE / AMS\", \"Deferred Crispification\", \"Qwen 3.8 sparring ≠ Archer\", \"pd.cut bins by equal width while jeval bins by quantile\", \"ECE 0.113 and ECE 0.076\", \"jeval drift is not implemented yet\", \"ranking ≠ calibration\", \"g0runmezadam/what-is-jev IS tunahansahin897/what-is-jev\", \"same GitHub id 1378007307\", \"947 repos scored\", \"A 273 · B 302 · C 372\", \"LLM rubric ≠ benches\", \"Probabilities are advisory, not calibrated guarantees\", \"light_cutoff_applied_to_combination 0\", \"AND: product (independence assumed and recorded in the trace)\", \"circuit-vl-4b ≠ Archer\", \"soundness theater / measurement theater\", \"hourly 0843 / notes.md §114\", Not a substitute for the official typesafe-ai skill (live Jev API contracts)." +description: "Use when placing typed probabilistic judgment (Jev-class System One / decision models) with mathematical, logical, or algorithmic mental models — in AI, software, business, knowledge work, or life, not only SWE; deciding where a fast cheap categorization/classification/scoring model belongs versus generation, exact policy/code, or proof; applying expected utility, selective classification/abstention, calibration, cost-sensitive thresholds, value of information, MCDA, signal detection, search/control substitutions, or Leveson-style org/safety; using NATM/snap-fit/Norman as design intuition; designing mixed architecture (decision model + LLM writing); auditing an existing system, PR, workflow, or non-software practice for judgment-shaped holes and code smells; debugging a question that hovers near 0.5, clusters mid-scale, or hides two judgments; placing agent self-supervision gates (pre-action, output judge, done-check, stuck-detector, context sieve); coupling a typed judge as an optimizer metric (Ax, DSPy); choosing among TypeSafe Jev, open heads (Laya, kev, openjev-lm, Nimble, encoder DeBERTa, LoRA distill, open multimodal RLCD / blackwood), announced open decision-model (Watch — still not landed), constrained-AR (TypeAR, pcdServer, decision-token LoRA), diffusion structured reads, GLiNER/GLiClass/GLiGuard encoder family (locate vs categorize vs safety-schema classify vs local multi-head), listwise rankers, or vision scorers; placing judgment beside TLA+/Alloy/Apalache/Dafny/DST (Antithesis, Resonate, PufferLib) without laundering a Noul as a proof; answering \"it's just classification\", \"is Jev probabilistic programming\" (marginals vs joint, not a PPL), \"low/medium/high entropy\" (allocator, not a meter), \"perception specialist then judgment vs shared multimodal System One\", \"wait for Archer vs open multimodal RLCD\", \"screenshot/DOM candidates → typed Choice\", \"eval path\", \"jevals\", \"Harbor taskset\", \"shared bake-off ECE/NLL/Brier\", \"LLM-as-judge is not the System One score\", \"pipeline / measure / hill-climb perception into a decision\", \"Ax vs DSPy\", \"held-out\", \"correctness is not confidence\", \"is this only for software?\", Alloy vs Apalache, GLiNER vs Jev, \"is GLiGuard Jev?\", LLM-as-judge, paraphrase brittleness, allowlist then judge (allowlist *proves*; fail-open cannot block), \"missing other → confident wrong Choice\", \"lint the Jev request\", \"training confronts Choice other / none-of-the-above\", \"S1 reflex keeps control / optional S2 one-use advice\", \"soft AGENTS.md rules vs linter\" (Abide / jev-pref), \"edit-phase vs turn-phase observation window\", banded confidence fail-open preference lint, \"extractive selection / pointer-not-generator\", \"encoder GLiNER compaction vs Jev Score compaction (same job; pointer not summarizer)\", \"fail-closed keep_full under mutation envelope\", \"CI flaky-vs-real merge gate\", \"fail-open VOI wake/resume (Horvitz)\", \"claim/evidence Stop integrity\", \"S1 extract + escalate-S2 indexer\", \"Harbor on/off routing\", \"policy-as-judgment PR marshal\", \"shadow-mode compaction rollout\", \"Jev Ultrafast vs GLiNER Ultrafast (observe-score-act backend-agnostic)\", \"hybrid local decide + remote fill\", \"DONE ≠ verified success\", \"observed a11y/DOM candidates vs screenshot multimodal\", \"evidence-preserving stdout prune (not summarize)\", \"hard token/format envelope then soft Noul\", \"fail-safe keep original on prune failure\", \"stdout prune vs session compaction\", \"specialist S1 computer-use (Cua-S1 form-v0; not TypeSafe Jev)\", \"plan ≠ execute / dry-run default\", \"observed-element option head (fill/check/click/skip)\", \"local /v1/systemone drop-in (stub until hf scorer)\", \"dataframe-native semantic columns\", \"route≠memory\", \"advisory sidecar receipts\", \"structure induction over bags\", \"AST ∩ semantic lint\", \"extractable-from-state / retrieve first\", \"decision-model vs constrained-LLM bake-off\", \"dual-process S1 decide / S2 generate\", \"combinatorial grid ≠ extractive\", TOCTOU-of-Noul, vacuous specs, open weights vs constrained decoding vs encoder vs LoRA vs kev, whether a decision needs a model at all (meta-VOI), env-break vs policy-break, sqlite-jev / in-engine vs CLI store index, hard safety envelope (Jev proposes, code clamps), host-adapter routing (not MCP), distill-to-device memory gate, \"uncalibrated local likelihoods vs Noul / CUDA replica\", \"decision-native RAG retrieve wide then decide then evidence set\", \"classify-first MCP / read selectively\", \"living applied-mappings atlas / class patterns not a 342 hit list\", \"draft-gate silence as safer / heartbeat\", \"robotics text-state vs pixels\", \"verbatim session ledger / scored recall\", \"judgment as language primitive / English-as-config\", \"pre-registered AMBIGUOUS eval / cascade sign-flip\", \"healthcare Harbor-shaped S1+S2\", \"pre-exec tool gate allow/block/review\", \"productized public primitive / judgment wall\", \"meaning-search without embeddings\", \"attention≠correctness PR review\", \"skills→oxlint / AST prove ∩ remainder\", \"session-sticky first-prompt routing\", \"measured RAG rerank vs generative rerank\", \"Stagehand extract pick-and-copy / judge\", \"harness observe-score-act productization\", or \"formally verify with Jev\", \"capability kernel / secrets never in the agent\", \"Jev is SENSOR not policy\", \"type-safe ≠ correct\", \"typed control plane around DSPy\", \"native-probability calibration / Brier/ECE arena\", \"fan-out as measurement economics\", \"engine owns truth / Jev owns judgment\", \"human-confirmed kill gate\", \"train specialist when downstream reads p vs few-shot hosted when only argmax\", \"decide→policy→LLM leftover cascade\", \"Noul 0.5 cannot-tell never rounded\", \"calibration ≠ sortable / ORDER BY over Jev probs\", \"pairwise inversion / Score ordinality / two-decimal ties\", \"wire-compat self-hosted /v1/systemone GLiFormer\", \"class-backend economics\", \"loopback gateway hosted + local OpenJev\", \"do not distill Jev as teacher of record\", \"active-learning triage / training-data VOI\", \"index-once ask-many / citable evidence packets\", \"meaning-grep AND/OR/NOT line Nouls\", \"closed-vote-only computer-use / no planner LLM\", \"Jev vs local MLX PCD Harbor\", \"PCD O(1) speed ≠ calibrated Noul\", \"host-owned handlers × System One\", \"OMP/pi fail-open acceptance gate\", \"permission vs probability / operator owns the safety bar\", \"judgment ≠ permission / Jev never grants access\", \"eval integrity / instrument not score / dinostomp jev-as-if\", \"constrained optimizer + S1 features / never sole hot-path gate\", \"privilege ≠ verdict / effect contracts not tokens\", \"attention filter / VOI for human review / never blocks / never green unless sure\", \"measurement owns endorsement / evidence-gated question packs\", \"Jev supplies evidence / code owns authority\", \"ranking ≠ calibration / never hard-threshold raw p as frequency\", \"hot-click CU / indexed element table / S1 on click path\", \"Jev judges relevance / code decides structure / never rewrite\", \"local rules first then remainder / never auto-train on model's own hides\", \"combinators / System One as control plane / not chat turns\", \"receipts not leaderboard / type-safe ≠ correct jaggedness\", \"VOI over skill library / skillranker abstention\", \"OOD calibration / AUC ≠ ECE / sign of miscalibration by type\", \"Jev vs thinking-budget small models / frontier-100\", \"turnstile / replayable evidence≠authority\", \"MLX one-pass schema→JSON / Apple Silicon replica economics\", \"memory leases ended by new evidence\", \"never confidently wrong / TLA+ compose with judgment / escalate instead of hard-gate\", \"no seal no advance / coverage ledger / mint ≠ product brain\", \"skill-broker sibling turnstile/skillranker / judgment ≠ permission\", \"sureness / CERTAIN|CONFIDENT|LEANING|TORN|CLUELESS / max_prob is generous\", \"JevBench Harbor/jevals practice / calibration not in Main Score\", \"CI typed gate before expensive review / ci-gatekeeper\", \"Codex MCP host adapter / jev_select_capability\", \"judgment as attention redirect not merge blocker / jev-preflight\", \"compress-before-first-send / dizk jev-lens vs rashed attention filter\", \"tools≠use / SessionStart over hoping the model recalls\", \"observational memory / keep-kind verbatim / pi-om\", \"open-Jev class / openvons / JevPick menu decode\", \"physical-world System One / HA-Jev / not for locks\", \"judgment outside the store / jevql CLI\", \"landed-script trust / headless≠auto-approve\", \"digital-design combinators / extended Router Loop Retry Fallback Memory\", \"VOI cache admission / same-intent skip LLM\", \"BM25 vs Jev skill routing Harbor harness\", \"zeroshot vs BERT / contamination DiD / label-equivalence\", \"typed escalate continue abort baton / inverted loop\", \"worth-your-attention VOI / ThinkyMiner Winnow vs kevinpita winnow\", \"Jev WHETHER Python HOW LLM WHAT\", \"conflict vs ignorance / named Choice escape\", \"Playwright executes Jev chooses / sample-from-distribution\", \"OpenJev /v1/decide not TypeSafe drop-in\", \"SemIf wire-compat runoff\", \"decision-as-memory flywheel\", \"record/replay CI / jevassert\", \"failure-finding arena / jevarena ≠ jev-arena\", \"BBQ stereotype/uncertainty/cost\", \"decider≠executor / jeffrey\", \"sentence-as-rule lint / jevlint ≠ JevLint\", \"VOI hunk prune / prune-review\", \"whole-repo intent VERIFIED/VIOLATION/UNKNOWN\", \"GLiNER2 System One spec ≠ replica\", \"Rust/WebGPU grande / Clojure Laya byte parity / CPU SemIf\", \"ONNX ModernBERT local-jev measured not equivalent\", \"persist constraints across compaction / pi-heed\", \"calibration+cost as first-class gates\", \"Harbor-shaped Jev vs schema-guided LLM-as-judge / jev-judge-bench ≠ jevarena ≠ jevbench\", \"hand no-text steps to Jev / jev-use / Vercel drops confidence / margin fallback\", \"Pi System-One control plane / pi-jev-control\", \"generation as tree of Choices / never free-generates / jev-gpt\", \"OpenRouter recipe atlas / samples not benches / jev-cookbook\", \"personal history feed / no social graph / jevfeed\", \"competing NAR claims / dual-channel ECE / claim-verification / openJev-verdict ≠ OpenJev\", \"empty compaction-proxy skip / IPECTER\", \"throughput ≠ latency / like-for-like ECE\", \"1-token logprob endpoint ≠ Noul / coverage ≠ correctness / chakuho\", \"open replica engine / jevinf / argmax-parity ≠ ECE\", \"unofficial Elixir SDK ≠ OTP peer / dannote/jev\", \"jevex rename + n=16 SWE VOI / files-to-read\", \"commit pre-review attention≠verdict / middle band never rounded / commitjev\", \"Hermes plugin is Agnes not TypeSafe\", \"pi-jev-compact ≠ pi-jev-compaction / verbatim summarizer replacement\", \"empty Codex-proxy skip / IPECTER runway\", \"decision-native inbox / mailordinal / humans own ambiguity\", \"unofficial jev-cli not ready / ≠ jevql\", \"laya-multilingual / English checkpoint confident-wrong OOD / ships uncalibrated\", \"schema-conditioned DeBERTa scorer / peaked ranking ≠ calibration\", \"HF 401 access / GitHub 404 Hub-only\", \"productized System One HTTP / classifier.dev / label+confidence public contract\", \"escalate-under-threshold / smart tier 0.7 / multi-label ignores tier\", \"silent-fallback FALLBACK marker / granite 0.546 vs advertised 0.800\", \"vs_jev tracked JSON not transcription / read eval/README before quoting\", \"choxos/jev-reviewer ≠ egma-ai / systematic-review pointer-not-generator\", \"two-pass Choice+Noul / relative which-line + absolute does-this-line\", \"not-found is an answer / no paraphrase invent\", \"human check as productized judgment / checked never overwritten\", \"githubnext/localjev ≠ kunchenguid/local-jev / prompted JSON ≠ structured logit read\", \"wire-compat ≠ logit-equiv / self-reported probs / entropy confidence\", \"institutional open-replica / GitHub Next /v1/systemone\", \"Harbor-shaped bake-off AG News BoolQ SST-5 / 1200-request caveats\", \"LM Studio runner gap / structured-read primitives for OpenJev parity\", \"NandhaKishorM/laya packaging ≠ Hub-only / Router script-before-p\", \"post-T ECE ≠ raw ECE / Banking77 token-budget / 0.85 still soft / not TypeSafe drop-in / external census ≠ scored bake-off / GLiNER2+routers class-boundary / incomplete vs watch / Harbor honesty watch / JevBench v1.2 geometric-mean I/C/S/K / cal now ON rank / weight sensitivity / option-order 72→21 / instruction models class-boundary / ×2 latency assumption / est. costs / Laya absent gap / Qwen3.8 27B ≠ Archer\", \"hourly already-folded watch / apply-the-five / skip thin noise\", \"hard-gate Noul as PR gate is soundness theater / totally-tim/jev-gate ≠ jev-gateway\", \"S1 never stalls waiting / S2 one-use advisory\", \"purple telemetry = consumed not arrived\", \"Local controller ≠ githubnext/localjev\", \"seed = geometry not async replay\", \"20% starting gate still soft / schema-safe ≠ correct\", \"no pixels to either provider / confidence ≠ selected probability\", \"experimental viz not a flight controller / S2 never grants\", \"OCR+AX observe-score-act / typesafe-computer-use\", \"never send screenshot to frontier for the decision\", \"overlapping CU options = false low confidence\", \"split kind/item/site / offscreen\", \"writer/decider split + post-type Noul still soft\", \"155× one-screenshot Harbor-shaped ≠ taskset\", \"AX never sole / Spotify 0\", \"decision ≠ answer-reader capture\", \"typesafe-computer-use ≠ jev-ultrafast ≠ cua-s1 ≠ camoufox\", \"ASR observe-score-act / jev-voice-browser\", \"partial-speech VOI / complete Noul / free-text waits\", \"spoken confirm ≠ hard auth\", \"numbered overlay disambiguate without another model\", \"moritzkremb/jev-voice-browser ≠ jev-voice-control ≠ nikolas-j\", \"wrap-as-execution / AgentGhost ALLOW ASK DENY\", \"rules first then Jev remainder / ASK throws / fail-closed\", \"reddpy/AgentGhost ≠ jwen5419807/agentghost ≠ vventirozos\", \"JP genre atlas / studio_yebisu / stars ephemeral ≠ eval\", \"Jev Clearly Explained / akshay_pachaar / LLM hammer\", \"schema-safe ≠ correct / 200× 400× TypeSafe ceiling\", \"questions-as-code / shadow first / not a TypeSafe how-to\", \"proposition ≠ embedding / contrast-set refund\", \"boolean composition of soft Nouls / AND OR NOT after threshold\", \"uehaj/jev-semgrep ≠ semgrep.dev\", \"meaning-grep dedicated fold / not a gate\", \"decision-validated UI / Jev never authors text / gram-render\", \"decision-as-assert / jevtest ambiguous band\", \"typed decisions drive UI / jev2ui\", \"hybrid S1 closed verb menu / anima3 / jeff confidently flat\", \"pointer-not-generator search / JevFind\", \"jev-frontier-bench ≠ frontier-100 / ChaosNLI JS\", \"product bakeoff ≠ architecture duel / jev-gliclass-bench\", \"four engines same questions / majority floor / calibration ≠ discrimination\", \"authorship named escape / not courtroom evidence\", \"ha-switchboard HA remains execution / ≠ HA-Jev\", \"n8n classify/route/score / Low Confidence abstention\", \"fast-jev-compaction-pi ≠ pi-jev-compact ≠ pi-jev-compaction\", \"jevloop full-distribution optimizer / no LLM in the loop\", \"laya-vision SmolVLM / score untrained / ≠ blackwood ≠ Archer\", \"Cerebellum-2B /v1/decide ≠ TypeSafe / wire-compat vs agent-routing\", \"laya-grounded not drop-in / phishing regress / Platt not temperature\", \"GestaltLabs/Jeff-1 ≠ logan-markewich/jeff / acc vs ECE n=9730\", \"stanley-code empty findings ≠ approval / human promote\", \"findme ≠ JevFind / NL memory beam-search FS\", \"jevsubrouter price workers not conversation / counts ≠ dollars\", \"feelings .feels() default 0.5 is Noul-0.5-never-rounded / ≠ hunch ≠ Probably\", \"apa-agent-harness ≠ AntonioCoppe/jev-harness / unpublished npm\", \"grok-bot-jev skill cannot force a bot that ignores it / A/B proxies not tokens\", \"Essentiel-Jev never authority / human every action\", \"enzo-mcp independently falsifiable claims / ≠ jev-sift\", \"pigeonhole OTHER skip / decision-as-filing\", \"jev-reliability Nothing about accuracy\", \"clduab11/jev-test ≠ realZachi/jevtest / Nothing runs yet\", \"jev-rag-benchmark Jev wins is not an assumption\", \"dairui1/jev-lab ≠ BrendanH18/jev-lab\", \"jevmail gmail.readonly / mailjay archive/trash\", \"ZHUBoer/ego-jev reserved __none__\", \"runWorkflow completed ≠ success\", \"jsort scores are relative\", \"Noul not Choice for scale\", \"groundedness-judge-bench native vs schema-guided\", \"implicit_true included in yes\", \"jev_playground 0 promotions\", \"routing-backtest 0.0447%\", \"yuyang2230/jev-agent-skill jev-1.13-free\", \"jev-techstack-classifier stack_config.json\", \"s1_ruby collapse late\", \"undecided? abstain\", \"2389-research/judgement license null\", \"confidence ≠ winner p\", \"typesafeai-sdk-community not a new species\", \"tpellet/hunch exit 3\", \"never-execute list\", \"jev-file-search scores not calibrated accuracy\", \"jev-linkmap Jev never sees S2 prose\", \"muhammedilyasy/jev-mail metadata only\", \"tidy none-of-folders stay\", \"tab-bouncer pinned/audio/current never closed\", \"lkclean Show fail-open\", \"jev-yt-time-saver Show anyway\", \"ORIGIN pause-if-no-Jev\", \"validResponse sums-to-1\", \"jev-crawlers risk bands never raw boolean\", \"jevbrain AUTO_ACT is not a Noul\", \"judgekit YAML classify/score/route/verify\", \"typed-judge-kit verdict-in-code\", \"alsoleg89/decide packing VOI\", \"0.8 ≠ 80% accuracy\", \"Jev-Calibration Platt ECE 0.117→0.052\", \"jev-calibration-arena never acts\", \"ctmx/openrouter-jev-mcp Decision-as-Plugin\", \"FrancoisChastel/jev-code ≠ npm jev-code\", \"claudecode-jev-marketplace fail-open not hot path\", \"pedroknigge/mcp_jev packs not ask_jev\", \"cyrusasco/typesafe-mcp noul deadband 0.35–0.65\", \"codaaiteam/jev-skill jevtypesafeai.com ≠ TypeSafe\", \"hermes-switchyard ≠ hermes-jev-router ≠ hermes-plugin-jev\", \"nanoprune 2.8MB ECE 2.58%\", \"smartdio/jev-browser-agent ≠ ZHUBoer/ego-jev\", \"Dakai/omp-jev-web DONE ≠ proof\", \"hari007sh/jev ≠ dannote/jev\", \"0thernet/system-one-skills deterministic verify\", \"typed-gate band [0.40,0.60] is refusal\", \"pi-jev-gate fail-closed; choice is the verdict\", \"Foq ~25ms/2.2GB local\", \"rev prefill-only + HF jev-0.5b\", \"robfrase/jev planning memo\", \"typesafe_agent_gates 27/27 / 31/31\", \"EpicEric/safe-sh static remainder\", \"pastepilot Confirm before act\", \"Jev-Reranker live Jev not yet measured\", \"sessionwise opt-in relevance\", \"jev-search pointer sieve\", \"400ms Salesforce WebMCP\", \"typesafe-scheduler-diagnostics advisory\", \"droidjev screenshot-free\", \"Tewoto1 jevcu planner still writes\", \"ha-conversation-jev Jev→Grok\", \"dsh-jev can only gate\", \"jev-classification-benchmark specified not run\", \"jev-luna-pagerduty p≥0.50\", \"meldltd/meldecision laya-go ONNX\", \"laya-doom never pixels\", \"logixism/laya-api empty README\", \"akpsahan/laya ≠ Archer\", \"choxos/jevchess engine owns truth\", \"jev-drive sim not AV\", \"story-arc Jev never authors\", \"jev-hs-assistant HS6\", \"golergka/jev-plays-starcraft-2 UI-verified ≠ API Victory\", \"awesome-jev-use-cases catalog\", \"Nibir1/typesafe-go ≠ official\", \"fingerprint after redact\", \"recall vs decide\", \"publish fingerprints+answers\", \"CI replay as Harbor cousin\", \"Cache hit ≠ correctness\", \"hyperspaceai/jevcache ≠ kushals256/jevcache\", \"human labels only\", \"score never auto-accepts\", \"production capture flywheel\", \"sutro-sh/jev-align ≠ caiovicentino/jev-align\", \"guidance ≠ hook\", \"catalysts ≠ summaries\", \"compile-time System One\", \"unofficial ≠ TypeSafe\", \"format_version modernbert-jev/1\", \"Argos1111/jev_local ≠ us/jev-local ≠ kunchenguid/local-jev\", \"LFM default ≠ ModernBERT backend\", \"Nemotron ≠ TypeSafe Jev\", \"not a calibrated replacement\", \"djev-dev complements djev-spark\", \"images as Choice options\", \"Laya essay numbers *theirs*\", \"Router/OOD confidence\", \"hosted bootstrap ≠ silent TypeSafe\", \"difficulty + policy thresholds + JSONL trace\", \"jev-codex-pilot model + reasoning depth\", \"keep/shadow/hybrid/reject\", \"quarry evidence projection\", \"Frank-ZY-Dou/awesome-jev robotics/3D/control\", \"one-dollar-tahoe TypeSafe Jev defense eval\", \"jevguard calibrator/cache/escape\", \"jev-ci-selector CI shadow mode\", \"llama-jev llama.cpp replica\", \"petercr/jev-orchestrator ≠ FleeexCorp/jev-orchestrator\", \"seb4ez/jevguard ≠ AseemPrasad/JevGuard ≠ pablozr/JevGuard\", \"webNeat/llama-jev ≠ WiktorB2004/llama-index-jev\", \"OpenCode jev-pruner context sieve\", \"observe→score-candidates→prune\", \"jev-zen / jev-1.13-free\", \"zen-chat ≠ Noul\", \"fail-open original\", \"keepScore >0.1 floor\", \"host port of tamaratran/jev-pruner\", \"indiejoseph/opencode-jev-pruner ≠ nrdz-labs/fast-jev-opencode\", \"jev-webagent-bench empty stub\", \"Kiln-AI/jev_jsonschema noul_threshold 0.5\", \"NSStudent/JevSwiftSDK unofficial\", \"GLiNER2 native Apple path\", \"unofficial Swift/Core ML GLiNER 2.5-small\", \"entity spans + confidence\", \"not Choice/Score/Noul\", \"not TypeSafe\", \"label descriptions as schema\", \"on-device ANE economics\", \"honesty locks\", \"shershah1024/gliner-native-runtime ≠ Fastino\", \"≠ gliner25-compaction ≠ gliner2-ultrafast ≠ Eran-BA/Jev_from_GLiNER2 ≠ NSStudent/JevSwiftSDK ≠ jevmlx\", \"default threshold 0.1 still soft\", \"Decision Graph Protocol frame→assess→commit\", \"app retains permissions/effects\", \"Jev-first assessor-neutral\", \"guarded commit / receipt/next frame\", \"assessment batching\", \"hard-gating DGP as safety theater\", \"numerous-com/dgp ≠ TypeSafe official\", \"jegrep calibrated path+range Nouls\", \"no embeddings/index/daemon\", \"~$0.01–0.03 typical\", \"agent --json\", \"can1357/jegrep ≠ Bentlybro/jevgrep ≠ uehaj/jev-semgrep\", \"Archer-arch fidelity\", \"kev family OOD 0.76–0.77 vs Jev 0.86\", \"block-causal isolation\", \"pointer/readout CE-trained\", \"/v1/systemone drop-in\", \"replica honesty\", \"cost-sensitive decision theory × System One probabilities → control flow\", \"thresholds derived from costs not hard-coded\", \"YES / NO / UNSURE from cost_false_yes / cost_false_no / cost_human\", \"auto-batching same-object questions\", \"Kungie/gut ≠ tpellet/hunch ≠ carldaws/hunch\", \"judgment vs generation\", \"deterministic execution after probabilistic judgment\", \"exactly one app-owned callback\", \"explicit uncertain branch\", \"Illusion47586/judge ≠ lexingtonhibiki/judgekit ≠ Ascurse/typed-judge-kit\", \"variable-N option scoring as the trainable object\", \"dynamic candidate bags not fixed label sets\", \"zwliJay/jev-forge ≠ NanoJev\", \"open replica economics / latency vs closed Jev\", \"NAR local drop-in\", \"wfzyx/von late-catch HIGH\", \"competing NAR claims / replica honesty\", \"typed judgments vs chat judges on guardrailing\", \"ishaannk/llm-vs-jev cross-note only\", \"deeper integrity fold is rh-guard\", \"nothing wins outright\", \"can be argued out of guarding\", \"Jev IS the if-statement\", \"judgments/probabilities drive branches\", \"text model only writes prose\", \"interpreter owns variables/loops/budgets/replay\", \"otherwise maybe / confidence gate\", \"chaos samples after the gate\", \"southpolesteve/probably ≠ carldaws/hunch ≠ feelings ≠ Kungie/gut ≠ Illusion47586/judge ≠ tidymodels/probably\", \"133★ / forks 10 live\", \"build calibrated classifiers from human feedback\", \"retrieve by relevance not resemblance\", \"one calibrated yes/no per memory in one request\", \"pointer mode 17/18 19/20 *theirs*\", \"embedding resemblance misses the allergy\", \"samdotmak/jev-recall ≠ jev-search ≠ jev-sift ≠ carryforward ≠ chopratejas/invalidate\", \"memory leases ended by new evidence\", \"six Nouls then fixed rules in code\", \"0 of 157 false invalidations\", \"questions/plans/directives are not evidence\", \"unsure → review queue\", \"host keeps the store\", \"name↔body / comment truth / test-claims\", \"mizchi/jev-lint is mizchi/jevlint rename\", \"no shipped rule has severity error\", \"~1 in 5 findings wrong *theirs*\", \"mizchi/jev-lint ≠ huntedman/JevLint ≠ MichitoSugawara/jev-lint\", \"JSON Schema → typed JSON via Jev\", \"noul_threshold 0.5 decoder not a proof\", \"IncompatibleSchemaError lists every bad property\", \"on-device Laya CoreML ANE\", \"~5 ms P50 short decisions\", \"189/189 FP16 checkpoint parity\", \"10× not achieved\", \"mizorewww/laya-coreml ≠ gliner-native-runtime ≠ jevmlx ≠ NandhaKishorM/laya\", \"softmax over allowed tokens ≠ Noul\", \"question-first cache\", \"Micha0827/snapjudge ≠ githubnext/localjev ≠ jevmlx ≠ cendress/SnapJudge\", \"Jev-first Pi agent loop\", \"slow-LLM fallback\", \"explicit action menu / CandidateSource unimplemented\", \"62 tests wiring not quality\", \"direwolfiy/JevPi ≠ standardagents/jevpilot ≠ pi-jev-control\", \"resume-screening bias audit methodology\", \"name×resume factorial independent Nouls\", \"callback determined by resume quality\", \"mean-probability name gaps operationally negligible\", \"natemoo-re/bias-bench ≠ BBQ\", \"Plan/PRD panel → code-owned pass|review|block\", \"cheerleading out of scope\", \"austindixson/planalyzer ≠ single-goodness Noul\", \"cost-aware multi-model routing/escalation\", \"decide vs do\", \"successful-task cost\", \"cannacre8ive/switchboard-ai ≠ ha-switchboard ≠ hermes-switchyard\", \"frozen-protocol zero-shot bench\", \"TypeSafe Jev vs PrismNLI vs Laya\", \"contamination caveat\", \"elcronos/jev-vs-open-decision-models ≠ JevBench ≠ DMB\", \"context-window admission control\", \"VOI gate which tokens are worth the expensive model\", \"fail polarity per lens\", \"on small inputs lenses lose money\", \"cvsgireesh/jevusher ≠ jev-sift ≠ winnow\", \"typed decision control plane\", \"receipt ≠ authorization\", \"historical-v0 zero retained cases\", \"MokiMeow/jev-fabric ≠ jev-forge ≠ dgp\", \"live 15-dim typed rubric re-score per pause\", \"scoring economics exemplar\", \"OpenJev/Codiv ≠ TypeSafe hosted\", \"jose-troche/live-rubric ~$0.000004 desc / ~$0.000006 README\", \"adversarial pre-registered Jev eval\", \"28 predictions before data\", \"123,805 requests\", \"confidence does not track ignorance\", \"polite injection 65% / crude 0%\", \"willkelly/jev-evaluation ≠ jevals ≠ jev-baselines-eval\", \"provider-neutral Elixir/BEAM Noul/Choice/Score SDK\", \"class infrastructure\", \"nshkrdotcom/system_one_sdk ≠ typesafe_sdk ≠ dannote/jev\", \"question-linting of Jev questions themselves\", \"nine jaggedness rules, no API key, no labelled data\", \"static lint ≠ measured separation\", \"yodablocks/jevq ≠ tenbin ≠ JevLint ≠ commitjev\", \"open-weights Laya as class exemplar (binding)\", \"Nx/Bumblebee runtime\", \"host chooses backend\", \"ChristianAlexander/laya_ex ≠ system_one_sdk ≠ dannote/jev ≠ NandhaKishorM/laya\", \"on-chain/edge Laya deploy\", \"parity_verified stays false\", \"model output never grants Tx\", \"humandebri/IC-Laya ≠ laya_ex\", \"auditable weekend replica\", \"Jev outputs never used for training\", \"soft human-vote distributions\", \"unpaired 0.577 vs 0.727\", \"agilabs-ai/jev48 ≠ JevBench ≠ Mapika/decider\", \"adversarial dual-judge / framing attack surface\", \"comparative framing is the usable judgment\", \"prior injection crowds out evidence\", \"copyleftdev/ember ≠ ember.js\", \"Laya specialist fine-tune pipeline\", \"training still GPU-pending\", \"PIXELZX0/XERON ≠ convaiinnovations/laya\", \"Hub Laya replica drop\", \"daliborsb/laya ≠ convaiinnovations/laya ≠ NandhaKishorM/laya\", \"System One student distillation corpus\", \"gold is programmatic\", \"teacher is closed-API clone\", \"do not distill Jev as teacher of record\", \"MagaBitmex/jev-4b-distill-data ≠ missing student checkpoint\", \"non-LLM VIN System One\", \"planning depth not chat\", \"lewislululu/jevon ≠ douglance/jevon\", \"source-bound evidence checks\", \"local quote mismatch needs no API\", \"exit 0 ≠ claim truth\", \"WaynezProg/jev-kit ≠ jonathanavis96/jev-kit (Airlock) ≠ jev-use ≠ jev-mcp\", \"independent System One evidence catalog\", \"scores not one leaderboard\", \"no external record currently reproduced\", \"TokenTrim no-Jev matched hybrid 62.4%\", \"reachjalil/system-one-bench ≠ mallahyari/system-one-benchmark\", \"21 tasks · 134 items · 208 questions\", \"scenes from public GitHub contracts, not production logs\", \"SivletLabs/jev-eval ≠ willkelly/jev-evaluation ≠ 4esv/jev-eval ≠ xxkuboxx/jev-eval ≠ onlyoneaman/jev-eval ≠ dayhaysoos/jevals\", \"option isolation (sibling-blind)\", \"permutation-equivariant\", \"Hub OWNER not published\", \"nafisazizir/hev ≠ jaredpalmer/kev\", \"frozen local LLM logits, no trained decision head\", \"residual-head 9,222-param decreased 73/96→67/96\", \"confidence = 1−normalized entropy, not P(correct)\", \"yuki-oshio/mini-jev ≠ r-ms/mini-jev\", \"Jev classifier as autoregressive next-token predictor\", \"ChatJev-style soundness theater\", \"erik-dunteman/ChatJev ≠ dannote/jev ≠ jev-gpt\", \"calibrated decision head × AlphaProof value head\", \"implementation-layer isomorphism, semantic difference\", \"timeout = censoring\", \"do not launder Noul as proof\", \"parallel rank-prediction vs serial selection\", \"independent questions can conflict\", \"zzzzzec/jevsort ≠ keltokhy/jsort\", \"curated open System One ecosystem catalog\", \"rupeshpoojary9/awesome-open-system-one ≠ AnotiaWang/awesome-jev\", \"arXiv paper radar with Jev relevance scoring\", \"ranking ≠ calibration / 0.5 still soft\", \"fail-open failed evals not marked seen\", \"train calibrated ~27M from scratch\", \"typed Q→prob dist / one forward pass / no LLM decode\", \"hyusi2003/MiniSystemOne ≠ Colvin0315/MiniSystemOne\", \"description-only stub / size 5\", \"ESCI hard probe fails four of six\", \"jev_bool ECE 0.242 inversion 0.255\", \"do not re-fold §60 six-gates as new\", \"jobbyjev one-request-per-company from batch-size result\", \"find/design/evaluate TypeSafe Jev decision loops\", \"karanb192/jev-architect ≠ samtay32/jev-system-architect\", \"Jairik/jev-distiller size 1\", \"distill-Jev UI stub / do not distill Jev as teacher of record\", \"post-launch scored use-case map / Jev self-scores then human curation\", \"licensedsaucer9-web/jev-opportunities\", \"Jev-inize a use case into classifier/router\", \"gavinHuang/jevinize → simple-jev not TypeSafe\", \"featherless-ai/simple-jev\", \"compare saved decisions / same label can still change the branch\", \"VihaanAgarwal/jev-diff ≠ Saik0s/diffusiongemma-jev-macos\", \"not tested with a live Jev API key\", \"constrained logprob + temp/Platt ≠ Noul\", \"OpenJevPro pastes openjev-sglang JevBench as own\", \"zhangcy122/OpenJevPro ≠ IamBusy/OpenJev ≠ ekzhang/openjev-sglang\", \"PolyForm Noncommercial\", \"SmolLM-135M / sub-70ms / 0 output tokens\", \"demo P(True) 0.5052 / Choice conf 0.2872 / Score conf 0.0055\", \"README claims MIT / GitHub license null / no LICENSE file\", \"patelvishwa112/jev-system-one-rlcd ≠ arnabgho/rlcd-lite ≠ blackwood-rlcd\", \"source-backed Awesome Jev radar / 306+ commit-pinned\", \"logicrw/awesome-jev-projects ≠ AnotiaWang/awesome-jev ≠ yibie/awesome-jev ≠ cobanov/awesome-jev ≠ rupeshpoojary9/awesome-open-system-one\", \"auto GitHub sync / Issue-only submissions\", \"hashed n-gram encoder / rival-aware attention\", \"olanotolu/jevbetter vs jevlike starter\", \"synthetic hard menus top-1 0.916 vs 0.873 / ECE 0.0182 vs 0.0367 / 40 vs 4608 menus/sec\", \"shuffled-context control 0.335\", \"structured probability readouts\", \"distribution > argmax\", \"Noul 0.5 midpoint\", \"score is expectation not integer\", \"bare HTTP not SDK\", \"Arohtea/jev-readout\", \"Jev-style Choice/Score/Noul from ordinary models\", \"optional DSH plugin\", \"schema-valid ≠ calibrated\", \"gulagala001/jevify ≠ Mintzs/jevify\", \"Laya RLCD benchmark\", \"40.3% below constant-answer\", \"open-weight measurement\", \"mourad-ghafiri/laya-rlcd-benchmark ≠ yibie/laya-jev-lab\", \"cheap fail-open semantic edge\", \"second signal not sole\", \"FastLoopError catch\", \"SupremeDreamZ/jev-fastloop ≠ jev-ultrafast\", \"asking more questions in one call\", \"0.980 at every N\", \"nearly not fully deterministic\", \"TheWebDevel/jev-fanout\", \"Qwen3-VL perception + Jev decisions train RL\", \"0 model calls at deployment\", \"VLM alone 1.7 vs +Jev 4.4\", \"harneet2512/reflexrl ≠ khordoo/jev-reflex-autonomy-lab\", \"independent Jev API vs Laya\", \"cascade 0.60 matches 78% at 1.8×\", \"noul facts not judgements\", \"yibie/laya-jev-lab ≠ dairui1/jev-lab ≠ BrendanH18/jev-lab\", \"GLiNER vs GLiFormer vs Laya vs Jev\", \"extractors ≠ decision engines\", \"Laya dict-instructions collapse 58.3%\", \"umstek/zero-shot-ie-bench\", \"decisions-per-minute & cost\", \"204 moves vs 73\", \"throughput not intelligence\", \"angelgalvisc/snake-arena-jev-vs-llms ≠ vtrivedy/jev-plays-games\", \"behavioral contracts\", \"pin expectations eval upgrades\", \"raw 0.94 is not a release\", \"sathariels/jevcheck ≠ dayhaysoos/jevals ≠ SivletLabs/jev-eval\", \"evidence-linked dependency upgrade\", \"Jev never generates filenames\", \"no_direct_evidence ≠ safe to merge\", \"GaneshVG18/upgrade-radar ≠ LYchoon/paper-radar-jev\", \"discography theme/mood/complexity\", \"five atomic questions one call\", \"lirantal/discoprint\", \"Turn any open LLM into System-One Jev\", \"uspraveen/Jevify ≠ Mintzs/jevify ≠ gulagala001/jevify\", \"Jevify-any-LLM architecture probe\", \"description-only stub / size 0\", \"Train encoder-only calibrated decision models from a task sentence\", \"Exu is a toolkit, not a method\", \"strictly proper scoring rule\", \"Pre-alpha\", \"Ruivalim/exu-base\", \"scratch-trained calibrated decision model\", \"typed Q → probability dists\", \"Colvin0315/MiniSystemOne ≠ hyusi2003/MiniSystemOne\", \"no published weights download URL\", \"90.5 seconds / 29.2% pipeline evidence\", \"p_i/p_j independent of other candidates\", \"Recipe for calibrated decision models — small model out\", \"init → synth → train → eval → serve\", \"91.1 % / ECE 0.022 *theirs*\", \"Jev zero-shot 75.1\", \"scienthoon/luce\", \"Put Jev's three headline claims on trial\", \"0.5B local GPU\", \"46x speedup / accuracy identical\", \"ECE 0.624 sentiment catastrophe\", \"bigger model worse calibration\", \"RichardoMrMu/jev-mini ≠ yuki-oshio/mini-jev ≠ r-ms/mini-jev\", \"System-1 decision engine for local LLMs\", \"structured choices only\", \"JSON parse of generated text ≠ Noul\", \"TypefAI JEV / Journal Entry Voucher\", \"tapsin/jev-local ≠ us/jev-local ≠ Argos1111/jev_local\", \"Jev 1.13 reward-model eval across 8 benchmark tracks\", \"40,940 examples / 0 API errors\", \"RewardBench v1 92.58%\", \"Precise IF 50.63%\", \"goya4140/jev-reward-model-evaluation\", \"Scaffolding in progress\", \"Jev vs LLM support-ticket routing\", \"static + live decision bench\", \"TypeSafe's own published benchmark\", \"illustrative simulations, not live API calls\", \"JevBench v1 — smart/cheap/fast/reliable\", \"I/C/S/K 25% geometric mean\", \"classifier.dev fast tier 84.8 is Jev behind its own API\", \"do not re-fold §78 v1.2 board as new\", \"Laya (421M) 70.1 now on board\", \"Zero-shot/few-shot LLM routing\", \"hard budget filter before Jev\", \"Jev never asked to perform budget arithmetic\", \"Jev judges the next state, XState enforces transitions\", \"simulation uses synthetic keyword fixtures\", \"catalog gravity\", \"v-modal/awesome-jev-tools\", \"★339 live REST\", \"curation is not endorsement\", \"crawler-maintained directory\", \"Daily GitHub + npm sweep, human-merged\", \"RadRebelSam/awesome-jev ≠ AnotiaWang ≠ yibie ≠ cobanov ≠ logicrw ≠ v-modal\", \"HF peft SPLADE/BGE reranker\", \"rdxtremity/jev-reranking ≠ carlaiau/jev-reranking\", \"query-side encoders, not a Jev replica\", \"ONNX System One Qwen3.5-4B scorer\", \"source:pngwn/system-one-qwen3.5-4b-scorer\", \"CC-BY-NC-4.0\", \"temperature 1.75\", \"transformers.js AutoModel cannot load this graph\", \"Consistency benchmark Space\", \"This Space contains no benchmark result yet\", \"12-case plumbing fixture\", \"Benchmark-driven Jev router and judge\", \"cheap alone is not success\", \"Jev does not write, sum prices, or claim accuracy %\", \"Sol 94.2 / Luna 83.9 / Jev path 89.7\", \"19.2% Sol / 62.3% cost save / 4.5pp miss of 2pp non-inferiority\", \"p50 latency worse than Sol due to routing overhead\", \"erendikmenn/jev-llm-router-benchmark ≠ jev-rag-benchmark ≠ ryantsai/jev-llm-router\", \"Express + node:sqlite\", \"mock and Jev decision engines\", \"previous_ticket_count >= 3 is code\", \"MIN_CONFIDENCE 0.6 still soft\", \"substring false positives\", \"aesaganda/jev-ticket-router ≠ SarathChandraBellam/jev-vs-llm-ticket-router\", \"Universal Figure & Diagram Router\", \"confidence ≥ 0.85 hard-gate is theater\", \"generative AI banned from scientific plots\", \"six visual branches\", \"hoangngochuong24947-gif/jev-figure-router\", \"human-labeled (state, question, label)\", \"166,054 rows / 22 configs\", \"soft_label for human uncertainty\", \"Praveenrajus/jev-bench ≠ fstandhartinger/jevbench\", \"ternary bonsai System One GGUF\", \"openjev's mechanism, Bonsai's weights\", \"Hub does not ship weights\", \"100/100 easy T/F is not Harbor\", \"label_mass ≠ correctness\", \"stock llama.cpp Q2_0 silently gibberish\", \"NicolaiMTLassen/open-bonzi-jev ≠ NicolaiLassen\", \"transformers.js DeBERTa ONNX\", \"source:com-kotobalabs/open-jev-deberta-v3-large\", \"temperature 1.05\", \"AutoModel from_pretrained works\", \"onnx-community/open-jev-deberta-v3-large-ONNX ≠ system-one-qwen3.5-4b-scorer-ONNX\", \"107★ densify\", \"GH 151M vs README 149.6M\", \"PR #1 now closed unmerged\", \"do not re-fold §71 claim-audit as a beat\", \"typed decisions, RLCD, confidence-gated routing\", \"structured ≠ correct\", \"mock not live API\", \"26 tests\", \"wjdjdakf17/jev-study ≠ baekenough/jev-study\", \"bonzi-27b-v2 / ternary-8b / 27b-v1 GGUF family densify\", \"Hub still does not ship weights\", \"WANLI-256 74.6% / 65.2% / 71.1% *theirs*\", \"Bonsai 1 27B Q1_0 runs on stock llama.cpp\", \"ternary still needs PrismML fork\", \"hf:heman10x/openJev-verdict-2.0 twin tokenizer-only\", \"OpenJev Vision image classification + uncertainty\", \"CLEVR-4 held-out joint 0%\", \"hfdataset:IamBusy/OpenJev-Vision-Research-v0.1 12,832\", \"294,912 derived targets not independent samples\", \"Laya multilingual ONNX WebGPU typed-decisions port\", \"63/63 selected answers / 5.1e-4 CPU / 1.2e-2 WebGPU\", \"UpHash-Network/mini-jev is yuki-oshio transfer\", \"jev-injection-bench 11,900 labelled prompts\", \"Jev best ranking / Haiku better ECE 0.021 vs 0.058\", \"0.5–0.9 band is where Jev's numbers do not mean what they say\", \"Prompt wording moves panic 28%\", \"manojlds/jev-dspy-bench ≠ dspachos/jev-dspy ≠ jmanhype/jev-dspy-lab\", \"Jev agreement is similarity, never ground truth\", \"no aggregate quality grade or merge gate\", \"AbstentionBench-on-Jev rank 1 of 20 vs 2025 field\", \"question-asymmetry\", \"forward-looking 0.465 never extreme\", \"openkev calibration layer not a runtime\", \"ECE vs coverage independent\", \"select_threshold returns inf\", \"escalation catches uncertainty not ignorance\", \"misakaikato/openkev ≠ jaredpalmer/kev\", \"pdf-race Docling→Jev vs Gemini\", \"parser owns the wall clock\", \"12/12 tie is a tie\", \"titles selected not generated\", \"flopcheck 16 calibrated tweet judgments\", \"mechanical tells in code\", \"ZeroX-01/jev-atlas ≠ Zaious/jev-capability-atlas ≠ gorock007/jev-atlas\", \"catalog not endorsement\", \"Laya calibration lab Gradio MCP\", \"T never changes argmax\", \"confidence ≠ top-label p\", \"easy probe set refused\", \"40–48 rows too small to ship T\", \"Gemma-4 26B-A4B jevify classification+calibration\", \"LoRA adapter twin not independent eval\", \"Gemma-4 E4B jevify\", \"E4B LoRA stub card\", \"kushalpatil/jevify-gemma4 ≠ Mintzs/jevify ≠ gulagala001/jevify ≠ uspraveen/Jevify\", \"GH kushalpatil07/jevify 404\", \"PAWS 0.580/ece 0.288 is the weak cell\", \"smaller E4B slightly better OOD ECE than 26B-A4B\", \"Hub jevify merged LoRA ships weights\", \"bonzi Bonsai-8B v1 GGUF densify\", \"Bonsai-1.7B v1\", \"Bonsai-4B v1\", \"WANLI-256 64.5% / 60.2% / 52.0% *theirs*\", \"rank #4 / #5 / #6 of 6\", \"JulesHuisman/jev-eval scaffolding / README SHA c356a584 (was empty e69de29b)\", \"JulesHuisman/jev-eval ≠ SivletLabs/jev-eval ≠ willkelly/jev-evaluation ≠ 4esv ≠ xxkuboxx ≠ onlyoneaman ≠ dayhaysoos/jevals\", \"7 bands 6/10 vs 40 bands 0/10\", \"source receipts + confidence slider re-policy without re-inference\", \"32/32 synthetic is smoke not production\", \"classify HF datasets across typed semantic dimensions\", \"roadus2 watch misspelling\", \"lock roadius2/ultra_laya\", \"ultra_laya REVIEW defects\", \"default branch claude/laya-jev-review-gg5ppo\", \"XNLI EN 88.3% ECE 0.032 → RU 77.3% ECE 0.096\", \"Δ −11.0 pp [−14.2,−7.8]\", \"ECE +0.063\", \"MASSIVE no detectable difference at n=600\", \"confidence is function of p_max (r=1.000)\", \"pointer-not-generator 400 human-authored responses\", \"proposed ≠ authorized\", \"FewRel 160: Jev 85.0% vs lexical 13.125%\", \"gated 100% (95/95) coverage 59.375%\", \"J++ composable semantic computation language\", \"judge-jev 0.5 still soft\", \"947 repos scored\", \"A 273 / B 302 / C 372\", \"LLM rubric ≠ benches\", \"No benchmark winner is claimed\", \"phishing: naive 62.6% vs regex 91.8%\", \"5-atomic + LR 95.0% *theirs*\", \"AITuber tension ±15\", \"README npm global\", \"repo is Rust\", \"git-confess code owns counting/blame/ratio\", \"httpx exhibit 11% (13/119) *theirs*\", \"90d trend +12.40% vs random +12.75% vs BH +41.71%\", \"5m win rate 25%\", \"Awesomejev 656 entries / 38,160 stars\", \"tracker likes 64 (+4) lastModified UNCHANGED\", \"Laya present\", \"Blackwood ABSENT\", \"Archer still promised_not_landed\", \"Blackwood tracker ABSENT; likes 2 gated manual\", \"r = c - p_a\", \"ECE 0.021; acc 0.807 vs warmup 0.746\", \"calibration beyond ~500 tokens unmeasured\", \"Independent primitive\", \"11.57s vs 54.10s · 4.67× · 120/128 *theirs*\", \"default path is pretrained Gemma probs not trained RLCD head\", \"GH Meanblock 404; lock leesk212/JEV-CPU\", \"softmax over letter slots ≠ Noul\", \"WANLI 0.741 vs openjev v2 0.77 *theirs*\", \"3-way NLI ≠ Noul\", \"priority 0.464 = majority floor\", \"banking77 contaminated\", \"raw margins not probabilities\", \"GH jev-haiku-benchmarking 404\", \"do not distill Jev as teacher of record (they distilled Haiku)\", \"“0.9 is not one number”\", \"ranking ≠ calibration\", \"banking77 0.8–0.9 stated 0.86 actual 0.73 over-confident *theirs*\", \"≠ Praveenrajus/jev-bench ≠ fstandhartinger/jevbench\", \"$0.0000153–$0.0000226 vs circulating $0.0004 (~20×)\", \"Score is 0..n-1 expectation not 0–1\", \"Noul has no confidence field\", \"TCP floor 198.8 ms\", \"type reliability is not a reason to choose Jev (json_schema 5/5)\", \"gateway tax not one number\", \"Function-only 5/8 vs hybrid 8/8\", \"4/8 without Jev\", \"8 designed cases not conversion lift\", \"≠ RadRebelSam/awesome-jev\", \"200-row pilot Jev 86.5% 173/200 vs Gemini Flash-Lite 86.0% 172/200 vs Pro 87.0% 174/200 *theirs*\", \"not a ranking\", \"NLI Tetris argmax P(entail)−P(contradict)\", \"情緒測謊器\", \"1q 396ms / 30q 567ms\", \"±0.03\", \"33q $0.000045 vs Gemini ~5× slower ~60× cost *theirs*\", \"≠ realZachi/jevtest\", \"8-example Jev vs GPT-5.6 Sol ~64× cost 5.4× latency *theirs*\", \"synthetic; no inference\", \"≠ JevBench v1.2 §78\", \"Judged 3317 / listed 2560\", \"Jev judges, code applies policy\", \"catalog ≠ endorsement\", \"APA “microsecond policy / zero hallucination” overclaim\", \"Client-side quiz; pointer from held docs; scanned-PDF warn\", \"CSP only api.typesafe.ai\", \"Jev judges / agent reasons / user decides\", \"selecting an option is not permission to implement\", \"degraded fallback\", \"pattern exact, judgement must clear floor\", \"no matching pattern → no model call\", \"not a correctness oracle\", \"$0.00022 vs chat $0.00306 *theirs*\", \"Spec vs artifact remainder\", \"treating 0.85 as 85% / minProbability hard-gate as Harbor\", \"VERIFY acquires discriminating evidence, never same-pool confidence-only rescoring\", \"fast/full/max are ceilings not sizes\", \"Solar writes, Jev chooses NEXT ACTION\", \"SemIf 2186★ (+20 vs §109 2166)\", \"jevlike 1038★ (+7 vs 1031)\", \"TypeAR 14★ flat\", \"AnotiaWang 96★ (+1 vs 95)\", \"yibie/awesome-jev 490★\", \"Laya likes 802 (was 783)\", \"tracker likes 64 flat, lastModified UNCHANGED\", \"do not reopen or amend PR #23/#24/#25/#26/#27.\", \"do not reopen or amend PR #23 or #24 or #25 or #26 or #27\", \"Calibration is not alpha\", \"NO CURRENT ALPHA CANDIDATE\", \"ΔR² approximately +0.00084\", \"Brier 0.2131387\", \"ECE 0.0421875\", \"Adding Jev probability to deterministic volatility improved Brier by only 1.4058e-05\", \"default 0.5 keeps zero non pinned\", \"keepResult median 0.14 to 0.17\", \"keepCall median 0.28 to 0.35\", \"usable range is about 0.10 to 0.25\", \"7.8% to 57.9%\", \"judges results it never sees\", \"task-finish eval not built yet\", \"$0.002 per compaction\", \"slavadubrov/sgr-judge-bench ≠ slavadubrov/jev-judge-bench\", \"Jev 108/120 $0.083 0.34 s\", \"Luna SGR 114/120\", \"paired Jev accuracy-difference intervals include zero\", \"not evidence of equivalence\", \"GLM SGR 26/120 93 format failures\", \"Terra-planned Jev hybrid 55/120\", \"rule-based by default, optionally Jev-backed\", \"empty README\", \"missing key cannot break the experience\", \"prefill plus exactly one decode\", \"softmax over A/B/C ≠ Noul\", \"BBQ 9,053/10,000 (90.53%)\", \"ECE 0.0890\", \"Mean confidence 0.9943\", \"overconfident\", \"score and noul not implemented\", \"DGUI 12 rows (was 6)\", \"INSTRUCT 119 rows likes 2\", \"encode the state once, decide everything in parallel\", \"0.740 accuracy against a 0.508 majority\", \"ECE 0.047\", \"fine-tune's advantage ends where its 384-token training data does\", \"jasonkneen/open-jev ≠ pngwn/open-jev\", \"same sha d41dc3cd\", \"Space does not call Jev\", \"recomputes routing from saved probabilities\", \"200-case Jev 97.0% / 100.0% / 95.0% / MAE 9.22\", \"synthetic repository benchmark\", \"Jev evaluations are advisory\", \"YehuiTang0316/jev-nlgrep ≠ Bentlybro/jevgrep ≠ can1357/jegrep ≠ uehaj/jev-semgrep\", \"default threshold 0.8 still soft\", \"40-line windows cannot prove whole function\", \"token-native sequential start/end Choice\", \"Gemini/Haiku stubs not configured yet\", \"handful of hand-written examples, not a benchmark\", \"Jev judged exactly what it was given\", \"laguagu/jev-skills ≠ laguagu/jev-evidence-lab ≠ Pleo2/awesome-jev-agent-skills\", \"contract_passed is not a claim of guaranteed factual truth\", \"Wilson lower bound 0.85 floor\", \"fixture mode no savings claim\", \"SemIf 2207★ (+21 vs §110 2186)\", \"jevlike 1043★ (+5 vs 1038)\", \"TypeAR 15★ (+1 vs 14)\", \"AnotiaWang 97★ (+1 vs 96)\", \"yibie/awesome-jev 506★ (+16 vs 490)\", \"Laya likes 822 (was 802)\", \"tracker likes 64 flat, lastModified UNCHANGED\", \"do not reopen or amend PR #23/#24/#25/#26/#27/#28\", \"people who compare Jev against GPT-5.6 has never fine-tuned BERTForXYZ for living and it shows\", \"zero shot classifiers\", \"scale them as much as decoder only models\", \"many problems solved with LLMs could have been solved with them, it was a skill issue\", \"opt for DeBERTa and ModernBERT ones\", \"BERTForXYZ → DeBERTa → ModernBERT\", \"Jev vs GPT-5.6 bakeoffs are a category error\", \"encoder / ZS classifiers\", \"institutional HF voice\", \"quote *theirs*\", \"do not invent accuracy numbers\", \"softmax/ZS scores still ≠ calibrated Noul\", \"soft scores ≠ hard gates\", \"@mervenoyann\", \"likes 421 / 189\", \"impressions 35498 / 9613\", \"multimodal image<>text ZS as perception front-end\", \"hf:MoritzLaurer/deberta-v3-large-zeroshot-v2.0 likes 139\", \"hf:MoritzLaurer/ModernBERT-large-zeroshot-v2.0 likes 72\", \"Bart, bert, deberta, modernbert, these are all LLMs\", \"Maziyar quoted\", \"Jev is exemplar not the mandate\", \"do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29\", \"Heman10x-NGU/Verdict-open-jev ≠ Heman10x-NGU/openJev-verdict-2.0\", \"TF-IDF + LogReg ECE 0.0207 vs Jev 0.1440\", \"Verdict-open-jev 48.07% vs Jev 90.80%\", \"abstention combined recall 10.00%\", \"p50 35.58 ms\", \"K=25 (maximum capacity) 72.00%\", \"0.85 coverage 84.60% selective risk 1.18%\", \"26.1× faster than standard Qwen JSON generation\", \"Jevify 90.0% / 167 ms CUDA graphs disabled\", \"Finding 1: Brier on stated confidence alone is a trap\", \"grpo_rlcr 0.78 / ECE 0.084\", \"reliability 0.007 but resolution 0.000\", \"27 900 schema-driven decisions\", \"13 600 / 13 600 questions\", \"candidate mass min 0.99999624\", \"22 configs · 166,054 rows · 4 calibration-gold\", \"sha a39eba3f\", \"Student B MAE 0.148 / Pearson 0.836 / 86.0%\", \"pngwn/open-jev-laya-bench README 404\", \"sha 9f69c742 likes 2\", \"HDFS 0.9933 (745/750) / retain 0.0084\", \"BGL ERROR/FATAL protection 1.0000\", \"2,479 / 2,500 HDFS uncertain\", \"cache hit 0.9648 (2412/2500)\", \"$0.153936 estimated\", \"E2 recomputes from saved probabilities\", \"Space sha eda59e0a\", \"MASSIVE English 0.783 / Khmer 0.033 / Hindi 0.133\", \"40–48 rows too small to ship T\", \"T never changes argmax\", \"siren2345/jev-single-decode-transformers ≠ siren2345/jev-single-decode\", \"Split Transformers experiment from llama.cpp runtime\", \"tanayvasishtha/jev-lab ≠ dairui1/jev-lab ≠ BrendanH18/jev-lab ≠ yibie/laya-jev-lab\", \"Four experiments stress-testing TypeSafe's Jev: calibration, bundle bias, label bias, and ensembling\", \"second pass must be $0.00 from cache\", \"The pages never call Jev\", \"Gemma 4 31B 77.0% / Jev 1.13.0 61.4% / Laya 322M 0.0%\", \"restriction state 95.0% against 84.4%\", \"None of the systems are particularly good at knowing when to stop and ask\", \"They skip the question and call a tool directly\", \"100% schema pass\", \"six-field joint 48.8% vs 72.8%\", \"ywchiu/jev_benchmark ≠ Running-Dolphins/jev-bench ≠ Praveenrajus/jev-bench\", \"ACT / REVIEW / FALLBACK\", \"A provider failure, timeout, malformed output, or missing answer is **not** a policy outcome\", \"confidence is descriptive provider output, not a substitute for probability\", \"Quality denominators include only valid scored answers\", \"an exact halfway tie chooses the lower level\", \"aiwithenoch/Jev-Skill ≠ simplosophy/jev-skill ≠ laguagu/jev-skills\", \"The local path does not claim to turn a smaller checkpoint into Jev\", \"Low support becomes decision: \"review\"\", \"MIT-0 SPDX NOASSERTION\", \"current-llm\", \"结构兼容,不是 Jev 模型能力\", \"altryne/jevify ≠ Mintzs/jevify ≠ gulagala001/jevify ≠ uspraveen/Jevify\", \"Find where Jev belongs. Design the questions. Measure the difference\", \"TypeAR-AI/TypeAR 301 → TypeLLM/TypeLLM\", \"TypeLLM/TypeLLM 16★\", \"SemIf 2241★ (+34 vs §111 2207)\", \"jevlike 1051★ (+8 vs 1043)\", \"AnotiaWang 98★ (+1 vs 97)\", \"yibie/awesome-jev 525★ (+19 vs 506)\", \"Laya likes 864 (was 822)\", \"tracker likes 67 (+3 vs 64)\", \"lastModified UNCHANGED `2026-09-20T04:29:16.000Z`\", \"do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#32\", \"soft Noul ≠ hard safety\" with a placement, not a stack replacement or a vendor how-to. Formal methods are one pillar., \"A hunch is a probability with a policy attached\", \"{ enter: 0.8, exit: 0.6 } is hysteresis\", \"replay a policy change without inference\", \"Decision models are providers, not the product\", \"huncho ≠ Kungie/gut ≠ carldaws/hunch ≠ tpellet/hunch\", \"pretrained Qwen2.5 base ECE 0.030 (0.5B) / 0.040 (7B)\", \"instruct 0.302 / 0.269\", \"70.9% → 70.0% mean conf 74.1% → 96.7%\", \"temperature scaling still matches it in-distribution\", \"No Jev API was called\", \"Qwen2.5 ≠ Archer\", \"Qwen/Qwen3.8-27B ≠ Archer\", \"calibration does not compose\", \"ECE has exactly zero statistical power to detect the failure mode that kills trajectories\", \"25–60× headline withdrawn\", \"P(all-correct): 0.0071 vs 0.0001\", \"TCE / AMS\", \"Deferred Crispification\", \"Qwen 3.8 sparring ≠ Archer\", \"pd.cut bins by equal width while jeval bins by quantile\", \"ECE 0.113 and ECE 0.076\", \"jeval drift is not implemented yet\", \"ranking ≠ calibration\", \"g0runmezadam/what-is-jev IS tunahansahin897/what-is-jev\", \"same GitHub id 1378007307\", \"947 repos scored\", \"A 273 · B 302 · C 372\", \"LLM rubric ≠ benches\", \"Probabilities are advisory, not calibrated guarantees\", \"light_cutoff_applied_to_combination 0\", \"AND: product (independence assumed and recorded in the trace)\", \"circuit-vl-4b ≠ Archer\", \"soundness theater / measurement theater\", \"hourly 0843 / notes.md §114\", , \"Jev Capability Resolver / NiazMorshed2007/jcr\", \"one tool to find documented deterministic commands in a nested capability tree\", \"returns context; **does not execute**\", \"skills = workflow+judgment; capabilities = individual operations\", \"format independent of Jev; proposed open standard exploration\", \"classify (Jev) → optional OpenAI decompose compound → beam search geometric mean of routing probs\", \"keep up to 3 paths ≥60% of best (JCR_BAND_RATIO 0.6)\", \"soft scores ≠ hard gates; 0.6 band is application policy\", \"routing ≠ permission; docs ≠ authority to run\", \"sol-vs-opus5-20 *theirs*; lookup+explain only; n=1 per cell; Not Harbor task-execution\", \"Claude Opus 5: agent input 108,585→15,819 (−85%)\", \"Codex GPT-5.6-Sol wall 25.3s→62.4s (Sol slower with JCR in 19/20)\", \"NiazMorshed2007/jcr ≠ skill-broker ≠ skillranker ≠ jev-sift ≠ jev-lens ≠ jevusher ≠ jev_select_capability\", \"do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33/#34\", \"notes.md §116\", \"soft Noul ≠ hard safety\" with a placement, not a stack replacement or a vendor how-to. Formal methods are one pillar. Not a substitute for the official typesafe-ai skill (live Jev API contracts)." license: MIT metadata: version: 0.4.0 @@ -54,7 +54,7 @@ classical method you already trust, substitute it, classify the win "paraphrase brittleness", "allowlist then judge", "TOCTOU-of-Noul", "Jev inside the database / sqlite-jev", "Jev picks bitrate / join order / the model", "wait for Archer", "lint the request / missing - other", "training confronts Choice other / none-of-the-above", "soft AGENTS.md rules vs the linter", "screenshot Choice / omni System One", "extractive quotes / pointer not generator", "compaction summarize vs pointer", "encoder vs Jev compaction backend", "shadow-mode compaction rollout", "CI flaky-vs-real merge gate", "fail-open VOI wake/resume", "claim vs session evidence", "S1 indexer escalate-S2", "Harbor on/off routing", "fail-open vs fail-closed wake vs CI gate", "encoder vs Jev computer-use backend", "hybrid local decide + remote fill", "DONE vs verified success", "stdout prune vs session compaction", "OpenCode jev-pruner vs Claude jev-pruner", "zen-chat vs jev-zen Noul", "hard envelope then Noul prune", "Cua-S1 vs TypeSafe Jev", "plan vs execute dry-run", "specialist computer-use vs general agent", "local drop-in vs stub scorer", "route vs memory", "when does it hold / extractable from state", "decision model vs constrained LLM", "dual-process S1/S2", "combinatorial grid vs extractive", "uncalibrated local likelihoods", "decision-native RAG", "classify-first / read selectively", "living applied-mappings atlas / class patterns", "silence as safer / draft-gate heartbeat", "robotics text-state vs pixels", "verbatim ledger vs summary", "judgment as language primitive", "Stagehand extract pick-and-copy", "harness observe-score-act vs demo loop", "public judgment wall / six parallel questions", "meaning-search without embeddings", "attention ≠ correctness", "skills→oxlint / AST prove ∩ remainder", "session-sticky first-prompt routing", "measured RAG rerank vs generative rerank", "capability kernel / secrets never in the agent", "Jev is SENSOR not policy", "type-safe ≠ correct", "typed control plane around DSPy", "native vs verbalized confidence", "engine owns truth / Jev owns judgment", "human-confirmed kill gate", "train specialist vs few-shot hosted", "decide→policy→LLM leftover", "Noul 0.5 cannot-tell never rounded", "calibration ≠ sortable / ORDER BY", "pairwise inversion / Score ordinality / two-decimal ties", "wire-compat GLiFormer /v1/systemone", "class-backend economics", "loopback gateway hosted + local", "do not distill Jev as teacher", "active-learning triage", "evidence-packet explorer", "meaning-grep AND/OR/NOT", "closed-vote-only / no planner LLM", "Jev vs PCD Harbor", "PCD O(1) ≠ Noul", "host-owned handlers × System One", "OMP/pi fail-open gate", "permission vs probability / operator owns thresholds", "judgment ≠ permission / Jev never grants access", "eval integrity / instrument not score", "constrained optimizer + S1 features / never sole hot-path gate", "privilege ≠ verdict / effect contracts not tokens", "attention filter / VOI for human review / never blocks / never green unless sure", "measurement owns endorsement / evidence-gated question packs", "Jev supplies evidence / code owns authority", "ranking ≠ calibration / never hard-threshold raw p as frequency", "hot-click CU / indexed element table", "Jev judges relevance / code decides structure", "local rules first then remainder / never auto-train on own hides", "combinators / System One as control plane", "receipts not leaderboard / type-safe ≠ correct jaggedness", "VOI over skill library / skillranker abstention", "OOD calibration / AUC ≠ ECE", "Jev vs thinking-budget small models", "turnstile / replayable evidence≠authority", "MLX one-pass schema→JSON / Apple Silicon replica economics", "memory leases ended by new evidence", "never confidently wrong / TLA+ compose / escalate instead of hard-gate", "no seal no advance / coverage ledger / mint ≠ product brain", "skill-broker sibling / judgment ≠ permission", "sureness bands / max_prob is generous", "JevBench / calibration not in Main Score", "CI typed gate before expensive review", "Codex MCP host adapter", "judgment as attention redirect / jev-preflight", "compress-before-first-send / dizk jev-lens", "tools≠use / SessionStart over hoping", "observational memory / pi-om keep-kind", "open-Jev class / openvons / JevPick", "physical-world System One / HA-Jev / not for locks", "judgment outside the store / jevql", "landed-script trust / headless≠auto-approve", "digital-design combinators / extended five", "VOI cache admission / same-intent skip LLM", "BM25 vs Jev skill routing Harbor harness", "zeroshot vs BERT / contamination DiD", "typed escalate continue abort baton / inverted loop", "worth-your-attention VOI / ThinkyMiner Winnow", "Jev WHETHER Python HOW LLM WHAT", "conflict vs ignorance / named Choice escape", "Playwright executes Jev chooses", "OpenJev /v1/decide not drop-in", "SemIf wire-compat runoff", "decision-as-memory flywheel", "record/replay CI / jevassert", "failure-finding arena / jevarena ≠ jev-arena", "BBQ not a bias cert", "decider≠executor", "sentence-as-rule lint / jevlint", "sentence-as-rule lint / jev-lint is jevlint rename", "VOI hunk prune", "whole-repo intent VERIFIED/VIOLATION/UNKNOWN", "GLiNER2 spec ≠ replica", "open replica substrates / grande / laya-jolt / JEV-CPU", "ONNX local-jev not equivalent", "persist constraints across compaction / pi-heed", "calibration+cost first-class gates", "Harbor-shaped Jev vs SGR LLM-as-judge / jev-judge-bench ≠ jevarena ≠ jevbench", "hand no-text steps / jev-use / Vercel drops confidence", "Pi System-One control plane / pi-jev-control", "never free-generates / jev-gpt tree of Choices", "OpenRouter recipe atlas / samples not benches", "personal history feed / jevfeed / no social graph", "competing NAR claims / dual-channel ECE / openJev-verdict ≠ OpenJev", "empty compaction-proxy skip / IPECTER", "throughput ≠ latency / like-for-like ECE", "1-token logprob endpoint ≠ Noul / coverage ≠ correctness", "open replica engine / jevinf", "unofficial Elixir SDK ≠ OTP peer", "jevex n=16 files-to-read VOI", "commit pre-review attention≠verdict / middle band", "Hermes plugin is Agnes not TypeSafe", "pi-jev-compact ≠ pi-jev-compaction", "empty Codex-proxy skip / IPECTER runway", "decision-native inbox / mailordinal", "unofficial jev-cli not ready / ≠ jevql", "laya-multilingual / English checkpoint confident-wrong OOD", "schema-scorer peaked ranking ≠ calibration", "HF 401 / GitHub 404 Hub-only", "productized System One HTTP / classifier.dev", "escalate-under-threshold / smart tier / multi-label ignores", "silent FALLBACK / granite 0.546 vs advertised 0.800", "vs_jev tracked JSON / read eval/README", "choxos/jev-reviewer ≠ egma-ai / systematic-review pointer", "two-pass Choice+Noul evidence extraction", "not-found is an answer", "human check as productized judgment", "githubnext/localjev ≠ kunchenguid/local-jev", "wire-compat ≠ logit-equiv / prompted JSON ≠ structured read", "self-reported probs / entropy confidence", "GitHub Next local /v1/systemone", "LM Studio runner gap / structured-read primitives", "NandhaKishorM/laya packaging ≠ Hub-only / Router script-before-p", "post-T ECE ≠ raw ECE / Banking77 token-budget", "0.85 still soft / not TypeSafe drop-in", "external census ≠ scored bake-off", "GLiNER2+routers class-boundary", "incomplete openjev census vs watch", "Harbor honesty watch / silent fallback", "JevBench v1.2 geometric mean / cal ON rank / weight sensitivity", "option-order 72→21 / instruction models in the class table", "self-host latency ×2 assumption / est. costs", "Laya absent is a gap not a named exclusion", "Qwen3.8 27B ≠ Archer", "hourly already-folded watch / apply-the-five / skip thin noise", "hard-gate Noul as PR/quality gate is soundness theater", "S1 never stalls waiting / S2 one-use advisory", "Local controller ≠ githubnext/localjev", "purple telemetry = consumed not arrived", "seed = geometry not async replay", "20% starting gate still soft", "no pixels to either provider", "OCR+AX observe-score-act / typesafe-computer-use", "never send screenshot to frontier for the decision", "overlapping CU options = false low confidence", "split kind/item/site", "155× one-screenshot ≠ Harbor taskset", "decision ≠ answer-reader capture", "ASR observe-score-act / jev-voice-browser", "partial-speech VOI / free-text waits", "spoken confirm ≠ hard auth", "numbered overlay without another model", "wrap-as-execution / AgentGhost ALLOW ASK DENY", "rules first then Jev remainder / ASK throws / fail-closed", "reddpy/AgentGhost ≠ jwen5419807/agentghost ≠ vventirozos", "JP genre atlas / studio_yebisu / stars ephemeral ≠ eval", "Jev Clearly Explained / akshay_pachaar / LLM hammer", "schema-safe ≠ correct / 200× 400× TypeSafe ceiling", "questions-as-code / shadow first / not a TypeSafe how-to", "proposition ≠ embedding / contrast-set", "boolean composition of soft Nouls / AND OR NOT", "uehaj/jev-semgrep ≠ semgrep.dev", "meaning-grep dedicated fold / not a gate", "decision-validated UI / Jev never authors text", "decision-as-assert / jevtest ambiguous band", "typed decisions drive UI / jev2ui", "hybrid S1 closed verb menu / anima3", "pointer-not-generator search / JevFind", "jev-frontier-bench ≠ frontier-100", "product bakeoff ≠ architecture duel / GLiClass", "four engines same questions / majority floor", "authorship named escape / not evidence", "ha-switchboard HA remains execution", "n8n classify/route/score / Low Confidence", "fast-jev-compaction-pi ≠ pi-jev-compact ≠ pi-jev-compaction", "jevloop full-distribution optimizer / no LLM in the loop", "laya-vision SmolVLM / score untrained", "Cerebellum-2B /v1/decide ≠ TypeSafe / wire-compat vs agent-routing", "laya-grounded not drop-in / Platt not temperature", "GestaltLabs/Jeff-1 ≠ logan-markewich/jeff / acc vs ECE n=9730", "stanley-code empty findings ≠ approval / human promote", "findme ≠ JevFind / NL memory beam-search FS", "jevsubrouter price workers not conversation / counts ≠ dollars", "feelings .feels() default 0.5 is Noul-0.5-never-rounded / ≠ hunch ≠ Probably", "apa-agent-harness ≠ AntonioCoppe/jev-harness / unpublished npm", "grok-bot-jev skill cannot force a bot that ignores it / A/B proxies not tokens", "Essentiel-Jev never authority / human every action", "enzo-mcp independently falsifiable claims / ≠ jev-sift", "pigeonhole OTHER skip / decision-as-filing", "jev-reliability Nothing about accuracy", "clduab11/jev-test ≠ realZachi/jevtest / Nothing runs yet", "jev-rag-benchmark Jev wins is not an assumption", "dairui1/jev-lab ≠ BrendanH18/jev-lab", "jevmail gmail.readonly / mailjay archive/trash", "ZHUBoer/ego-jev reserved __none__", "runWorkflow completed ≠ success", "jsort scores are relative", "Noul not Choice for scale", "groundedness-judge-bench native vs schema-guided", "implicit_true included in yes", "jev_playground 0 promotions", "routing-backtest 0.0447%", "yuyang2230/jev-agent-skill jev-1.13-free", "jev-techstack-classifier stack_config.json", "s1_ruby collapse late", "undecided? abstain", "2389-research/judgement license null", "confidence ≠ winner p", "typesafeai-sdk-community not a new species", "tpellet/hunch exit 3", "never-execute list", "jev-file-search scores not calibrated accuracy", "jev-linkmap Jev never sees S2 prose", "muhammedilyasy/jev-mail metadata only", "tidy none-of-folders stay", "tab-bouncer pinned/audio/current never closed", "lkclean Show fail-open", "jev-yt-time-saver Show anyway", "ORIGIN pause-if-no-Jev", "validResponse sums-to-1", "jev-crawlers risk bands never raw boolean", "jevbrain AUTO_ACT is not a Noul", "judgekit YAML classify/score/route/verify", "typed-judge-kit verdict-in-code", "alsoleg89/decide packing VOI", "0.8 ≠ 80% accuracy", "Jev-Calibration Platt ECE 0.117→0.052", "jev-calibration-arena never acts", "ctmx/openrouter-jev-mcp Decision-as-Plugin", "FrancoisChastel/jev-code ≠ npm jev-code", "claudecode-jev-marketplace fail-open not hot path", "pedroknigge/mcp_jev packs not ask_jev", "cyrusasco/typesafe-mcp noul deadband 0.35–0.65", "codaaiteam/jev-skill jevtypesafeai.com ≠ TypeSafe", "hermes-switchyard ≠ hermes-jev-router ≠ hermes-plugin-jev", "nanoprune 2.8MB ECE 2.58%", "smartdio/jev-browser-agent ≠ ZHUBoer/ego-jev", "Dakai/omp-jev-web DONE ≠ proof", "hari007sh/jev ≠ dannote/jev", "0thernet/system-one-skills deterministic verify", "typed-gate band [0.40,0.60] is refusal", "pi-jev-gate fail-closed; choice is the verdict", "Foq ~25ms/2.2GB local", "rev prefill-only + HF jev-0.5b", "robfrase/jev planning memo", "typesafe_agent_gates 27/27 / 31/31", "EpicEric/safe-sh static remainder", "pastepilot Confirm before act", "Jev-Reranker live Jev not yet measured", "sessionwise opt-in relevance", "jev-search pointer sieve", "400ms Salesforce WebMCP", "typesafe-scheduler-diagnostics advisory", "droidjev screenshot-free", "Tewoto1 jevcu planner still writes", "ha-conversation-jev Jev→Grok", "dsh-jev can only gate", "jev-classification-benchmark specified not run", "jev-luna-pagerduty p≥0.50", "meldltd/meldecision laya-go ONNX", "laya-doom never pixels", "logixism/laya-api empty README", "akpsahan/laya ≠ Archer", "choxos/jevchess engine owns truth", "jev-drive sim not AV", "story-arc Jev never authors", "jev-hs-assistant HS6", "golergka/jev-plays-starcraft-2 UI-verified ≠ API Victory", "awesome-jev-use-cases catalog", "Nibir1/typesafe-go ≠ official", "fingerprint after redact", "recall vs decide", "publish fingerprints+answers", "CI replay as Harbor cousin", "Cache hit ≠ correctness", "hyperspaceai/jevcache ≠ kushals256/jevcache", "human labels only", "score never auto-accepts", "production capture flywheel", "sutro-sh/jev-align ≠ caiovicentino/jev-align", "guidance ≠ hook", "catalysts ≠ summaries", "compile-time System One", "unofficial ≠ TypeSafe", "format_version modernbert-jev/1", "Argos1111/jev_local ≠ us/jev-local ≠ kunchenguid/local-jev", "LFM default ≠ ModernBERT backend", "Nemotron ≠ TypeSafe Jev", "not a calibrated replacement", "djev-dev complements djev-spark", "images as Choice options", "Laya essay numbers *theirs*", "Router/OOD confidence", "hosted bootstrap ≠ silent TypeSafe", "difficulty + policy thresholds + JSONL trace", "jev-codex-pilot model + reasoning depth", "keep/shadow/hybrid/reject", "quarry evidence projection", "Frank-ZY-Dou/awesome-jev robotics/3D/control", "one-dollar-tahoe TypeSafe Jev defense eval", "jevguard calibrator/cache/escape", "jev-ci-selector CI shadow mode", "llama-jev llama.cpp replica", "petercr/jev-orchestrator ≠ FleeexCorp/jev-orchestrator", "seb4ez/jevguard ≠ AseemPrasad/JevGuard ≠ pablozr/JevGuard", "webNeat/llama-jev ≠ WiktorB2004/llama-index-jev", "OpenCode jev-pruner context sieve", "observe→score-candidates→prune", "jev-zen / jev-1.13-free", "zen-chat ≠ Noul", "fail-open original", "keepScore >0.1 floor", "host port of tamaratran/jev-pruner", "indiejoseph/opencode-jev-pruner ≠ nrdz-labs/fast-jev-opencode", "jev-webagent-bench empty stub", "Kiln-AI/jev_jsonschema noul_threshold 0.5", "NSStudent/JevSwiftSDK unofficial", "GLiNER2 native Apple path", "unofficial Swift/Core ML GLiNER 2.5-small", "entity spans + confidence", "not Choice/Score/Noul", "not TypeSafe", "label descriptions as schema", "on-device ANE economics", "honesty locks", "shershah1024/gliner-native-runtime ≠ Fastino", "≠ gliner25-compaction ≠ gliner2-ultrafast ≠ Eran-BA/Jev_from_GLiNER2 ≠ NSStudent/JevSwiftSDK ≠ jevmlx", "default threshold 0.1 still soft", "soft Noul ≠ hard safety", "Decision Graph Protocol frame→assess→commit", "app retains permissions/effects", "Jev-first assessor-neutral", "guarded commit / receipt/next frame", "assessment batching", "hard-gating DGP as safety theater", "numerous-com/dgp ≠ TypeSafe official", "jegrep calibrated path+range Nouls", "no embeddings/index/daemon", "~$0.01–0.03 typical", "agent --json", "can1357/jegrep ≠ Bentlybro/jevgrep ≠ uehaj/jev-semgrep", "Archer-arch fidelity", "kev family OOD 0.76–0.77 vs Jev 0.86", "block-causal isolation", "pointer/readout CE-trained", "/v1/systemone drop-in", "replica honesty", "cost-sensitive decision theory × System One probabilities → control flow", "thresholds derived from costs not hard-coded", "YES / NO / UNSURE from cost_false_yes / cost_false_no / cost_human", "auto-batching same-object questions", "Kungie/gut ≠ tpellet/hunch ≠ carldaws/hunch", "judgment vs generation", "deterministic execution after probabilistic judgment", "exactly one app-owned callback", "explicit uncertain branch", "Illusion47586/judge ≠ lexingtonhibiki/judgekit ≠ Ascurse/typed-judge-kit", "variable-N option scoring as the trainable object", "dynamic candidate bags not fixed label sets", "zwliJay/jev-forge ≠ NanoJev", "open replica economics / latency vs closed Jev", "NAR local drop-in", "wfzyx/von late-catch HIGH", "competing NAR claims / replica honesty", "typed judgments vs chat judges on guardrailing", "ishaannk/llm-vs-jev cross-note only", "deeper integrity fold is rh-guard", "nothing wins outright", "can be argued out of guarding"", "Jev IS the if-statement", "judgments/probabilities drive branches", "text model only writes prose", "interpreter owns variables/loops/budgets/replay", "otherwise maybe / confidence gate", "chaos samples after the gate", "southpolesteve/probably ≠ carldaws/hunch ≠ feelings ≠ Kungie/gut ≠ Illusion47586/judge ≠ tidymodels/probably", "133★ / forks 10 live", "build calibrated classifiers from human feedback", "retrieve by relevance not resemblance", "one calibrated yes/no per memory in one request", "pointer mode 17/18 19/20 *theirs*", "embedding resemblance misses the allergy", "samdotmak/jev-recall ≠ jev-search ≠ jev-sift ≠ carryforward ≠ chopratejas/invalidate", "memory leases ended by new evidence", "six Nouls then fixed rules in code", "0 of 157 false invalidations", "questions/plans/directives are not evidence", "unsure → review queue", "host keeps the store", "name↔body / comment truth / test-claims", "mizchi/jev-lint is mizchi/jevlint rename", "no shipped rule has severity error", "~1 in 5 findings wrong *theirs*", "mizchi/jev-lint ≠ huntedman/JevLint ≠ MichitoSugawara/jev-lint", "JSON Schema → typed JSON via Jev", "noul_threshold 0.5 decoder not a proof", "IncompatibleSchemaError lists every bad property", "on-device Laya CoreML ANE", "~5 ms P50 short decisions", "189/189 FP16 checkpoint parity", "10× not achieved", "mizorewww/laya-coreml ≠ gliner-native-runtime ≠ jevmlx ≠ NandhaKishorM/laya", "softmax over allowed tokens ≠ Noul", "question-first cache", "Micha0827/snapjudge ≠ githubnext/localjev ≠ jevmlx ≠ cendress/SnapJudge", "Jev-first Pi agent loop", "slow-LLM fallback", "explicit action menu / CandidateSource unimplemented", "62 tests wiring not quality", "direwolfiy/JevPi ≠ standardagents/jevpilot ≠ pi-jev-control", "resume-screening bias audit methodology", "name×resume factorial independent Nouls", "callback determined by resume quality", "mean-probability name gaps operationally negligible", "natemoo-re/bias-bench ≠ BBQ", "Plan/PRD panel → code-owned pass|review|block", "cheerleading out of scope", "austindixson/planalyzer ≠ single-goodness Noul", "cost-aware multi-model routing/escalation", "decide vs do", "successful-task cost", "cannacre8ive/switchboard-ai ≠ ha-switchboard ≠ hermes-switchyard", "frozen-protocol zero-shot bench", "TypeSafe Jev vs PrismNLI vs Laya", "contamination caveat", "elcronos/jev-vs-open-decision-models ≠ JevBench ≠ DMB", "context-window admission control", "VOI gate which tokens are worth the expensive model", "fail polarity per lens", "on small inputs lenses lose money", "cvsgireesh/jevusher ≠ jev-sift ≠ winnow", "typed decision control plane", "receipt ≠ authorization", "historical-v0 zero retained cases", "MokiMeow/jev-fabric ≠ jev-forge ≠ dgp", "live 15-dim typed rubric re-score per pause", "scoring economics exemplar", "OpenJev/Codiv ≠ TypeSafe hosted", "jose-troche/live-rubric ~$0.000004 desc / ~$0.000006 README", "adversarial pre-registered Jev eval", "28 predictions before data", "123,805 requests", "confidence does not track ignorance", "polite injection 65% / crude 0%", "willkelly/jev-evaluation ≠ jevals ≠ jev-baselines-eval", "provider-neutral Elixir/BEAM Noul/Choice/Score SDK", "class infrastructure", "nshkrdotcom/system_one_sdk ≠ typesafe_sdk ≠ dannote/jev", "question-linting of Jev questions themselves", "nine jaggedness rules, no API key, no labelled data", "static lint ≠ measured separation", "yodablocks/jevq ≠ tenbin ≠ JevLint ≠ commitjev", "open-weights Laya as class exemplar (binding)", "Nx/Bumblebee runtime", "host chooses backend", "ChristianAlexander/laya_ex ≠ system_one_sdk ≠ dannote/jev ≠ NandhaKishorM/laya", "on-chain/edge Laya deploy", "parity_verified stays false", "model output never grants Tx", "humandebri/IC-Laya ≠ laya_ex", "auditable weekend replica", "Jev outputs never used for training", "soft human-vote distributions", "unpaired 0.577 vs 0.727", "agilabs-ai/jev48 ≠ JevBench ≠ Mapika/decider", "adversarial dual-judge / framing attack surface", "comparative framing is the usable judgment", "prior injection crowds out evidence", "copyleftdev/ember ≠ ember.js", "Laya specialist fine-tune pipeline", "training still GPU-pending", "PIXELZX0/XERON ≠ convaiinnovations/laya", "Hub Laya replica drop", "daliborsb/laya ≠ convaiinnovations/laya ≠ NandhaKishorM/laya", "System One student distillation corpus", "gold is programmatic", "teacher is closed-API clone", "do not distill Jev as teacher of record", "MagaBitmex/jev-4b-distill-data ≠ missing student checkpoint", "non-LLM VIN System One", "planning depth not chat", "lewislululu/jevon ≠ douglance/jevon", "source-bound evidence checks", "local quote mismatch needs no API", "exit 0 ≠ claim truth", "WaynezProg/jev-kit ≠ jonathanavis96/jev-kit (Airlock) ≠ jev-use ≠ jev-mcp", "independent System One evidence catalog", "scores not one leaderboard", "no external record currently reproduced", "TokenTrim no-Jev matched hybrid 62.4%", "reachjalil/system-one-bench ≠ mallahyari/system-one-benchmark", "21 tasks · 134 items · 208 questions", "scenes from public GitHub contracts, not production logs", "SivletLabs/jev-eval ≠ willkelly/jev-evaluation ≠ 4esv/jev-eval ≠ xxkuboxx/jev-eval ≠ onlyoneaman/jev-eval ≠ dayhaysoos/jevals", "option isolation (sibling-blind)", "permutation-equivariant", "Hub OWNER not published", "nafisazizir/hev ≠ jaredpalmer/kev", "frozen local LLM logits, no trained decision head", "residual-head 9,222-param decreased 73/96→67/96", "confidence = 1−normalized entropy, not P(correct)", "yuki-oshio/mini-jev ≠ r-ms/mini-jev", "Jev classifier as autoregressive next-token predictor", "ChatJev-style soundness theater", "erik-dunteman/ChatJev ≠ dannote/jev ≠ jev-gpt", "calibrated decision head × AlphaProof value head", "implementation-layer isomorphism, semantic difference", "timeout = censoring", "do not launder Noul as proof", "parallel rank-prediction vs serial selection", "independent questions can conflict", "zzzzzec/jevsort ≠ keltokhy/jsort", "curated open System One ecosystem catalog", "rupeshpoojary9/awesome-open-system-one ≠ AnotiaWang/awesome-jev", "arXiv paper radar with Jev relevance scoring", "ranking ≠ calibration / 0.5 still soft", "fail-open failed evals not marked seen", "train calibrated ~27M from scratch", "typed Q→prob dist / one forward pass / no LLM decode", "hyusi2003/MiniSystemOne ≠ Colvin0315/MiniSystemOne", "description-only stub / size 5", "ESCI hard probe fails four of six", "jev_bool ECE 0.242 inversion 0.255", "do not re-fold §60 six-gates as new", "jobbyjev one-request-per-company from batch-size result", "find/design/evaluate TypeSafe Jev decision loops", "karanb192/jev-architect ≠ samtay32/jev-system-architect", "Jairik/jev-distiller size 1", "distill-Jev UI stub / do not distill Jev as teacher of record", "post-launch scored use-case map / Jev self-scores then human curation", "licensedsaucer9-web/jev-opportunities", "Jev-inize a use case into classifier/router", "gavinHuang/jevinize → simple-jev not TypeSafe", "featherless-ai/simple-jev", "compare saved decisions / same label can still change the branch", "VihaanAgarwal/jev-diff ≠ Saik0s/diffusiongemma-jev-macos", "not tested with a live Jev API key", "constrained logprob + temp/Platt ≠ Noul", "OpenJevPro pastes openjev-sglang JevBench as own", "zhangcy122/OpenJevPro ≠ IamBusy/OpenJev ≠ ekzhang/openjev-sglang", "PolyForm Noncommercial", "SmolLM-135M / sub-70ms / 0 output tokens", "demo P(True) 0.5052 / Choice conf 0.2872 / Score conf 0.0055", "README claims MIT / GitHub license null / no LICENSE file", "patelvishwa112/jev-system-one-rlcd ≠ arnabgho/rlcd-lite ≠ blackwood-rlcd", "source-backed Awesome Jev radar / 306+ commit-pinned", "logicrw/awesome-jev-projects ≠ AnotiaWang/awesome-jev ≠ yibie/awesome-jev ≠ cobanov/awesome-jev ≠ rupeshpoojary9/awesome-open-system-one", "auto GitHub sync / Issue-only submissions", "hashed n-gram encoder / rival-aware attention", "olanotolu/jevbetter vs jevlike starter", "synthetic hard menus top-1 0.916 vs 0.873 / ECE 0.0182 vs 0.0367 / 40 vs 4608 menus/sec", "shuffled-context control 0.335", "Turn any open LLM into System-One Jev", "uspraveen/Jevify ≠ Mintzs/jevify ≠ gulagala001/jevify", "Jevify-any-LLM architecture probe", "description-only stub / size 0", "Train encoder-only calibrated decision models from a task sentence", "Exu is a toolkit, not a method", "strictly proper scoring rule", "Pre-alpha", "Ruivalim/exu-base", "scratch-trained calibrated decision model", "typed Q → probability dists", "Colvin0315/MiniSystemOne ≠ hyusi2003/MiniSystemOne", "no published weights download URL", "90.5 seconds / 29.2% pipeline evidence", "p_i/p_j independent of other candidates", "Recipe for calibrated decision models — small model out", "init → synth → train → eval → serve", "91.1 % / ECE 0.022 *theirs*", "Jev zero-shot 75.1", "scienthoon/luce", "Put Jev's three headline claims on trial", "0.5B local GPU", "46x speedup / accuracy identical", "ECE 0.624 sentiment catastrophe", "bigger model worse calibration", "RichardoMrMu/jev-mini ≠ yuki-oshio/mini-jev ≠ r-ms/mini-jev", "System-1 decision engine for local LLMs", "structured choices only", "JSON parse of generated text ≠ Noul", "TypefAI JEV / Journal Entry Voucher", "tapsin/jev-local ≠ us/jev-local ≠ Argos1111/jev_local", "Jev 1.13 reward-model eval across 8 benchmark tracks", "40,940 examples / 0 API errors", "RewardBench v1 92.58%", "Precise IF 50.63%", "goya4140/jev-reward-model-evaluation", "Scaffolding in progress", "Jev vs LLM support-ticket routing", "static + live decision bench", "TypeSafe's own published benchmark", "illustrative simulations, not live API calls", "JevBench v1 — smart/cheap/fast/reliable", "I/C/S/K 25% geometric mean", "classifier.dev fast tier 84.8 is Jev behind its own API", "do not re-fold §78 v1.2 board as new", "Laya (421M) 70.1 now on board", "Zero-shot/few-shot LLM routing", "hard budget filter before Jev", "Jev never asked to perform budget arithmetic", "Jev judges the next state, XState enforces transitions", "simulation uses synthetic keyword fixtures", "catalog gravity", "v-modal/awesome-jev-tools", "★339 live REST", "curation is not endorsement", "crawler-maintained directory", "Daily GitHub + npm sweep, human-merged", "RadRebelSam/awesome-jev ≠ AnotiaWang ≠ yibie ≠ cobanov ≠ logicrw ≠ v-modal", "HF peft SPLADE/BGE reranker", "rdxtremity/jev-reranking ≠ carlaiau/jev-reranking", "query-side encoders, not a Jev replica", "ONNX System One Qwen3.5-4B scorer", "source:pngwn/system-one-qwen3.5-4b-scorer", "CC-BY-NC-4.0", "temperature 1.75", "transformers.js AutoModel cannot load this graph", "Consistency benchmark Space", "This Space contains no benchmark result yet", "12-case plumbing fixture", "Benchmark-driven Jev router and judge", "cheap alone is not success", "Jev does not write, sum prices, or claim accuracy %", "Sol 94.2 / Luna 83.9 / Jev path 89.7", "19.2% Sol / 62.3% cost save / 4.5pp miss of 2pp non-inferiority", "p50 latency worse than Sol due to routing overhead", "erendikmenn/jev-llm-router-benchmark ≠ jev-rag-benchmark ≠ ryantsai/jev-llm-router", "Express + node:sqlite", "mock and Jev decision engines", "previous_ticket_count >= 3 is code", "MIN_CONFIDENCE 0.6 still soft", "substring false positives", "aesaganda/jev-ticket-router ≠ SarathChandraBellam/jev-vs-llm-ticket-router", "Universal Figure & Diagram Router", "confidence ≥ 0.85 hard-gate is theater", "generative AI banned from scientific plots", "six visual branches", "hoangngochuong24947-gif/jev-figure-router", "human-labeled (state, question, label)", "166,054 rows / 22 configs", "soft_label for human uncertainty", "Praveenrajus/jev-bench ≠ fstandhartinger/jevbench", "ternary bonsai System One GGUF", "openjev's mechanism, Bonsai's weights", "Hub does not ship weights", "100/100 easy T/F is not Harbor", "label_mass ≠ correctness", "stock llama.cpp Q2_0 silently gibberish", "NicolaiMTLassen/open-bonzi-jev ≠ NicolaiLassen", "transformers.js DeBERTa ONNX", "source:com-kotobalabs/open-jev-deberta-v3-large", "temperature 1.05", "AutoModel from_pretrained works", "onnx-community/open-jev-deberta-v3-large-ONNX ≠ system-one-qwen3.5-4b-scorer-ONNX", "107★ densify", "GH 151M vs README 149.6M", "PR #1 now closed unmerged", "do not re-fold §71 claim-audit as a beat", "typed decisions, RLCD, confidence-gated routing", "structured ≠ correct", "mock not live API", "26 tests", "wjdjdakf17/jev-study ≠ baekenough/jev-study", "bonzi-27b-v2 / ternary-8b / 27b-v1 GGUF family densify", "WANLI-256 74.6% / 65.2% / 71.1% *theirs*", "Bonsai 1 27B Q1_0 runs on stock llama.cpp", "ternary still needs PrismML fork", "hf:heman10x/openJev-verdict-2.0 twin tokenizer-only", "OpenJev Vision image classification + uncertainty", "CLEVR-4 held-out joint 0%", "hfdataset:IamBusy/OpenJev-Vision-Research-v0.1 12,832", "294,912 derived targets not independent samples", "Laya multilingual ONNX WebGPU typed-decisions port", "63/63 selected answers / 5.1e-4 CPU / 1.2e-2 WebGPU", "UpHash-Network/mini-jev is yuki-oshio transfer", "jev-injection-bench 11,900 labelled prompts", "Jev best ranking / Haiku better ECE 0.021 vs 0.058", "0.5–0.9 band is where Jev's numbers do not mean what they say", "Prompt wording moves panic 28%", "manojlds/jev-dspy-bench ≠ dspachos/jev-dspy ≠ jmanhype/jev-dspy-lab", "Jev agreement is similarity, never ground truth", "no aggregate quality grade or merge gate", "AbstentionBench-on-Jev rank 1 of 20 vs 2025 field", "question-asymmetry", "forward-looking 0.465 never extreme", "openkev calibration layer not a runtime", "ECE vs coverage independent", "select_threshold returns inf", "escalation catches uncertainty not ignorance", "misakaikato/openkev ≠ jaredpalmer/kev", "pdf-race Docling→Jev vs Gemini", "parser owns the wall clock", "12/12 tie is a tie", "titles selected not generated", "flopcheck 16 calibrated tweet judgments", "mechanical tells in code", "ZeroX-01/jev-atlas ≠ Zaious/jev-capability-atlas ≠ gorock007/jev-atlas", "Laya calibration lab Gradio MCP", "T never changes argmax", "confidence ≠ top-label p", "easy probe set refused", "40–48 rows too small to ship T", "Gemma-4 26B-A4B jevify classification+calibration", "LoRA adapter twin not independent eval", "Gemma-4 E4B jevify", "E4B LoRA stub card", "kushalpatil/jevify-gemma4 ≠ Mintzs/jevify ≠ gulagala001/jevify ≠ uspraveen/Jevify", "GH kushalpatil07/jevify 404", "PAWS 0.580/ece 0.288 is the weak cell", "smaller E4B slightly better OOD ECE than 26B-A4B", "Hub jevify merged LoRA ships weights", "bonzi Bonsai-8B v1 GGUF densify", "Bonsai-1.7B v1", "Bonsai-4B v1", "WANLI-256 64.5% / 60.2% / 52.0% *theirs*", "rank #4 / #5 / #6 of 6", "JulesHuisman/jev-eval scaffolding / README SHA c356a584 (was empty e69de29b)", "JulesHuisman/jev-eval ≠ SivletLabs/jev-eval ≠ willkelly/jev-evaluation ≠ 4esv ≠ xxkuboxx ≠ onlyoneaman ≠ dayhaysoos/jevals", "7 bands 6/10 vs 40 bands 0/10", "source receipts + confidence slider re-policy without re-inference", "32/32 synthetic is smoke not production", "classify HF datasets across typed semantic dimensions", "roadus2 watch misspelling; lock roadius2/ultra_laya", "ultra_laya REVIEW defects", "default branch claude/laya-jev-review-gg5ppo", "XNLI EN 88.3% ECE 0.032 → RU 77.3% ECE 0.096", "Δ −11.0 pp [−14.2,−7.8]; ECE +0.063", "MASSIVE no detectable difference at n=600", "confidence is function of p_max (r=1.000)", "pointer-not-generator 400 human-authored responses", "proposed ≠ authorized", "FewRel 160: Jev 85.0% vs lexical 13.125%", "gated 100% (95/95) coverage 59.375%", "J++ composable semantic computation language", "judge-jev 0.5 still soft", "947 repos scored; A 273 / B 302 / C 372", "LLM rubric ≠ benches", "No benchmark winner is claimed", "phishing: naive 62.6% vs regex 91.8%; 5-atomic + LR 95.0% *theirs*", "AITuber tension ±15", "README npm global; repo is Rust", "git-confess code owns counting/blame/ratio", "httpx exhibit 11% (13/119) *theirs*", "90d trend +12.40% vs random +12.75% vs BH +41.71%", "5m win rate 25%", "Awesomejev 656 entries / 38,160 stars", "tracker likes 64 (+4) lastModified UNCHANGED", "Laya present; Blackwood ABSENT; Archer still promised_not_landed", "Blackwood tracker ABSENT; likes 2 gated manual", "r = c - p_a", "ECE 0.021; acc 0.807 vs warmup 0.746", "Independent primitive", "11.57s vs 54.10s · 4.67× · 120/128 *theirs*", "default path is pretrained Gemma probs not trained RLCD head", "GH Meanblock 404; lock leesk212/JEV-CPU", "softmax over letter slots ≠ Noul", "WANLI 0.741 vs openjev v2 0.77 *theirs*", "3-way NLI ≠ Noul", "priority 0.464 = majority floor", "banking77 contaminated", "raw margins not probabilities", "do not distill Jev as teacher of record (they distilled Haiku)", "“0.9 is not one number”", "ranking ≠ calibration", "banking77 0.8–0.9 stated 0.86 actual 0.73 over-confident *theirs*", "≠ Praveenrajus/jev-bench ≠ fstandhartinger/jevbench", "$0.0000153–$0.0000226 vs circulating $0.0004 (~20×)", "Score is 0..n-1 expectation not 0–1", "Noul has no confidence field", "TCP floor 198.8 ms", "type reliability is not a reason to choose Jev (json_schema 5/5)", "gateway tax not one number", "Function-only 5/8 vs hybrid 8/8", "4/8 without Jev", "8 designed cases not conversion lift", "200-row pilot Jev 86.5% 173/200 vs Gemini Flash-Lite 86.0% 172/200 vs Pro 87.0% 174/200 *theirs*", "not a ranking", "情緒測謊器", "8-example Jev vs GPT-5.6 Sol ~64× cost 5.4× latency *theirs*", "synthetic; no inference", "≠ JevBench v1.2 §78", "Judged 3317 / listed 2560", "Jev judges, code applies policy", "APA “microsecond policy / zero hallucination” overclaim", "Client-side quiz; pointer from held docs; scanned-PDF warn", "Jev judges / agent reasons / user decides", "selecting an option is not permission to implement", "pattern exact, judgement must clear floor", "no matching pattern → no model call", "not a correctness oracle", "Spec vs artifact remainder", "treating 0.85 as 85% / minProbability hard-gate as Harbor", "VERIFY acquires discriminating evidence, never same-pool confidence-only rescoring", "fast/full/max are ceilings not sizes", "Solar writes, Jev chooses NEXT ACTION", "do not reopen or amend PR #23 or #24 or #25 or #26 or #27", , "Calibration is not alpha", "NO CURRENT ALPHA CANDIDATE", "ΔR² approximately +0.00084", "Brier 0.2131387", "ECE 0.0421875", "Adding Jev probability to deterministic volatility improved Brier by only 1.4058e-05", "default 0.5 keeps zero non pinned", "keepResult median 0.14 to 0.17", "keepCall median 0.28 to 0.35", "usable range is about 0.10 to 0.25", "7.8% to 57.9%", "judges results it never sees", "task-finish eval not built yet", "$0.002 per compaction", "slavadubrov/sgr-judge-bench ≠ slavadubrov/jev-judge-bench", "Jev 108/120 $0.083 0.34 s", "Luna SGR 114/120", "paired Jev accuracy-difference intervals include zero", "not evidence of equivalence", "GLM SGR 26/120 93 format failures", "Terra-planned Jev hybrid 55/120", "rule-based by default, optionally Jev-backed", "empty README", "missing key cannot break the experience", "prefill plus exactly one decode", "softmax over A/B/C ≠ Noul", "BBQ 9,053/10,000 (90.53%)", "ECE 0.0890", "Mean confidence 0.9943", "overconfident", "score and noul not implemented", "DGUI 12 rows (was 6)", "INSTRUCT 119 rows likes 2", "encode the state once, decide everything in parallel", "0.740 accuracy against a 0.508 majority", "ECE 0.047", "fine-tune's advantage ends where its 384-token training data does", "jasonkneen/open-jev ≠ pngwn/open-jev", "same sha d41dc3cd", "Space does not call Jev", "recomputes routing from saved probabilities", "200-case Jev 97.0% / 100.0% / 95.0% / MAE 9.22", "synthetic repository benchmark", "Jev evaluations are advisory", "YehuiTang0316/jev-nlgrep ≠ Bentlybro/jevgrep ≠ can1357/jegrep ≠ uehaj/jev-semgrep", "default threshold 0.8 still soft", "40-line windows cannot prove whole function", "token-native sequential start/end Choice", "Gemini/Haiku stubs not configured yet", "handful of hand-written examples, not a benchmark", "Jev judged exactly what it was given", "laguagu/jev-skills ≠ laguagu/jev-evidence-lab ≠ Pleo2/awesome-jev-agent-skills", "contract_passed is not a claim of guaranteed factual truth", "Wilson lower bound 0.85 floor", "fixture mode no savings claim", "SemIf 2207★ (+21 vs §110 2186)", "jevlike 1043★ (+5 vs 1038)", "TypeAR 15★ (+1 vs 14)", "AnotiaWang 97★ (+1 vs 96)", "yibie/awesome-jev 506★ (+16 vs 490)", "Laya likes 822 (was 802)", "tracker likes 64 flat, lastModified UNCHANGED", "do not reopen or amend PR #23/#24/#25/#26/#27/#28", "Heman10x-NGU/Verdict-open-jev ≠ Heman10x-NGU/openJev-verdict-2.0", "TF-IDF + LogReg ECE 0.0207 vs Jev 0.1440", "Verdict-open-jev 48.07% vs Jev 90.80%", "abstention combined recall 10.00%", "p50 35.58 ms", "K=25 (maximum capacity) 72.00%", "0.85 coverage 84.60% selective risk 1.18%", "26.1× faster than standard Qwen JSON generation", "Jevify 90.0% / 167 ms CUDA graphs disabled", "Finding 1: Brier on stated confidence alone is a trap", "grpo_rlcr 0.78 / ECE 0.084", "reliability 0.007 but resolution 0.000", "27 900 schema-driven decisions", "13 600 / 13 600 questions", "candidate mass min 0.99999624", "22 configs · 166,054 rows · 4 calibration-gold", "sha a39eba3f", "Student B MAE 0.148 / Pearson 0.836 / 86.0%", "pngwn/open-jev-laya-bench README 404", "sha 9f69c742 likes 2", "HDFS 0.9933 (745/750) / retain 0.0084", "BGL ERROR/FATAL protection 1.0000", "2,479 / 2,500 HDFS uncertain", "cache hit 0.9648 (2412/2500)", "$0.153936 estimated", "E2 recomputes from saved probabilities", "Space sha eda59e0a", "MASSIVE English 0.783 / Khmer 0.033 / Hindi 0.133", "40–48 rows too small to ship T", "T never changes argmax", "siren2345/jev-single-decode-transformers ≠ siren2345/jev-single-decode", "Split Transformers experiment from llama.cpp runtime", "tanayvasishtha/jev-lab ≠ dairui1/jev-lab ≠ BrendanH18/jev-lab ≠ yibie/laya-jev-lab", "Four experiments stress-testing TypeSafe's Jev: calibration, bundle bias, label bias, and ensembling", "second pass must be $0.00 from cache", "The pages never call Jev", "Gemma 4 31B 77.0% / Jev 1.13.0 61.4% / Laya 322M 0.0%", "restriction state 95.0% against 84.4%", "None of the systems are particularly good at knowing when to stop and ask", "They skip the question and call a tool directly", "100% schema pass", "six-field joint 48.8% vs 72.8%", "ywchiu/jev_benchmark ≠ Running-Dolphins/jev-bench ≠ Praveenrajus/jev-bench", "ACT / REVIEW / FALLBACK", "A provider failure, timeout, malformed output, or missing answer is **not** a policy outcome", "confidence is descriptive provider output, not a substitute for probability", "Quality denominators include only valid scored answers", "an exact halfway tie chooses the lower level", "aiwithenoch/Jev-Skill ≠ simplosophy/jev-skill ≠ laguagu/jev-skills", "The local path does not claim to turn a smaller checkpoint into Jev", "Low support becomes decision: \"review\"", "MIT-0 SPDX NOASSERTION", "current-llm", "结构兼容,不是 Jev 模型能力", "altryne/jevify ≠ Mintzs/jevify ≠ gulagala001/jevify ≠ uspraveen/Jevify", "Find where Jev belongs. Design the questions. Measure the difference", "TypeAR-AI/TypeAR 301 → TypeLLM/TypeLLM", "TypeLLM/TypeLLM 16★", "SemIf 2241★ (+34 vs §111 2207)", "jevlike 1051★ (+8 vs 1043)", "AnotiaWang 98★ (+1 vs 97)", "yibie/awesome-jev 525★ (+19 vs 506)", "Laya likes 864 (was 822)", "tracker likes 67 (+3 vs 64)", "lastModified UNCHANGED `2026-09-20T04:29:16.000Z`", "do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#32", "hysteresis enter/exit / replay policy without inference", "calibration does not compose / hop-ECE permutation-invariant", "equal-width vs quantile ECE / ranking ≠ calibration", "Qwen2.5 ≠ Archer / Qwen 3.8 sparring ≠ Archer / Qwen/Qwen3.8-27B ≠ Archer", "Deferred Crispification / TCE / AMS", "g0runmezadam/what-is-jev IS tunahansahin897/what-is-jev", "pd.cut equal-width vs jeval quantile", "A hunch is a probability with a policy attached", "soundness theater / measurement theater / hourly 0843", or "cascade sign-flip / calibration theater": read `references/faq.md`, + other", "training confronts Choice other / none-of-the-above", "soft AGENTS.md rules vs the linter", "screenshot Choice / omni System One", "extractive quotes / pointer not generator", "compaction summarize vs pointer", "encoder vs Jev compaction backend", "shadow-mode compaction rollout", "CI flaky-vs-real merge gate", "fail-open VOI wake/resume", "claim vs session evidence", "S1 indexer escalate-S2", "Harbor on/off routing", "fail-open vs fail-closed wake vs CI gate", "encoder vs Jev computer-use backend", "hybrid local decide + remote fill", "DONE vs verified success", "stdout prune vs session compaction", "OpenCode jev-pruner vs Claude jev-pruner", "zen-chat vs jev-zen Noul", "hard envelope then Noul prune", "Cua-S1 vs TypeSafe Jev", "plan vs execute dry-run", "specialist computer-use vs general agent", "local drop-in vs stub scorer", "route vs memory", "when does it hold / extractable from state", "decision model vs constrained LLM", "dual-process S1/S2", "combinatorial grid vs extractive", "uncalibrated local likelihoods", "decision-native RAG", "classify-first / read selectively", "living applied-mappings atlas / class patterns", "silence as safer / draft-gate heartbeat", "robotics text-state vs pixels", "verbatim ledger vs summary", "judgment as language primitive", "Stagehand extract pick-and-copy", "harness observe-score-act vs demo loop", "public judgment wall / six parallel questions", "meaning-search without embeddings", "attention ≠ correctness", "skills→oxlint / AST prove ∩ remainder", "session-sticky first-prompt routing", "measured RAG rerank vs generative rerank", "capability kernel / secrets never in the agent", "Jev is SENSOR not policy", "type-safe ≠ correct", "typed control plane around DSPy", "native vs verbalized confidence", "engine owns truth / Jev owns judgment", "human-confirmed kill gate", "train specialist vs few-shot hosted", "decide→policy→LLM leftover", "Noul 0.5 cannot-tell never rounded", "calibration ≠ sortable / ORDER BY", "pairwise inversion / Score ordinality / two-decimal ties", "wire-compat GLiFormer /v1/systemone", "class-backend economics", "loopback gateway hosted + local", "do not distill Jev as teacher", "active-learning triage", "evidence-packet explorer", "meaning-grep AND/OR/NOT", "closed-vote-only / no planner LLM", "Jev vs PCD Harbor", "PCD O(1) ≠ Noul", "host-owned handlers × System One", "OMP/pi fail-open gate", "permission vs probability / operator owns thresholds", "judgment ≠ permission / Jev never grants access", "eval integrity / instrument not score", "constrained optimizer + S1 features / never sole hot-path gate", "privilege ≠ verdict / effect contracts not tokens", "attention filter / VOI for human review / never blocks / never green unless sure", "measurement owns endorsement / evidence-gated question packs", "Jev supplies evidence / code owns authority", "ranking ≠ calibration / never hard-threshold raw p as frequency", "hot-click CU / indexed element table", "Jev judges relevance / code decides structure", "local rules first then remainder / never auto-train on own hides", "combinators / System One as control plane", "receipts not leaderboard / type-safe ≠ correct jaggedness", "VOI over skill library / skillranker abstention", "OOD calibration / AUC ≠ ECE", "Jev vs thinking-budget small models", "turnstile / replayable evidence≠authority", "MLX one-pass schema→JSON / Apple Silicon replica economics", "memory leases ended by new evidence", "never confidently wrong / TLA+ compose / escalate instead of hard-gate", "no seal no advance / coverage ledger / mint ≠ product brain", "skill-broker sibling / judgment ≠ permission", "sureness bands / max_prob is generous", "JevBench / calibration not in Main Score", "CI typed gate before expensive review", "Codex MCP host adapter", "judgment as attention redirect / jev-preflight", "compress-before-first-send / dizk jev-lens", "tools≠use / SessionStart over hoping", "observational memory / pi-om keep-kind", "open-Jev class / openvons / JevPick", "physical-world System One / HA-Jev / not for locks", "judgment outside the store / jevql", "landed-script trust / headless≠auto-approve", "digital-design combinators / extended five", "VOI cache admission / same-intent skip LLM", "BM25 vs Jev skill routing Harbor harness", "zeroshot vs BERT / contamination DiD", "typed escalate continue abort baton / inverted loop", "worth-your-attention VOI / ThinkyMiner Winnow", "Jev WHETHER Python HOW LLM WHAT", "conflict vs ignorance / named Choice escape", "Playwright executes Jev chooses", "OpenJev /v1/decide not drop-in", "SemIf wire-compat runoff", "decision-as-memory flywheel", "record/replay CI / jevassert", "failure-finding arena / jevarena ≠ jev-arena", "BBQ not a bias cert", "decider≠executor", "sentence-as-rule lint / jevlint", "sentence-as-rule lint / jev-lint is jevlint rename", "VOI hunk prune", "whole-repo intent VERIFIED/VIOLATION/UNKNOWN", "GLiNER2 spec ≠ replica", "open replica substrates / grande / laya-jolt / JEV-CPU", "ONNX local-jev not equivalent", "persist constraints across compaction / pi-heed", "calibration+cost first-class gates", "Harbor-shaped Jev vs SGR LLM-as-judge / jev-judge-bench ≠ jevarena ≠ jevbench", "hand no-text steps / jev-use / Vercel drops confidence", "Pi System-One control plane / pi-jev-control", "never free-generates / jev-gpt tree of Choices", "OpenRouter recipe atlas / samples not benches", "personal history feed / jevfeed / no social graph", "competing NAR claims / dual-channel ECE / openJev-verdict ≠ OpenJev", "empty compaction-proxy skip / IPECTER", "throughput ≠ latency / like-for-like ECE", "1-token logprob endpoint ≠ Noul / coverage ≠ correctness", "open replica engine / jevinf", "unofficial Elixir SDK ≠ OTP peer", "jevex n=16 files-to-read VOI", "commit pre-review attention≠verdict / middle band", "Hermes plugin is Agnes not TypeSafe", "pi-jev-compact ≠ pi-jev-compaction", "empty Codex-proxy skip / IPECTER runway", "decision-native inbox / mailordinal", "unofficial jev-cli not ready / ≠ jevql", "laya-multilingual / English checkpoint confident-wrong OOD", "schema-scorer peaked ranking ≠ calibration", "HF 401 / GitHub 404 Hub-only", "productized System One HTTP / classifier.dev", "escalate-under-threshold / smart tier / multi-label ignores", "silent FALLBACK / granite 0.546 vs advertised 0.800", "vs_jev tracked JSON / read eval/README", "choxos/jev-reviewer ≠ egma-ai / systematic-review pointer", "two-pass Choice+Noul evidence extraction", "not-found is an answer", "human check as productized judgment", "githubnext/localjev ≠ kunchenguid/local-jev", "wire-compat ≠ logit-equiv / prompted JSON ≠ structured read", "self-reported probs / entropy confidence", "GitHub Next local /v1/systemone", "LM Studio runner gap / structured-read primitives", "NandhaKishorM/laya packaging ≠ Hub-only / Router script-before-p", "post-T ECE ≠ raw ECE / Banking77 token-budget", "0.85 still soft / not TypeSafe drop-in", "external census ≠ scored bake-off", "GLiNER2+routers class-boundary", "incomplete openjev census vs watch", "Harbor honesty watch / silent fallback", "JevBench v1.2 geometric mean / cal ON rank / weight sensitivity", "option-order 72→21 / instruction models in the class table", "self-host latency ×2 assumption / est. costs", "Laya absent is a gap not a named exclusion", "Qwen3.8 27B ≠ Archer", "hourly already-folded watch / apply-the-five / skip thin noise", "hard-gate Noul as PR/quality gate is soundness theater", "S1 never stalls waiting / S2 one-use advisory", "Local controller ≠ githubnext/localjev", "purple telemetry = consumed not arrived", "seed = geometry not async replay", "20% starting gate still soft", "no pixels to either provider", "OCR+AX observe-score-act / typesafe-computer-use", "never send screenshot to frontier for the decision", "overlapping CU options = false low confidence", "split kind/item/site", "155× one-screenshot ≠ Harbor taskset", "decision ≠ answer-reader capture", "ASR observe-score-act / jev-voice-browser", "partial-speech VOI / free-text waits", "spoken confirm ≠ hard auth", "numbered overlay without another model", "wrap-as-execution / AgentGhost ALLOW ASK DENY", "rules first then Jev remainder / ASK throws / fail-closed", "reddpy/AgentGhost ≠ jwen5419807/agentghost ≠ vventirozos", "JP genre atlas / studio_yebisu / stars ephemeral ≠ eval", "Jev Clearly Explained / akshay_pachaar / LLM hammer", "schema-safe ≠ correct / 200× 400× TypeSafe ceiling", "questions-as-code / shadow first / not a TypeSafe how-to", "proposition ≠ embedding / contrast-set", "boolean composition of soft Nouls / AND OR NOT", "uehaj/jev-semgrep ≠ semgrep.dev", "meaning-grep dedicated fold / not a gate", "decision-validated UI / Jev never authors text", "decision-as-assert / jevtest ambiguous band", "typed decisions drive UI / jev2ui", "hybrid S1 closed verb menu / anima3", "pointer-not-generator search / JevFind", "jev-frontier-bench ≠ frontier-100", "product bakeoff ≠ architecture duel / GLiClass", "four engines same questions / majority floor", "authorship named escape / not evidence", "ha-switchboard HA remains execution", "n8n classify/route/score / Low Confidence", "fast-jev-compaction-pi ≠ pi-jev-compact ≠ pi-jev-compaction", "jevloop full-distribution optimizer / no LLM in the loop", "laya-vision SmolVLM / score untrained", "Cerebellum-2B /v1/decide ≠ TypeSafe / wire-compat vs agent-routing", "laya-grounded not drop-in / Platt not temperature", "GestaltLabs/Jeff-1 ≠ logan-markewich/jeff / acc vs ECE n=9730", "stanley-code empty findings ≠ approval / human promote", "findme ≠ JevFind / NL memory beam-search FS", "jevsubrouter price workers not conversation / counts ≠ dollars", "feelings .feels() default 0.5 is Noul-0.5-never-rounded / ≠ hunch ≠ Probably", "apa-agent-harness ≠ AntonioCoppe/jev-harness / unpublished npm", "grok-bot-jev skill cannot force a bot that ignores it / A/B proxies not tokens", "Essentiel-Jev never authority / human every action", "enzo-mcp independently falsifiable claims / ≠ jev-sift", "pigeonhole OTHER skip / decision-as-filing", "jev-reliability Nothing about accuracy", "clduab11/jev-test ≠ realZachi/jevtest / Nothing runs yet", "jev-rag-benchmark Jev wins is not an assumption", "dairui1/jev-lab ≠ BrendanH18/jev-lab", "jevmail gmail.readonly / mailjay archive/trash", "ZHUBoer/ego-jev reserved __none__", "runWorkflow completed ≠ success", "jsort scores are relative", "Noul not Choice for scale", "groundedness-judge-bench native vs schema-guided", "implicit_true included in yes", "jev_playground 0 promotions", "routing-backtest 0.0447%", "yuyang2230/jev-agent-skill jev-1.13-free", "jev-techstack-classifier stack_config.json", "s1_ruby collapse late", "undecided? abstain", "2389-research/judgement license null", "confidence ≠ winner p", "typesafeai-sdk-community not a new species", "tpellet/hunch exit 3", "never-execute list", "jev-file-search scores not calibrated accuracy", "jev-linkmap Jev never sees S2 prose", "muhammedilyasy/jev-mail metadata only", "tidy none-of-folders stay", "tab-bouncer pinned/audio/current never closed", "lkclean Show fail-open", "jev-yt-time-saver Show anyway", "ORIGIN pause-if-no-Jev", "validResponse sums-to-1", "jev-crawlers risk bands never raw boolean", "jevbrain AUTO_ACT is not a Noul", "judgekit YAML classify/score/route/verify", "typed-judge-kit verdict-in-code", "alsoleg89/decide packing VOI", "0.8 ≠ 80% accuracy", "Jev-Calibration Platt ECE 0.117→0.052", "jev-calibration-arena never acts", "ctmx/openrouter-jev-mcp Decision-as-Plugin", "FrancoisChastel/jev-code ≠ npm jev-code", "claudecode-jev-marketplace fail-open not hot path", "pedroknigge/mcp_jev packs not ask_jev", "cyrusasco/typesafe-mcp noul deadband 0.35–0.65", "codaaiteam/jev-skill jevtypesafeai.com ≠ TypeSafe", "hermes-switchyard ≠ hermes-jev-router ≠ hermes-plugin-jev", "nanoprune 2.8MB ECE 2.58%", "smartdio/jev-browser-agent ≠ ZHUBoer/ego-jev", "Dakai/omp-jev-web DONE ≠ proof", "hari007sh/jev ≠ dannote/jev", "0thernet/system-one-skills deterministic verify", "typed-gate band [0.40,0.60] is refusal", "pi-jev-gate fail-closed; choice is the verdict", "Foq ~25ms/2.2GB local", "rev prefill-only + HF jev-0.5b", "robfrase/jev planning memo", "typesafe_agent_gates 27/27 / 31/31", "EpicEric/safe-sh static remainder", "pastepilot Confirm before act", "Jev-Reranker live Jev not yet measured", "sessionwise opt-in relevance", "jev-search pointer sieve", "400ms Salesforce WebMCP", "typesafe-scheduler-diagnostics advisory", "droidjev screenshot-free", "Tewoto1 jevcu planner still writes", "ha-conversation-jev Jev→Grok", "dsh-jev can only gate", "jev-classification-benchmark specified not run", "jev-luna-pagerduty p≥0.50", "meldltd/meldecision laya-go ONNX", "laya-doom never pixels", "logixism/laya-api empty README", "akpsahan/laya ≠ Archer", "choxos/jevchess engine owns truth", "jev-drive sim not AV", "story-arc Jev never authors", "jev-hs-assistant HS6", "golergka/jev-plays-starcraft-2 UI-verified ≠ API Victory", "awesome-jev-use-cases catalog", "Nibir1/typesafe-go ≠ official", "fingerprint after redact", "recall vs decide", "publish fingerprints+answers", "CI replay as Harbor cousin", "Cache hit ≠ correctness", "hyperspaceai/jevcache ≠ kushals256/jevcache", "human labels only", "score never auto-accepts", "production capture flywheel", "sutro-sh/jev-align ≠ caiovicentino/jev-align", "guidance ≠ hook", "catalysts ≠ summaries", "compile-time System One", "unofficial ≠ TypeSafe", "format_version modernbert-jev/1", "Argos1111/jev_local ≠ us/jev-local ≠ kunchenguid/local-jev", "LFM default ≠ ModernBERT backend", "Nemotron ≠ TypeSafe Jev", "not a calibrated replacement", "djev-dev complements djev-spark", "images as Choice options", "Laya essay numbers *theirs*", "Router/OOD confidence", "hosted bootstrap ≠ silent TypeSafe", "difficulty + policy thresholds + JSONL trace", "jev-codex-pilot model + reasoning depth", "keep/shadow/hybrid/reject", "quarry evidence projection", "Frank-ZY-Dou/awesome-jev robotics/3D/control", "one-dollar-tahoe TypeSafe Jev defense eval", "jevguard calibrator/cache/escape", "jev-ci-selector CI shadow mode", "llama-jev llama.cpp replica", "petercr/jev-orchestrator ≠ FleeexCorp/jev-orchestrator", "seb4ez/jevguard ≠ AseemPrasad/JevGuard ≠ pablozr/JevGuard", "webNeat/llama-jev ≠ WiktorB2004/llama-index-jev", "OpenCode jev-pruner context sieve", "observe→score-candidates→prune", "jev-zen / jev-1.13-free", "zen-chat ≠ Noul", "fail-open original", "keepScore >0.1 floor", "host port of tamaratran/jev-pruner", "indiejoseph/opencode-jev-pruner ≠ nrdz-labs/fast-jev-opencode", "jev-webagent-bench empty stub", "Kiln-AI/jev_jsonschema noul_threshold 0.5", "NSStudent/JevSwiftSDK unofficial", "GLiNER2 native Apple path", "unofficial Swift/Core ML GLiNER 2.5-small", "entity spans + confidence", "not Choice/Score/Noul", "not TypeSafe", "label descriptions as schema", "on-device ANE economics", "honesty locks", "shershah1024/gliner-native-runtime ≠ Fastino", "≠ gliner25-compaction ≠ gliner2-ultrafast ≠ Eran-BA/Jev_from_GLiNER2 ≠ NSStudent/JevSwiftSDK ≠ jevmlx", "default threshold 0.1 still soft", "soft Noul ≠ hard safety", "Decision Graph Protocol frame→assess→commit", "app retains permissions/effects", "Jev-first assessor-neutral", "guarded commit / receipt/next frame", "assessment batching", "hard-gating DGP as safety theater", "numerous-com/dgp ≠ TypeSafe official", "jegrep calibrated path+range Nouls", "no embeddings/index/daemon", "~$0.01–0.03 typical", "agent --json", "can1357/jegrep ≠ Bentlybro/jevgrep ≠ uehaj/jev-semgrep", "Archer-arch fidelity", "kev family OOD 0.76–0.77 vs Jev 0.86", "block-causal isolation", "pointer/readout CE-trained", "/v1/systemone drop-in", "replica honesty", "cost-sensitive decision theory × System One probabilities → control flow", "thresholds derived from costs not hard-coded", "YES / NO / UNSURE from cost_false_yes / cost_false_no / cost_human", "auto-batching same-object questions", "Kungie/gut ≠ tpellet/hunch ≠ carldaws/hunch", "judgment vs generation", "deterministic execution after probabilistic judgment", "exactly one app-owned callback", "explicit uncertain branch", "Illusion47586/judge ≠ lexingtonhibiki/judgekit ≠ Ascurse/typed-judge-kit", "variable-N option scoring as the trainable object", "dynamic candidate bags not fixed label sets", "zwliJay/jev-forge ≠ NanoJev", "open replica economics / latency vs closed Jev", "NAR local drop-in", "wfzyx/von late-catch HIGH", "competing NAR claims / replica honesty", "typed judgments vs chat judges on guardrailing", "ishaannk/llm-vs-jev cross-note only", "deeper integrity fold is rh-guard", "nothing wins outright", "can be argued out of guarding"", "Jev IS the if-statement", "judgments/probabilities drive branches", "text model only writes prose", "interpreter owns variables/loops/budgets/replay", "otherwise maybe / confidence gate", "chaos samples after the gate", "southpolesteve/probably ≠ carldaws/hunch ≠ feelings ≠ Kungie/gut ≠ Illusion47586/judge ≠ tidymodels/probably", "133★ / forks 10 live", "build calibrated classifiers from human feedback", "retrieve by relevance not resemblance", "one calibrated yes/no per memory in one request", "pointer mode 17/18 19/20 *theirs*", "embedding resemblance misses the allergy", "samdotmak/jev-recall ≠ jev-search ≠ jev-sift ≠ carryforward ≠ chopratejas/invalidate", "memory leases ended by new evidence", "six Nouls then fixed rules in code", "0 of 157 false invalidations", "questions/plans/directives are not evidence", "unsure → review queue", "host keeps the store", "name↔body / comment truth / test-claims", "mizchi/jev-lint is mizchi/jevlint rename", "no shipped rule has severity error", "~1 in 5 findings wrong *theirs*", "mizchi/jev-lint ≠ huntedman/JevLint ≠ MichitoSugawara/jev-lint", "JSON Schema → typed JSON via Jev", "noul_threshold 0.5 decoder not a proof", "IncompatibleSchemaError lists every bad property", "on-device Laya CoreML ANE", "~5 ms P50 short decisions", "189/189 FP16 checkpoint parity", "10× not achieved", "mizorewww/laya-coreml ≠ gliner-native-runtime ≠ jevmlx ≠ NandhaKishorM/laya", "softmax over allowed tokens ≠ Noul", "question-first cache", "Micha0827/snapjudge ≠ githubnext/localjev ≠ jevmlx ≠ cendress/SnapJudge", "Jev-first Pi agent loop", "slow-LLM fallback", "explicit action menu / CandidateSource unimplemented", "62 tests wiring not quality", "direwolfiy/JevPi ≠ standardagents/jevpilot ≠ pi-jev-control", "resume-screening bias audit methodology", "name×resume factorial independent Nouls", "callback determined by resume quality", "mean-probability name gaps operationally negligible", "natemoo-re/bias-bench ≠ BBQ", "Plan/PRD panel → code-owned pass|review|block", "cheerleading out of scope", "austindixson/planalyzer ≠ single-goodness Noul", "cost-aware multi-model routing/escalation", "decide vs do", "successful-task cost", "cannacre8ive/switchboard-ai ≠ ha-switchboard ≠ hermes-switchyard", "frozen-protocol zero-shot bench", "TypeSafe Jev vs PrismNLI vs Laya", "contamination caveat", "elcronos/jev-vs-open-decision-models ≠ JevBench ≠ DMB", "context-window admission control", "VOI gate which tokens are worth the expensive model", "fail polarity per lens", "on small inputs lenses lose money", "cvsgireesh/jevusher ≠ jev-sift ≠ winnow", "typed decision control plane", "receipt ≠ authorization", "historical-v0 zero retained cases", "MokiMeow/jev-fabric ≠ jev-forge ≠ dgp", "live 15-dim typed rubric re-score per pause", "scoring economics exemplar", "OpenJev/Codiv ≠ TypeSafe hosted", "jose-troche/live-rubric ~$0.000004 desc / ~$0.000006 README", "adversarial pre-registered Jev eval", "28 predictions before data", "123,805 requests", "confidence does not track ignorance", "polite injection 65% / crude 0%", "willkelly/jev-evaluation ≠ jevals ≠ jev-baselines-eval", "provider-neutral Elixir/BEAM Noul/Choice/Score SDK", "class infrastructure", "nshkrdotcom/system_one_sdk ≠ typesafe_sdk ≠ dannote/jev", "question-linting of Jev questions themselves", "nine jaggedness rules, no API key, no labelled data", "static lint ≠ measured separation", "yodablocks/jevq ≠ tenbin ≠ JevLint ≠ commitjev", "open-weights Laya as class exemplar (binding)", "Nx/Bumblebee runtime", "host chooses backend", "ChristianAlexander/laya_ex ≠ system_one_sdk ≠ dannote/jev ≠ NandhaKishorM/laya", "on-chain/edge Laya deploy", "parity_verified stays false", "model output never grants Tx", "humandebri/IC-Laya ≠ laya_ex", "auditable weekend replica", "Jev outputs never used for training", "soft human-vote distributions", "unpaired 0.577 vs 0.727", "agilabs-ai/jev48 ≠ JevBench ≠ Mapika/decider", "adversarial dual-judge / framing attack surface", "comparative framing is the usable judgment", "prior injection crowds out evidence", "copyleftdev/ember ≠ ember.js", "Laya specialist fine-tune pipeline", "training still GPU-pending", "PIXELZX0/XERON ≠ convaiinnovations/laya", "Hub Laya replica drop", "daliborsb/laya ≠ convaiinnovations/laya ≠ NandhaKishorM/laya", "System One student distillation corpus", "gold is programmatic", "teacher is closed-API clone", "do not distill Jev as teacher of record", "MagaBitmex/jev-4b-distill-data ≠ missing student checkpoint", "non-LLM VIN System One", "planning depth not chat", "lewislululu/jevon ≠ douglance/jevon", "source-bound evidence checks", "local quote mismatch needs no API", "exit 0 ≠ claim truth", "WaynezProg/jev-kit ≠ jonathanavis96/jev-kit (Airlock) ≠ jev-use ≠ jev-mcp", "independent System One evidence catalog", "scores not one leaderboard", "no external record currently reproduced", "TokenTrim no-Jev matched hybrid 62.4%", "reachjalil/system-one-bench ≠ mallahyari/system-one-benchmark", "21 tasks · 134 items · 208 questions", "scenes from public GitHub contracts, not production logs", "SivletLabs/jev-eval ≠ willkelly/jev-evaluation ≠ 4esv/jev-eval ≠ xxkuboxx/jev-eval ≠ onlyoneaman/jev-eval ≠ dayhaysoos/jevals", "option isolation (sibling-blind)", "permutation-equivariant", "Hub OWNER not published", "nafisazizir/hev ≠ jaredpalmer/kev", "frozen local LLM logits, no trained decision head", "residual-head 9,222-param decreased 73/96→67/96", "confidence = 1−normalized entropy, not P(correct)", "yuki-oshio/mini-jev ≠ r-ms/mini-jev", "Jev classifier as autoregressive next-token predictor", "ChatJev-style soundness theater", "erik-dunteman/ChatJev ≠ dannote/jev ≠ jev-gpt", "calibrated decision head × AlphaProof value head", "implementation-layer isomorphism, semantic difference", "timeout = censoring", "do not launder Noul as proof", "parallel rank-prediction vs serial selection", "independent questions can conflict", "zzzzzec/jevsort ≠ keltokhy/jsort", "curated open System One ecosystem catalog", "rupeshpoojary9/awesome-open-system-one ≠ AnotiaWang/awesome-jev", "arXiv paper radar with Jev relevance scoring", "ranking ≠ calibration / 0.5 still soft", "fail-open failed evals not marked seen", "train calibrated ~27M from scratch", "typed Q→prob dist / one forward pass / no LLM decode", "hyusi2003/MiniSystemOne ≠ Colvin0315/MiniSystemOne", "description-only stub / size 5", "ESCI hard probe fails four of six", "jev_bool ECE 0.242 inversion 0.255", "do not re-fold §60 six-gates as new", "jobbyjev one-request-per-company from batch-size result", "find/design/evaluate TypeSafe Jev decision loops", "karanb192/jev-architect ≠ samtay32/jev-system-architect", "Jairik/jev-distiller size 1", "distill-Jev UI stub / do not distill Jev as teacher of record", "post-launch scored use-case map / Jev self-scores then human curation", "licensedsaucer9-web/jev-opportunities", "Jev-inize a use case into classifier/router", "gavinHuang/jevinize → simple-jev not TypeSafe", "featherless-ai/simple-jev", "compare saved decisions / same label can still change the branch", "VihaanAgarwal/jev-diff ≠ Saik0s/diffusiongemma-jev-macos", "not tested with a live Jev API key", "constrained logprob + temp/Platt ≠ Noul", "OpenJevPro pastes openjev-sglang JevBench as own", "zhangcy122/OpenJevPro ≠ IamBusy/OpenJev ≠ ekzhang/openjev-sglang", "PolyForm Noncommercial", "SmolLM-135M / sub-70ms / 0 output tokens", "demo P(True) 0.5052 / Choice conf 0.2872 / Score conf 0.0055", "README claims MIT / GitHub license null / no LICENSE file", "patelvishwa112/jev-system-one-rlcd ≠ arnabgho/rlcd-lite ≠ blackwood-rlcd", "source-backed Awesome Jev radar / 306+ commit-pinned", "logicrw/awesome-jev-projects ≠ AnotiaWang/awesome-jev ≠ yibie/awesome-jev ≠ cobanov/awesome-jev ≠ rupeshpoojary9/awesome-open-system-one", "auto GitHub sync / Issue-only submissions", "hashed n-gram encoder / rival-aware attention", "olanotolu/jevbetter vs jevlike starter", "synthetic hard menus top-1 0.916 vs 0.873 / ECE 0.0182 vs 0.0367 / 40 vs 4608 menus/sec", "shuffled-context control 0.335", "Turn any open LLM into System-One Jev", "uspraveen/Jevify ≠ Mintzs/jevify ≠ gulagala001/jevify", "Jevify-any-LLM architecture probe", "description-only stub / size 0", "Train encoder-only calibrated decision models from a task sentence", "Exu is a toolkit, not a method", "strictly proper scoring rule", "Pre-alpha", "Ruivalim/exu-base", "scratch-trained calibrated decision model", "typed Q → probability dists", "Colvin0315/MiniSystemOne ≠ hyusi2003/MiniSystemOne", "no published weights download URL", "90.5 seconds / 29.2% pipeline evidence", "p_i/p_j independent of other candidates", "Recipe for calibrated decision models — small model out", "init → synth → train → eval → serve", "91.1 % / ECE 0.022 *theirs*", "Jev zero-shot 75.1", "scienthoon/luce", "Put Jev's three headline claims on trial", "0.5B local GPU", "46x speedup / accuracy identical", "ECE 0.624 sentiment catastrophe", "bigger model worse calibration", "RichardoMrMu/jev-mini ≠ yuki-oshio/mini-jev ≠ r-ms/mini-jev", "System-1 decision engine for local LLMs", "structured choices only", "JSON parse of generated text ≠ Noul", "TypefAI JEV / Journal Entry Voucher", "tapsin/jev-local ≠ us/jev-local ≠ Argos1111/jev_local", "Jev 1.13 reward-model eval across 8 benchmark tracks", "40,940 examples / 0 API errors", "RewardBench v1 92.58%", "Precise IF 50.63%", "goya4140/jev-reward-model-evaluation", "Scaffolding in progress", "Jev vs LLM support-ticket routing", "static + live decision bench", "TypeSafe's own published benchmark", "illustrative simulations, not live API calls", "JevBench v1 — smart/cheap/fast/reliable", "I/C/S/K 25% geometric mean", "classifier.dev fast tier 84.8 is Jev behind its own API", "do not re-fold §78 v1.2 board as new", "Laya (421M) 70.1 now on board", "Zero-shot/few-shot LLM routing", "hard budget filter before Jev", "Jev never asked to perform budget arithmetic", "Jev judges the next state, XState enforces transitions", "simulation uses synthetic keyword fixtures", "catalog gravity", "v-modal/awesome-jev-tools", "★339 live REST", "curation is not endorsement", "crawler-maintained directory", "Daily GitHub + npm sweep, human-merged", "RadRebelSam/awesome-jev ≠ AnotiaWang ≠ yibie ≠ cobanov ≠ logicrw ≠ v-modal", "HF peft SPLADE/BGE reranker", "rdxtremity/jev-reranking ≠ carlaiau/jev-reranking", "query-side encoders, not a Jev replica", "ONNX System One Qwen3.5-4B scorer", "source:pngwn/system-one-qwen3.5-4b-scorer", "CC-BY-NC-4.0", "temperature 1.75", "transformers.js AutoModel cannot load this graph", "Consistency benchmark Space", "This Space contains no benchmark result yet", "12-case plumbing fixture", "Benchmark-driven Jev router and judge", "cheap alone is not success", "Jev does not write, sum prices, or claim accuracy %", "Sol 94.2 / Luna 83.9 / Jev path 89.7", "19.2% Sol / 62.3% cost save / 4.5pp miss of 2pp non-inferiority", "p50 latency worse than Sol due to routing overhead", "erendikmenn/jev-llm-router-benchmark ≠ jev-rag-benchmark ≠ ryantsai/jev-llm-router", "Express + node:sqlite", "mock and Jev decision engines", "previous_ticket_count >= 3 is code", "MIN_CONFIDENCE 0.6 still soft", "substring false positives", "aesaganda/jev-ticket-router ≠ SarathChandraBellam/jev-vs-llm-ticket-router", "Universal Figure & Diagram Router", "confidence ≥ 0.85 hard-gate is theater", "generative AI banned from scientific plots", "six visual branches", "hoangngochuong24947-gif/jev-figure-router", "human-labeled (state, question, label)", "166,054 rows / 22 configs", "soft_label for human uncertainty", "Praveenrajus/jev-bench ≠ fstandhartinger/jevbench", "ternary bonsai System One GGUF", "openjev's mechanism, Bonsai's weights", "Hub does not ship weights", "100/100 easy T/F is not Harbor", "label_mass ≠ correctness", "stock llama.cpp Q2_0 silently gibberish", "NicolaiMTLassen/open-bonzi-jev ≠ NicolaiLassen", "transformers.js DeBERTa ONNX", "source:com-kotobalabs/open-jev-deberta-v3-large", "temperature 1.05", "AutoModel from_pretrained works", "onnx-community/open-jev-deberta-v3-large-ONNX ≠ system-one-qwen3.5-4b-scorer-ONNX", "107★ densify", "GH 151M vs README 149.6M", "PR #1 now closed unmerged", "do not re-fold §71 claim-audit as a beat", "typed decisions, RLCD, confidence-gated routing", "structured ≠ correct", "mock not live API", "26 tests", "wjdjdakf17/jev-study ≠ baekenough/jev-study", "bonzi-27b-v2 / ternary-8b / 27b-v1 GGUF family densify", "WANLI-256 74.6% / 65.2% / 71.1% *theirs*", "Bonsai 1 27B Q1_0 runs on stock llama.cpp", "ternary still needs PrismML fork", "hf:heman10x/openJev-verdict-2.0 twin tokenizer-only", "OpenJev Vision image classification + uncertainty", "CLEVR-4 held-out joint 0%", "hfdataset:IamBusy/OpenJev-Vision-Research-v0.1 12,832", "294,912 derived targets not independent samples", "Laya multilingual ONNX WebGPU typed-decisions port", "63/63 selected answers / 5.1e-4 CPU / 1.2e-2 WebGPU", "UpHash-Network/mini-jev is yuki-oshio transfer", "jev-injection-bench 11,900 labelled prompts", "Jev best ranking / Haiku better ECE 0.021 vs 0.058", "0.5–0.9 band is where Jev's numbers do not mean what they say", "Prompt wording moves panic 28%", "manojlds/jev-dspy-bench ≠ dspachos/jev-dspy ≠ jmanhype/jev-dspy-lab", "Jev agreement is similarity, never ground truth", "no aggregate quality grade or merge gate", "AbstentionBench-on-Jev rank 1 of 20 vs 2025 field", "question-asymmetry", "forward-looking 0.465 never extreme", "openkev calibration layer not a runtime", "ECE vs coverage independent", "select_threshold returns inf", "escalation catches uncertainty not ignorance", "misakaikato/openkev ≠ jaredpalmer/kev", "pdf-race Docling→Jev vs Gemini", "parser owns the wall clock", "12/12 tie is a tie", "titles selected not generated", "flopcheck 16 calibrated tweet judgments", "mechanical tells in code", "ZeroX-01/jev-atlas ≠ Zaious/jev-capability-atlas ≠ gorock007/jev-atlas", "Laya calibration lab Gradio MCP", "T never changes argmax", "confidence ≠ top-label p", "easy probe set refused", "40–48 rows too small to ship T", "Gemma-4 26B-A4B jevify classification+calibration", "LoRA adapter twin not independent eval", "Gemma-4 E4B jevify", "E4B LoRA stub card", "kushalpatil/jevify-gemma4 ≠ Mintzs/jevify ≠ gulagala001/jevify ≠ uspraveen/Jevify", "GH kushalpatil07/jevify 404", "PAWS 0.580/ece 0.288 is the weak cell", "smaller E4B slightly better OOD ECE than 26B-A4B", "Hub jevify merged LoRA ships weights", "bonzi Bonsai-8B v1 GGUF densify", "Bonsai-1.7B v1", "Bonsai-4B v1", "WANLI-256 64.5% / 60.2% / 52.0% *theirs*", "rank #4 / #5 / #6 of 6", "JulesHuisman/jev-eval scaffolding / README SHA c356a584 (was empty e69de29b)", "JulesHuisman/jev-eval ≠ SivletLabs/jev-eval ≠ willkelly/jev-evaluation ≠ 4esv ≠ xxkuboxx ≠ onlyoneaman ≠ dayhaysoos/jevals", "7 bands 6/10 vs 40 bands 0/10", "source receipts + confidence slider re-policy without re-inference", "32/32 synthetic is smoke not production", "classify HF datasets across typed semantic dimensions", "roadus2 watch misspelling; lock roadius2/ultra_laya", "ultra_laya REVIEW defects", "default branch claude/laya-jev-review-gg5ppo", "XNLI EN 88.3% ECE 0.032 → RU 77.3% ECE 0.096", "Δ −11.0 pp [−14.2,−7.8]; ECE +0.063", "MASSIVE no detectable difference at n=600", "confidence is function of p_max (r=1.000)", "pointer-not-generator 400 human-authored responses", "proposed ≠ authorized", "FewRel 160: Jev 85.0% vs lexical 13.125%", "gated 100% (95/95) coverage 59.375%", "J++ composable semantic computation language", "judge-jev 0.5 still soft", "947 repos scored; A 273 / B 302 / C 372", "LLM rubric ≠ benches", "No benchmark winner is claimed", "phishing: naive 62.6% vs regex 91.8%; 5-atomic + LR 95.0% *theirs*", "AITuber tension ±15", "README npm global; repo is Rust", "git-confess code owns counting/blame/ratio", "httpx exhibit 11% (13/119) *theirs*", "90d trend +12.40% vs random +12.75% vs BH +41.71%", "5m win rate 25%", "Awesomejev 656 entries / 38,160 stars", "tracker likes 64 (+4) lastModified UNCHANGED", "Laya present; Blackwood ABSENT; Archer still promised_not_landed", "Blackwood tracker ABSENT; likes 2 gated manual", "r = c - p_a", "ECE 0.021; acc 0.807 vs warmup 0.746", "Independent primitive", "11.57s vs 54.10s · 4.67× · 120/128 *theirs*", "default path is pretrained Gemma probs not trained RLCD head", "GH Meanblock 404; lock leesk212/JEV-CPU", "softmax over letter slots ≠ Noul", "WANLI 0.741 vs openjev v2 0.77 *theirs*", "3-way NLI ≠ Noul", "priority 0.464 = majority floor", "banking77 contaminated", "raw margins not probabilities", "do not distill Jev as teacher of record (they distilled Haiku)", "“0.9 is not one number”", "ranking ≠ calibration", "banking77 0.8–0.9 stated 0.86 actual 0.73 over-confident *theirs*", "≠ Praveenrajus/jev-bench ≠ fstandhartinger/jevbench", "$0.0000153–$0.0000226 vs circulating $0.0004 (~20×)", "Score is 0..n-1 expectation not 0–1", "Noul has no confidence field", "TCP floor 198.8 ms", "type reliability is not a reason to choose Jev (json_schema 5/5)", "gateway tax not one number", "Function-only 5/8 vs hybrid 8/8", "4/8 without Jev", "8 designed cases not conversion lift", "200-row pilot Jev 86.5% 173/200 vs Gemini Flash-Lite 86.0% 172/200 vs Pro 87.0% 174/200 *theirs*", "not a ranking", "情緒測謊器", "8-example Jev vs GPT-5.6 Sol ~64× cost 5.4× latency *theirs*", "synthetic; no inference", "≠ JevBench v1.2 §78", "Judged 3317 / listed 2560", "Jev judges, code applies policy", "APA “microsecond policy / zero hallucination” overclaim", "Client-side quiz; pointer from held docs; scanned-PDF warn", "Jev judges / agent reasons / user decides", "selecting an option is not permission to implement", "pattern exact, judgement must clear floor", "no matching pattern → no model call", "not a correctness oracle", "Spec vs artifact remainder", "treating 0.85 as 85% / minProbability hard-gate as Harbor", "VERIFY acquires discriminating evidence, never same-pool confidence-only rescoring", "fast/full/max are ceilings not sizes", "Solar writes, Jev chooses NEXT ACTION", "do not reopen or amend PR #23 or #24 or #25 or #26 or #27", , "Calibration is not alpha", "NO CURRENT ALPHA CANDIDATE", "ΔR² approximately +0.00084", "Brier 0.2131387", "ECE 0.0421875", "Adding Jev probability to deterministic volatility improved Brier by only 1.4058e-05", "default 0.5 keeps zero non pinned", "keepResult median 0.14 to 0.17", "keepCall median 0.28 to 0.35", "usable range is about 0.10 to 0.25", "7.8% to 57.9%", "judges results it never sees", "task-finish eval not built yet", "$0.002 per compaction", "slavadubrov/sgr-judge-bench ≠ slavadubrov/jev-judge-bench", "Jev 108/120 $0.083 0.34 s", "Luna SGR 114/120", "paired Jev accuracy-difference intervals include zero", "not evidence of equivalence", "GLM SGR 26/120 93 format failures", "Terra-planned Jev hybrid 55/120", "rule-based by default, optionally Jev-backed", "empty README", "missing key cannot break the experience", "prefill plus exactly one decode", "softmax over A/B/C ≠ Noul", "BBQ 9,053/10,000 (90.53%)", "ECE 0.0890", "Mean confidence 0.9943", "overconfident", "score and noul not implemented", "DGUI 12 rows (was 6)", "INSTRUCT 119 rows likes 2", "encode the state once, decide everything in parallel", "0.740 accuracy against a 0.508 majority", "ECE 0.047", "fine-tune's advantage ends where its 384-token training data does", "jasonkneen/open-jev ≠ pngwn/open-jev", "same sha d41dc3cd", "Space does not call Jev", "recomputes routing from saved probabilities", "200-case Jev 97.0% / 100.0% / 95.0% / MAE 9.22", "synthetic repository benchmark", "Jev evaluations are advisory", "YehuiTang0316/jev-nlgrep ≠ Bentlybro/jevgrep ≠ can1357/jegrep ≠ uehaj/jev-semgrep", "default threshold 0.8 still soft", "40-line windows cannot prove whole function", "token-native sequential start/end Choice", "Gemini/Haiku stubs not configured yet", "handful of hand-written examples, not a benchmark", "Jev judged exactly what it was given", "laguagu/jev-skills ≠ laguagu/jev-evidence-lab ≠ Pleo2/awesome-jev-agent-skills", "contract_passed is not a claim of guaranteed factual truth", "Wilson lower bound 0.85 floor", "fixture mode no savings claim", "SemIf 2207★ (+21 vs §110 2186)", "jevlike 1043★ (+5 vs 1038)", "TypeAR 15★ (+1 vs 14)", "AnotiaWang 97★ (+1 vs 96)", "yibie/awesome-jev 506★ (+16 vs 490)", "Laya likes 822 (was 802)", "tracker likes 64 flat, lastModified UNCHANGED", "do not reopen or amend PR #23/#24/#25/#26/#27/#28", "Heman10x-NGU/Verdict-open-jev ≠ Heman10x-NGU/openJev-verdict-2.0", "TF-IDF + LogReg ECE 0.0207 vs Jev 0.1440", "Verdict-open-jev 48.07% vs Jev 90.80%", "abstention combined recall 10.00%", "p50 35.58 ms", "K=25 (maximum capacity) 72.00%", "0.85 coverage 84.60% selective risk 1.18%", "26.1× faster than standard Qwen JSON generation", "Jevify 90.0% / 167 ms CUDA graphs disabled", "Finding 1: Brier on stated confidence alone is a trap", "grpo_rlcr 0.78 / ECE 0.084", "reliability 0.007 but resolution 0.000", "27 900 schema-driven decisions", "13 600 / 13 600 questions", "candidate mass min 0.99999624", "22 configs · 166,054 rows · 4 calibration-gold", "sha a39eba3f", "Student B MAE 0.148 / Pearson 0.836 / 86.0%", "pngwn/open-jev-laya-bench README 404", "sha 9f69c742 likes 2", "HDFS 0.9933 (745/750) / retain 0.0084", "BGL ERROR/FATAL protection 1.0000", "2,479 / 2,500 HDFS uncertain", "cache hit 0.9648 (2412/2500)", "$0.153936 estimated", "E2 recomputes from saved probabilities", "Space sha eda59e0a", "MASSIVE English 0.783 / Khmer 0.033 / Hindi 0.133", "40–48 rows too small to ship T", "T never changes argmax", "siren2345/jev-single-decode-transformers ≠ siren2345/jev-single-decode", "Split Transformers experiment from llama.cpp runtime", "tanayvasishtha/jev-lab ≠ dairui1/jev-lab ≠ BrendanH18/jev-lab ≠ yibie/laya-jev-lab", "Four experiments stress-testing TypeSafe's Jev: calibration, bundle bias, label bias, and ensembling", "second pass must be $0.00 from cache", "The pages never call Jev", "Gemma 4 31B 77.0% / Jev 1.13.0 61.4% / Laya 322M 0.0%", "restriction state 95.0% against 84.4%", "None of the systems are particularly good at knowing when to stop and ask", "They skip the question and call a tool directly", "100% schema pass", "six-field joint 48.8% vs 72.8%", "ywchiu/jev_benchmark ≠ Running-Dolphins/jev-bench ≠ Praveenrajus/jev-bench", "ACT / REVIEW / FALLBACK", "A provider failure, timeout, malformed output, or missing answer is **not** a policy outcome", "confidence is descriptive provider output, not a substitute for probability", "Quality denominators include only valid scored answers", "an exact halfway tie chooses the lower level", "aiwithenoch/Jev-Skill ≠ simplosophy/jev-skill ≠ laguagu/jev-skills", "The local path does not claim to turn a smaller checkpoint into Jev", "Low support becomes decision: \"review\"", "MIT-0 SPDX NOASSERTION", "current-llm", "结构兼容,不是 Jev 模型能力", "altryne/jevify ≠ Mintzs/jevify ≠ gulagala001/jevify ≠ uspraveen/Jevify", "Find where Jev belongs. Design the questions. Measure the difference", "TypeAR-AI/TypeAR 301 → TypeLLM/TypeLLM", "TypeLLM/TypeLLM 16★", "SemIf 2241★ (+34 vs §111 2207)", "jevlike 1051★ (+8 vs 1043)", "AnotiaWang 98★ (+1 vs 97)", "yibie/awesome-jev 525★ (+19 vs 506)", "Laya likes 864 (was 822)", "tracker likes 67 (+3 vs 64)", "lastModified UNCHANGED `2026-09-20T04:29:16.000Z`", "do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#32", "hysteresis enter/exit / replay policy without inference", "calibration does not compose / hop-ECE permutation-invariant", "equal-width vs quantile ECE / ranking ≠ calibration", "Qwen2.5 ≠ Archer / Qwen 3.8 sparring ≠ Archer / Qwen/Qwen3.8-27B ≠ Archer", "Deferred Crispification / TCE / AMS", "g0runmezadam/what-is-jev IS tunahansahin897/what-is-jev", "pd.cut equal-width vs jeval quantile", "A hunch is a probability with a policy attached", "soundness theater / measurement theater / hourly 0843", , "Jev Capability Resolver / NiazMorshed2007/jcr", "one tool nested capability tree / returns context / does not execute", "skills vs capabilities / workflow+judgment vs operations", "format independent of Jev / proposed open standard", "JCR_BAND_RATIO 0.6 is application policy / soft scores ≠ hard gates", "routing ≠ permission / docs ≠ authority to run", "sol-vs-opus5-20 lookup+explain / n=1 / Not Harbor task-execution", "wall-time mixed / Sol slower with JCR in 19/20", "NiazMorshed2007/jcr ≠ skill-broker ≠ skillranker ≠ jev-sift ≠ jev-lens ≠ jevusher ≠ jev_select_capability", "do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33/#34", "notes.md §116", or "cascade sign-flip / calibration theater": read `references/faq.md`, then `references/mental-models.md`, then `references/mixed-architecture.md`, then `references/judgment-class.md` before any mapping. Proof, @@ -115,16 +115,16 @@ classical method you already trust, substitute it, classify the win | Familiar method | Judgment shape | Detail | |---|---|---| -| Mental models across domains (not SWE-only) | EU, abstention, VOI, MCDA, SDT, search/control, Leveson, NATM/Norman/snap-fit; **extractable-from-state boundary map** (self-contained vs needs outside knowledge). **Apply queued 1047 follow-ons** (`notes.md` §88): **Replica honesty** / **Empty ≠ approve** / **Beam as control** / **Cache is the exact envelope**. Apply 1144 (`notes.md` §89): **Typed if** / **Shadow then honor** / **Human every action** / **Atom then sense** / **File by Choice** / **Question preflight** / **Inbox read-only vs write**. Apply 1241 (`notes.md` §90): **Observe→score→act (namesake lock)** / **Decision-as-ranking** / **Native vs schema-guided Harbor** / **0 promotions / authored vs real** / **Offload + classifier-not-generator** / **Collapse late** / **Unofficial toolbelt** / **Pointer shell** / **Preview-first VOI / rubric rewrite** / **Life fail-open covers** / **S1 decide / S2 plan** / **Seed/expand/judge/verify + local daemon ≠ Jev**. Apply 1347 (`notes.md` §91): **Judge harness as control API** / **Batch packing VOI** / **Calibration as product** / **Decision-as-Plugin** / **Policy-constrained skill select** / **Tiny local econ pruner** / **Observe→score→act cousins** / **Deterministic verify ≠ System One** / **Soft-score vs hard-argmax**. Apply 1441 (`notes.md` §92): **Self-hosted econ** / **Soft-judgment gate integrity** / **Retrieval as calibrated decision space** / **Enterprise reflexes** / **Screenshot-free CU** / **Hybrid S1/S2** / **Harbor-jevals / SRE** / **Laya densifies** / **Demos / unofficial toolbelt**. Apply SIGNAL jevcache/jev-align (`notes.md` §93): **Decision ledger / memoization** / **GEPA alignment loop** (fingerprint after redact; recall vs decide; publish fingerprints+answers; CI replay as Harbor cousin; Cache hit ≠ correctness; hyperspaceai/jevcache ≠ kushals256/jevcache; human labels only; score never auto-accepts; production capture flywheel; sutro-sh/jev-align ≠ caiovicentino/jev-align). Apply SIGNAL enzyme/JA ModernBERT/Gemma (`notes.md` §94): **Compile-time System One / questions-as-index** / **Unofficial JA ModernBERT cross-encoder** / **NAR class legitimacy / multimodal observe→decide / Router-OOD** (guidance ≠ hook; catalysts ≠ summaries; compile-time System One; unofficial ≠ TypeSafe; format_version modernbert-jev/1; Argos1111/jev_local ≠ us/jev-local ≠ kunchenguid/local-jev; LFM default ≠ ModernBERT backend; Nemotron ≠ TypeSafe Jev; not a calibrated replacement; djev-dev complements djev-spark; images as Choice options; Laya essay numbers *theirs*; Router/OOD confidence; hosted bootstrap ≠ silent TypeSafe). Apply 1541 (`notes.md` §95): **Decision-as-plugin** (difficulty + policy thresholds + JSONL trace; jev-codex-pilot model + reasoning depth; keep/shadow/hybrid/reject) / **Evidence projection** (quarry evidence projection) / **Soft judgment integrity** (jevguard calibrator/cache/escape; jev-ci-selector CI shadow mode) / **Physical/control** (Frank-ZY-Dou/awesome-jev robotics/3D/control) / **Harbor-jevals injection-firewall** (one-dollar-tahoe TypeSafe Jev defense eval) / **llama.cpp replica** (llama-jev llama.cpp replica; petercr/jev-orchestrator ≠ FleeexCorp/jev-orchestrator; seb4ez/jevguard ≠ AseemPrasad/JevGuard ≠ pablozr/JevGuard; webNeat/llama-jev ≠ WiktorB2004/llama-index-jev) Apply 1639 (`notes.md` §96): **OpenCode stdout-prune host port** (OpenCode jev-pruner context sieve; observe→score-candidates→prune; jev-zen / jev-1.13-free; zen-chat ≠ Noul; fail-open original; keepScore >0.1 floor; host port of tamaratran/jev-pruner; indiejoseph/opencode-jev-pruner ≠ nrdz-labs/fast-jev-opencode; jev-webagent-bench empty stub; Kiln-AI/jev_jsonschema noul_threshold 0.5; NSStudent/JevSwiftSDK unofficial). Apply SIGNAL gliner-native-runtime (`notes.md` §97): **GLiNER2 native Apple path** (unofficial Swift/Core ML GLiNER 2.5-small; entity spans + confidence; not Choice/Score/Noul; not TypeSafe; label descriptions as schema; on-device ANE economics; honesty locks; shershah1024/gliner-native-runtime ≠ Fastino; ≠ gliner25-compaction ≠ gliner2-ultrafast ≠ Eran-BA/Jev_from_GLiNER2 ≠ NSStudent/JevSwiftSDK ≠ jevmlx; default threshold 0.1 still soft). Apply 1740 (`notes.md` §98): **Decision Graph Protocol envelope** (Decision Graph Protocol frame→assess→commit; app retains permissions/effects; Jev-first assessor-neutral; guarded commit / receipt/next frame; assessment batching; hard-gating DGP as safety theater; numerous-com/dgp ≠ TypeSafe official) / **Calibrated meaning-grep live tree** (jegrep calibrated path+range Nouls; no embeddings/index/daemon; ~$0.01–0.03 typical; agent --json; can1357/jegrep ≠ Bentlybro/jevgrep ≠ uehaj/jev-semgrep) / **Archer-arch fidelity + measured calibration gap** (Archer-arch fidelity; kev family OOD 0.76–0.77 vs Jev 0.86; block-causal isolation; pointer/readout CE-trained; /v1/systemone drop-in; replica honesty). Apply 1843 (`notes.md` §99): **Cost-derived YES/NO/UNSURE control flow** (cost-sensitive decision theory × System One probabilities → control flow; thresholds derived from costs not hard-coded; YES / NO / UNSURE from cost_false_yes / cost_false_no / cost_human; auto-batching same-object questions; Kungie/gut ≠ tpellet/hunch ≠ carldaws/hunch) / **Typed-callback control flow** (judgment vs generation; deterministic execution after probabilistic judgment; exactly one app-owned callback; explicit uncertain branch; Illusion47586/judge ≠ lexingtonhibiki/judgekit ≠ Ascurse/typed-judge-kit) / **Variable-N option scoring as the trainable object** (dynamic candidate bags not fixed label sets; zwliJay/jev-forge ≠ NanoJev; not a new class-table species) / **Open NAR replica economics** (NAR local drop-in; open replica economics / latency vs closed Jev; wfzyx/von late-catch HIGH; competing NAR claims / replica honesty) / **Typed vs chat judges on guardrailing** (typed judgments vs chat judges on guardrailing; ishaannk/llm-vs-jev cross-note only; deeper integrity fold is rh-guard; nothing wins outright; can be argued out of guarding). Apply 1943 (`notes.md` §100): **Jev IS the if-statement** (judgments/probabilities drive branches; text model only writes prose; interpreter owns variables/loops/budgets/replay; otherwise maybe / confidence gate; chaos samples after the gate; southpolesteve/probably ≠ carldaws/hunch ≠ feelings ≠ Kungie/gut ≠ Illusion47586/judge ≠ tidymodels/probably) / **GEPA live delta** (133★ / forks 10 live; build calibrated classifiers from human feedback; HEAD/README SHA unchanged vs §93) / **Memory retrieve vs lease** (retrieve by relevance not resemblance; one calibrated yes/no per memory in one request; pointer mode 17/18 19/20 *theirs*; embedding resemblance misses the allergy; samdotmak/jev-recall ≠ jev-search ≠ jev-sift ≠ carryforward ≠ chopratejas/invalidate; memory leases ended by new evidence; six Nouls then fixed rules in code; 0 of 157 false invalidations; questions/plans/directives are not evidence; unsure → review queue; host keeps the store) / **Contract-of-artifact lint rename** (name↔body / comment truth / test-claims; mizchi/jev-lint is mizchi/jevlint rename; no shipped rule has severity error; ~1 in 5 findings wrong *theirs*; mizchi/jev-lint ≠ huntedman/JevLint ≠ MichitoSugawara/jev-lint) / **JSON Schema question compiler** (JSON Schema → typed JSON via Jev; noul_threshold 0.5 decoder not a proof; IncompatibleSchemaError lists every bad property) / **Local System One economics** (on-device Laya CoreML ANE; ~5 ms P50 short decisions; 189/189 FP16 checkpoint parity; 10× not achieved; mizorewww/laya-coreml ≠ gliner-native-runtime ≠ jevmlx ≠ NandhaKishorM/laya; softmax over allowed tokens ≠ Noul; question-first cache; Micha0827/snapjudge ≠ githubnext/localjev ≠ jevmlx ≠ cendress/SnapJudge; Jev-first Pi agent loop; slow-LLM fallback; explicit action menu / CandidateSource unimplemented; 62 tests wiring not quality; direwolfiy/JevPi ≠ standardagents/jevpilot ≠ pi-jev-control). Apply 2041 (`notes.md` §101): **Resume-screening bias audit** (resume-screening bias audit methodology; name×resume factorial independent Nouls; callback determined by resume quality; mean-probability name gaps operationally negligible; natemoo-re/bias-bench ≠ BBQ) / **MCDA panel code-owned verdict** (Plan/PRD panel → code-owned pass|review|block; cheerleading out of scope; austindixson/planalyzer ≠ single-goodness Noul) / **EU cost-aware routing** (cost-aware multi-model routing/escalation; decide vs do; successful-task cost; cannacre8ive/switchboard-ai ≠ ha-switchboard ≠ hermes-switchyard) / **Frozen-protocol bake-off** (frozen-protocol zero-shot bench; TypeSafe Jev vs PrismNLI vs Laya; contamination caveat; elcronos/jev-vs-open-decision-models ≠ JevBench ≠ DMB) / **VOI admission** (context-window admission control; VOI gate which tokens are worth the expensive model; fail polarity per lens; on small inputs lenses lose money; cvsgireesh/jevusher ≠ jev-sift ≠ winnow) / **Leveson control plane** (typed decision control plane; receipt ≠ authorization; historical-v0 zero retained cases; MokiMeow/jev-fabric ≠ jev-forge ≠ dgp) / **Scoring economics** (live 15-dim typed rubric re-score per pause; scoring economics exemplar; OpenJev/Codiv ≠ TypeSafe hosted; jose-troche/live-rubric ~$0.000004 desc / ~$0.000006 README) / **Pre-registered calibration science** (adversarial pre-registered Jev eval; 28 predictions before data; 123,805 requests; confidence does not track ignorance; polite injection 65% / crude 0%; willkelly/jev-evaluation ≠ jevals ≠ jev-baselines-eval) / **Class infrastructure SDK** (provider-neutral Elixir/BEAM Noul/Choice/Score SDK; class infrastructure; nshkrdotcom/system_one_sdk ≠ typesafe_sdk ≠ dannote/jev). Apply 2145 (`notes.md` §102): **Question-linting of Jev questions themselves** (question-linting of Jev questions themselves; nine jaggedness rules, no API key, no labelled data; static lint ≠ measured separation; yodablocks/jevq ≠ tenbin ≠ JevLint ≠ commitjev) / **Open-weights Laya as class exemplar (binding)** (Nx/Bumblebee runtime; host chooses backend; ChristianAlexander/laya_ex ≠ system_one_sdk ≠ dannote/jev ≠ NandhaKishorM/laya) / **On-chain/edge Laya deploy** (parity_verified stays false; model output never grants Tx; humandebri/IC-Laya ≠ laya_ex) / **Auditable weekend replica** (Jev outputs never used for training; unpaired 0.577 vs 0.727; agilabs-ai/jev48 ≠ JevBench ≠ Mapika/decider) / **Adversarial dual-judge / framing** (comparative framing is the usable judgment; prior injection crowds out evidence; copyleftdev/ember ≠ ember.js) / **Laya specialist + Hub replica** (training still GPU-pending; PIXELZX0/XERON ≠ convaiinnovations/laya; daliborsb/laya ≠ convaiinnovations/laya ≠ NandhaKishorM/laya) / **Distillation economics** (gold is programmatic; teacher is closed-API clone; do not distill Jev as teacher of record; MagaBitmex/jev-4b-distill-data ≠ missing student checkpoint) / **Non-LLM VIN System One** (planning depth not chat; lewislululu/jevon ≠ douglance/jevon) / **Source-bound evidence** (local quote mismatch needs no API; exit 0 ≠ claim truth; WaynezProg/jev-kit ≠ jonathanavis96/jev-kit (Airlock) ≠ jev-use ≠ jev-mcp) / Apply 2246 (`notes.md` §103): **Independent System One evidence catalog** (independent System One evidence catalog; 19 reviewed records; scores not one leaderboard; no external record currently reproduced; TokenTrim no-Jev matched hybrid 62.4%; reachjalil/system-one-bench ≠ mallahyari/system-one-benchmark) / **Typed eval freeze** (21 tasks · 134 items · 208 questions; scenes from public GitHub contracts, not production logs; SivletLabs/jev-eval ≠ willkelly/jev-evaluation ≠ 4esv/jev-eval ≠ xxkuboxx/jev-eval ≠ onlyoneaman/jev-eval ≠ dayhaysoos/jevals) / **Option-isolated tiny replica** (option isolation (sibling-blind); permutation-equivariant; Hub OWNER not published; nafisazizir/hev ≠ jaredpalmer/kev) / **Frozen-LLM typed decisions** (frozen local LLM logits, no trained decision head; residual-head 9,222-param decreased 73/96→67/96; confidence = 1−normalized entropy, not P(correct); yuki-oshio/mini-jev ≠ r-ms/mini-jev) / **AR next-token anti-pattern** (Jev classifier as autoregressive next-token predictor; ChatJev-style soundness theater; erik-dunteman/ChatJev ≠ dannote/jev ≠ jev-gpt) / **Formal compose with scoring** (calibrated decision head × AlphaProof value head; implementation-layer isomorphism, semantic difference; timeout = censoring; do not launder Noul as proof) / **Parallel rank vs serial selection** (parallel rank-prediction vs serial selection; independent questions can conflict; zzzzzec/jevsort ≠ keltokhy/jsort) / **Open-side ecosystem catalog** (curated open System One ecosystem catalog; rupeshpoojary9/awesome-open-system-one ≠ AnotiaWang/awesome-jev) / **Knowledge-work paper radar** (arXiv paper radar with Jev relevance scoring; ranking ≠ calibration / 0.5 still soft; fail-open failed evals not marked seen) / Apply 2340 (`notes.md` §104): **From-scratch calibrated decision model** (train calibrated ~27M from scratch; typed Q→prob dist / one forward pass / no LLM decode; hyusi2003/MiniSystemOne ≠ Colvin0315/MiniSystemOne; description-only stub / size 5) / **ORDER BY ranking upgrade** (ESCI hard probe fails four of six; jev_bool ECE 0.242 inversion 0.255; do not re-fold §60 six-gates as new; jobbyjev one-request-per-company from batch-size result) / **Find/design/evaluate decision loops** (find/design/evaluate TypeSafe Jev decision loops; karanb192/jev-architect ≠ samtay32/jev-system-architect) / **Distill-Jev UI stub** (Jairik/jev-distiller size 1; distill-Jev UI stub / do not distill Jev as teacher of record) / **Post-launch scored opportunity map** (post-launch scored use-case map / Jev self-scores then human curation; licensedsaucer9-web/jev-opportunities) / **Jev-inize a use case** (Jev-inize a use case into classifier/router; gavinHuang/jevinize → simple-jev not TypeSafe; featherless-ai/simple-jev) / **Saved-decision regression** (compare saved decisions / same label can still change the branch; VihaanAgarwal/jev-diff ≠ Saik0s/diffusiongemma-jev-macos; not tested with a live Jev API key) / **Constrained-logprob API** (constrained logprob + temp/Platt ≠ Noul; OpenJevPro pastes openjev-sglang JevBench as own; zhangcy122/OpenJevPro ≠ IamBusy/OpenJev ≠ ekzhang/openjev-sglang; PolyForm Noncommercial) / **SmolLM RLCD reproduction** (SmolLM-135M / sub-70ms / 0 output tokens; demo P(True) 0.5052 / Choice conf 0.2872 / Score conf 0.0055; README claims MIT / GitHub license null / no LICENSE file; patelvishwa112/jev-system-one-rlcd ≠ arnabgho/rlcd-lite ≠ blackwood-rlcd) / **Source-backed Awesome radar** (source-backed Awesome Jev radar / 306+ commit-pinned; logicrw/awesome-jev-projects ≠ AnotiaWang/awesome-jev ≠ yibie/awesome-jev ≠ cobanov/awesome-jev ≠ rupeshpoojary9/awesome-open-system-one; auto GitHub sync / Issue-only submissions) / **Rival-aware one-pass scorer** (hashed n-gram encoder / rival-aware attention; olanotolu/jevbetter vs jevlike starter; synthetic hard menus top-1 0.916 vs 0.873 / ECE 0.0182 vs 0.0367 / 40 vs 4608 menus/sec; shuffled-context control 0.335) | `references/mental-models.md` Apply 0042 (`notes.md` §105): structured probability readouts; distribution > argmax; Noul 0.5 midpoint; score is expectation not integer; bare HTTP not SDK; Arohtea/jev-readout) / **Ordinary-model Jev-shape** (Jev-style Choice/Score/Noul from ordinary models; optional DSH plugin; schema-valid ≠ calibrated; gulagala001/jevify ≠ Mintzs/jevify) / **Open-weight Laya measurement** (Laya RLCD benchmark; 40.3% below constant-answer; open-weight measurement; mourad-ghafiri/laya-rlcd-benchmark ≠ yibie/laya-jev-lab) / **Cheap fail-open semantic edge** (cheap fail-open semantic edge; second signal not sole; FastLoopError catch; SupremeDreamZ/jev-fastloop ≠ jev-ultrafast) / **Fan-out measurement** (asking more questions in one call; 0.980 at every N; nearly not fully deterministic; TheWebDevel/jev-fanout) / **VLM+Jev RL teacher** (Qwen3-VL perception + Jev decisions train RL; 0 model calls at deployment; VLM alone 1.7 vs +Jev 4.4; harneet2512/reflexrl ≠ khordoo/jev-reflex-autonomy-lab) / **Independent Jev API vs Laya** (independent Jev API vs Laya; cascade 0.60 matches 78% at 1.8×; noul facts not judgements; yibie/laya-jev-lab ≠ dairui1/jev-lab ≠ BrendanH18/jev-lab) / **Locate vs decide** (GLiNER vs GLiFormer vs Laya vs Jev; extractors ≠ decision engines; Laya dict-instructions collapse 58.3%; umstek/zero-shot-ie-bench) / **Throughput arena** (decisions-per-minute & cost; 204 moves vs 73; throughput not intelligence; angelgalvisc/snake-arena-jev-vs-llms ≠ vtrivedy/jev-plays-games) / **Behavioral contracts** (behavioral contracts; pin expectations eval upgrades; raw 0.94 is not a release; sathariels/jevcheck ≠ dayhaysoos/jevals ≠ SivletLabs/jev-eval) / **Evidence-linked upgrade review** (evidence-linked dependency upgrade; Jev never generates filenames; no_direct_evidence ≠ safe to merge; GaneshVG18/upgrade-radar ≠ LYchoon/paper-radar-jev) / **Knowledge-work discography** (discography theme/mood/complexity; five atomic questions one call; lirantal/discoprint) / Apply 0145 (`notes.md` §106): **Architecture probes PRIMARY** (Turn any open LLM into System-One Jev; uspraveen/Jevify ≠ Mintzs/jevify ≠ gulagala001/jevify; Jevify-any-LLM architecture probe; description-only stub / size 0; Train encoder-only calibrated decision models from a task sentence; Exu is a toolkit, not a method; strictly proper scoring rule; Pre-alpha; Ruivalim/exu-base; scratch-trained calibrated decision model; typed Q → probability dists; Colvin0315/MiniSystemOne ≠ hyusi2003/MiniSystemOne; no published weights download URL; 90.5 seconds / 29.2% pipeline evidence; p_i/p_j independent of other candidates; Recipe for calibrated decision models — small model out; init → synth → train → eval → serve; 91.1 % / ECE 0.022 *theirs*; Jev zero-shot 75.1; scienthoon/luce; Put Jev's three headline claims on trial; 0.5B local GPU; 46x speedup / accuracy identical; ECE 0.624 sentiment catastrophe; bigger model worse calibration; RichardoMrMu/jev-mini ≠ yuki-oshio/mini-jev ≠ r-ms/mini-jev; System-1 decision engine for local LLMs; structured choices only; JSON parse of generated text ≠ Noul; TypefAI JEV / Journal Entry Voucher; tapsin/jev-local ≠ us/jev-local ≠ Argos1111/jev_local) / **Measurement densifies** (Jev 1.13 reward-model eval across 8 benchmark tracks; 40,940 examples / 0 API errors; RewardBench v1 92.58%; Precise IF 50.63%; goya4140/jev-reward-model-evaluation; Scaffolding in progress; Jev vs LLM support-ticket routing; static + live decision bench; TypeSafe's own published benchmark; illustrative simulations, not live API calls; JevBench v1 — smart/cheap/fast/reliable; I/C/S/K 25% geometric mean; classifier.dev fast tier 84.8 is Jev behind its own API; do not re-fold §78 v1.2 board as new; Laya (421M) 70.1 now on board; Zero-shot/few-shot LLM routing; hard budget filter before Jev; Jev never asked to perform budget arithmetic; Jev judges the next state, XState enforces transitions; simulation uses synthetic keyword fixtures; Consistency benchmark Space; This Space contains no benchmark result yet; 12-case plumbing fixture) / **Catalog gravity** (catalog gravity; v-modal/awesome-jev-tools; ★339 live REST; curation is not endorsement; crawler-maintained directory; Daily GitHub + npm sweep, human-merged; RadRebelSam/awesome-jev ≠ AnotiaWang ≠ yibie ≠ cobanov ≠ logicrw ≠ v-modal) / **HF class ports** (HF peft SPLADE/BGE reranker; rdxtremity/jev-reranking ≠ carlaiau/jev-reranking; query-side encoders, not a Jev replica; ONNX System One Qwen3.5-4B scorer; source:pngwn/system-one-qwen3.5-4b-scorer; CC-BY-NC-4.0; temperature 1.75; transformers.js AutoModel cannot load this graph) / do not reopen or amend PR #23 / Apply 0243: Benchmark-driven Jev router and judge; cheap alone is not success; Jev does not write, sum prices, or claim accuracy %; Sol 94.2 / Luna 83.9 / Jev path 89.7; 19.2% Sol / 62.3% cost save / 4.5pp miss of 2pp non-inferiority; p50 latency worse than Sol due to routing overhead; erendikmenn/jev-llm-router-benchmark ≠ jev-rag-benchmark ≠ ryantsai/jev-llm-router; Express + node:sqlite; mock and Jev decision engines; previous_ticket_count >= 3 is code; MIN_CONFIDENCE 0.6 still soft; substring false positives; aesaganda/jev-ticket-router ≠ SarathChandraBellam/jev-vs-llm-ticket-router; Universal Figure & Diagram Router; confidence ≥ 0.85 hard-gate is theater; generative AI banned from scientific plots; six visual branches; human-labeled (state, question, label); 166,054 rows / 22 configs; soft_label for human uncertainty; Praveenrajus/jev-bench ≠ fstandhartinger/jevbench; ternary bonsai System One GGUF; Hub does not ship weights; 100/100 easy T/F is not Harbor; label_mass ≠ correctness; stock llama.cpp Q2_0 silently gibberish; NicolaiMTLassen/open-bonzi-jev ≠ NicolaiLassen; transformers.js DeBERTa ONNX; temperature 1.05; AutoModel from_pretrained works; onnx-community/open-jev-deberta-v3-large-ONNX ≠ system-one-qwen3.5-4b-scorer-ONNX; 107★ densify; GH 151M vs README 149.6M; PR #1 now closed unmerged; do not re-fold §71 claim-audit as a beat; typed decisions, RLCD, confidence-gated routing; structured ≠ correct; mock not live API; 26 tests; wjdjdakf17/jev-study ≠ baekenough/jev-study / Apply 0345 (`notes.md` §108): **Open reproduction densifies / measurement densifies PRIMARY** (bonzi-27b-v2 / ternary-8b / 27b-v1 GGUF family densify; WANLI-256 74.6% / 65.2% / 71.1% *theirs*; Bonsai 1 27B Q1_0 runs on stock llama.cpp; ternary still needs PrismML fork; hf:heman10x/openJev-verdict-2.0 twin tokenizer-only; OpenJev Vision image classification + uncertainty; CLEVR-4 held-out joint 0%; hfdataset:IamBusy/OpenJev-Vision-Research-v0.1 12,832; 294,912 derived targets not independent samples; Laya multilingual ONNX WebGPU typed-decisions port; 63/63 selected answers / 5.1e-4 CPU / 1.2e-2 WebGPU; UpHash-Network/mini-jev is yuki-oshio transfer; jev-injection-bench 11,900 labelled prompts; Jev best ranking / Haiku better ECE 0.021 vs 0.058; 0.5–0.9 band is where Jev's numbers do not mean what they say; Prompt wording moves panic 28%; manojlds/jev-dspy-bench ≠ dspachos/jev-dspy ≠ jmanhype/jev-dspy-lab; Jev agreement is similarity, never ground truth; no aggregate quality grade or merge gate; AbstentionBench-on-Jev rank 1 of 20 vs 2025 field; question-asymmetry; forward-looking 0.465 never extreme; openkev calibration layer not a runtime; ECE vs coverage independent; select_threshold returns inf; escalation catches uncertainty not ignorance; misakaikato/openkev ≠ jaredpalmer/kev; pdf-race Docling→Jev vs Gemini; parser owns the wall clock; 12/12 tie is a tie; titles selected not generated; flopcheck 16 calibrated tweet judgments; mechanical tells in code; ZeroX-01/jev-atlas ≠ Zaious/jev-capability-atlas ≠ gorock007/jev-atlas; Laya calibration lab Gradio MCP; T never changes argmax; confidence ≠ top-label p; easy probe set refused; 40–48 rows too small to ship T; do not reopen or amend PR #23 or #24 or #25. Apply 0439 (`notes.md` §109): Gemma-4 26B-A4B jevify classification+calibration; LoRA adapter twin not independent eval; Gemma-4 E4B jevify; E4B LoRA stub card; kushalpatil/jevify-gemma4 ≠ Mintzs/jevify ≠ gulagala001/jevify ≠ uspraveen/Jevify; GH kushalpatil07/jevify 404; PAWS 0.580/ece 0.288 is the weak cell; smaller E4B slightly better OOD ECE than 26B-A4B; Hub jevify merged LoRA ships weights; bonzi Bonsai-8B v1 GGUF densify; Bonsai-1.7B v1; Bonsai-4B v1; WANLI-256 64.5% / 60.2% / 52.0% *theirs*; rank #4 / #5 / #6 of 6; JulesHuisman/jev-eval scaffolding / README SHA c356a584 (was empty e69de29b); JulesHuisman/jev-eval ≠ SivletLabs/jev-eval ≠ willkelly/jev-evaluation ≠ 4esv ≠ xxkuboxx ≠ onlyoneaman ≠ dayhaysoos/jevals; 7 bands 6/10 vs 40 bands 0/10; source receipts + confidence slider re-policy without re-inference; 32/32 synthetic is smoke not production; classify HF datasets across typed semantic dimensions; roadus2 watch misspelling; lock roadius2/ultra_laya; ultra_laya REVIEW defects; default branch claude/laya-jev-review-gg5ppo; XNLI EN 88.3% ECE 0.032 → RU 77.3% ECE 0.096; Δ −11.0 pp [−14.2,−7.8]; ECE +0.063; MASSIVE no detectable difference at n=600; confidence is function of p_max (r=1.000); pointer-not-generator 400 human-authored responses; proposed ≠ authorized; FewRel 160: Jev 85.0% vs lexical 13.125%; gated 100% (95/95) coverage 59.375%; J++ composable semantic computation language; judge-jev 0.5 still soft; 947 repos scored; A 273 / B 302 / C 372; LLM rubric ≠ benches; No benchmark winner is claimed; phishing: naive 62.6% vs regex 91.8%; 5-atomic + LR 95.0% *theirs*; AITuber tension ±15; README npm global; repo is Rust; git-confess code owns counting/blame/ratio; httpx exhibit 11% (13/119) *theirs*; 90d trend +12.40% vs random +12.75% vs BH +41.71%; 5m win rate 25%; Awesomejev 656 entries / 38,160 stars; tracker likes 64 (+4) lastModified UNCHANGED; Laya present; Blackwood ABSENT; Archer still promised_not_landed; do not reopen or amend PR #23/#24/#25/#26. Apply 0541 (`notes.md` §110): Blackwood tracker ABSENT; likes 2 gated manual; r = c - p_a; ECE 0.021; acc 0.807 vs warmup 0.746; calibration beyond ~500 tokens unmeasured; Independent primitive; 11.57s vs 54.10s · 4.67× · 120/128 *theirs*; default path is pretrained Gemma probs not trained RLCD head; GH Meanblock 404; lock leesk212/JEV-CPU; softmax over letter slots ≠ Noul; WANLI 0.741 vs openjev v2 0.77 *theirs*; 3-way NLI ≠ Noul; priority 0.464 = majority floor; banking77 contaminated; raw margins not probabilities; GH jev-haiku-benchmarking 404; do not distill Jev as teacher of record (they distilled Haiku); “0.9 is not one number”; ranking ≠ calibration; banking77 0.8–0.9 stated 0.86 actual 0.73 over-confident *theirs*; ≠ Praveenrajus/jev-bench ≠ fstandhartinger/jevbench; $0.0000153–$0.0000226 vs circulating $0.0004 (~20×); Score is 0..n-1 expectation not 0–1; Noul has no confidence field; TCP floor 198.8 ms; type reliability is not a reason to choose Jev (json_schema 5/5); gateway tax not one number; Function-only 5/8 vs hybrid 8/8; 4/8 without Jev; 8 designed cases not conversion lift; ≠ RadRebelSam/awesome-jev; 200-row pilot Jev 86.5% 173/200 vs Gemini Flash-Lite 86.0% 172/200 vs Pro 87.0% 174/200 *theirs*; not a ranking; NLI Tetris argmax P(entail)−P(contradict); 情緒測謊器; 1q 396ms / 30q 567ms; ±0.03; 33q $0.000045 vs Gemini ~5× slower ~60× cost *theirs*; ≠ realZachi/jevtest; 8-example Jev vs GPT-5.6 Sol ~64× cost 5.4× latency *theirs*; synthetic; no inference; ≠ JevBench v1.2 §78; Judged 3317 / listed 2560; Jev judges, code applies policy; catalog ≠ endorsement; APA “microsecond policy / zero hallucination” overclaim; Client-side quiz; pointer from held docs; scanned-PDF warn; CSP only api.typesafe.ai; Jev judges / agent reasons / user decides; selecting an option is not permission to implement; degraded fallback; pattern exact, judgement must clear floor; no matching pattern → no model call; not a correctness oracle; $0.00022 vs chat $0.00306 *theirs*; Spec vs artifact remainder; treating 0.85 as 85% / minProbability hard-gate as Harbor; VERIFY acquires discriminating evidence, never same-pool confidence-only rescoring; fast/full/max are ceilings not sizes; Solar writes, Jev chooses NEXT ACTION; SemIf 2186★ (+20 vs §109 2166); jevlike 1038★ (+7 vs 1031); TypeAR 14★ flat; AnotiaWang 96★ (+1 vs 95); yibie/awesome-jev 490★; Laya likes 802 (was 783); tracker likes 64 flat, lastModified UNCHANGED; do not reopen or amend PR #23/#24/#25/#26/#27. Apply 0743 (`notes.md` §113): Hourly 0743 uniqueness lock: Heman10x-NGU/Verdict-open-jev ≠ Heman10x-NGU/openJev-verdict-2.0; TF-IDF + LogReg ECE 0.0207 vs Jev 0.1440; Verdict-open-jev 48.07% vs Jev 90.80%; abstention combined recall 10.00%; p50 35.58 ms; K=25 (maximum capacity) 72.00%; 0.85 coverage 84.60% selective risk 1.18%; 26.1× faster than standard Qwen JSON generation; Jevify 90.0% / 167 ms CUDA graphs disabled; Finding 1: Brier on stated confidence alone is a trap; grpo_rlcr 0.78 / ECE 0.084; reliability 0.007 but resolution 0.000; 27 900 schema-driven decisions; 13 600 / 13 600 questions; candidate mass min 0.99999624; 22 configs · 166,054 rows · 4 calibration-gold; sha a39eba3f; Student B MAE 0.148 / Pearson 0.836 / 86.0%; pngwn/open-jev-laya-bench README 404; sha 9f69c742 likes 2; HDFS 0.9933 (745/750) / retain 0.0084; BGL ERROR/FATAL protection 1.0000; 2,479 / 2,500 HDFS uncertain; cache hit 0.9648 (2412/2500); $0.153936 estimated; E2 recomputes from saved probabilities; Space sha eda59e0a; MASSIVE English 0.783 / Khmer 0.033 / Hindi 0.133; 40–48 rows too small to ship T; T never changes argmax; siren2345/jev-single-decode-transformers ≠ siren2345/jev-single-decode; Split Transformers experiment from llama.cpp runtime; tanayvasishtha/jev-lab ≠ dairui1/jev-lab ≠ BrendanH18/jev-lab ≠ yibie/laya-jev-lab; Four experiments stress-testing TypeSafe's Jev: calibration, bundle bias, label bias, and ensembling; second pass must be $0.00 from cache; The pages never call Jev; Gemma 4 31B 77.0% / Jev 1.13.0 61.4% / Laya 322M 0.0%; restriction state 95.0% against 84.4%; None of the systems are particularly good at knowing when to stop and ask; They skip the question and call a tool directly; 100% schema pass; six-field joint 48.8% vs 72.8%; ywchiu/jev_benchmark ≠ Running-Dolphins/jev-bench ≠ Praveenrajus/jev-bench; ACT / REVIEW / FALLBACK; A provider failure, timeout, malformed output, or missing answer is **not** a policy outcome; confidence is descriptive provider output, not a substitute for probability; Quality denominators include only valid scored answers; an exact halfway tie chooses the lower level; aiwithenoch/Jev-Skill ≠ simplosophy/jev-skill ≠ laguagu/jev-skills; The local path does not claim to turn a smaller checkpoint into Jev; Low support becomes decision: "review"; MIT-0 SPDX NOASSERTION; current-llm; 结构兼容,不是 Jev 模型能力; altryne/jevify ≠ Mintzs/jevify ≠ gulagala001/jevify ≠ uspraveen/Jevify; Find where Jev belongs. Design the questions. Measure the difference; TypeAR-AI/TypeAR 301 → TypeLLM/TypeLLM; TypeLLM/TypeLLM 16★; SemIf 2241★ (+34 vs §111 2207); jevlike 1051★ (+8 vs 1043); AnotiaWang 98★ (+1 vs 97); yibie/awesome-jev 525★ (+19 vs 506); Laya likes 864 (was 822); tracker likes 67 (+3 vs 64); lastModified UNCHANGED `2026-09-20T04:29:16.000Z`; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#32/#33; notes.md §113 Apply 0646 (`notes.md` §111): Calibration is not alpha; NO CURRENT ALPHA CANDIDATE; ΔR² approximately +0.00084; Brier 0.2131387; ECE 0.0421875; Adding Jev probability to deterministic volatility improved Brier by only 1.4058e-05; default 0.5 keeps zero non pinned; keepResult median 0.14 to 0.17; keepCall median 0.28 to 0.35; usable range is about 0.10 to 0.25; 7.8% to 57.9%; judges results it never sees; task-finish eval not built yet; $0.002 per compaction; slavadubrov/sgr-judge-bench ≠ slavadubrov/jev-judge-bench; Jev 108/120 $0.083 0.34 s; Luna SGR 114/120; paired Jev accuracy-difference intervals include zero; not evidence of equivalence; GLM SGR 26/120 93 format failures; Terra-planned Jev hybrid 55/120; rule-based by default, optionally Jev-backed; empty README; missing key cannot break the experience; prefill plus exactly one decode; softmax over A/B/C ≠ Noul; BBQ 9,053/10,000 (90.53%); ECE 0.0890; Mean confidence 0.9943; overconfident; score and noul not implemented; DGUI 12 rows (was 6); INSTRUCT 119 rows likes 2; encode the state once, decide everything in parallel; 0.740 accuracy against a 0.508 majority; ECE 0.047; fine-tune's advantage ends where its 384-token training data does; jasonkneen/open-jev ≠ pngwn/open-jev; same sha d41dc3cd; Space does not call Jev; recomputes routing from saved probabilities; 200-case Jev 97.0% / 100.0% / 95.0% / MAE 9.22; synthetic repository benchmark; Jev evaluations are advisory; YehuiTang0316/jev-nlgrep ≠ Bentlybro/jevgrep ≠ can1357/jegrep ≠ uehaj/jev-semgrep; default threshold 0.8 still soft; 40-line windows cannot prove whole function; token-native sequential start/end Choice; Gemini/Haiku stubs not configured yet; handful of hand-written examples, not a benchmark; Jev judged exactly what it was given; laguagu/jev-skills ≠ laguagu/jev-evidence-lab ≠ Pleo2/awesome-jev-agent-skills; contract_passed is not a claim of guaranteed factual truth; Wilson lower bound 0.85 floor; fixture mode no savings claim; SemIf 2207★ (+21 vs §110 2186); jevlike 1043★ (+5 vs 1038); TypeAR 15★ (+1 vs 14); AnotiaWang 97★ (+1 vs 96); yibie/awesome-jev 506★ (+16 vs 490); Laya likes 822 (was 802); tracker likes 64 flat, lastModified UNCHANGED; do not reopen or amend PR #23/#24/#25/#26/#27/#28.) | +| Mental models across domains (not SWE-only) | EU, abstention, VOI, MCDA, SDT, search/control, Leveson, NATM/Norman/snap-fit; **extractable-from-state boundary map** (self-contained vs needs outside knowledge). **Apply queued 1047 follow-ons** (`notes.md` §88): **Replica honesty** / **Empty ≠ approve** / **Beam as control** / **Cache is the exact envelope**. Apply 1144 (`notes.md` §89): **Typed if** / **Shadow then honor** / **Human every action** / **Atom then sense** / **File by Choice** / **Question preflight** / **Inbox read-only vs write**. Apply 1241 (`notes.md` §90): **Observe→score→act (namesake lock)** / **Decision-as-ranking** / **Native vs schema-guided Harbor** / **0 promotions / authored vs real** / **Offload + classifier-not-generator** / **Collapse late** / **Unofficial toolbelt** / **Pointer shell** / **Preview-first VOI / rubric rewrite** / **Life fail-open covers** / **S1 decide / S2 plan** / **Seed/expand/judge/verify + local daemon ≠ Jev**. Apply 1347 (`notes.md` §91): **Judge harness as control API** / **Batch packing VOI** / **Calibration as product** / **Decision-as-Plugin** / **Policy-constrained skill select** / **Tiny local econ pruner** / **Observe→score→act cousins** / **Deterministic verify ≠ System One** / **Soft-score vs hard-argmax**. Apply 1441 (`notes.md` §92): **Self-hosted econ** / **Soft-judgment gate integrity** / **Retrieval as calibrated decision space** / **Enterprise reflexes** / **Screenshot-free CU** / **Hybrid S1/S2** / **Harbor-jevals / SRE** / **Laya densifies** / **Demos / unofficial toolbelt**. Apply SIGNAL jevcache/jev-align (`notes.md` §93): **Decision ledger / memoization** / **GEPA alignment loop** (fingerprint after redact; recall vs decide; publish fingerprints+answers; CI replay as Harbor cousin; Cache hit ≠ correctness; hyperspaceai/jevcache ≠ kushals256/jevcache; human labels only; score never auto-accepts; production capture flywheel; sutro-sh/jev-align ≠ caiovicentino/jev-align). Apply SIGNAL enzyme/JA ModernBERT/Gemma (`notes.md` §94): **Compile-time System One / questions-as-index** / **Unofficial JA ModernBERT cross-encoder** / **NAR class legitimacy / multimodal observe→decide / Router-OOD** (guidance ≠ hook; catalysts ≠ summaries; compile-time System One; unofficial ≠ TypeSafe; format_version modernbert-jev/1; Argos1111/jev_local ≠ us/jev-local ≠ kunchenguid/local-jev; LFM default ≠ ModernBERT backend; Nemotron ≠ TypeSafe Jev; not a calibrated replacement; djev-dev complements djev-spark; images as Choice options; Laya essay numbers *theirs*; Router/OOD confidence; hosted bootstrap ≠ silent TypeSafe). Apply 1541 (`notes.md` §95): **Decision-as-plugin** (difficulty + policy thresholds + JSONL trace; jev-codex-pilot model + reasoning depth; keep/shadow/hybrid/reject) / **Evidence projection** (quarry evidence projection) / **Soft judgment integrity** (jevguard calibrator/cache/escape; jev-ci-selector CI shadow mode) / **Physical/control** (Frank-ZY-Dou/awesome-jev robotics/3D/control) / **Harbor-jevals injection-firewall** (one-dollar-tahoe TypeSafe Jev defense eval) / **llama.cpp replica** (llama-jev llama.cpp replica; petercr/jev-orchestrator ≠ FleeexCorp/jev-orchestrator; seb4ez/jevguard ≠ AseemPrasad/JevGuard ≠ pablozr/JevGuard; webNeat/llama-jev ≠ WiktorB2004/llama-index-jev) Apply 1639 (`notes.md` §96): **OpenCode stdout-prune host port** (OpenCode jev-pruner context sieve; observe→score-candidates→prune; jev-zen / jev-1.13-free; zen-chat ≠ Noul; fail-open original; keepScore >0.1 floor; host port of tamaratran/jev-pruner; indiejoseph/opencode-jev-pruner ≠ nrdz-labs/fast-jev-opencode; jev-webagent-bench empty stub; Kiln-AI/jev_jsonschema noul_threshold 0.5; NSStudent/JevSwiftSDK unofficial). Apply SIGNAL gliner-native-runtime (`notes.md` §97): **GLiNER2 native Apple path** (unofficial Swift/Core ML GLiNER 2.5-small; entity spans + confidence; not Choice/Score/Noul; not TypeSafe; label descriptions as schema; on-device ANE economics; honesty locks; shershah1024/gliner-native-runtime ≠ Fastino; ≠ gliner25-compaction ≠ gliner2-ultrafast ≠ Eran-BA/Jev_from_GLiNER2 ≠ NSStudent/JevSwiftSDK ≠ jevmlx; default threshold 0.1 still soft). Apply 1740 (`notes.md` §98): **Decision Graph Protocol envelope** (Decision Graph Protocol frame→assess→commit; app retains permissions/effects; Jev-first assessor-neutral; guarded commit / receipt/next frame; assessment batching; hard-gating DGP as safety theater; numerous-com/dgp ≠ TypeSafe official) / **Calibrated meaning-grep live tree** (jegrep calibrated path+range Nouls; no embeddings/index/daemon; ~$0.01–0.03 typical; agent --json; can1357/jegrep ≠ Bentlybro/jevgrep ≠ uehaj/jev-semgrep) / **Archer-arch fidelity + measured calibration gap** (Archer-arch fidelity; kev family OOD 0.76–0.77 vs Jev 0.86; block-causal isolation; pointer/readout CE-trained; /v1/systemone drop-in; replica honesty). Apply 1843 (`notes.md` §99): **Cost-derived YES/NO/UNSURE control flow** (cost-sensitive decision theory × System One probabilities → control flow; thresholds derived from costs not hard-coded; YES / NO / UNSURE from cost_false_yes / cost_false_no / cost_human; auto-batching same-object questions; Kungie/gut ≠ tpellet/hunch ≠ carldaws/hunch) / **Typed-callback control flow** (judgment vs generation; deterministic execution after probabilistic judgment; exactly one app-owned callback; explicit uncertain branch; Illusion47586/judge ≠ lexingtonhibiki/judgekit ≠ Ascurse/typed-judge-kit) / **Variable-N option scoring as the trainable object** (dynamic candidate bags not fixed label sets; zwliJay/jev-forge ≠ NanoJev; not a new class-table species) / **Open NAR replica economics** (NAR local drop-in; open replica economics / latency vs closed Jev; wfzyx/von late-catch HIGH; competing NAR claims / replica honesty) / **Typed vs chat judges on guardrailing** (typed judgments vs chat judges on guardrailing; ishaannk/llm-vs-jev cross-note only; deeper integrity fold is rh-guard; nothing wins outright; can be argued out of guarding). Apply 1943 (`notes.md` §100): **Jev IS the if-statement** (judgments/probabilities drive branches; text model only writes prose; interpreter owns variables/loops/budgets/replay; otherwise maybe / confidence gate; chaos samples after the gate; southpolesteve/probably ≠ carldaws/hunch ≠ feelings ≠ Kungie/gut ≠ Illusion47586/judge ≠ tidymodels/probably) / **GEPA live delta** (133★ / forks 10 live; build calibrated classifiers from human feedback; HEAD/README SHA unchanged vs §93) / **Memory retrieve vs lease** (retrieve by relevance not resemblance; one calibrated yes/no per memory in one request; pointer mode 17/18 19/20 *theirs*; embedding resemblance misses the allergy; samdotmak/jev-recall ≠ jev-search ≠ jev-sift ≠ carryforward ≠ chopratejas/invalidate; memory leases ended by new evidence; six Nouls then fixed rules in code; 0 of 157 false invalidations; questions/plans/directives are not evidence; unsure → review queue; host keeps the store) / **Contract-of-artifact lint rename** (name↔body / comment truth / test-claims; mizchi/jev-lint is mizchi/jevlint rename; no shipped rule has severity error; ~1 in 5 findings wrong *theirs*; mizchi/jev-lint ≠ huntedman/JevLint ≠ MichitoSugawara/jev-lint) / **JSON Schema question compiler** (JSON Schema → typed JSON via Jev; noul_threshold 0.5 decoder not a proof; IncompatibleSchemaError lists every bad property) / **Local System One economics** (on-device Laya CoreML ANE; ~5 ms P50 short decisions; 189/189 FP16 checkpoint parity; 10× not achieved; mizorewww/laya-coreml ≠ gliner-native-runtime ≠ jevmlx ≠ NandhaKishorM/laya; softmax over allowed tokens ≠ Noul; question-first cache; Micha0827/snapjudge ≠ githubnext/localjev ≠ jevmlx ≠ cendress/SnapJudge; Jev-first Pi agent loop; slow-LLM fallback; explicit action menu / CandidateSource unimplemented; 62 tests wiring not quality; direwolfiy/JevPi ≠ standardagents/jevpilot ≠ pi-jev-control). Apply 2041 (`notes.md` §101): **Resume-screening bias audit** (resume-screening bias audit methodology; name×resume factorial independent Nouls; callback determined by resume quality; mean-probability name gaps operationally negligible; natemoo-re/bias-bench ≠ BBQ) / **MCDA panel code-owned verdict** (Plan/PRD panel → code-owned pass|review|block; cheerleading out of scope; austindixson/planalyzer ≠ single-goodness Noul) / **EU cost-aware routing** (cost-aware multi-model routing/escalation; decide vs do; successful-task cost; cannacre8ive/switchboard-ai ≠ ha-switchboard ≠ hermes-switchyard) / **Frozen-protocol bake-off** (frozen-protocol zero-shot bench; TypeSafe Jev vs PrismNLI vs Laya; contamination caveat; elcronos/jev-vs-open-decision-models ≠ JevBench ≠ DMB) / **VOI admission** (context-window admission control; VOI gate which tokens are worth the expensive model; fail polarity per lens; on small inputs lenses lose money; cvsgireesh/jevusher ≠ jev-sift ≠ winnow) / **Leveson control plane** (typed decision control plane; receipt ≠ authorization; historical-v0 zero retained cases; MokiMeow/jev-fabric ≠ jev-forge ≠ dgp) / **Scoring economics** (live 15-dim typed rubric re-score per pause; scoring economics exemplar; OpenJev/Codiv ≠ TypeSafe hosted; jose-troche/live-rubric ~$0.000004 desc / ~$0.000006 README) / **Pre-registered calibration science** (adversarial pre-registered Jev eval; 28 predictions before data; 123,805 requests; confidence does not track ignorance; polite injection 65% / crude 0%; willkelly/jev-evaluation ≠ jevals ≠ jev-baselines-eval) / **Class infrastructure SDK** (provider-neutral Elixir/BEAM Noul/Choice/Score SDK; class infrastructure; nshkrdotcom/system_one_sdk ≠ typesafe_sdk ≠ dannote/jev). Apply 2145 (`notes.md` §102): **Question-linting of Jev questions themselves** (question-linting of Jev questions themselves; nine jaggedness rules, no API key, no labelled data; static lint ≠ measured separation; yodablocks/jevq ≠ tenbin ≠ JevLint ≠ commitjev) / **Open-weights Laya as class exemplar (binding)** (Nx/Bumblebee runtime; host chooses backend; ChristianAlexander/laya_ex ≠ system_one_sdk ≠ dannote/jev ≠ NandhaKishorM/laya) / **On-chain/edge Laya deploy** (parity_verified stays false; model output never grants Tx; humandebri/IC-Laya ≠ laya_ex) / **Auditable weekend replica** (Jev outputs never used for training; unpaired 0.577 vs 0.727; agilabs-ai/jev48 ≠ JevBench ≠ Mapika/decider) / **Adversarial dual-judge / framing** (comparative framing is the usable judgment; prior injection crowds out evidence; copyleftdev/ember ≠ ember.js) / **Laya specialist + Hub replica** (training still GPU-pending; PIXELZX0/XERON ≠ convaiinnovations/laya; daliborsb/laya ≠ convaiinnovations/laya ≠ NandhaKishorM/laya) / **Distillation economics** (gold is programmatic; teacher is closed-API clone; do not distill Jev as teacher of record; MagaBitmex/jev-4b-distill-data ≠ missing student checkpoint) / **Non-LLM VIN System One** (planning depth not chat; lewislululu/jevon ≠ douglance/jevon) / **Source-bound evidence** (local quote mismatch needs no API; exit 0 ≠ claim truth; WaynezProg/jev-kit ≠ jonathanavis96/jev-kit (Airlock) ≠ jev-use ≠ jev-mcp) / Apply 2246 (`notes.md` §103): **Independent System One evidence catalog** (independent System One evidence catalog; 19 reviewed records; scores not one leaderboard; no external record currently reproduced; TokenTrim no-Jev matched hybrid 62.4%; reachjalil/system-one-bench ≠ mallahyari/system-one-benchmark) / **Typed eval freeze** (21 tasks · 134 items · 208 questions; scenes from public GitHub contracts, not production logs; SivletLabs/jev-eval ≠ willkelly/jev-evaluation ≠ 4esv/jev-eval ≠ xxkuboxx/jev-eval ≠ onlyoneaman/jev-eval ≠ dayhaysoos/jevals) / **Option-isolated tiny replica** (option isolation (sibling-blind); permutation-equivariant; Hub OWNER not published; nafisazizir/hev ≠ jaredpalmer/kev) / **Frozen-LLM typed decisions** (frozen local LLM logits, no trained decision head; residual-head 9,222-param decreased 73/96→67/96; confidence = 1−normalized entropy, not P(correct); yuki-oshio/mini-jev ≠ r-ms/mini-jev) / **AR next-token anti-pattern** (Jev classifier as autoregressive next-token predictor; ChatJev-style soundness theater; erik-dunteman/ChatJev ≠ dannote/jev ≠ jev-gpt) / **Formal compose with scoring** (calibrated decision head × AlphaProof value head; implementation-layer isomorphism, semantic difference; timeout = censoring; do not launder Noul as proof) / **Parallel rank vs serial selection** (parallel rank-prediction vs serial selection; independent questions can conflict; zzzzzec/jevsort ≠ keltokhy/jsort) / **Open-side ecosystem catalog** (curated open System One ecosystem catalog; rupeshpoojary9/awesome-open-system-one ≠ AnotiaWang/awesome-jev) / **Knowledge-work paper radar** (arXiv paper radar with Jev relevance scoring; ranking ≠ calibration / 0.5 still soft; fail-open failed evals not marked seen) / Apply 2340 (`notes.md` §104): **From-scratch calibrated decision model** (train calibrated ~27M from scratch; typed Q→prob dist / one forward pass / no LLM decode; hyusi2003/MiniSystemOne ≠ Colvin0315/MiniSystemOne; description-only stub / size 5) / **ORDER BY ranking upgrade** (ESCI hard probe fails four of six; jev_bool ECE 0.242 inversion 0.255; do not re-fold §60 six-gates as new; jobbyjev one-request-per-company from batch-size result) / **Find/design/evaluate decision loops** (find/design/evaluate TypeSafe Jev decision loops; karanb192/jev-architect ≠ samtay32/jev-system-architect) / **Distill-Jev UI stub** (Jairik/jev-distiller size 1; distill-Jev UI stub / do not distill Jev as teacher of record) / **Post-launch scored opportunity map** (post-launch scored use-case map / Jev self-scores then human curation; licensedsaucer9-web/jev-opportunities) / **Jev-inize a use case** (Jev-inize a use case into classifier/router; gavinHuang/jevinize → simple-jev not TypeSafe; featherless-ai/simple-jev) / **Saved-decision regression** (compare saved decisions / same label can still change the branch; VihaanAgarwal/jev-diff ≠ Saik0s/diffusiongemma-jev-macos; not tested with a live Jev API key) / **Constrained-logprob API** (constrained logprob + temp/Platt ≠ Noul; OpenJevPro pastes openjev-sglang JevBench as own; zhangcy122/OpenJevPro ≠ IamBusy/OpenJev ≠ ekzhang/openjev-sglang; PolyForm Noncommercial) / **SmolLM RLCD reproduction** (SmolLM-135M / sub-70ms / 0 output tokens; demo P(True) 0.5052 / Choice conf 0.2872 / Score conf 0.0055; README claims MIT / GitHub license null / no LICENSE file; patelvishwa112/jev-system-one-rlcd ≠ arnabgho/rlcd-lite ≠ blackwood-rlcd) / **Source-backed Awesome radar** (source-backed Awesome Jev radar / 306+ commit-pinned; logicrw/awesome-jev-projects ≠ AnotiaWang/awesome-jev ≠ yibie/awesome-jev ≠ cobanov/awesome-jev ≠ rupeshpoojary9/awesome-open-system-one; auto GitHub sync / Issue-only submissions) / **Rival-aware one-pass scorer** (hashed n-gram encoder / rival-aware attention; olanotolu/jevbetter vs jevlike starter; synthetic hard menus top-1 0.916 vs 0.873 / ECE 0.0182 vs 0.0367 / 40 vs 4608 menus/sec; shuffled-context control 0.335) | `references/mental-models.md` Apply 0042 (`notes.md` §105): structured probability readouts; distribution > argmax; Noul 0.5 midpoint; score is expectation not integer; bare HTTP not SDK; Arohtea/jev-readout) / **Ordinary-model Jev-shape** (Jev-style Choice/Score/Noul from ordinary models; optional DSH plugin; schema-valid ≠ calibrated; gulagala001/jevify ≠ Mintzs/jevify) / **Open-weight Laya measurement** (Laya RLCD benchmark; 40.3% below constant-answer; open-weight measurement; mourad-ghafiri/laya-rlcd-benchmark ≠ yibie/laya-jev-lab) / **Cheap fail-open semantic edge** (cheap fail-open semantic edge; second signal not sole; FastLoopError catch; SupremeDreamZ/jev-fastloop ≠ jev-ultrafast) / **Fan-out measurement** (asking more questions in one call; 0.980 at every N; nearly not fully deterministic; TheWebDevel/jev-fanout) / **VLM+Jev RL teacher** (Qwen3-VL perception + Jev decisions train RL; 0 model calls at deployment; VLM alone 1.7 vs +Jev 4.4; harneet2512/reflexrl ≠ khordoo/jev-reflex-autonomy-lab) / **Independent Jev API vs Laya** (independent Jev API vs Laya; cascade 0.60 matches 78% at 1.8×; noul facts not judgements; yibie/laya-jev-lab ≠ dairui1/jev-lab ≠ BrendanH18/jev-lab) / **Locate vs decide** (GLiNER vs GLiFormer vs Laya vs Jev; extractors ≠ decision engines; Laya dict-instructions collapse 58.3%; umstek/zero-shot-ie-bench) / **Throughput arena** (decisions-per-minute & cost; 204 moves vs 73; throughput not intelligence; angelgalvisc/snake-arena-jev-vs-llms ≠ vtrivedy/jev-plays-games) / **Behavioral contracts** (behavioral contracts; pin expectations eval upgrades; raw 0.94 is not a release; sathariels/jevcheck ≠ dayhaysoos/jevals ≠ SivletLabs/jev-eval) / **Evidence-linked upgrade review** (evidence-linked dependency upgrade; Jev never generates filenames; no_direct_evidence ≠ safe to merge; GaneshVG18/upgrade-radar ≠ LYchoon/paper-radar-jev) / **Knowledge-work discography** (discography theme/mood/complexity; five atomic questions one call; lirantal/discoprint) / Apply 0145 (`notes.md` §106): **Architecture probes PRIMARY** (Turn any open LLM into System-One Jev; uspraveen/Jevify ≠ Mintzs/jevify ≠ gulagala001/jevify; Jevify-any-LLM architecture probe; description-only stub / size 0; Train encoder-only calibrated decision models from a task sentence; Exu is a toolkit, not a method; strictly proper scoring rule; Pre-alpha; Ruivalim/exu-base; scratch-trained calibrated decision model; typed Q → probability dists; Colvin0315/MiniSystemOne ≠ hyusi2003/MiniSystemOne; no published weights download URL; 90.5 seconds / 29.2% pipeline evidence; p_i/p_j independent of other candidates; Recipe for calibrated decision models — small model out; init → synth → train → eval → serve; 91.1 % / ECE 0.022 *theirs*; Jev zero-shot 75.1; scienthoon/luce; Put Jev's three headline claims on trial; 0.5B local GPU; 46x speedup / accuracy identical; ECE 0.624 sentiment catastrophe; bigger model worse calibration; RichardoMrMu/jev-mini ≠ yuki-oshio/mini-jev ≠ r-ms/mini-jev; System-1 decision engine for local LLMs; structured choices only; JSON parse of generated text ≠ Noul; TypefAI JEV / Journal Entry Voucher; tapsin/jev-local ≠ us/jev-local ≠ Argos1111/jev_local) / **Measurement densifies** (Jev 1.13 reward-model eval across 8 benchmark tracks; 40,940 examples / 0 API errors; RewardBench v1 92.58%; Precise IF 50.63%; goya4140/jev-reward-model-evaluation; Scaffolding in progress; Jev vs LLM support-ticket routing; static + live decision bench; TypeSafe's own published benchmark; illustrative simulations, not live API calls; JevBench v1 — smart/cheap/fast/reliable; I/C/S/K 25% geometric mean; classifier.dev fast tier 84.8 is Jev behind its own API; do not re-fold §78 v1.2 board as new; Laya (421M) 70.1 now on board; Zero-shot/few-shot LLM routing; hard budget filter before Jev; Jev never asked to perform budget arithmetic; Jev judges the next state, XState enforces transitions; simulation uses synthetic keyword fixtures; Consistency benchmark Space; This Space contains no benchmark result yet; 12-case plumbing fixture) / **Catalog gravity** (catalog gravity; v-modal/awesome-jev-tools; ★339 live REST; curation is not endorsement; crawler-maintained directory; Daily GitHub + npm sweep, human-merged; RadRebelSam/awesome-jev ≠ AnotiaWang ≠ yibie ≠ cobanov ≠ logicrw ≠ v-modal) / **HF class ports** (HF peft SPLADE/BGE reranker; rdxtremity/jev-reranking ≠ carlaiau/jev-reranking; query-side encoders, not a Jev replica; ONNX System One Qwen3.5-4B scorer; source:pngwn/system-one-qwen3.5-4b-scorer; CC-BY-NC-4.0; temperature 1.75; transformers.js AutoModel cannot load this graph) / do not reopen or amend PR #23 / Apply 0243: Benchmark-driven Jev router and judge; cheap alone is not success; Jev does not write, sum prices, or claim accuracy %; Sol 94.2 / Luna 83.9 / Jev path 89.7; 19.2% Sol / 62.3% cost save / 4.5pp miss of 2pp non-inferiority; p50 latency worse than Sol due to routing overhead; erendikmenn/jev-llm-router-benchmark ≠ jev-rag-benchmark ≠ ryantsai/jev-llm-router; Express + node:sqlite; mock and Jev decision engines; previous_ticket_count >= 3 is code; MIN_CONFIDENCE 0.6 still soft; substring false positives; aesaganda/jev-ticket-router ≠ SarathChandraBellam/jev-vs-llm-ticket-router; Universal Figure & Diagram Router; confidence ≥ 0.85 hard-gate is theater; generative AI banned from scientific plots; six visual branches; human-labeled (state, question, label); 166,054 rows / 22 configs; soft_label for human uncertainty; Praveenrajus/jev-bench ≠ fstandhartinger/jevbench; ternary bonsai System One GGUF; Hub does not ship weights; 100/100 easy T/F is not Harbor; label_mass ≠ correctness; stock llama.cpp Q2_0 silently gibberish; NicolaiMTLassen/open-bonzi-jev ≠ NicolaiLassen; transformers.js DeBERTa ONNX; temperature 1.05; AutoModel from_pretrained works; onnx-community/open-jev-deberta-v3-large-ONNX ≠ system-one-qwen3.5-4b-scorer-ONNX; 107★ densify; GH 151M vs README 149.6M; PR #1 now closed unmerged; do not re-fold §71 claim-audit as a beat; typed decisions, RLCD, confidence-gated routing; structured ≠ correct; mock not live API; 26 tests; wjdjdakf17/jev-study ≠ baekenough/jev-study / Apply 0345 (`notes.md` §108): **Open reproduction densifies / measurement densifies PRIMARY** (bonzi-27b-v2 / ternary-8b / 27b-v1 GGUF family densify; WANLI-256 74.6% / 65.2% / 71.1% *theirs*; Bonsai 1 27B Q1_0 runs on stock llama.cpp; ternary still needs PrismML fork; hf:heman10x/openJev-verdict-2.0 twin tokenizer-only; OpenJev Vision image classification + uncertainty; CLEVR-4 held-out joint 0%; hfdataset:IamBusy/OpenJev-Vision-Research-v0.1 12,832; 294,912 derived targets not independent samples; Laya multilingual ONNX WebGPU typed-decisions port; 63/63 selected answers / 5.1e-4 CPU / 1.2e-2 WebGPU; UpHash-Network/mini-jev is yuki-oshio transfer; jev-injection-bench 11,900 labelled prompts; Jev best ranking / Haiku better ECE 0.021 vs 0.058; 0.5–0.9 band is where Jev's numbers do not mean what they say; Prompt wording moves panic 28%; manojlds/jev-dspy-bench ≠ dspachos/jev-dspy ≠ jmanhype/jev-dspy-lab; Jev agreement is similarity, never ground truth; no aggregate quality grade or merge gate; AbstentionBench-on-Jev rank 1 of 20 vs 2025 field; question-asymmetry; forward-looking 0.465 never extreme; openkev calibration layer not a runtime; ECE vs coverage independent; select_threshold returns inf; escalation catches uncertainty not ignorance; misakaikato/openkev ≠ jaredpalmer/kev; pdf-race Docling→Jev vs Gemini; parser owns the wall clock; 12/12 tie is a tie; titles selected not generated; flopcheck 16 calibrated tweet judgments; mechanical tells in code; ZeroX-01/jev-atlas ≠ Zaious/jev-capability-atlas ≠ gorock007/jev-atlas; Laya calibration lab Gradio MCP; T never changes argmax; confidence ≠ top-label p; easy probe set refused; 40–48 rows too small to ship T; do not reopen or amend PR #23 or #24 or #25. Apply 0439 (`notes.md` §109): Gemma-4 26B-A4B jevify classification+calibration; LoRA adapter twin not independent eval; Gemma-4 E4B jevify; E4B LoRA stub card; kushalpatil/jevify-gemma4 ≠ Mintzs/jevify ≠ gulagala001/jevify ≠ uspraveen/Jevify; GH kushalpatil07/jevify 404; PAWS 0.580/ece 0.288 is the weak cell; smaller E4B slightly better OOD ECE than 26B-A4B; Hub jevify merged LoRA ships weights; bonzi Bonsai-8B v1 GGUF densify; Bonsai-1.7B v1; Bonsai-4B v1; WANLI-256 64.5% / 60.2% / 52.0% *theirs*; rank #4 / #5 / #6 of 6; JulesHuisman/jev-eval scaffolding / README SHA c356a584 (was empty e69de29b); JulesHuisman/jev-eval ≠ SivletLabs/jev-eval ≠ willkelly/jev-evaluation ≠ 4esv ≠ xxkuboxx ≠ onlyoneaman ≠ dayhaysoos/jevals; 7 bands 6/10 vs 40 bands 0/10; source receipts + confidence slider re-policy without re-inference; 32/32 synthetic is smoke not production; classify HF datasets across typed semantic dimensions; roadus2 watch misspelling; lock roadius2/ultra_laya; ultra_laya REVIEW defects; default branch claude/laya-jev-review-gg5ppo; XNLI EN 88.3% ECE 0.032 → RU 77.3% ECE 0.096; Δ −11.0 pp [−14.2,−7.8]; ECE +0.063; MASSIVE no detectable difference at n=600; confidence is function of p_max (r=1.000); pointer-not-generator 400 human-authored responses; proposed ≠ authorized; FewRel 160: Jev 85.0% vs lexical 13.125%; gated 100% (95/95) coverage 59.375%; J++ composable semantic computation language; judge-jev 0.5 still soft; 947 repos scored; A 273 / B 302 / C 372; LLM rubric ≠ benches; No benchmark winner is claimed; phishing: naive 62.6% vs regex 91.8%; 5-atomic + LR 95.0% *theirs*; AITuber tension ±15; README npm global; repo is Rust; git-confess code owns counting/blame/ratio; httpx exhibit 11% (13/119) *theirs*; 90d trend +12.40% vs random +12.75% vs BH +41.71%; 5m win rate 25%; Awesomejev 656 entries / 38,160 stars; tracker likes 64 (+4) lastModified UNCHANGED; Laya present; Blackwood ABSENT; Archer still promised_not_landed; do not reopen or amend PR #23/#24/#25/#26. Apply 0541 (`notes.md` §110): Blackwood tracker ABSENT; likes 2 gated manual; r = c - p_a; ECE 0.021; acc 0.807 vs warmup 0.746; calibration beyond ~500 tokens unmeasured; Independent primitive; 11.57s vs 54.10s · 4.67× · 120/128 *theirs*; default path is pretrained Gemma probs not trained RLCD head; GH Meanblock 404; lock leesk212/JEV-CPU; softmax over letter slots ≠ Noul; WANLI 0.741 vs openjev v2 0.77 *theirs*; 3-way NLI ≠ Noul; priority 0.464 = majority floor; banking77 contaminated; raw margins not probabilities; GH jev-haiku-benchmarking 404; do not distill Jev as teacher of record (they distilled Haiku); “0.9 is not one number”; ranking ≠ calibration; banking77 0.8–0.9 stated 0.86 actual 0.73 over-confident *theirs*; ≠ Praveenrajus/jev-bench ≠ fstandhartinger/jevbench; $0.0000153–$0.0000226 vs circulating $0.0004 (~20×); Score is 0..n-1 expectation not 0–1; Noul has no confidence field; TCP floor 198.8 ms; type reliability is not a reason to choose Jev (json_schema 5/5); gateway tax not one number; Function-only 5/8 vs hybrid 8/8; 4/8 without Jev; 8 designed cases not conversion lift; ≠ RadRebelSam/awesome-jev; 200-row pilot Jev 86.5% 173/200 vs Gemini Flash-Lite 86.0% 172/200 vs Pro 87.0% 174/200 *theirs*; not a ranking; NLI Tetris argmax P(entail)−P(contradict); 情緒測謊器; 1q 396ms / 30q 567ms; ±0.03; 33q $0.000045 vs Gemini ~5× slower ~60× cost *theirs*; ≠ realZachi/jevtest; 8-example Jev vs GPT-5.6 Sol ~64× cost 5.4× latency *theirs*; synthetic; no inference; ≠ JevBench v1.2 §78; Judged 3317 / listed 2560; Jev judges, code applies policy; catalog ≠ endorsement; APA “microsecond policy / zero hallucination” overclaim; Client-side quiz; pointer from held docs; scanned-PDF warn; CSP only api.typesafe.ai; Jev judges / agent reasons / user decides; selecting an option is not permission to implement; degraded fallback; pattern exact, judgement must clear floor; no matching pattern → no model call; not a correctness oracle; $0.00022 vs chat $0.00306 *theirs*; Spec vs artifact remainder; treating 0.85 as 85% / minProbability hard-gate as Harbor; VERIFY acquires discriminating evidence, never same-pool confidence-only rescoring; fast/full/max are ceilings not sizes; Solar writes, Jev chooses NEXT ACTION; SemIf 2186★ (+20 vs §109 2166); jevlike 1038★ (+7 vs 1031); TypeAR 14★ flat; AnotiaWang 96★ (+1 vs 95); yibie/awesome-jev 490★; Laya likes 802 (was 783); tracker likes 64 flat, lastModified UNCHANGED; do not reopen or amend PR #23/#24/#25/#26/#27. Apply 0920 jcr (`notes.md` §116): **Retrieve-wide→decide→evidence-set** / **Skills vs capability catalogs** / **VOI of context admission** / **Measurement honesty (wall-time mixed)**. Apply 0743 (`notes.md` §113): Hourly 0743 uniqueness lock: Heman10x-NGU/Verdict-open-jev ≠ Heman10x-NGU/openJev-verdict-2.0; TF-IDF + LogReg ECE 0.0207 vs Jev 0.1440; Verdict-open-jev 48.07% vs Jev 90.80%; abstention combined recall 10.00%; p50 35.58 ms; K=25 (maximum capacity) 72.00%; 0.85 coverage 84.60% selective risk 1.18%; 26.1× faster than standard Qwen JSON generation; Jevify 90.0% / 167 ms CUDA graphs disabled; Finding 1: Brier on stated confidence alone is a trap; grpo_rlcr 0.78 / ECE 0.084; reliability 0.007 but resolution 0.000; 27 900 schema-driven decisions; 13 600 / 13 600 questions; candidate mass min 0.99999624; 22 configs · 166,054 rows · 4 calibration-gold; sha a39eba3f; Student B MAE 0.148 / Pearson 0.836 / 86.0%; pngwn/open-jev-laya-bench README 404; sha 9f69c742 likes 2; HDFS 0.9933 (745/750) / retain 0.0084; BGL ERROR/FATAL protection 1.0000; 2,479 / 2,500 HDFS uncertain; cache hit 0.9648 (2412/2500); $0.153936 estimated; E2 recomputes from saved probabilities; Space sha eda59e0a; MASSIVE English 0.783 / Khmer 0.033 / Hindi 0.133; 40–48 rows too small to ship T; T never changes argmax; siren2345/jev-single-decode-transformers ≠ siren2345/jev-single-decode; Split Transformers experiment from llama.cpp runtime; tanayvasishtha/jev-lab ≠ dairui1/jev-lab ≠ BrendanH18/jev-lab ≠ yibie/laya-jev-lab; Four experiments stress-testing TypeSafe's Jev: calibration, bundle bias, label bias, and ensembling; second pass must be $0.00 from cache; The pages never call Jev; Gemma 4 31B 77.0% / Jev 1.13.0 61.4% / Laya 322M 0.0%; restriction state 95.0% against 84.4%; None of the systems are particularly good at knowing when to stop and ask; They skip the question and call a tool directly; 100% schema pass; six-field joint 48.8% vs 72.8%; ywchiu/jev_benchmark ≠ Running-Dolphins/jev-bench ≠ Praveenrajus/jev-bench; ACT / REVIEW / FALLBACK; A provider failure, timeout, malformed output, or missing answer is **not** a policy outcome; confidence is descriptive provider output, not a substitute for probability; Quality denominators include only valid scored answers; an exact halfway tie chooses the lower level; aiwithenoch/Jev-Skill ≠ simplosophy/jev-skill ≠ laguagu/jev-skills; The local path does not claim to turn a smaller checkpoint into Jev; Low support becomes decision: "review"; MIT-0 SPDX NOASSERTION; current-llm; 结构兼容,不是 Jev 模型能力; altryne/jevify ≠ Mintzs/jevify ≠ gulagala001/jevify ≠ uspraveen/Jevify; Find where Jev belongs. Design the questions. Measure the difference; TypeAR-AI/TypeAR 301 → TypeLLM/TypeLLM; TypeLLM/TypeLLM 16★; SemIf 2241★ (+34 vs §111 2207); jevlike 1051★ (+8 vs 1043); AnotiaWang 98★ (+1 vs 97); yibie/awesome-jev 525★ (+19 vs 506); Laya likes 864 (was 822); tracker likes 67 (+3 vs 64); lastModified UNCHANGED `2026-09-20T04:29:16.000Z`; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#32/#33; notes.md §113 Apply 0646 (`notes.md` §111): Calibration is not alpha; NO CURRENT ALPHA CANDIDATE; ΔR² approximately +0.00084; Brier 0.2131387; ECE 0.0421875; Adding Jev probability to deterministic volatility improved Brier by only 1.4058e-05; default 0.5 keeps zero non pinned; keepResult median 0.14 to 0.17; keepCall median 0.28 to 0.35; usable range is about 0.10 to 0.25; 7.8% to 57.9%; judges results it never sees; task-finish eval not built yet; $0.002 per compaction; slavadubrov/sgr-judge-bench ≠ slavadubrov/jev-judge-bench; Jev 108/120 $0.083 0.34 s; Luna SGR 114/120; paired Jev accuracy-difference intervals include zero; not evidence of equivalence; GLM SGR 26/120 93 format failures; Terra-planned Jev hybrid 55/120; rule-based by default, optionally Jev-backed; empty README; missing key cannot break the experience; prefill plus exactly one decode; softmax over A/B/C ≠ Noul; BBQ 9,053/10,000 (90.53%); ECE 0.0890; Mean confidence 0.9943; overconfident; score and noul not implemented; DGUI 12 rows (was 6); INSTRUCT 119 rows likes 2; encode the state once, decide everything in parallel; 0.740 accuracy against a 0.508 majority; ECE 0.047; fine-tune's advantage ends where its 384-token training data does; jasonkneen/open-jev ≠ pngwn/open-jev; same sha d41dc3cd; Space does not call Jev; recomputes routing from saved probabilities; 200-case Jev 97.0% / 100.0% / 95.0% / MAE 9.22; synthetic repository benchmark; Jev evaluations are advisory; YehuiTang0316/jev-nlgrep ≠ Bentlybro/jevgrep ≠ can1357/jegrep ≠ uehaj/jev-semgrep; default threshold 0.8 still soft; 40-line windows cannot prove whole function; token-native sequential start/end Choice; Gemini/Haiku stubs not configured yet; handful of hand-written examples, not a benchmark; Jev judged exactly what it was given; laguagu/jev-skills ≠ laguagu/jev-evidence-lab ≠ Pleo2/awesome-jev-agent-skills; contract_passed is not a claim of guaranteed factual truth; Wilson lower bound 0.85 floor; fixture mode no savings claim; SemIf 2207★ (+21 vs §110 2186); jevlike 1043★ (+5 vs 1038); TypeAR 15★ (+1 vs 14); AnotiaWang 97★ (+1 vs 96); yibie/awesome-jev 506★ (+16 vs 490); Laya likes 822 (was 802); tracker likes 64 flat, lastModified UNCHANGED; do not reopen or amend PR #23/#24/#25/#26/#27/#28.) | | Judgment-model class (Jev is exemplar, not monopoly) | Species: decide / locate (GLiNER) / categorize (GLiClass) / rank / perceive; open heads include encoder DeBERTa, LoRA distill, **domain specialist LoRA on independent gold**, openjev-lm, kev. Compaction job is backend-agnostic (Jev Score/Noul vs GLiNER2.5 encoder). Indexer cousin: GLiNER extract + escalate-S2 (10–50× unfilled). Computer-use observe→score-among-candidates→code-acts is backend-agnostic (Jev Ultrafast ↔ GLiNER2 Ultrafast ↔ Cua-S1 specialist ↔ Stagehand experimental Jev stack; **OCR+AX desktop:** typesafe-computer-use hosted Jev, never ships a screenshot for the *decision*; Cua-S1 is not TypeSafe Jev; Stagehand pick is a fast path, not a replacement; **≠** jev-macos-loop OmniParser **≠** camoufox; **ASR voice-browser:** jev-voice-browser hosted Jev, never ships a waveform). GLiFormer encoder serving `/v1/systemone` is a class-backend (jeff), not a Jev replica. Local MLX PCD is O(1) constrained-AR speed, **not** a calibrated Noul (system-one-benchmark Brier 0.3884 vs Jev 0.1096). **jevmlx** is the productized Apple Silicon one-pass schema→JSON+probs library (softmax ≠ Noul; no local leaderboard yet). **openvons** is an independent open-Jev class (LM/vision/voice finite-choice+prob; JevPick 3.2–4.8×; Flutter on-device; `/v1/systemone` wire-compat, not a TypeSafe replica). **OpenJev** (IamBusy) is a local 0.6B LoRA+scalar head on `/v1/decide` (45/60 *theirs*; **not** a TypeSafe drop-in; distinct from hraness/sysone OpenJev runners). **semif-serve** puts SemIf behind `/v1/systemone` (runoff, no option ceiling; 1164 vs 178 ms *theirs*; wire-compat ≠ replica). **grande** Rust/WebGPU kev-shaped branches (JGLUE *theirs* JNLI 0.614 ECE→0.088 / JCQA 0.853; 270M 0.710/0.710; isolation 0.098/0.996). **laya-jolt** Clojure/Jolt byte parity vs Python Laya. **JEV-CPU** SemIf on CPU (leesk212; Meanblock 404). **local-jev** ONNX ModernBERT measured not-equivalent (done 30% / shape 57% *theirs*). **GLiNER2→Choice/Score/Noul spec** (Eran-BA; design only, ≠ jeff GLiFormer). **openJev-verdict-2.0** competing NAR claims as **audit object not endorsement** (77.10%/0.0636/0.0144 *theirs*; PR #1; ≠ IamBusy/OpenJev `/v1/decide`). **chakuho** 1-token logprob local `/v1/systemone` (softmax ≠ Noul; coverage ≠ correctness; GUI 336 *theirs* 27B 95%/92% vs Jev 89%/82%). **jevinf** open replica engine (NanoJev/decider-2b/Laya; 2.57×/2.27× 100% argmax; MPS only). **laya-multilingual** mmBERT-base 322M (MASSIVE 0.366/0.387 vs English laya 0.227/0.733; Khmer 0.000@0.952 conf; ships uncalibrated). **schema-scorer** DeBERTa-v3-large scalar head (Hub; GitHub 404; v2 Choice 0.841 *theirs*; peaked ranking ≠ calibration). **githubnext/localjev** prompted-JSON `/v1/systemone` (MIT **261★**; TypeSafe SDK drop-in; wire-compat ≠ logit-equiv vs razorback16 structured-read; **≠** kunchenguid/local-jev; 1,200-req bake-off *theirs* Qwen3.6 76.7% / Gemma 26B 75.0% / DiffusionGemma 74.2% short; not calibrated). **NandhaKishorM/laya** PyPI+Router packaging of Hub Laya (Apache-2.0; **710★**; not a new species; T4 32.8 ms *theirs*; post-T ECE 0.081 vs Jev 0.246; Banking77 0.425 vs Jev 0.870; 0.766 is fine-tune not zero-shot; 0.85 still soft; **≠** TypeSafe `/v1/systemone`). **external openjev census** (@airesearch12 tweet ≠ jevbench v1.1; GLiNER2+routers class-boundary; incomplete vs watch). **JevBench v1.2 scored board** (geo-mean I/C/S/K 25% each; Jev 75.3 / SemIf 74.6 *theirs*; cal ON rank; instruction models in the table; ≠ v1.1 87.6; ≠ tweet census; Laya absent gap; Qwen3.8 27B ≠ Archer). **Hourly 0842:** already-folded class as a recipe (wire≠logit · product+FALLBACK · packaging honesty · pointer-not-generator · leaderboard VOI); skip thin noise. **Open LoRA replica, different jeff:** GestaltLabs/Jeff-1 LoRA Qwen3-4B ≠ logan-markewich/jeff GLiFormer; acc/ECE tradeoff n=9730 *theirs*; set reused. **Hourly 1241:** typesafeai-sdk-community not a new species; 2389-research/judgement license null; confidence ≠ winner p; jevbrain AUTO_ACT is not a Noul. **Hourly 1347:** FrancoisChastel/jev-code ≠ npm jev-code; ctmx/openrouter-jev-mcp Decision-as-Plugin; claudecode-jev-marketplace fail-open not hot path; pedroknigge/mcp_jev packs not ask_jev; 0thernet/system-one-skills deterministic verify; hari007sh/jev ≠ dannote/jev; nanoprune 2.8MB ECE 2.58%. **Hourly 1441:** Foq ~25ms/2.2GB local; rev prefill-only + HF jev-0.5b; robfrase/jev planning memo; meldltd/meldecision laya-go ONNX; laya-doom never pixels; logixism/laya-api empty README; akpsahan/laya ≠ Archer; Nibir1/typesafe-go ≠ official. **SIGNAL §94:** unofficial ≠ TypeSafe JA ModernBERT (argos1111/modernbert-ja-310m-jev; format_version modernbert-jev/1); Argos1111/jev_local ≠ us/jev-local ≠ kunchenguid/local-jev; LFM default ≠ ModernBERT backend; Nemotron ≠ TypeSafe Jev; djev-dev complements djev-spark (images as Choice options); Laya essay ≠ new species; hosted bootstrap ≠ silent TypeSafe. **Hourly 1541:** llama-jev llama.cpp replica; softmax ≠ Noul; **≠** TypeSafe **≠** pcdServer. **SIGNAL §97:** GLiNER2 native Apple path; unofficial Swift/Core ML GLiNER 2.5-small; entity spans + confidence; not Choice/Score/Noul; not TypeSafe; label descriptions as schema; on-device ANE economics; honesty locks; shershah1024/gliner-native-runtime ≠ Fastino; ≠ gliner25-compaction ≠ gliner2-ultrafast ≠ Eran-BA/Jev_from_GLiNER2 ≠ NSStudent/JevSwiftSDK ≠ jevmlx; default threshold 0.1 still soft. **Hourly 1740:** Decision Graph Protocol frame→assess→commit; app retains permissions/effects; Jev-first assessor-neutral; guarded commit / receipt/next frame; assessment batching; hard-gating DGP as safety theater; numerous-com/dgp ≠ TypeSafe official; jegrep calibrated path+range Nouls; no embeddings/index/daemon; ~$0.01–0.03 typical; agent --json; can1357/jegrep ≠ Bentlybro/jevgrep ≠ uehaj/jev-semgrep; Archer-arch fidelity; kev family OOD 0.76–0.77 vs Jev 0.86; block-causal isolation; pointer/readout CE-trained; /v1/systemone drop-in; replica honesty. **Hourly 1843:** variable-N option scoring as the trainable object; dynamic candidate bags not fixed label sets; zwliJay/jev-forge ≠ NanoJev; not a new class-table species; NAR local drop-in; wfzyx/von late-catch HIGH; open replica economics / latency vs closed Jev; competing NAR claims / replica honesty; gut/judge are control-flow overlays not species. **Hourly 1943:** softmax over allowed tokens ≠ Noul; on-device Laya CoreML ANE; JSON Schema → typed JSON via Jev; noul_threshold 0.5 decoder not a proof; mizorewww/laya-coreml ≠ gliner-native-runtime ≠ jevmlx ≠ NandhaKishorM/laya; Micha0827/snapjudge ≠ githubnext/localjev ≠ jevmlx ≠ cendress/SnapJudge; Jev IS the if-statement is a language primitive not a new species **Hourly 2041:** resume-screening bias audit methodology; Plan/PRD panel → code-owned pass|review|block; cost-aware multi-model routing/escalation; frozen-protocol zero-shot bench; context-window admission control; typed decision control plane; receipt ≠ authorization; live 15-dim typed rubric re-score per pause; adversarial pre-registered Jev eval; confidence does not track ignorance; provider-neutral Elixir/BEAM Noul/Choice/Score SDK **Hourly 2145:** question-linting of Jev questions themselves; nine jaggedness rules, no API key, no labelled data; static lint ≠ measured separation; yodablocks/jevq ≠ tenbin ≠ JevLint ≠ commitjev; open-weights Laya as class exemplar (binding); Nx/Bumblebee runtime; host chooses backend; ChristianAlexander/laya_ex ≠ system_one_sdk ≠ dannote/jev ≠ NandhaKishorM/laya; on-chain/edge Laya deploy; parity_verified stays false; model output never grants Tx; humandebri/IC-Laya ≠ laya_ex; auditable weekend replica; Jev outputs never used for training; unpaired 0.577 vs 0.727; agilabs-ai/jev48 ≠ JevBench ≠ Mapika/decider; adversarial dual-judge / framing attack surface; comparative framing is the usable judgment; prior injection crowds out evidence; copyleftdev/ember ≠ ember.js; Laya specialist fine-tune pipeline; training still GPU-pending; PIXELZX0/XERON ≠ convaiinnovations/laya; Hub Laya replica drop; daliborsb/laya ≠ convaiinnovations/laya ≠ NandhaKishorM/laya; System One student distillation corpus; gold is programmatic; teacher is closed-API clone; MagaBitmex/jev-4b-distill-data ≠ missing student checkpoint; non-LLM VIN System One; planning depth not chat; lewislululu/jevon ≠ douglance/jevon; source-bound evidence checks; local quote mismatch needs no API; exit 0 ≠ claim truth; WaynezProg/jev-kit ≠ jonathanavis96/jev-kit (Airlock) ≠ jev-use ≠ jev-mcp **Hourly 2246:** independent System One evidence catalog; 19 reviewed records; scores not one leaderboard; no external record currently reproduced; TokenTrim no-Jev matched hybrid 62.4%; reachjalil/system-one-bench ≠ mallahyari/system-one-benchmark; 21 tasks · 134 items · 208 questions; scenes from public GitHub contracts, not production logs; SivletLabs/jev-eval ≠ willkelly/jev-evaluation ≠ 4esv/jev-eval ≠ xxkuboxx/jev-eval ≠ onlyoneaman/jev-eval ≠ dayhaysoos/jevals; option isolation (sibling-blind); permutation-equivariant; Hub OWNER not published; nafisazizir/hev ≠ jaredpalmer/kev; frozen local LLM logits, no trained decision head; residual-head 9,222-param decreased 73/96→67/96; confidence = 1−normalized entropy, not P(correct); yuki-oshio/mini-jev ≠ r-ms/mini-jev; Jev classifier as autoregressive next-token predictor; ChatJev-style soundness theater; erik-dunteman/ChatJev ≠ dannote/jev ≠ jev-gpt; calibrated decision head × AlphaProof value head; implementation-layer isomorphism, semantic difference; timeout = censoring; do not launder Noul as proof; parallel rank-prediction vs serial selection; independent questions can conflict; zzzzzec/jevsort ≠ keltokhy/jsort; curated open System One ecosystem catalog; rupeshpoojary9/awesome-open-system-one ≠ AnotiaWang/awesome-jev; arXiv paper radar with Jev relevance scoring; ranking ≠ calibration / 0.5 still soft; fail-open failed evals not marked seen **Hourly 2340:** train calibrated ~27M from scratch; typed Q→prob dist / one forward pass / no LLM decode; hyusi2003/MiniSystemOne ≠ Colvin0315/MiniSystemOne; description-only stub / size 5; ESCI hard probe fails four of six; jev_bool ECE 0.242 inversion 0.255; do not re-fold §60 six-gates as new; jobbyjev one-request-per-company from batch-size result; find/design/evaluate TypeSafe Jev decision loops; karanb192/jev-architect ≠ samtay32/jev-system-architect; Jairik/jev-distiller size 1; distill-Jev UI stub / do not distill Jev as teacher of record; post-launch scored use-case map / Jev self-scores then human curation; licensedsaucer9-web/jev-opportunities; Jev-inize a use case into classifier/router; gavinHuang/jevinize → simple-jev not TypeSafe; featherless-ai/simple-jev; compare saved decisions / same label can still change the branch; VihaanAgarwal/jev-diff ≠ Saik0s/diffusiongemma-jev-macos; not tested with a live Jev API key; constrained logprob + temp/Platt ≠ Noul; OpenJevPro pastes openjev-sglang JevBench as own; zhangcy122/OpenJevPro ≠ IamBusy/OpenJev ≠ ekzhang/openjev-sglang; PolyForm Noncommercial; SmolLM-135M / sub-70ms / 0 output tokens; demo P(True) 0.5052 / Choice conf 0.2872 / Score conf 0.0055; README claims MIT / GitHub license null / no LICENSE file; patelvishwa112/jev-system-one-rlcd ≠ arnabgho/rlcd-lite ≠ blackwood-rlcd; source-backed Awesome Jev radar / 306+ commit-pinned; logicrw/awesome-jev-projects ≠ AnotiaWang/awesome-jev ≠ yibie/awesome-jev ≠ cobanov/awesome-jev ≠ rupeshpoojary9/awesome-open-system-one; auto GitHub sync / Issue-only submissions; hashed n-gram encoder / rival-aware attention; olanotolu/jevbetter vs jevlike starter; synthetic hard menus top-1 0.916 vs 0.873 / ECE 0.0182 vs 0.0367 / 40 vs 4608 menus/sec; shuffled-context control 0.335 | `references/judgment-class.md` Apply 0042 (`notes.md` §105): structured probability readouts; distribution > argmax; Noul 0.5 midpoint; score is expectation not integer; bare HTTP not SDK; Arohtea/jev-readout; Jev-style Choice/Score/Noul from ordinary models; optional DSH plugin; schema-valid ≠ calibrated; gulagala001/jevify ≠ Mintzs/jevify; Laya RLCD benchmark; 40.3% below constant-answer; open-weight measurement; mourad-ghafiri/laya-rlcd-benchmark ≠ yibie/laya-jev-lab; cheap fail-open semantic edge; second signal not sole; FastLoopError catch; SupremeDreamZ/jev-fastloop ≠ jev-ultrafast; asking more questions in one call; 0.980 at every N; nearly not fully deterministic; TheWebDevel/jev-fanout; Qwen3-VL perception + Jev decisions train RL; 0 model calls at deployment; VLM alone 1.7 vs +Jev 4.4; harneet2512/reflexrl ≠ khordoo/jev-reflex-autonomy-lab; independent Jev API vs Laya; cascade 0.60 matches 78% at 1.8×; noul facts not judgements; yibie/laya-jev-lab ≠ dairui1/jev-lab ≠ BrendanH18/jev-lab; GLiNER vs GLiFormer vs Laya vs Jev; extractors ≠ decision engines; Laya dict-instructions collapse 58.3%; umstek/zero-shot-ie-bench; decisions-per-minute & cost; 204 moves vs 73; throughput not intelligence; angelgalvisc/snake-arena-jev-vs-llms ≠ vtrivedy/jev-plays-games; behavioral contracts; pin expectations eval upgrades; raw 0.94 is not a release; sathariels/jevcheck ≠ dayhaysoos/jevals ≠ SivletLabs/jev-eval; evidence-linked dependency upgrade; Jev never generates filenames; no_direct_evidence ≠ safe to merge; GaneshVG18/upgrade-radar ≠ LYchoon/paper-radar-jev; discography theme/mood/complexity; five atomic questions one call; lirantal/discoprint / Apply 0145 (`notes.md` §106): Turn any open LLM into System-One Jev; description-only stub / size 0; Exu is a toolkit, not a method; typed Q → probability dists; JSON parse of generated text ≠ Noul; classifier.dev fast tier 84.8 is Jev behind its own API; hard budget filter before Jev; Jev judges the next state, XState enforces transitions; catalog gravity; ★339 live REST; query-side encoders, not a Jev replica; transformers.js AutoModel cannot load this graph; This Space contains no benchmark result yet; do not reopen or amend PR #23. / Apply 0243: Benchmark-driven Jev router and judge; cheap alone is not success; Jev does not write, sum prices, or claim accuracy %; Sol 94.2 / Luna 83.9 / Jev path 89.7; 19.2% Sol / 62.3% cost save / 4.5pp miss of 2pp non-inferiority; p50 latency worse than Sol due to routing overhead; erendikmenn/jev-llm-router-benchmark ≠ jev-rag-benchmark ≠ ryantsai/jev-llm-router; Express + node:sqlite; mock and Jev decision engines; previous_ticket_count >= 3 is code; MIN_CONFIDENCE 0.6 still soft; substring false positives; aesaganda/jev-ticket-router ≠ SarathChandraBellam/jev-vs-llm-ticket-router; Universal Figure & Diagram Router; confidence ≥ 0.85 hard-gate is theater; generative AI banned from scientific plots; six visual branches; human-labeled (state, question, label); 166,054 rows / 22 configs; soft_label for human uncertainty; Praveenrajus/jev-bench ≠ fstandhartinger/jevbench; ternary bonsai System One GGUF; Hub does not ship weights; 100/100 easy T/F is not Harbor; label_mass ≠ correctness; stock llama.cpp Q2_0 silently gibberish; NicolaiMTLassen/open-bonzi-jev ≠ NicolaiLassen; transformers.js DeBERTa ONNX; temperature 1.05; AutoModel from_pretrained works; onnx-community/open-jev-deberta-v3-large-ONNX ≠ system-one-qwen3.5-4b-scorer-ONNX; 107★ densify; GH 151M vs README 149.6M; PR #1 now closed unmerged; do not re-fold §71 claim-audit as a beat; typed decisions, RLCD, confidence-gated routing; structured ≠ correct; mock not live API; 26 tests; wjdjdakf17/jev-study ≠ baekenough/jev-study / Apply 0345 (`notes.md` §108): **Open reproduction densifies / measurement densifies PRIMARY** (bonzi-27b-v2 / ternary-8b / 27b-v1 GGUF family densify; WANLI-256 74.6% / 65.2% / 71.1% *theirs*; Bonsai 1 27B Q1_0 runs on stock llama.cpp; ternary still needs PrismML fork; hf:heman10x/openJev-verdict-2.0 twin tokenizer-only; OpenJev Vision image classification + uncertainty; CLEVR-4 held-out joint 0%; hfdataset:IamBusy/OpenJev-Vision-Research-v0.1 12,832; 294,912 derived targets not independent samples; Laya multilingual ONNX WebGPU typed-decisions port; 63/63 selected answers / 5.1e-4 CPU / 1.2e-2 WebGPU; UpHash-Network/mini-jev is yuki-oshio transfer; jev-injection-bench 11,900 labelled prompts; Jev best ranking / Haiku better ECE 0.021 vs 0.058; 0.5–0.9 band is where Jev's numbers do not mean what they say; Prompt wording moves panic 28%; manojlds/jev-dspy-bench ≠ dspachos/jev-dspy ≠ jmanhype/jev-dspy-lab; Jev agreement is similarity, never ground truth; no aggregate quality grade or merge gate; AbstentionBench-on-Jev rank 1 of 20 vs 2025 field; question-asymmetry; forward-looking 0.465 never extreme; openkev calibration layer not a runtime; ECE vs coverage independent; select_threshold returns inf; escalation catches uncertainty not ignorance; misakaikato/openkev ≠ jaredpalmer/kev; pdf-race Docling→Jev vs Gemini; parser owns the wall clock; 12/12 tie is a tie; titles selected not generated; flopcheck 16 calibrated tweet judgments; mechanical tells in code; ZeroX-01/jev-atlas ≠ Zaious/jev-capability-atlas ≠ gorock007/jev-atlas; Laya calibration lab Gradio MCP; T never changes argmax; confidence ≠ top-label p; easy probe set refused; 40–48 rows too small to ship T; do not reopen or amend PR #23 or #24 or #25. Apply 0439 (`notes.md` §109): Gemma-4 26B-A4B jevify classification+calibration; LoRA adapter twin not independent eval; Gemma-4 E4B jevify; E4B LoRA stub card; kushalpatil/jevify-gemma4 ≠ Mintzs/jevify ≠ gulagala001/jevify ≠ uspraveen/Jevify; GH kushalpatil07/jevify 404; PAWS 0.580/ece 0.288 is the weak cell; smaller E4B slightly better OOD ECE than 26B-A4B; Hub jevify merged LoRA ships weights; bonzi Bonsai-8B v1 GGUF densify; Bonsai-1.7B v1; Bonsai-4B v1; WANLI-256 64.5% / 60.2% / 52.0% *theirs*; rank #4 / #5 / #6 of 6; JulesHuisman/jev-eval scaffolding / README SHA c356a584 (was empty e69de29b); JulesHuisman/jev-eval ≠ SivletLabs/jev-eval ≠ willkelly/jev-evaluation ≠ 4esv ≠ xxkuboxx ≠ onlyoneaman ≠ dayhaysoos/jevals; 7 bands 6/10 vs 40 bands 0/10; source receipts + confidence slider re-policy without re-inference; 32/32 synthetic is smoke not production; classify HF datasets across typed semantic dimensions; roadus2 watch misspelling; lock roadius2/ultra_laya; ultra_laya REVIEW defects; default branch claude/laya-jev-review-gg5ppo; XNLI EN 88.3% ECE 0.032 → RU 77.3% ECE 0.096; Δ −11.0 pp [−14.2,−7.8]; ECE +0.063; MASSIVE no detectable difference at n=600; confidence is function of p_max (r=1.000); pointer-not-generator 400 human-authored responses; proposed ≠ authorized; FewRel 160: Jev 85.0% vs lexical 13.125%; gated 100% (95/95) coverage 59.375%; J++ composable semantic computation language; judge-jev 0.5 still soft; 947 repos scored; A 273 / B 302 / C 372; LLM rubric ≠ benches; No benchmark winner is claimed; phishing: naive 62.6% vs regex 91.8%; 5-atomic + LR 95.0% *theirs*; AITuber tension ±15; README npm global; repo is Rust; git-confess code owns counting/blame/ratio; httpx exhibit 11% (13/119) *theirs*; 90d trend +12.40% vs random +12.75% vs BH +41.71%; 5m win rate 25%; Awesomejev 656 entries / 38,160 stars; tracker likes 64 (+4) lastModified UNCHANGED; Laya present; Blackwood ABSENT; Archer still promised_not_landed; do not reopen or amend PR #23/#24/#25/#26. Apply 0541 (`notes.md` §110): Blackwood tracker ABSENT; likes 2 gated manual; r = c - p_a; ECE 0.021; acc 0.807 vs warmup 0.746; calibration beyond ~500 tokens unmeasured; Independent primitive; 11.57s vs 54.10s · 4.67× · 120/128 *theirs*; default path is pretrained Gemma probs not trained RLCD head; GH Meanblock 404; lock leesk212/JEV-CPU; softmax over letter slots ≠ Noul; WANLI 0.741 vs openjev v2 0.77 *theirs*; 3-way NLI ≠ Noul; priority 0.464 = majority floor; banking77 contaminated; raw margins not probabilities; GH jev-haiku-benchmarking 404; do not distill Jev as teacher of record (they distilled Haiku); “0.9 is not one number”; ranking ≠ calibration; banking77 0.8–0.9 stated 0.86 actual 0.73 over-confident *theirs*; ≠ Praveenrajus/jev-bench ≠ fstandhartinger/jevbench; $0.0000153–$0.0000226 vs circulating $0.0004 (~20×); Score is 0..n-1 expectation not 0–1; Noul has no confidence field; TCP floor 198.8 ms; type reliability is not a reason to choose Jev (json_schema 5/5); gateway tax not one number; Function-only 5/8 vs hybrid 8/8; 4/8 without Jev; 8 designed cases not conversion lift; ≠ RadRebelSam/awesome-jev; 200-row pilot Jev 86.5% 173/200 vs Gemini Flash-Lite 86.0% 172/200 vs Pro 87.0% 174/200 *theirs*; not a ranking; NLI Tetris argmax P(entail)−P(contradict); 情緒測謊器; 1q 396ms / 30q 567ms; ±0.03; 33q $0.000045 vs Gemini ~5× slower ~60× cost *theirs*; ≠ realZachi/jevtest; 8-example Jev vs GPT-5.6 Sol ~64× cost 5.4× latency *theirs*; synthetic; no inference; ≠ JevBench v1.2 §78; Judged 3317 / listed 2560; Jev judges, code applies policy; catalog ≠ endorsement; APA “microsecond policy / zero hallucination” overclaim; Client-side quiz; pointer from held docs; scanned-PDF warn; CSP only api.typesafe.ai; Jev judges / agent reasons / user decides; selecting an option is not permission to implement; degraded fallback; pattern exact, judgement must clear floor; no matching pattern → no model call; not a correctness oracle; $0.00022 vs chat $0.00306 *theirs*; Spec vs artifact remainder; treating 0.85 as 85% / minProbability hard-gate as Harbor; VERIFY acquires discriminating evidence, never same-pool confidence-only rescoring; fast/full/max are ceilings not sizes; Solar writes, Jev chooses NEXT ACTION; SemIf 2186★ (+20 vs §109 2166); jevlike 1038★ (+7 vs 1031); TypeAR 14★ flat; AnotiaWang 96★ (+1 vs 95); yibie/awesome-jev 490★; Laya likes 802 (was 783); tracker likes 64 flat, lastModified UNCHANGED; do not reopen or amend PR #23/#24/#25/#26/#27. Apply 0743 (`notes.md` §113): Hourly 0743 uniqueness lock: Heman10x-NGU/Verdict-open-jev ≠ Heman10x-NGU/openJev-verdict-2.0; TF-IDF + LogReg ECE 0.0207 vs Jev 0.1440; Verdict-open-jev 48.07% vs Jev 90.80%; abstention combined recall 10.00%; p50 35.58 ms; K=25 (maximum capacity) 72.00%; 0.85 coverage 84.60% selective risk 1.18%; 26.1× faster than standard Qwen JSON generation; Jevify 90.0% / 167 ms CUDA graphs disabled; Finding 1: Brier on stated confidence alone is a trap; grpo_rlcr 0.78 / ECE 0.084; reliability 0.007 but resolution 0.000; 27 900 schema-driven decisions; 13 600 / 13 600 questions; candidate mass min 0.99999624; 22 configs · 166,054 rows · 4 calibration-gold; sha a39eba3f; Student B MAE 0.148 / Pearson 0.836 / 86.0%; pngwn/open-jev-laya-bench README 404; sha 9f69c742 likes 2; HDFS 0.9933 (745/750) / retain 0.0084; BGL ERROR/FATAL protection 1.0000; 2,479 / 2,500 HDFS uncertain; cache hit 0.9648 (2412/2500); $0.153936 estimated; E2 recomputes from saved probabilities; Space sha eda59e0a; MASSIVE English 0.783 / Khmer 0.033 / Hindi 0.133; 40–48 rows too small to ship T; T never changes argmax; siren2345/jev-single-decode-transformers ≠ siren2345/jev-single-decode; Split Transformers experiment from llama.cpp runtime; tanayvasishtha/jev-lab ≠ dairui1/jev-lab ≠ BrendanH18/jev-lab ≠ yibie/laya-jev-lab; Four experiments stress-testing TypeSafe's Jev: calibration, bundle bias, label bias, and ensembling; second pass must be $0.00 from cache; The pages never call Jev; Gemma 4 31B 77.0% / Jev 1.13.0 61.4% / Laya 322M 0.0%; restriction state 95.0% against 84.4%; None of the systems are particularly good at knowing when to stop and ask; They skip the question and call a tool directly; 100% schema pass; six-field joint 48.8% vs 72.8%; ywchiu/jev_benchmark ≠ Running-Dolphins/jev-bench ≠ Praveenrajus/jev-bench; ACT / REVIEW / FALLBACK; A provider failure, timeout, malformed output, or missing answer is **not** a policy outcome; confidence is descriptive provider output, not a substitute for probability; Quality denominators include only valid scored answers; an exact halfway tie chooses the lower level; aiwithenoch/Jev-Skill ≠ simplosophy/jev-skill ≠ laguagu/jev-skills; The local path does not claim to turn a smaller checkpoint into Jev; Low support becomes decision: "review"; MIT-0 SPDX NOASSERTION; current-llm; 结构兼容,不是 Jev 模型能力; altryne/jevify ≠ Mintzs/jevify ≠ gulagala001/jevify ≠ uspraveen/Jevify; Find where Jev belongs. Design the questions. Measure the difference; TypeAR-AI/TypeAR 301 → TypeLLM/TypeLLM; TypeLLM/TypeLLM 16★; SemIf 2241★ (+34 vs §111 2207); jevlike 1051★ (+8 vs 1043); AnotiaWang 98★ (+1 vs 97); yibie/awesome-jev 525★ (+19 vs 506); Laya likes 864 (was 822); tracker likes 67 (+3 vs 64); lastModified UNCHANGED `2026-09-20T04:29:16.000Z`; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#32/#33; notes.md §113 Apply 0646 (`notes.md` §111): Calibration is not alpha; NO CURRENT ALPHA CANDIDATE; ΔR² approximately +0.00084; Brier 0.2131387; ECE 0.0421875; Adding Jev probability to deterministic volatility improved Brier by only 1.4058e-05; default 0.5 keeps zero non pinned; keepResult median 0.14 to 0.17; keepCall median 0.28 to 0.35; usable range is about 0.10 to 0.25; 7.8% to 57.9%; judges results it never sees; task-finish eval not built yet; $0.002 per compaction; slavadubrov/sgr-judge-bench ≠ slavadubrov/jev-judge-bench; Jev 108/120 $0.083 0.34 s; Luna SGR 114/120; paired Jev accuracy-difference intervals include zero; not evidence of equivalence; GLM SGR 26/120 93 format failures; Terra-planned Jev hybrid 55/120; rule-based by default, optionally Jev-backed; empty README; missing key cannot break the experience; prefill plus exactly one decode; softmax over A/B/C ≠ Noul; BBQ 9,053/10,000 (90.53%); ECE 0.0890; Mean confidence 0.9943; overconfident; score and noul not implemented; DGUI 12 rows (was 6); INSTRUCT 119 rows likes 2; encode the state once, decide everything in parallel; 0.740 accuracy against a 0.508 majority; ECE 0.047; fine-tune's advantage ends where its 384-token training data does; jasonkneen/open-jev ≠ pngwn/open-jev; same sha d41dc3cd; Space does not call Jev; recomputes routing from saved probabilities; 200-case Jev 97.0% / 100.0% / 95.0% / MAE 9.22; synthetic repository benchmark; Jev evaluations are advisory; YehuiTang0316/jev-nlgrep ≠ Bentlybro/jevgrep ≠ can1357/jegrep ≠ uehaj/jev-semgrep; default threshold 0.8 still soft; 40-line windows cannot prove whole function; token-native sequential start/end Choice; Gemini/Haiku stubs not configured yet; handful of hand-written examples, not a benchmark; Jev judged exactly what it was given; laguagu/jev-skills ≠ laguagu/jev-evidence-lab ≠ Pleo2/awesome-jev-agent-skills; contract_passed is not a claim of guaranteed factual truth; Wilson lower bound 0.85 floor; fixture mode no savings claim; SemIf 2207★ (+21 vs §110 2186); jevlike 1043★ (+5 vs 1038); TypeAR 15★ (+1 vs 14); AnotiaWang 97★ (+1 vs 96); yibie/awesome-jev 506★ (+16 vs 490); Laya likes 822 (was 802); tracker likes 64 flat, lastModified UNCHANGED; do not reopen or amend PR #23/#24/#25/#26/#27/#28.) | | Open weights vs constrained decoding vs proprietary API | Three open paths: encoder open-jev / AR constrained decode (TypeAR + pcdServer; decision-token LoRA; packed one-forward logprob on open LLMs; **CUDA/PyTorch local replica** jevify — uncalibrated likelihoods ≠ Noul) / trained decision-only (Laya + ONNX port, Nimble, kev, **blackwood-rlcd** multimodal now, Archer Watch still Watch). **Domain LoRA specialist on independent gold** (not a Jev teacher-copy): train when downstream reads p; few-shot hosted when only argmax. Local `/v1/systemone` surfaces: jev-local (stub until `hf`), kev (trained pointer), von (§49 Needle SAN snapshot ≠ this-pass 395M / n=78; not a replica), **jeff** (GLiFormer-400M encoder, typesafe-sdk drop-in — not a Jev replica), **local-jev** (ModernBERT approximation — not equivalence). Laya ONNX: Mattepiu port vs **gqgs** complete browser int8 (distinct). Loopback **gateway** (sysone) routes hosted + local; does not run weights. Softmax over allowed tokens ≠ Noul. **jevmlx:** MLX one-pass schema→JSON+probs (Apple Silicon replica economics; not a Jev replica). **openvons:** independent open-Jev class (Apache-2.0 code; GitHub SPDX NOASSERTION); wire-compat `/v1/systemone`; NOTA + execute/confirm/reject. **OpenJev** `/v1/decide` ≠ TypeSafe. **semif-serve:** SemIf runoff wire (MIT pyproject / GitHub SPDX null). **grande** Rust/WebGPU `/v1/systemone` (license null; softmax ≠ Noul until T). **laya-jolt** Clojure Apache-2.0 byte-parity Laya. **JEV-CPU** CPU SemIf. **local-jev** ONNX NLI approximation (confidence omitted). **chakuho:** 1-token logprob constrained-AR endpoint (uncalibrated; coverage is format-mass). **jevinf:** Jev-kind engine + wire (not a replica). **laya-multilingual:** mmBERT-base for non-English; route by script. **githubnext/localjev:** Bun Chat Completions bridge; prompted JSON + entropy confidence; **≠** kunchenguid/local-jev; **≠** razorback16 logits. **NandhaKishorM/laya:** PyPI `laya` + Router over the three Hub ckpts; packaging ≠ new species; Jev still leads >20 options / soft-acc / raw ECE. **@airesearch12 census:** list ≠ rank; GLiNER2/routers counted as openjevs are a class-boundary. **JevBench v1.2:** same class table includes Luna/Gemini/DeepSeek/Qwen3.8; OpenJev on board = razorback16 DiffusionGemma ≠ IamBusy; SemIf formerly OpenJev. **≠** GestaltLabs/Jeff-1 Qwen3-4B LoRA replica. **SIGNAL §94:** unofficial JA ModernBERT cross-encoder (format_version modernbert-jev/1; unofficial ≠ TypeSafe; LFM default ≠ ModernBERT backend); Nemotron ≠ TypeSafe Jev / not a calibrated replacement; djev-dev complements djev-spark (native image / images as Choice options); Laya essay numbers *theirs* / Router/OOD confidence. **Hourly 1541:** llama-jev llama.cpp replica; softmax ≠ Noul; **≠** TypeSafe **≠** pcdServer. **Hourly 1740:** Archer-arch fidelity; kev family OOD 0.76–0.77 vs Jev 0.86; block-causal isolation; pointer/readout CE-trained; /v1/systemone drop-in; replica honesty. **Hourly 1843:** NAR local drop-in; wfzyx/von late-catch HIGH; open replica economics / latency vs closed Jev; competing NAR claims / replica honesty; variable-N option scoring as the trainable object; dynamic candidate bags not fixed label sets. candidate bags not fixed label sets; zwliJay/jev-forge ≠ NanoJev. **Hourly 1943:** on-device Laya CoreML ANE; ~5 ms P50 short decisions; 189/189 FP16 checkpoint parity; 10× not achieved; softmax over allowed tokens ≠ Noul; question-first cache **Hourly 2041:** resume-screening bias audit methodology; Plan/PRD panel → code-owned pass|review|block; cost-aware multi-model routing/escalation; frozen-protocol zero-shot bench; context-window admission control; typed decision control plane; receipt ≠ authorization; live 15-dim typed rubric re-score per pause; adversarial pre-registered Jev eval; confidence does not track ignorance; provider-neutral Elixir/BEAM Noul/Choice/Score SDK **Hourly 2145:** question-linting of Jev questions themselves; nine jaggedness rules, no API key, no labelled data; static lint ≠ measured separation; yodablocks/jevq ≠ tenbin ≠ JevLint ≠ commitjev; open-weights Laya as class exemplar (binding); Nx/Bumblebee runtime; host chooses backend; ChristianAlexander/laya_ex ≠ system_one_sdk ≠ dannote/jev ≠ NandhaKishorM/laya; on-chain/edge Laya deploy; parity_verified stays false; model output never grants Tx; humandebri/IC-Laya ≠ laya_ex; auditable weekend replica; Jev outputs never used for training; unpaired 0.577 vs 0.727; agilabs-ai/jev48 ≠ JevBench ≠ Mapika/decider; adversarial dual-judge / framing attack surface; comparative framing is the usable judgment; prior injection crowds out evidence; copyleftdev/ember ≠ ember.js; Laya specialist fine-tune pipeline; training still GPU-pending; PIXELZX0/XERON ≠ convaiinnovations/laya; Hub Laya replica drop; daliborsb/laya ≠ convaiinnovations/laya ≠ NandhaKishorM/laya; System One student distillation corpus; gold is programmatic; teacher is closed-API clone; MagaBitmex/jev-4b-distill-data ≠ missing student checkpoint; non-LLM VIN System One; planning depth not chat; lewislululu/jevon ≠ douglance/jevon; source-bound evidence checks; local quote mismatch needs no API; exit 0 ≠ claim truth; WaynezProg/jev-kit ≠ jonathanavis96/jev-kit (Airlock) ≠ jev-use ≠ jev-mcp **Hourly 2246:** independent System One evidence catalog; 19 reviewed records; scores not one leaderboard; no external record currently reproduced; TokenTrim no-Jev matched hybrid 62.4%; reachjalil/system-one-bench ≠ mallahyari/system-one-benchmark; 21 tasks · 134 items · 208 questions; scenes from public GitHub contracts, not production logs; SivletLabs/jev-eval ≠ willkelly/jev-evaluation ≠ 4esv/jev-eval ≠ xxkuboxx/jev-eval ≠ onlyoneaman/jev-eval ≠ dayhaysoos/jevals; option isolation (sibling-blind); permutation-equivariant; Hub OWNER not published; nafisazizir/hev ≠ jaredpalmer/kev; frozen local LLM logits, no trained decision head; residual-head 9,222-param decreased 73/96→67/96; confidence = 1−normalized entropy, not P(correct); yuki-oshio/mini-jev ≠ r-ms/mini-jev; Jev classifier as autoregressive next-token predictor; ChatJev-style soundness theater; erik-dunteman/ChatJev ≠ dannote/jev ≠ jev-gpt; calibrated decision head × AlphaProof value head; implementation-layer isomorphism, semantic difference; timeout = censoring; do not launder Noul as proof; parallel rank-prediction vs serial selection; independent questions can conflict; zzzzzec/jevsort ≠ keltokhy/jsort; curated open System One ecosystem catalog; rupeshpoojary9/awesome-open-system-one ≠ AnotiaWang/awesome-jev; arXiv paper radar with Jev relevance scoring; ranking ≠ calibration / 0.5 still soft; fail-open failed evals not marked seen **Hourly 2340:** train calibrated ~27M from scratch; typed Q→prob dist / one forward pass / no LLM decode; hyusi2003/MiniSystemOne ≠ Colvin0315/MiniSystemOne; description-only stub / size 5; ESCI hard probe fails four of six; jev_bool ECE 0.242 inversion 0.255; do not re-fold §60 six-gates as new; jobbyjev one-request-per-company from batch-size result; find/design/evaluate TypeSafe Jev decision loops; karanb192/jev-architect ≠ samtay32/jev-system-architect; Jairik/jev-distiller size 1; distill-Jev UI stub / do not distill Jev as teacher of record; post-launch scored use-case map / Jev self-scores then human curation; licensedsaucer9-web/jev-opportunities; Jev-inize a use case into classifier/router; gavinHuang/jevinize → simple-jev not TypeSafe; featherless-ai/simple-jev; compare saved decisions / same label can still change the branch; VihaanAgarwal/jev-diff ≠ Saik0s/diffusiongemma-jev-macos; not tested with a live Jev API key; constrained logprob + temp/Platt ≠ Noul; OpenJevPro pastes openjev-sglang JevBench as own; zhangcy122/OpenJevPro ≠ IamBusy/OpenJev ≠ ekzhang/openjev-sglang; PolyForm Noncommercial; SmolLM-135M / sub-70ms / 0 output tokens; demo P(True) 0.5052 / Choice conf 0.2872 / Score conf 0.0055; README claims MIT / GitHub license null / no LICENSE file; patelvishwa112/jev-system-one-rlcd ≠ arnabgho/rlcd-lite ≠ blackwood-rlcd; source-backed Awesome Jev radar / 306+ commit-pinned; logicrw/awesome-jev-projects ≠ AnotiaWang/awesome-jev ≠ yibie/awesome-jev ≠ cobanov/awesome-jev ≠ rupeshpoojary9/awesome-open-system-one; auto GitHub sync / Issue-only submissions; hashed n-gram encoder / rival-aware attention; olanotolu/jevbetter vs jevlike starter; synthetic hard menus top-1 0.916 vs 0.873 / ECE 0.0182 vs 0.0367 / 40 vs 4608 menus/sec; shuffled-context control 0.335 | `references/judgment-class.md` (when-to-use table) Apply 0042 (`notes.md` §105): structured probability readouts; distribution > argmax; Noul 0.5 midpoint; score is expectation not integer; bare HTTP not SDK; Arohtea/jev-readout; Jev-style Choice/Score/Noul from ordinary models; optional DSH plugin; schema-valid ≠ calibrated; gulagala001/jevify ≠ Mintzs/jevify; Laya RLCD benchmark; 40.3% below constant-answer; open-weight measurement; mourad-ghafiri/laya-rlcd-benchmark ≠ yibie/laya-jev-lab; cheap fail-open semantic edge; second signal not sole; FastLoopError catch; SupremeDreamZ/jev-fastloop ≠ jev-ultrafast; asking more questions in one call; 0.980 at every N; nearly not fully deterministic; TheWebDevel/jev-fanout; Qwen3-VL perception + Jev decisions train RL; 0 model calls at deployment; VLM alone 1.7 vs +Jev 4.4; harneet2512/reflexrl ≠ khordoo/jev-reflex-autonomy-lab; independent Jev API vs Laya; cascade 0.60 matches 78% at 1.8×; noul facts not judgements; yibie/laya-jev-lab ≠ dairui1/jev-lab ≠ BrendanH18/jev-lab; GLiNER vs GLiFormer vs Laya vs Jev; extractors ≠ decision engines; Laya dict-instructions collapse 58.3%; umstek/zero-shot-ie-bench; decisions-per-minute & cost; 204 moves vs 73; throughput not intelligence; angelgalvisc/snake-arena-jev-vs-llms ≠ vtrivedy/jev-plays-games; behavioral contracts; pin expectations eval upgrades; raw 0.94 is not a release; sathariels/jevcheck ≠ dayhaysoos/jevals ≠ SivletLabs/jev-eval; evidence-linked dependency upgrade; Jev never generates filenames; no_direct_evidence ≠ safe to merge; GaneshVG18/upgrade-radar ≠ LYchoon/paper-radar-jev; discography theme/mood/complexity; five atomic questions one call; lirantal/discoprint / Apply 0145 (`notes.md` §106): Turn any open LLM into System-One Jev; description-only stub / size 0; Exu is a toolkit, not a method; typed Q → probability dists; JSON parse of generated text ≠ Noul; classifier.dev fast tier 84.8 is Jev behind its own API; hard budget filter before Jev; Jev judges the next state, XState enforces transitions; catalog gravity; ★339 live REST; query-side encoders, not a Jev replica; transformers.js AutoModel cannot load this graph; This Space contains no benchmark result yet; do not reopen or amend PR #23. / Apply 0243: Benchmark-driven Jev router and judge; cheap alone is not success; Jev does not write, sum prices, or claim accuracy %; Sol 94.2 / Luna 83.9 / Jev path 89.7; 19.2% Sol / 62.3% cost save / 4.5pp miss of 2pp non-inferiority; p50 latency worse than Sol due to routing overhead; erendikmenn/jev-llm-router-benchmark ≠ jev-rag-benchmark ≠ ryantsai/jev-llm-router; Express + node:sqlite; mock and Jev decision engines; previous_ticket_count >= 3 is code; MIN_CONFIDENCE 0.6 still soft; substring false positives; aesaganda/jev-ticket-router ≠ SarathChandraBellam/jev-vs-llm-ticket-router; Universal Figure & Diagram Router; confidence ≥ 0.85 hard-gate is theater; generative AI banned from scientific plots; six visual branches; human-labeled (state, question, label); 166,054 rows / 22 configs; soft_label for human uncertainty; Praveenrajus/jev-bench ≠ fstandhartinger/jevbench; ternary bonsai System One GGUF; Hub does not ship weights; 100/100 easy T/F is not Harbor; label_mass ≠ correctness; stock llama.cpp Q2_0 silently gibberish; NicolaiMTLassen/open-bonzi-jev ≠ NicolaiLassen; transformers.js DeBERTa ONNX; temperature 1.05; AutoModel from_pretrained works; onnx-community/open-jev-deberta-v3-large-ONNX ≠ system-one-qwen3.5-4b-scorer-ONNX; 107★ densify; GH 151M vs README 149.6M; PR #1 now closed unmerged; do not re-fold §71 claim-audit as a beat; typed decisions, RLCD, confidence-gated routing; structured ≠ correct; mock not live API; 26 tests; wjdjdakf17/jev-study ≠ baekenough/jev-study / Apply 0345 (`notes.md` §108): **Open reproduction densifies / measurement densifies PRIMARY** (bonzi-27b-v2 / ternary-8b / 27b-v1 GGUF family densify; WANLI-256 74.6% / 65.2% / 71.1% *theirs*; Bonsai 1 27B Q1_0 runs on stock llama.cpp; ternary still needs PrismML fork; hf:heman10x/openJev-verdict-2.0 twin tokenizer-only; OpenJev Vision image classification + uncertainty; CLEVR-4 held-out joint 0%; hfdataset:IamBusy/OpenJev-Vision-Research-v0.1 12,832; 294,912 derived targets not independent samples; Laya multilingual ONNX WebGPU typed-decisions port; 63/63 selected answers / 5.1e-4 CPU / 1.2e-2 WebGPU; UpHash-Network/mini-jev is yuki-oshio transfer; jev-injection-bench 11,900 labelled prompts; Jev best ranking / Haiku better ECE 0.021 vs 0.058; 0.5–0.9 band is where Jev's numbers do not mean what they say; Prompt wording moves panic 28%; manojlds/jev-dspy-bench ≠ dspachos/jev-dspy ≠ jmanhype/jev-dspy-lab; Jev agreement is similarity, never ground truth; no aggregate quality grade or merge gate; AbstentionBench-on-Jev rank 1 of 20 vs 2025 field; question-asymmetry; forward-looking 0.465 never extreme; openkev calibration layer not a runtime; ECE vs coverage independent; select_threshold returns inf; escalation catches uncertainty not ignorance; misakaikato/openkev ≠ jaredpalmer/kev; pdf-race Docling→Jev vs Gemini; parser owns the wall clock; 12/12 tie is a tie; titles selected not generated; flopcheck 16 calibrated tweet judgments; mechanical tells in code; ZeroX-01/jev-atlas ≠ Zaious/jev-capability-atlas ≠ gorock007/jev-atlas; Laya calibration lab Gradio MCP; T never changes argmax; confidence ≠ top-label p; easy probe set refused; 40–48 rows too small to ship T; do not reopen or amend PR #23 or #24 or #25. Apply 0439 (`notes.md` §109): Gemma-4 26B-A4B jevify classification+calibration; LoRA adapter twin not independent eval; Gemma-4 E4B jevify; E4B LoRA stub card; kushalpatil/jevify-gemma4 ≠ Mintzs/jevify ≠ gulagala001/jevify ≠ uspraveen/Jevify; GH kushalpatil07/jevify 404; PAWS 0.580/ece 0.288 is the weak cell; smaller E4B slightly better OOD ECE than 26B-A4B; Hub jevify merged LoRA ships weights; bonzi Bonsai-8B v1 GGUF densify; Bonsai-1.7B v1; Bonsai-4B v1; WANLI-256 64.5% / 60.2% / 52.0% *theirs*; rank #4 / #5 / #6 of 6; JulesHuisman/jev-eval scaffolding / README SHA c356a584 (was empty e69de29b); JulesHuisman/jev-eval ≠ SivletLabs/jev-eval ≠ willkelly/jev-evaluation ≠ 4esv ≠ xxkuboxx ≠ onlyoneaman ≠ dayhaysoos/jevals; 7 bands 6/10 vs 40 bands 0/10; source receipts + confidence slider re-policy without re-inference; 32/32 synthetic is smoke not production; classify HF datasets across typed semantic dimensions; roadus2 watch misspelling; lock roadius2/ultra_laya; ultra_laya REVIEW defects; default branch claude/laya-jev-review-gg5ppo; XNLI EN 88.3% ECE 0.032 → RU 77.3% ECE 0.096; Δ −11.0 pp [−14.2,−7.8]; ECE +0.063; MASSIVE no detectable difference at n=600; confidence is function of p_max (r=1.000); pointer-not-generator 400 human-authored responses; proposed ≠ authorized; FewRel 160: Jev 85.0% vs lexical 13.125%; gated 100% (95/95) coverage 59.375%; J++ composable semantic computation language; judge-jev 0.5 still soft; 947 repos scored; A 273 / B 302 / C 372; LLM rubric ≠ benches; No benchmark winner is claimed; phishing: naive 62.6% vs regex 91.8%; 5-atomic + LR 95.0% *theirs*; AITuber tension ±15; README npm global; repo is Rust; git-confess code owns counting/blame/ratio; httpx exhibit 11% (13/119) *theirs*; 90d trend +12.40% vs random +12.75% vs BH +41.71%; 5m win rate 25%; Awesomejev 656 entries / 38,160 stars; tracker likes 64 (+4) lastModified UNCHANGED; Laya present; Blackwood ABSENT; Archer still promised_not_landed; do not reopen or amend PR #23/#24/#25/#26. Apply 0541 (`notes.md` §110): Blackwood tracker ABSENT; likes 2 gated manual; r = c - p_a; ECE 0.021; acc 0.807 vs warmup 0.746; calibration beyond ~500 tokens unmeasured; Independent primitive; 11.57s vs 54.10s · 4.67× · 120/128 *theirs*; default path is pretrained Gemma probs not trained RLCD head; GH Meanblock 404; lock leesk212/JEV-CPU; softmax over letter slots ≠ Noul; WANLI 0.741 vs openjev v2 0.77 *theirs*; 3-way NLI ≠ Noul; priority 0.464 = majority floor; banking77 contaminated; raw margins not probabilities; GH jev-haiku-benchmarking 404; do not distill Jev as teacher of record (they distilled Haiku); “0.9 is not one number”; ranking ≠ calibration; banking77 0.8–0.9 stated 0.86 actual 0.73 over-confident *theirs*; ≠ Praveenrajus/jev-bench ≠ fstandhartinger/jevbench; $0.0000153–$0.0000226 vs circulating $0.0004 (~20×); Score is 0..n-1 expectation not 0–1; Noul has no confidence field; TCP floor 198.8 ms; type reliability is not a reason to choose Jev (json_schema 5/5); gateway tax not one number; Function-only 5/8 vs hybrid 8/8; 4/8 without Jev; 8 designed cases not conversion lift; ≠ RadRebelSam/awesome-jev; 200-row pilot Jev 86.5% 173/200 vs Gemini Flash-Lite 86.0% 172/200 vs Pro 87.0% 174/200 *theirs*; not a ranking; NLI Tetris argmax P(entail)−P(contradict); 情緒測謊器; 1q 396ms / 30q 567ms; ±0.03; 33q $0.000045 vs Gemini ~5× slower ~60× cost *theirs*; ≠ realZachi/jevtest; 8-example Jev vs GPT-5.6 Sol ~64× cost 5.4× latency *theirs*; synthetic; no inference; ≠ JevBench v1.2 §78; Judged 3317 / listed 2560; Jev judges, code applies policy; catalog ≠ endorsement; APA “microsecond policy / zero hallucination” overclaim; Client-side quiz; pointer from held docs; scanned-PDF warn; CSP only api.typesafe.ai; Jev judges / agent reasons / user decides; selecting an option is not permission to implement; degraded fallback; pattern exact, judgement must clear floor; no matching pattern → no model call; not a correctness oracle; $0.00022 vs chat $0.00306 *theirs*; Spec vs artifact remainder; treating 0.85 as 85% / minProbability hard-gate as Harbor; VERIFY acquires discriminating evidence, never same-pool confidence-only rescoring; fast/full/max are ceilings not sizes; Solar writes, Jev chooses NEXT ACTION; SemIf 2186★ (+20 vs §109 2166); jevlike 1038★ (+7 vs 1031); TypeAR 14★ flat; AnotiaWang 96★ (+1 vs 95); yibie/awesome-jev 490★; Laya likes 802 (was 783); tracker likes 64 flat, lastModified UNCHANGED; do not reopen or amend PR #23/#24/#25/#26/#27. Apply 0743 (`notes.md` §113): Hourly 0743 uniqueness lock: Heman10x-NGU/Verdict-open-jev ≠ Heman10x-NGU/openJev-verdict-2.0; TF-IDF + LogReg ECE 0.0207 vs Jev 0.1440; Verdict-open-jev 48.07% vs Jev 90.80%; abstention combined recall 10.00%; p50 35.58 ms; K=25 (maximum capacity) 72.00%; 0.85 coverage 84.60% selective risk 1.18%; 26.1× faster than standard Qwen JSON generation; Jevify 90.0% / 167 ms CUDA graphs disabled; Finding 1: Brier on stated confidence alone is a trap; grpo_rlcr 0.78 / ECE 0.084; reliability 0.007 but resolution 0.000; 27 900 schema-driven decisions; 13 600 / 13 600 questions; candidate mass min 0.99999624; 22 configs · 166,054 rows · 4 calibration-gold; sha a39eba3f; Student B MAE 0.148 / Pearson 0.836 / 86.0%; pngwn/open-jev-laya-bench README 404; sha 9f69c742 likes 2; HDFS 0.9933 (745/750) / retain 0.0084; BGL ERROR/FATAL protection 1.0000; 2,479 / 2,500 HDFS uncertain; cache hit 0.9648 (2412/2500); $0.153936 estimated; E2 recomputes from saved probabilities; Space sha eda59e0a; MASSIVE English 0.783 / Khmer 0.033 / Hindi 0.133; 40–48 rows too small to ship T; T never changes argmax; siren2345/jev-single-decode-transformers ≠ siren2345/jev-single-decode; Split Transformers experiment from llama.cpp runtime; tanayvasishtha/jev-lab ≠ dairui1/jev-lab ≠ BrendanH18/jev-lab ≠ yibie/laya-jev-lab; Four experiments stress-testing TypeSafe's Jev: calibration, bundle bias, label bias, and ensembling; second pass must be $0.00 from cache; The pages never call Jev; Gemma 4 31B 77.0% / Jev 1.13.0 61.4% / Laya 322M 0.0%; restriction state 95.0% against 84.4%; None of the systems are particularly good at knowing when to stop and ask; They skip the question and call a tool directly; 100% schema pass; six-field joint 48.8% vs 72.8%; ywchiu/jev_benchmark ≠ Running-Dolphins/jev-bench ≠ Praveenrajus/jev-bench; ACT / REVIEW / FALLBACK; A provider failure, timeout, malformed output, or missing answer is **not** a policy outcome; confidence is descriptive provider output, not a substitute for probability; Quality denominators include only valid scored answers; an exact halfway tie chooses the lower level; aiwithenoch/Jev-Skill ≠ simplosophy/jev-skill ≠ laguagu/jev-skills; The local path does not claim to turn a smaller checkpoint into Jev; Low support becomes decision: "review"; MIT-0 SPDX NOASSERTION; current-llm; 结构兼容,不是 Jev 模型能力; altryne/jevify ≠ Mintzs/jevify ≠ gulagala001/jevify ≠ uspraveen/Jevify; Find where Jev belongs. Design the questions. Measure the difference; TypeAR-AI/TypeAR 301 → TypeLLM/TypeLLM; TypeLLM/TypeLLM 16★; SemIf 2241★ (+34 vs §111 2207); jevlike 1051★ (+8 vs 1043); AnotiaWang 98★ (+1 vs 97); yibie/awesome-jev 525★ (+19 vs 506); Laya likes 864 (was 822); tracker likes 67 (+3 vs 64); lastModified UNCHANGED `2026-09-20T04:29:16.000Z`; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#32/#33; notes.md §113 Apply 0646 (`notes.md` §111): Calibration is not alpha; NO CURRENT ALPHA CANDIDATE; ΔR² approximately +0.00084; Brier 0.2131387; ECE 0.0421875; Adding Jev probability to deterministic volatility improved Brier by only 1.4058e-05; default 0.5 keeps zero non pinned; keepResult median 0.14 to 0.17; keepCall median 0.28 to 0.35; usable range is about 0.10 to 0.25; 7.8% to 57.9%; judges results it never sees; task-finish eval not built yet; $0.002 per compaction; slavadubrov/sgr-judge-bench ≠ slavadubrov/jev-judge-bench; Jev 108/120 $0.083 0.34 s; Luna SGR 114/120; paired Jev accuracy-difference intervals include zero; not evidence of equivalence; GLM SGR 26/120 93 format failures; Terra-planned Jev hybrid 55/120; rule-based by default, optionally Jev-backed; empty README; missing key cannot break the experience; prefill plus exactly one decode; softmax over A/B/C ≠ Noul; BBQ 9,053/10,000 (90.53%); ECE 0.0890; Mean confidence 0.9943; overconfident; score and noul not implemented; DGUI 12 rows (was 6); INSTRUCT 119 rows likes 2; encode the state once, decide everything in parallel; 0.740 accuracy against a 0.508 majority; ECE 0.047; fine-tune's advantage ends where its 384-token training data does; jasonkneen/open-jev ≠ pngwn/open-jev; same sha d41dc3cd; Space does not call Jev; recomputes routing from saved probabilities; 200-case Jev 97.0% / 100.0% / 95.0% / MAE 9.22; synthetic repository benchmark; Jev evaluations are advisory; YehuiTang0316/jev-nlgrep ≠ Bentlybro/jevgrep ≠ can1357/jegrep ≠ uehaj/jev-semgrep; default threshold 0.8 still soft; 40-line windows cannot prove whole function; token-native sequential start/end Choice; Gemini/Haiku stubs not configured yet; handful of hand-written examples, not a benchmark; Jev judged exactly what it was given; laguagu/jev-skills ≠ laguagu/jev-evidence-lab ≠ Pleo2/awesome-jev-agent-skills; contract_passed is not a claim of guaranteed factual truth; Wilson lower bound 0.85 floor; fixture mode no savings claim; SemIf 2207★ (+21 vs §110 2186); jevlike 1043★ (+5 vs 1038); TypeAR 15★ (+1 vs 14); AnotiaWang 97★ (+1 vs 96); yibie/awesome-jev 506★ (+16 vs 490); Laya likes 822 (was 802); tracker likes 64 flat, lastModified UNCHANGED; do not reopen or amend PR #23/#24/#25/#26/#27/#28.) | Apply 0843 (`notes.md` §114): **Hysteresis is policy** / **Instruct-tuning honesty collapse** / **Equal-width ≠ quantile ECE** / **Calibration does not compose** / **Ranking ≠ calibration** / **Qwen2.5 ≠ Archer** / **Deferred Crispification** / **g0runmezadam IS tunahansahin897**. **Category error** (Jev vs GPT-5.6 bakeoffs)** / **Skill-issue thesis** / **opt for DeBERTa and ModernBERT ones** / **multimodal ZS perception front-end** / **softmax/ZS ≠ Noul**. people who compare Jev against GPT-5.6 has never fine-tuned BERTForXYZ for living and it shows; zero shot classifiers; scale them as much as decoder only models; many problems solved with LLMs could have been solved with them, it was a skill issue; opt for DeBERTa and ModernBERT ones; BERTForXYZ → DeBERTa → ModernBERT; Jev vs GPT-5.6 bakeoffs are a category error; encoder / ZS classifiers; institutional HF voice; quote *theirs*; do not invent accuracy numbers; softmax/ZS scores still ≠ calibrated Noul; soft scores ≠ hard gates; @mervenoyann; likes 421 / 189; impressions 35498 / 9613; multimodal image<>text ZS as perception front-end; hf:MoritzLaurer/deberta-v3-large-zeroshot-v2.0 likes 139; hf:MoritzLaurer/ModernBERT-large-zeroshot-v2.0 likes 72; Bart, bert, deberta, modernbert, these are all LLMs; Maziyar quoted; Jev is exemplar not the mandate; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29.)| Entropy as allocator (low / medium / high) | Typed low+medium decisions → System One marginals; high-entropy synthesis → frontier decoder. Product rhetoric, not a meter. **Hypothesis** | `references/judgment-class.md` | | Formal / semi-formal (proof vs judgment) | Sensor vs constraint vs searchlight; Alloy vs Apalache; DST trio; TOCTOU-of-Noul, AI×FM. Skills→oxlint: AST/precheck compose with remainder judgment without hard-gating a Noul as a proof. PR attention ≠ correctness (anti-soundness-theater). Engine owns truth / Jev owns judgment (Stockfish+Jev chess coach). Capability kernel: type-safe ≠ correct; irreversible behind threshold AND human. Eval integrity: check the instrument, not just the score (`dinostomp jev` tests a question like an if-statement). Effect contracts, not surface tokens (construct-auto-classifier; privilege ≠ verdict). **Jev supplies evidence, code owns authority** (actiongate-jev; a positive score never overrides a deterministic security failure). **Turnstile clone:** deterministic policy + Jev remainder + replay (evidence ≠ authority). **Type-safe ≠ correct as jaggedness receipts** (atlas; schema-valid ≠ picked-right). **TLA+ compose with a Jev-class oracle:** never confidently wrong; escalate is the safety valve (jev-labs; inverse of soundness theater is hard-gating without escalate). **Hourly 1347:** typed-gate band [0.40,0.60] is refusal; pi-jev-gate fail-closed; choice is the verdict; jev-calibration-arena never acts. **Hourly 1441:** typesafe_agent_gates 27/27 / 31/31; EpicEric/safe-sh static remainder; pastepilot Confirm before act; typesafe-scheduler-diagnostics advisory; choxos/jevchess engine owns truth. **Advance/coverage ledger:** Jev answers questions; SEAL answers whether the world may change (coverage.path auto|code|human|escalate; mint ≠ product brain). **Conflict ≠ ignorance:** Noul collapses both; named Choice escape separates (typed-evaluation-collapse; schema-as-interface). **Sentence-as-rule lint:** ast-grep matcher silent × Jev `ask:` loud (mizchi/jev-lint is mizchi/jevlint rename; ≠ huntedman/JevLint). **Whole-repo intent:** VERIFIED/VIOLATION/UNKNOWN; empty search ≠ proof (jev-intent-review). **Decision-as-assert:** meaning Noul vs exact `toContain`; ambiguous band fails both polarities (jevtest; 0.85 still soft; record/replay; hard-gating a matcher as a merge seal is soundness theater). **Authorship named escape:** `human`/`ai_generated`/`uncertain`; not courtroom evidence. **Empty findings as approval, or auto-promoting an agent-written workflow, is the same theater** (stanley-code `notChecked`; Soft Noul ≠ hard safety). **feelings `.feels()`** default 0.5 never rounded; **Essentiel-Jev never authority**; **enzo-mcp UNKNOWN**; **pigeonhole OTHER skip**. **Hourly 1241:** ZHUBoer/ego-jev reserved `__none__`; runWorkflow completed ≠ success; ORIGIN pause-if-no-Jev; jevbrain AUTO_ACT is not a Noul. **SIGNAL §93:** Cache hit ≠ correctness; score never auto-accepts. **SIGNAL §94:** guidance ≠ hook; unofficial ≠ TypeSafe; hosted bootstrap ≠ silent TypeSafe; Nemotron ≠ TypeSafe Jev; not a calibrated replacement; Router/OOD confidence. **Hourly 1541:** jevguard calibrator/cache/escape; jev-ci-selector CI shadow mode; one-dollar-tahoe TypeSafe Jev defense eval (rh-guard owns the gate cousin). **Hourly 1740:** Decision Graph Protocol frame→assess→commit; app retains permissions/effects; Jev-first assessor-neutral; guarded commit / receipt/next frame; assessment batching; hard-gating DGP as safety theater; numerous-com/dgp ≠ TypeSafe official. **Hourly 1843:** thresholds derived from costs not hard-coded; cost-sensitive decision theory × System One probabilities → control flow; hard-gating 0.038 as safety theater; "guaranteeing" calibration is theater; deeper integrity fold is rh-guard. **Hourly 1943:** noul_threshold 0.5 decoder not a proof; no shipped rule has severity error; 0 of 157 false invalidations; questions/plans/directives are not evidence; IncompatibleSchemaError lists every bad property **Hourly 2041:** resume-screening bias audit methodology; Plan/PRD panel → code-owned pass|review|block; cost-aware multi-model routing/escalation; frozen-protocol zero-shot bench; context-window admission control; typed decision control plane; receipt ≠ authorization; live 15-dim typed rubric re-score per pause; adversarial pre-registered Jev eval; confidence does not track ignorance; provider-neutral Elixir/BEAM Noul/Choice/Score SDK **Hourly 2145:** question-linting of Jev questions themselves; nine jaggedness rules, no API key, no labelled data; static lint ≠ measured separation; yodablocks/jevq ≠ tenbin ≠ JevLint ≠ commitjev; open-weights Laya as class exemplar (binding); Nx/Bumblebee runtime; host chooses backend; ChristianAlexander/laya_ex ≠ system_one_sdk ≠ dannote/jev ≠ NandhaKishorM/laya; on-chain/edge Laya deploy; parity_verified stays false; model output never grants Tx; humandebri/IC-Laya ≠ laya_ex; auditable weekend replica; Jev outputs never used for training; unpaired 0.577 vs 0.727; agilabs-ai/jev48 ≠ JevBench ≠ Mapika/decider; adversarial dual-judge / framing attack surface; comparative framing is the usable judgment; prior injection crowds out evidence; copyleftdev/ember ≠ ember.js; Laya specialist fine-tune pipeline; training still GPU-pending; PIXELZX0/XERON ≠ convaiinnovations/laya; Hub Laya replica drop; daliborsb/laya ≠ convaiinnovations/laya ≠ NandhaKishorM/laya; System One student distillation corpus; gold is programmatic; teacher is closed-API clone; MagaBitmex/jev-4b-distill-data ≠ missing student checkpoint; non-LLM VIN System One; planning depth not chat; lewislululu/jevon ≠ douglance/jevon; source-bound evidence checks; local quote mismatch needs no API; exit 0 ≠ claim truth; WaynezProg/jev-kit ≠ jonathanavis96/jev-kit (Airlock) ≠ jev-use ≠ jev-mcp **Hourly 2246:** independent System One evidence catalog; 19 reviewed records; scores not one leaderboard; no external record currently reproduced; TokenTrim no-Jev matched hybrid 62.4%; reachjalil/system-one-bench ≠ mallahyari/system-one-benchmark; 21 tasks · 134 items · 208 questions; scenes from public GitHub contracts, not production logs; SivletLabs/jev-eval ≠ willkelly/jev-evaluation ≠ 4esv/jev-eval ≠ xxkuboxx/jev-eval ≠ onlyoneaman/jev-eval ≠ dayhaysoos/jevals; option isolation (sibling-blind); permutation-equivariant; Hub OWNER not published; nafisazizir/hev ≠ jaredpalmer/kev; frozen local LLM logits, no trained decision head; residual-head 9,222-param decreased 73/96→67/96; confidence = 1−normalized entropy, not P(correct); yuki-oshio/mini-jev ≠ r-ms/mini-jev; Jev classifier as autoregressive next-token predictor; ChatJev-style soundness theater; erik-dunteman/ChatJev ≠ dannote/jev ≠ jev-gpt; calibrated decision head × AlphaProof value head; implementation-layer isomorphism, semantic difference; timeout = censoring; do not launder Noul as proof; parallel rank-prediction vs serial selection; independent questions can conflict; zzzzzec/jevsort ≠ keltokhy/jsort; curated open System One ecosystem catalog; rupeshpoojary9/awesome-open-system-one ≠ AnotiaWang/awesome-jev; arXiv paper radar with Jev relevance scoring; ranking ≠ calibration / 0.5 still soft; fail-open failed evals not marked seen **Hourly 2340:** train calibrated ~27M from scratch; typed Q→prob dist / one forward pass / no LLM decode; hyusi2003/MiniSystemOne ≠ Colvin0315/MiniSystemOne; description-only stub / size 5; ESCI hard probe fails four of six; jev_bool ECE 0.242 inversion 0.255; do not re-fold §60 six-gates as new; jobbyjev one-request-per-company from batch-size result; find/design/evaluate TypeSafe Jev decision loops; karanb192/jev-architect ≠ samtay32/jev-system-architect; Jairik/jev-distiller size 1; distill-Jev UI stub / do not distill Jev as teacher of record; post-launch scored use-case map / Jev self-scores then human curation; licensedsaucer9-web/jev-opportunities; Jev-inize a use case into classifier/router; gavinHuang/jevinize → simple-jev not TypeSafe; featherless-ai/simple-jev; compare saved decisions / same label can still change the branch; VihaanAgarwal/jev-diff ≠ Saik0s/diffusiongemma-jev-macos; not tested with a live Jev API key; constrained logprob + temp/Platt ≠ Noul; OpenJevPro pastes openjev-sglang JevBench as own; zhangcy122/OpenJevPro ≠ IamBusy/OpenJev ≠ ekzhang/openjev-sglang; PolyForm Noncommercial; SmolLM-135M / sub-70ms / 0 output tokens; demo P(True) 0.5052 / Choice conf 0.2872 / Score conf 0.0055; README claims MIT / GitHub license null / no LICENSE file; patelvishwa112/jev-system-one-rlcd ≠ arnabgho/rlcd-lite ≠ blackwood-rlcd; source-backed Awesome Jev radar / 306+ commit-pinned; logicrw/awesome-jev-projects ≠ AnotiaWang/awesome-jev ≠ yibie/awesome-jev ≠ cobanov/awesome-jev ≠ rupeshpoojary9/awesome-open-system-one; auto GitHub sync / Issue-only submissions; hashed n-gram encoder / rival-aware attention; olanotolu/jevbetter vs jevlike starter; synthetic hard menus top-1 0.916 vs 0.873 / ECE 0.0182 vs 0.0367 / 40 vs 4608 menus/sec; shuffled-context control 0.335 | `references/formal-methods.md` (one-screen: `references/formal-semi-formal.md`) Apply 0042 (`notes.md` §105): structured probability readouts; distribution > argmax; Noul 0.5 midpoint; score is expectation not integer; bare HTTP not SDK; Arohtea/jev-readout; Jev-style Choice/Score/Noul from ordinary models; optional DSH plugin; schema-valid ≠ calibrated; gulagala001/jevify ≠ Mintzs/jevify; Laya RLCD benchmark; 40.3% below constant-answer; open-weight measurement; mourad-ghafiri/laya-rlcd-benchmark ≠ yibie/laya-jev-lab; cheap fail-open semantic edge; second signal not sole; FastLoopError catch; SupremeDreamZ/jev-fastloop ≠ jev-ultrafast; asking more questions in one call; 0.980 at every N; nearly not fully deterministic; TheWebDevel/jev-fanout; Qwen3-VL perception + Jev decisions train RL; 0 model calls at deployment; VLM alone 1.7 vs +Jev 4.4; harneet2512/reflexrl ≠ khordoo/jev-reflex-autonomy-lab; independent Jev API vs Laya; cascade 0.60 matches 78% at 1.8×; noul facts not judgements; yibie/laya-jev-lab ≠ dairui1/jev-lab ≠ BrendanH18/jev-lab; GLiNER vs GLiFormer vs Laya vs Jev; extractors ≠ decision engines; Laya dict-instructions collapse 58.3%; umstek/zero-shot-ie-bench; decisions-per-minute & cost; 204 moves vs 73; throughput not intelligence; angelgalvisc/snake-arena-jev-vs-llms ≠ vtrivedy/jev-plays-games; behavioral contracts; pin expectations eval upgrades; raw 0.94 is not a release; sathariels/jevcheck ≠ dayhaysoos/jevals ≠ SivletLabs/jev-eval; evidence-linked dependency upgrade; Jev never generates filenames; no_direct_evidence ≠ safe to merge; GaneshVG18/upgrade-radar ≠ LYchoon/paper-radar-jev; discography theme/mood/complexity; five atomic questions one call; lirantal/discoprint / Apply 0145: hard budget filter before Jev; Jev never asked to perform budget arithmetic; Jev judges the next state, XState enforces transitions; do not reopen or amend PR #23. / Apply 0243: Benchmark-driven Jev router and judge; cheap alone is not success; Jev does not write, sum prices, or claim accuracy %; Sol 94.2 / Luna 83.9 / Jev path 89.7; 19.2% Sol / 62.3% cost save / 4.5pp miss of 2pp non-inferiority; p50 latency worse than Sol due to routing overhead; erendikmenn/jev-llm-router-benchmark ≠ jev-rag-benchmark ≠ ryantsai/jev-llm-router; Express + node:sqlite; mock and Jev decision engines; previous_ticket_count >= 3 is code; MIN_CONFIDENCE 0.6 still soft; substring false positives; aesaganda/jev-ticket-router ≠ SarathChandraBellam/jev-vs-llm-ticket-router; Universal Figure & Diagram Router; confidence ≥ 0.85 hard-gate is theater; generative AI banned from scientific plots; six visual branches; human-labeled (state, question, label); 166,054 rows / 22 configs; soft_label for human uncertainty; Praveenrajus/jev-bench ≠ fstandhartinger/jevbench; ternary bonsai System One GGUF; Hub does not ship weights; 100/100 easy T/F is not Harbor; label_mass ≠ correctness; stock llama.cpp Q2_0 silently gibberish; NicolaiMTLassen/open-bonzi-jev ≠ NicolaiLassen; transformers.js DeBERTa ONNX; temperature 1.05; AutoModel from_pretrained works; onnx-community/open-jev-deberta-v3-large-ONNX ≠ system-one-qwen3.5-4b-scorer-ONNX; 107★ densify; GH 151M vs README 149.6M; PR #1 now closed unmerged; do not re-fold §71 claim-audit as a beat; typed decisions, RLCD, confidence-gated routing; structured ≠ correct; mock not live API; 26 tests; wjdjdakf17/jev-study ≠ baekenough/jev-study / Apply 0345 (`notes.md` §108): **Open reproduction densifies / measurement densifies PRIMARY** (bonzi-27b-v2 / ternary-8b / 27b-v1 GGUF family densify; WANLI-256 74.6% / 65.2% / 71.1% *theirs*; Bonsai 1 27B Q1_0 runs on stock llama.cpp; ternary still needs PrismML fork; hf:heman10x/openJev-verdict-2.0 twin tokenizer-only; OpenJev Vision image classification + uncertainty; CLEVR-4 held-out joint 0%; hfdataset:IamBusy/OpenJev-Vision-Research-v0.1 12,832; 294,912 derived targets not independent samples; Laya multilingual ONNX WebGPU typed-decisions port; 63/63 selected answers / 5.1e-4 CPU / 1.2e-2 WebGPU; UpHash-Network/mini-jev is yuki-oshio transfer; jev-injection-bench 11,900 labelled prompts; Jev best ranking / Haiku better ECE 0.021 vs 0.058; 0.5–0.9 band is where Jev's numbers do not mean what they say; Prompt wording moves panic 28%; manojlds/jev-dspy-bench ≠ dspachos/jev-dspy ≠ jmanhype/jev-dspy-lab; Jev agreement is similarity, never ground truth; no aggregate quality grade or merge gate; AbstentionBench-on-Jev rank 1 of 20 vs 2025 field; question-asymmetry; forward-looking 0.465 never extreme; openkev calibration layer not a runtime; ECE vs coverage independent; select_threshold returns inf; escalation catches uncertainty not ignorance; misakaikato/openkev ≠ jaredpalmer/kev; pdf-race Docling→Jev vs Gemini; parser owns the wall clock; 12/12 tie is a tie; titles selected not generated; flopcheck 16 calibrated tweet judgments; mechanical tells in code; ZeroX-01/jev-atlas ≠ Zaious/jev-capability-atlas ≠ gorock007/jev-atlas; Laya calibration lab Gradio MCP; T never changes argmax; confidence ≠ top-label p; easy probe set refused; 40–48 rows too small to ship T; do not reopen or amend PR #23 or #24 or #25. Apply 0439 (`notes.md` §109): Gemma-4 26B-A4B jevify classification+calibration; LoRA adapter twin not independent eval; Gemma-4 E4B jevify; E4B LoRA stub card; kushalpatil/jevify-gemma4 ≠ Mintzs/jevify ≠ gulagala001/jevify ≠ uspraveen/Jevify; GH kushalpatil07/jevify 404; PAWS 0.580/ece 0.288 is the weak cell; smaller E4B slightly better OOD ECE than 26B-A4B; Hub jevify merged LoRA ships weights; bonzi Bonsai-8B v1 GGUF densify; Bonsai-1.7B v1; Bonsai-4B v1; WANLI-256 64.5% / 60.2% / 52.0% *theirs*; rank #4 / #5 / #6 of 6; JulesHuisman/jev-eval scaffolding / README SHA c356a584 (was empty e69de29b); JulesHuisman/jev-eval ≠ SivletLabs/jev-eval ≠ willkelly/jev-evaluation ≠ 4esv ≠ xxkuboxx ≠ onlyoneaman ≠ dayhaysoos/jevals; 7 bands 6/10 vs 40 bands 0/10; source receipts + confidence slider re-policy without re-inference; 32/32 synthetic is smoke not production; classify HF datasets across typed semantic dimensions; roadus2 watch misspelling; lock roadius2/ultra_laya; ultra_laya REVIEW defects; default branch claude/laya-jev-review-gg5ppo; XNLI EN 88.3% ECE 0.032 → RU 77.3% ECE 0.096; Δ −11.0 pp [−14.2,−7.8]; ECE +0.063; MASSIVE no detectable difference at n=600; confidence is function of p_max (r=1.000); pointer-not-generator 400 human-authored responses; proposed ≠ authorized; FewRel 160: Jev 85.0% vs lexical 13.125%; gated 100% (95/95) coverage 59.375%; J++ composable semantic computation language; judge-jev 0.5 still soft; 947 repos scored; A 273 / B 302 / C 372; LLM rubric ≠ benches; No benchmark winner is claimed; phishing: naive 62.6% vs regex 91.8%; 5-atomic + LR 95.0% *theirs*; AITuber tension ±15; README npm global; repo is Rust; git-confess code owns counting/blame/ratio; httpx exhibit 11% (13/119) *theirs*; 90d trend +12.40% vs random +12.75% vs BH +41.71%; 5m win rate 25%; Awesomejev 656 entries / 38,160 stars; tracker likes 64 (+4) lastModified UNCHANGED; Laya present; Blackwood ABSENT; Archer still promised_not_landed; do not reopen or amend PR #23/#24/#25/#26. Apply 0541 (`notes.md` §110): Blackwood tracker ABSENT; likes 2 gated manual; r = c - p_a; ECE 0.021; acc 0.807 vs warmup 0.746; calibration beyond ~500 tokens unmeasured; Independent primitive; 11.57s vs 54.10s · 4.67× · 120/128 *theirs*; default path is pretrained Gemma probs not trained RLCD head; GH Meanblock 404; lock leesk212/JEV-CPU; softmax over letter slots ≠ Noul; WANLI 0.741 vs openjev v2 0.77 *theirs*; 3-way NLI ≠ Noul; priority 0.464 = majority floor; banking77 contaminated; raw margins not probabilities; GH jev-haiku-benchmarking 404; do not distill Jev as teacher of record (they distilled Haiku); “0.9 is not one number”; ranking ≠ calibration; banking77 0.8–0.9 stated 0.86 actual 0.73 over-confident *theirs*; ≠ Praveenrajus/jev-bench ≠ fstandhartinger/jevbench; $0.0000153–$0.0000226 vs circulating $0.0004 (~20×); Score is 0..n-1 expectation not 0–1; Noul has no confidence field; TCP floor 198.8 ms; type reliability is not a reason to choose Jev (json_schema 5/5); gateway tax not one number; Function-only 5/8 vs hybrid 8/8; 4/8 without Jev; 8 designed cases not conversion lift; ≠ RadRebelSam/awesome-jev; 200-row pilot Jev 86.5% 173/200 vs Gemini Flash-Lite 86.0% 172/200 vs Pro 87.0% 174/200 *theirs*; not a ranking; NLI Tetris argmax P(entail)−P(contradict); 情緒測謊器; 1q 396ms / 30q 567ms; ±0.03; 33q $0.000045 vs Gemini ~5× slower ~60× cost *theirs*; ≠ realZachi/jevtest; 8-example Jev vs GPT-5.6 Sol ~64× cost 5.4× latency *theirs*; synthetic; no inference; ≠ JevBench v1.2 §78; Judged 3317 / listed 2560; Jev judges, code applies policy; catalog ≠ endorsement; APA “microsecond policy / zero hallucination” overclaim; Client-side quiz; pointer from held docs; scanned-PDF warn; CSP only api.typesafe.ai; Jev judges / agent reasons / user decides; selecting an option is not permission to implement; degraded fallback; pattern exact, judgement must clear floor; no matching pattern → no model call; not a correctness oracle; $0.00022 vs chat $0.00306 *theirs*; Spec vs artifact remainder; treating 0.85 as 85% / minProbability hard-gate as Harbor; VERIFY acquires discriminating evidence, never same-pool confidence-only rescoring; fast/full/max are ceilings not sizes; Solar writes, Jev chooses NEXT ACTION; SemIf 2186★ (+20 vs §109 2166); jevlike 1038★ (+7 vs 1031); TypeAR 14★ flat; AnotiaWang 96★ (+1 vs 95); yibie/awesome-jev 490★; Laya likes 802 (was 783); tracker likes 64 flat, lastModified UNCHANGED; do not reopen or amend PR #23/#24/#25/#26/#27. Apply 0743 (`notes.md` §113): Hourly 0743 uniqueness lock: Heman10x-NGU/Verdict-open-jev ≠ Heman10x-NGU/openJev-verdict-2.0; TF-IDF + LogReg ECE 0.0207 vs Jev 0.1440; Verdict-open-jev 48.07% vs Jev 90.80%; abstention combined recall 10.00%; p50 35.58 ms; K=25 (maximum capacity) 72.00%; 0.85 coverage 84.60% selective risk 1.18%; 26.1× faster than standard Qwen JSON generation; Jevify 90.0% / 167 ms CUDA graphs disabled; Finding 1: Brier on stated confidence alone is a trap; grpo_rlcr 0.78 / ECE 0.084; reliability 0.007 but resolution 0.000; 27 900 schema-driven decisions; 13 600 / 13 600 questions; candidate mass min 0.99999624; 22 configs · 166,054 rows · 4 calibration-gold; sha a39eba3f; Student B MAE 0.148 / Pearson 0.836 / 86.0%; pngwn/open-jev-laya-bench README 404; sha 9f69c742 likes 2; HDFS 0.9933 (745/750) / retain 0.0084; BGL ERROR/FATAL protection 1.0000; 2,479 / 2,500 HDFS uncertain; cache hit 0.9648 (2412/2500); $0.153936 estimated; E2 recomputes from saved probabilities; Space sha eda59e0a; MASSIVE English 0.783 / Khmer 0.033 / Hindi 0.133; 40–48 rows too small to ship T; T never changes argmax; siren2345/jev-single-decode-transformers ≠ siren2345/jev-single-decode; Split Transformers experiment from llama.cpp runtime; tanayvasishtha/jev-lab ≠ dairui1/jev-lab ≠ BrendanH18/jev-lab ≠ yibie/laya-jev-lab; Four experiments stress-testing TypeSafe's Jev: calibration, bundle bias, label bias, and ensembling; second pass must be $0.00 from cache; The pages never call Jev; Gemma 4 31B 77.0% / Jev 1.13.0 61.4% / Laya 322M 0.0%; restriction state 95.0% against 84.4%; None of the systems are particularly good at knowing when to stop and ask; They skip the question and call a tool directly; 100% schema pass; six-field joint 48.8% vs 72.8%; ywchiu/jev_benchmark ≠ Running-Dolphins/jev-bench ≠ Praveenrajus/jev-bench; ACT / REVIEW / FALLBACK; A provider failure, timeout, malformed output, or missing answer is **not** a policy outcome; confidence is descriptive provider output, not a substitute for probability; Quality denominators include only valid scored answers; an exact halfway tie chooses the lower level; aiwithenoch/Jev-Skill ≠ simplosophy/jev-skill ≠ laguagu/jev-skills; The local path does not claim to turn a smaller checkpoint into Jev; Low support becomes decision: "review"; MIT-0 SPDX NOASSERTION; current-llm; 结构兼容,不是 Jev 模型能力; altryne/jevify ≠ Mintzs/jevify ≠ gulagala001/jevify ≠ uspraveen/Jevify; Find where Jev belongs. Design the questions. Measure the difference; TypeAR-AI/TypeAR 301 → TypeLLM/TypeLLM; TypeLLM/TypeLLM 16★; SemIf 2241★ (+34 vs §111 2207); jevlike 1051★ (+8 vs 1043); AnotiaWang 98★ (+1 vs 97); yibie/awesome-jev 525★ (+19 vs 506); Laya likes 864 (was 822); tracker likes 67 (+3 vs 64); lastModified UNCHANGED `2026-09-20T04:29:16.000Z`; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#32/#33; notes.md §113 Apply 0646 (`notes.md` §111): Calibration is not alpha; NO CURRENT ALPHA CANDIDATE; ΔR² approximately +0.00084; Brier 0.2131387; ECE 0.0421875; Adding Jev probability to deterministic volatility improved Brier by only 1.4058e-05; default 0.5 keeps zero non pinned; keepResult median 0.14 to 0.17; keepCall median 0.28 to 0.35; usable range is about 0.10 to 0.25; 7.8% to 57.9%; judges results it never sees; task-finish eval not built yet; $0.002 per compaction; slavadubrov/sgr-judge-bench ≠ slavadubrov/jev-judge-bench; Jev 108/120 $0.083 0.34 s; Luna SGR 114/120; paired Jev accuracy-difference intervals include zero; not evidence of equivalence; GLM SGR 26/120 93 format failures; Terra-planned Jev hybrid 55/120; rule-based by default, optionally Jev-backed; empty README; missing key cannot break the experience; prefill plus exactly one decode; softmax over A/B/C ≠ Noul; BBQ 9,053/10,000 (90.53%); ECE 0.0890; Mean confidence 0.9943; overconfident; score and noul not implemented; DGUI 12 rows (was 6); INSTRUCT 119 rows likes 2; encode the state once, decide everything in parallel; 0.740 accuracy against a 0.508 majority; ECE 0.047; fine-tune's advantage ends where its 384-token training data does; jasonkneen/open-jev ≠ pngwn/open-jev; same sha d41dc3cd; Space does not call Jev; recomputes routing from saved probabilities; 200-case Jev 97.0% / 100.0% / 95.0% / MAE 9.22; synthetic repository benchmark; Jev evaluations are advisory; YehuiTang0316/jev-nlgrep ≠ Bentlybro/jevgrep ≠ can1357/jegrep ≠ uehaj/jev-semgrep; default threshold 0.8 still soft; 40-line windows cannot prove whole function; token-native sequential start/end Choice; Gemini/Haiku stubs not configured yet; handful of hand-written examples, not a benchmark; Jev judged exactly what it was given; laguagu/jev-skills ≠ laguagu/jev-evidence-lab ≠ Pleo2/awesome-jev-agent-skills; contract_passed is not a claim of guaranteed factual truth; Wilson lower bound 0.85 floor; fixture mode no savings claim; SemIf 2207★ (+21 vs §110 2186); jevlike 1043★ (+5 vs 1038); TypeAR 15★ (+1 vs 14); AnotiaWang 97★ (+1 vs 96); yibie/awesome-jev 506★ (+16 vs 490); Laya likes 822 (was 802); tracker likes 64 flat, lastModified UNCHANGED; do not reopen or amend PR #23/#24/#25/#26/#27/#28.) | | Mixed architecture (judgment model + LLM) | Provider judges, LLM writes, code owns control; not a stack replacement. Advisory sidecar never changes host routing. Dual-process: S1 decides, S2 generates (routing accuracy unmeasured); **Harbor-shaped cousin:** decide→policy→LLM leftover (shared Answer schema; jev vs gen-json vs gen-logprob; Noul 0.5 never rounded). **Closed-vote CU:** no planner LLM; code builds options, Jev only picks (JevOnly). **Host-owned product:** app retains handlers/permissions; Jev over live typed actions (waymode). S1 specialists + S2 coordinator is the same split (description-only greenfield). Internals ≠ FSM; placement is a **component node**. Browser-use strength = DOM-as-text + speculative fan-out, not vision. Fail polarity is per act: skip-wake fail-open vs merge-gate BLOCK fail-closed; OMP/pi acceptance+route **fail-open** (contrast pi-jev-approver fail-closed). OMP prompt suppression: operator owns the bar; not a sandbox (omp-greenlight). Skill-broker outline: Jev never grants access. Hybrid local decide + remote fill; `DONE` ≠ verified success. Specialist computer-use: plan ≠ execute, dry-run default (Cua-S1; not TypeSafe Jev). Judgment as a language primitive (Ruby `almost_certain?`/`pick`/`rate`). Decision-native RAG: retrieve wide → decide → evidence set → LLM. Classify-first MCP (topology A): content to the judge without entering main agent context first. Generative UI: model decides, compiler emits. Draft-gate silence ≠ safer (heartbeat). Living class-pattern atlas (not a 342-title dump). Stagehand experimental Jev: pick-and-copy extract + act tree + observe/cache-check; LLM fallback; draft stack. Public judgment wall (six parallel questions; policy-in-code; cost-to-1M). PR attention ≠ correctness. Session-sticky first-prompt route (fail-closed fallback). Capability kernel (LLM ring 3 / Interlock ring 0; secrets never in agent; Jev SENSOR; policy.py BLOCK/ASK/ALLOW; type-safe ≠ correct). Typed control plane around DSPy (drafts AFTER route+action). Engine owns truth / Jev owns judgment. Human-confirmed kill (mapped explanations; identity re-check). Wire-compat encoder backend (GLiFormer `/v1/systemone` drop-in; cheaper, less accurate on reasoning-heavy). Loopback gateway routes hosted + local (not a model). **Constrained optimizer + S1 features** (slo-router: Jev never the sole hot-path gate; fail-open local features). **Effect-based shell gate** (construct: privilege ≠ verdict; fail-closed). **Attention filter / VOI** (jev-lens: never blocks the agent; never green unless sure). **Jev supplies evidence, code owns authority** (actiongate-jev). **Measurement owns endorsement** (jev-packs evidence-gated). **Ranking ≠ calibration** (does-jev-confidence; never hard-threshold raw p). **Hot-click CU** (ego-jev: indexed table → operation+target; text model only for type). **Jev judges relevance, code decides structure** (jev-compactor; never rewrite; regex floor). **Local rules first / never auto-train on own hides** (x-reply-filter). **Control-plane combinators** (Then/Gate/Vote/Cascade/Weighted; not chat turns). **Skill VOI / abstention** (skillranker; hook fail-open). **Receipts not leaderboard** (atlas + frontier-100 + OOD). **Turnstile** evidence≠authority + replay. **MLX one-pass replica economics** (jevmlx; softmax ≠ Noul). **TLA+ consensus kernel** (jev-labs; never confidently wrong; escalate). **SEAL advance/coverage** (no seal, no advance; exception queue visible). **Sureness bands** (how-sure-is-jev; max_prob is generous). **JevBench v1.1** (calibration reported, not scored). **CI typed gate** (ci-gatekeeper before expensive review). **Codex MCP adapter** (jev-in-codex; ranking unbenchmarked; lexical fallback). **Stop-hook attention redirect** (jev-preflight; eight axes; assist=one reinspect; fail-open; not a merge blocker). **Pre-send view selection** (dizk/jev-lens; 79% fewer tokens *theirs*; compress-before-first-send; distinct from rashedInt32/jev-lens). **Observational memory** (pi-om; keep/kind verbatim; model-free compact). **tools≠use** (carryforward 0/4 recall; SessionStart > hoping). **Physical-world S1** (HA-Jev; sensors from typed answers; not for locks/heaters). **Judgment outside the store** (jevql CLI; DB never sees `jev()`). **Landed-script / headless≠auto-approve** (construct); **digital-design combinators** (jev-combinators rename + Router/Loop/Retry/Fallback/Memory; metaphor ≠ literal AND/OR); **VOI cache admission** (jevcache same-intent; 0 FP/100 *theirs*; fail-open); **worth-your-attention VOI** (ThinkyMiner/Winnow 80%/90%; ≠ kevinpita/winnow); **Jev WHETHER / Python HOW / LLM WHAT** (hermes-jev-router; license null); **typed escalate/continue/abort baton** (jev-handoff; inverted loop; gate never grants; fail-open); **Playwright executes, Jev chooses** (browser-jev; sample-from-distribution); **OpenJev `/v1/decide` ≠ drop-in** + **SemIf runoff wire**; **conflict ≠ ignorance** (named Choice escape); **decision-as-memory flywheel** (DGUI_HYPERMEM-JEV 6-row schema); **record/replay CI** (jevassert landed; accuracy+ECE+cost gates offline); **failure-finding arena** (chenmingtang830/jevarena ≠ meetr1912/jev-arena); **BBQ** 97.28%/0.04/0.34/$0.3429 *theirs*; **decider≠executor** (jeffrey: Jev next-tool, LLM fills args); **sentence-as-rule** (mizchi/jev-lint is jevlint rename); **VOI hunk prune** (prune-review ~20% target; 1.18% with outlier); **persist constraints across compaction** (pi-heed); **Harbor SGR-judge contract** (jev-judge-bench; canaries ≠ quality; ≠ jevarena/jevbench); **hand no-text steps** (jev-use; Vercel drops confidence; margin 0.4; fail-open gate; ≠ jev-ultrafast); **Pi System-One control plane** (pi-jev-control; GUI never force-click; compaction never writes session); **never free-generates** (jev-gpt tree of Choices; 400 calls / 75 s / 2¢ *theirs*); **OpenRouter recipe atlas** (jev-cookbook; 16–36 item samples not benches; 425 calls / $0.015); **personal-history feed** (jevfeed; no social graph; one request per batch of ten); **competing NAR claim-audit** (openJev-verdict-2.0; dual-channel ECE; PR #1; ≠ IamBusy/OpenJev); **empty compaction-proxy skip** (IPECTER context-pruner **and** jev-runway LICENSE-only); **1-token logprob endpoint ≠ Noul** (chakuho; coverage ≠ correctness); **open replica engine** (jevinf argmax-parity); **unofficial Elixir SDK ≠ OTP peer** (typesafe-elixir-sdk ≠ dannote/jev); **jevex n=16 files-to-read VOI** (rename of jev-semantic-explorer); **commit pre-review attention≠verdict** (commitjev; middle band never rounded); **Hermes plugin is Agnes not TypeSafe** (hermes-plugin-jev); **pi-jev-compact ≠ pi-jev-compaction** (verbatim summarizer replacement); **decision-native inbox** (mailordinal; humans own ambiguity); **unofficial jev-cli not ready** (≠ jevql); **laya-multilingual** English checkpoint confident-wrong OOD; **schema-scorer peaked ranking ≠ calibration**; **productized System One HTTP** (classifier.dev; label+p; batch ~1000; Jev primary / LLM fallback); **escalate-under-threshold** (smart single-label <0.7; multi-label ignores); **silent FALLBACK** (granite 0.546 vs advertised 0.800; rh-guard owns the gate); **systematic-review pointer (choxos/jev-reviewer ≠ egma-ai)** two-pass Choice+Noul; *Not found* is an answer; human check is the product; **githubnext/localjev** prompted JSON ≠ structured-read logits (wire-compat ≠ logit-equiv; **≠** kunchenguid/local-jev; GitHub Next **261★**; 1,200-req caveats *theirs*); **NandhaKishorM/laya packaging** Router script-before-p; post-T ECE ≠ raw ECE; 0.85 still soft; Banking77 token-budget; **≠** TypeSafe drop-in; **external census ≠ scored bake-off** (@airesearch12; GLiNER2+routers class-boundary; incomplete vs Laya/localjev/kev; Harbor honesty watch); **JevBench v1.2 geometric-mean product** (I/C/S/K 25% each; cal ON rank; Luna I=96.8 rank #7; ×2/est. Harbor honesty; option-order 72→21; Laya absent gap; Qwen3.8 27B ≠ Archer); **hourly 0842 apply-the-five** (already §73–§78; do not re-card); **skip thin noise**; **hard-gate a Noul as a PR/quality gate is soundness theater** (totally-tim/jev-gate / claude-jev-warden; ≠ jev-gateway / MongLong0214/jev-gate / jev-gate-student-b); **S1 keeps flying / S2 one-use** (khordoo delta: escalate without stall; purple = consumed; Local controller ≠ githubnext/localjev; seed = geometry; 20% still soft; no pixels; S2 never grants). **OCR+AX desktop CU:** typesafe-computer-use (hosted Jev; never screenshot-to-frontier for the decision; overlapping options = doubt; writer/decider; 155× *theirs* one screenshot; 0.4/0.5 still soft; **≠** jev-ultrafast **≠** cua-s1 **≠** camoufox). **ASR voice-browser CU:** jev-voice-browser (partial-speech VOI; pointer spans; spoken confirm ≠ auth; 27/27 *theirs* fixtures; **≠** jev-voice-control **≠** nikolas-j **≠** OCR desktop). **Wrap-as-execution ALLOW/ASK/DENY:** wrap *is* the tool function; rules first; ASK throws; fail-closed (AgentGhost; **≠** actiongate **≠** jev-use fail-open; rh-guard owns the gate cousin). **JP genre atlas:** apps by hole; stars research-time; not verified evals (@studio_yebisu; **≠** class census §77 **≠** v1.2). **External pedagogy:** Akshay “Jev Clearly Explained”; LLM hammer; schema-safe ≠ correct; 200×/400× TypeSafe ceiling; shadow + questions-as-code; **≠** official docs **≠** Flavio **≠** AgentGhost. **Meaning-grep dedicated:** proposition≠embedding; boolean composition of thresholded Nouls; Semgrep.dev collision; not a gate (jev-semgrep §86). **Hourly 1047 + deferred 0945:** decision-validated UI (gram-render never authors text; jev2ui Jev decides / Gemini writes); decision-as-assert (jevtest ambiguous band 0.15–0.85 fails both; 0.85 still soft); hybrid S1 (anima3 closed verb menu + hard safety first; jeff confidently flat on magnitude; Qwen logprob default — do not invent Laya as a backend); pointer search (JevFind path then window); Harbor bake-offs three shapes (frontier-bench ≠ frontier-100; GLiClass product bakeoff not architecture duel; four engines / majority floor / calibration ≠ discrimination); authorship named escape (not evidence); non-SWE (ha-switchboard HA remains execution ≠ HA-Jev; n8n Low Confidence abstention); compaction delta (fast-jev-compaction-pi ≠ pi-jev-compact ≠ pi-jev-compaction); full-distribution optimizer (jevloop UCB1+CEM; no LLM in the loop; mock default); deferred class (laya-vision SmolVLM `score` untrained ≠ blackwood ≠ Archer; Cerebellum-2B `/v1/decide` ≠ TypeSafe — wire-compat vs agent-routing as separate Harbor axes, competing NAR not endorsement; laya-grounded not drop-in / phishing 0.611→0.512 / Platt not temperature). Hard-gating a Noul as test/PR/HA write/authorship seal is soundness theater. **Queued SIGNALs:** Collapse GestaltLabs/Jeff-1 into logan-markewich/jeff; Quote Jeff-1 ECE as “better than Jev”; Treat empty stanley findings as approval; Treat findme beam score as file-identity; Price workers, not the conversation. Soft Noul ≠ hard safety. **Hourly 1144:** Treat `.feels()` 0.5 as a bool if; Collapse apa-agent-harness into AntonioCoppe/jev-harness; Quote apa "mathematically fulfilled" / npm @aipersona; Treat grok-bot-jev A/B as token savings; Let Essentiel Jev send / skip human approve; Collapse enzo-mcp into jev-sift / skip UNKNOWN; Treat pigeonhole OTHER as a move / 0.6 as Harbor τ; Treat the HF playground as live Jev / collapse into classifier.dev; Quote jev-reliability as accuracy; Collapse clduab11/jev-test into realZachi/jevtest / paste bars as results; Paste "Jev wins" from jev-rag-benchmark; Collapse dairui1/jev-lab into BrendanH18/jev-lab / re-card jev-desktop; Treat jevmail as mailordinal / mailjay as read-only. **Hourly 1241:** Collapse ZHUBoer/ego-jev into jiangkoumo / treat `completed` as success; Treat jsort logits as frequencies / Choice as the scale; Paste groundedness Macro-F1 as “Jev wins quality”; Quote jev_playground 83% / promote from authored bars; Copy `jev-latest` on Zen; Treat techstack ranks as a generated stack; Collapse s1_ruby into hunch/feelings; Treat judgement as jevql / confidence as winner p; Treat the Rust community SDK as official / a new species; Collapse tpellet/hunch into carldaws/hunch / skip exit 3; Quote file-search 15 matches as recall; Treat linkmap referee as gold / let Jev see S2 prose; Treat jev-mail as jevmail / tidy OTHER as a move; Close pinned/audio/current tabs / skip Show; Let ORIGIN LLM decide / continue without Jev; Gate crawlers on raw `bug_likely`; Sell jevbrain AUTO_ACT as a Noul. **Hourly 1347:** hermes-switchyard ≠ hermes-jev-router ≠ hermes-plugin-jev; cyrusasco/typesafe-mcp noul deadband 0.35–0.65; smartdio/jev-browser-agent ≠ ZHUBoer/ego-jev; Dakai/omp-jev-web DONE ≠ proof; typed-gate band [0.40,0.60] is refusal; pi-jev-gate fail-closed; choice is the verdict; Jev-Calibration Platt ECE 0.117→0.052. **Hourly 1441:** Jev-Reranker live Jev not yet measured; sessionwise opt-in relevance; jev-search pointer sieve; savka777/jev-search ≠ kazuhideoki/jev-search ≠ superagents-lab/jev-search; 400ms Salesforce WebMCP; droidjev screenshot-free; Tewoto1 jevcu planner still writes; ha-conversation-jev Jev→Grok; dsh-jev can only gate; jev-classification-benchmark specified not run; jev-luna-pagerduty p≥0.50; jev-drive sim not AV; story-arc Jev never authors; jev-hs-assistant HS6; golergka/jev-plays-starcraft-2 UI-verified ≠ API Victory; awesome-jev-use-cases catalog. **SIGNAL §93:** fingerprint after redact; recall vs decide; publish fingerprints+answers; CI replay as Harbor cousin; Cache hit ≠ correctness; hyperspaceai/jevcache ≠ kushals256/jevcache; human labels only; score never auto-accepts; production capture flywheel; sutro-sh/jev-align ≠ caiovicentino/jev-align. **SIGNAL §94:** guidance ≠ hook; catalysts ≠ summaries; compile-time System One; unofficial ≠ TypeSafe; format_version modernbert-jev/1; Argos1111/jev_local ≠ us/jev-local ≠ kunchenguid/local-jev; LFM default ≠ ModernBERT backend; Nemotron ≠ TypeSafe Jev; not a calibrated replacement; djev-dev complements djev-spark; images as Choice options; Laya essay numbers *theirs*; Router/OOD confidence; hosted bootstrap ≠ silent TypeSafe. **Hourly 1541:** difficulty + policy thresholds + JSONL trace; jev-codex-pilot model + reasoning depth; keep/shadow/hybrid/reject; quarry evidence projection; Frank-ZY-Dou/awesome-jev robotics/3D/control; one-dollar-tahoe TypeSafe Jev defense eval; jevguard calibrator/cache/escape; jev-ci-selector CI shadow mode; llama-jev llama.cpp replica; petercr/jev-orchestrator ≠ FleeexCorp/jev-orchestrator; seb4ez/jevguard ≠ AseemPrasad/JevGuard ≠ pablozr/JevGuard; webNeat/llama-jev ≠ WiktorB2004/llama-index-jev. **Hourly 1639:** OpenCode jev-pruner context sieve; observe→score-candidates→prune; jev-zen / jev-1.13-free; zen-chat ≠ Noul; fail-open original; keepScore >0.1 floor; host port of tamaratran/jev-pruner; indiejoseph/opencode-jev-pruner ≠ nrdz-labs/fast-jev-opencode; jev-webagent-bench empty stub; Kiln-AI/jev_jsonschema noul_threshold 0.5; NSStudent/JevSwiftSDK unofficial. **SIGNAL §97:** GLiNER2 native Apple path; unofficial Swift/Core ML GLiNER 2.5-small; entity spans + confidence; not Choice/Score/Noul; not TypeSafe; label descriptions as schema; on-device ANE economics; honesty locks; shershah1024/gliner-native-runtime ≠ Fastino; ≠ gliner25-compaction ≠ gliner2-ultrafast ≠ Eran-BA/Jev_from_GLiNER2 ≠ NSStudent/JevSwiftSDK ≠ jevmlx; default threshold 0.1 still soft. **Hourly 1740:** Decision Graph Protocol frame→assess→commit; app retains permissions/effects; Jev-first assessor-neutral; guarded commit / receipt/next frame; assessment batching; hard-gating DGP as safety theater; numerous-com/dgp ≠ TypeSafe official; jegrep calibrated path+range Nouls; no embeddings/index/daemon; ~$0.01–0.03 typical; agent --json; can1357/jegrep ≠ Bentlybro/jevgrep ≠ uehaj/jev-semgrep; Archer-arch fidelity; kev family OOD 0.76–0.77 vs Jev 0.86; block-causal isolation; pointer/readout CE-trained; /v1/systemone drop-in; replica honesty. **Hourly 1843:** judgment vs generation; deterministic execution after probabilistic judgment; exactly one app-owned callback; explicit uncertain branch; YES / NO / UNSURE from cost_false_yes / cost_false_no / cost_human; auto-batching same-object questions; Kungie/gut ≠ tpellet/hunch ≠ carldaws/hunch; Illusion47586/judge ≠ lexingtonhibiki/judgekit ≠ Ascurse/typed-judge-kit. **Hourly 1943:** Jev IS the if-statement; judgments/probabilities drive branches; text model only writes prose; interpreter owns variables/loops/budgets/replay; otherwise maybe / confidence gate; chaos samples after the gate; Jev-first Pi agent loop; slow-LLM fallback **Hourly 2041:** resume-screening bias audit methodology; Plan/PRD panel → code-owned pass|review|block; cost-aware multi-model routing/escalation; frozen-protocol zero-shot bench; context-window admission control; typed decision control plane; receipt ≠ authorization; live 15-dim typed rubric re-score per pause; adversarial pre-registered Jev eval; confidence does not track ignorance; provider-neutral Elixir/BEAM Noul/Choice/Score SDK **Hourly 2145:** question-linting of Jev questions themselves; nine jaggedness rules, no API key, no labelled data; static lint ≠ measured separation; yodablocks/jevq ≠ tenbin ≠ JevLint ≠ commitjev; open-weights Laya as class exemplar (binding); Nx/Bumblebee runtime; host chooses backend; ChristianAlexander/laya_ex ≠ system_one_sdk ≠ dannote/jev ≠ NandhaKishorM/laya; on-chain/edge Laya deploy; parity_verified stays false; model output never grants Tx; humandebri/IC-Laya ≠ laya_ex; auditable weekend replica; Jev outputs never used for training; unpaired 0.577 vs 0.727; agilabs-ai/jev48 ≠ JevBench ≠ Mapika/decider; adversarial dual-judge / framing attack surface; comparative framing is the usable judgment; prior injection crowds out evidence; copyleftdev/ember ≠ ember.js; Laya specialist fine-tune pipeline; training still GPU-pending; PIXELZX0/XERON ≠ convaiinnovations/laya; Hub Laya replica drop; daliborsb/laya ≠ convaiinnovations/laya ≠ NandhaKishorM/laya; System One student distillation corpus; gold is programmatic; teacher is closed-API clone; MagaBitmex/jev-4b-distill-data ≠ missing student checkpoint; non-LLM VIN System One; planning depth not chat; lewislululu/jevon ≠ douglance/jevon; source-bound evidence checks; local quote mismatch needs no API; exit 0 ≠ claim truth; WaynezProg/jev-kit ≠ jonathanavis96/jev-kit (Airlock) ≠ jev-use ≠ jev-mcp **Hourly 2246:** independent System One evidence catalog; 19 reviewed records; scores not one leaderboard; no external record currently reproduced; TokenTrim no-Jev matched hybrid 62.4%; reachjalil/system-one-bench ≠ mallahyari/system-one-benchmark; 21 tasks · 134 items · 208 questions; scenes from public GitHub contracts, not production logs; SivletLabs/jev-eval ≠ willkelly/jev-evaluation ≠ 4esv/jev-eval ≠ xxkuboxx/jev-eval ≠ onlyoneaman/jev-eval ≠ dayhaysoos/jevals; option isolation (sibling-blind); permutation-equivariant; Hub OWNER not published; nafisazizir/hev ≠ jaredpalmer/kev; frozen local LLM logits, no trained decision head; residual-head 9,222-param decreased 73/96→67/96; confidence = 1−normalized entropy, not P(correct); yuki-oshio/mini-jev ≠ r-ms/mini-jev; Jev classifier as autoregressive next-token predictor; ChatJev-style soundness theater; erik-dunteman/ChatJev ≠ dannote/jev ≠ jev-gpt; calibrated decision head × AlphaProof value head; implementation-layer isomorphism, semantic difference; timeout = censoring; do not launder Noul as proof; parallel rank-prediction vs serial selection; independent questions can conflict; zzzzzec/jevsort ≠ keltokhy/jsort; curated open System One ecosystem catalog; rupeshpoojary9/awesome-open-system-one ≠ AnotiaWang/awesome-jev; arXiv paper radar with Jev relevance scoring; ranking ≠ calibration / 0.5 still soft; fail-open failed evals not marked seen **Hourly 2340:** train calibrated ~27M from scratch; typed Q→prob dist / one forward pass / no LLM decode; hyusi2003/MiniSystemOne ≠ Colvin0315/MiniSystemOne; description-only stub / size 5; ESCI hard probe fails four of six; jev_bool ECE 0.242 inversion 0.255; do not re-fold §60 six-gates as new; jobbyjev one-request-per-company from batch-size result; find/design/evaluate TypeSafe Jev decision loops; karanb192/jev-architect ≠ samtay32/jev-system-architect; Jairik/jev-distiller size 1; distill-Jev UI stub / do not distill Jev as teacher of record; post-launch scored use-case map / Jev self-scores then human curation; licensedsaucer9-web/jev-opportunities; Jev-inize a use case into classifier/router; gavinHuang/jevinize → simple-jev not TypeSafe; featherless-ai/simple-jev; compare saved decisions / same label can still change the branch; VihaanAgarwal/jev-diff ≠ Saik0s/diffusiongemma-jev-macos; not tested with a live Jev API key; constrained logprob + temp/Platt ≠ Noul; OpenJevPro pastes openjev-sglang JevBench as own; zhangcy122/OpenJevPro ≠ IamBusy/OpenJev ≠ ekzhang/openjev-sglang; PolyForm Noncommercial; SmolLM-135M / sub-70ms / 0 output tokens; demo P(True) 0.5052 / Choice conf 0.2872 / Score conf 0.0055; README claims MIT / GitHub license null / no LICENSE file; patelvishwa112/jev-system-one-rlcd ≠ arnabgho/rlcd-lite ≠ blackwood-rlcd; source-backed Awesome Jev radar / 306+ commit-pinned; logicrw/awesome-jev-projects ≠ AnotiaWang/awesome-jev ≠ yibie/awesome-jev ≠ cobanov/awesome-jev ≠ rupeshpoojary9/awesome-open-system-one; auto GitHub sync / Issue-only submissions; hashed n-gram encoder / rival-aware attention; olanotolu/jevbetter vs jevlike starter; synthetic hard menus top-1 0.916 vs 0.873 / ECE 0.0182 vs 0.0367 / 40 vs 4608 menus/sec; shuffled-context control 0.335 | `references/mixed-architecture.md` Apply 0042 (`notes.md` §105): structured probability readouts; distribution > argmax; Noul 0.5 midpoint; score is expectation not integer; bare HTTP not SDK; Arohtea/jev-readout; Jev-style Choice/Score/Noul from ordinary models; optional DSH plugin; schema-valid ≠ calibrated; gulagala001/jevify ≠ Mintzs/jevify; Laya RLCD benchmark; 40.3% below constant-answer; open-weight measurement; mourad-ghafiri/laya-rlcd-benchmark ≠ yibie/laya-jev-lab; cheap fail-open semantic edge; second signal not sole; FastLoopError catch; SupremeDreamZ/jev-fastloop ≠ jev-ultrafast; asking more questions in one call; 0.980 at every N; nearly not fully deterministic; TheWebDevel/jev-fanout; Qwen3-VL perception + Jev decisions train RL; 0 model calls at deployment; VLM alone 1.7 vs +Jev 4.4; harneet2512/reflexrl ≠ khordoo/jev-reflex-autonomy-lab; independent Jev API vs Laya; cascade 0.60 matches 78% at 1.8×; noul facts not judgements; yibie/laya-jev-lab ≠ dairui1/jev-lab ≠ BrendanH18/jev-lab; GLiNER vs GLiFormer vs Laya vs Jev; extractors ≠ decision engines; Laya dict-instructions collapse 58.3%; umstek/zero-shot-ie-bench; decisions-per-minute & cost; 204 moves vs 73; throughput not intelligence; angelgalvisc/snake-arena-jev-vs-llms ≠ vtrivedy/jev-plays-games; behavioral contracts; pin expectations eval upgrades; raw 0.94 is not a release; sathariels/jevcheck ≠ dayhaysoos/jevals ≠ SivletLabs/jev-eval; evidence-linked dependency upgrade; Jev never generates filenames; no_direct_evidence ≠ safe to merge; GaneshVG18/upgrade-radar ≠ LYchoon/paper-radar-jev; discography theme/mood/complexity; five atomic questions one call; lirantal/discoprint / Apply 0145 (`notes.md` §106): Turn any open LLM into System-One Jev; description-only stub / size 0; Exu is a toolkit, not a method; typed Q → probability dists; JSON parse of generated text ≠ Noul; classifier.dev fast tier 84.8 is Jev behind its own API; hard budget filter before Jev; Jev judges the next state, XState enforces transitions; catalog gravity; ★339 live REST; query-side encoders, not a Jev replica; transformers.js AutoModel cannot load this graph; This Space contains no benchmark result yet; do not reopen or amend PR #23. / Apply 0243: Benchmark-driven Jev router and judge; cheap alone is not success; Jev does not write, sum prices, or claim accuracy %; Sol 94.2 / Luna 83.9 / Jev path 89.7; 19.2% Sol / 62.3% cost save / 4.5pp miss of 2pp non-inferiority; p50 latency worse than Sol due to routing overhead; erendikmenn/jev-llm-router-benchmark ≠ jev-rag-benchmark ≠ ryantsai/jev-llm-router; Express + node:sqlite; mock and Jev decision engines; previous_ticket_count >= 3 is code; MIN_CONFIDENCE 0.6 still soft; substring false positives; aesaganda/jev-ticket-router ≠ SarathChandraBellam/jev-vs-llm-ticket-router; Universal Figure & Diagram Router; confidence ≥ 0.85 hard-gate is theater; generative AI banned from scientific plots; six visual branches; human-labeled (state, question, label); 166,054 rows / 22 configs; soft_label for human uncertainty; Praveenrajus/jev-bench ≠ fstandhartinger/jevbench; ternary bonsai System One GGUF; Hub does not ship weights; 100/100 easy T/F is not Harbor; label_mass ≠ correctness; stock llama.cpp Q2_0 silently gibberish; NicolaiMTLassen/open-bonzi-jev ≠ NicolaiLassen; transformers.js DeBERTa ONNX; temperature 1.05; AutoModel from_pretrained works; onnx-community/open-jev-deberta-v3-large-ONNX ≠ system-one-qwen3.5-4b-scorer-ONNX; 107★ densify; GH 151M vs README 149.6M; PR #1 now closed unmerged; do not re-fold §71 claim-audit as a beat; typed decisions, RLCD, confidence-gated routing; structured ≠ correct; mock not live API; 26 tests; wjdjdakf17/jev-study ≠ baekenough/jev-study / Apply 0345 (`notes.md` §108): **Open reproduction densifies / measurement densifies PRIMARY** (bonzi-27b-v2 / ternary-8b / 27b-v1 GGUF family densify; WANLI-256 74.6% / 65.2% / 71.1% *theirs*; Bonsai 1 27B Q1_0 runs on stock llama.cpp; ternary still needs PrismML fork; hf:heman10x/openJev-verdict-2.0 twin tokenizer-only; OpenJev Vision image classification + uncertainty; CLEVR-4 held-out joint 0%; hfdataset:IamBusy/OpenJev-Vision-Research-v0.1 12,832; 294,912 derived targets not independent samples; Laya multilingual ONNX WebGPU typed-decisions port; 63/63 selected answers / 5.1e-4 CPU / 1.2e-2 WebGPU; UpHash-Network/mini-jev is yuki-oshio transfer; jev-injection-bench 11,900 labelled prompts; Jev best ranking / Haiku better ECE 0.021 vs 0.058; 0.5–0.9 band is where Jev's numbers do not mean what they say; Prompt wording moves panic 28%; manojlds/jev-dspy-bench ≠ dspachos/jev-dspy ≠ jmanhype/jev-dspy-lab; Jev agreement is similarity, never ground truth; no aggregate quality grade or merge gate; AbstentionBench-on-Jev rank 1 of 20 vs 2025 field; question-asymmetry; forward-looking 0.465 never extreme; openkev calibration layer not a runtime; ECE vs coverage independent; select_threshold returns inf; escalation catches uncertainty not ignorance; misakaikato/openkev ≠ jaredpalmer/kev; pdf-race Docling→Jev vs Gemini; parser owns the wall clock; 12/12 tie is a tie; titles selected not generated; flopcheck 16 calibrated tweet judgments; mechanical tells in code; ZeroX-01/jev-atlas ≠ Zaious/jev-capability-atlas ≠ gorock007/jev-atlas; Laya calibration lab Gradio MCP; T never changes argmax; confidence ≠ top-label p; easy probe set refused; 40–48 rows too small to ship T; do not reopen or amend PR #23 or #24 or #25. Apply 0439 (`notes.md` §109): Gemma-4 26B-A4B jevify classification+calibration; LoRA adapter twin not independent eval; Gemma-4 E4B jevify; E4B LoRA stub card; kushalpatil/jevify-gemma4 ≠ Mintzs/jevify ≠ gulagala001/jevify ≠ uspraveen/Jevify; GH kushalpatil07/jevify 404; PAWS 0.580/ece 0.288 is the weak cell; smaller E4B slightly better OOD ECE than 26B-A4B; Hub jevify merged LoRA ships weights; bonzi Bonsai-8B v1 GGUF densify; Bonsai-1.7B v1; Bonsai-4B v1; WANLI-256 64.5% / 60.2% / 52.0% *theirs*; rank #4 / #5 / #6 of 6; JulesHuisman/jev-eval scaffolding / README SHA c356a584 (was empty e69de29b); JulesHuisman/jev-eval ≠ SivletLabs/jev-eval ≠ willkelly/jev-evaluation ≠ 4esv ≠ xxkuboxx ≠ onlyoneaman ≠ dayhaysoos/jevals; 7 bands 6/10 vs 40 bands 0/10; source receipts + confidence slider re-policy without re-inference; 32/32 synthetic is smoke not production; classify HF datasets across typed semantic dimensions; roadus2 watch misspelling; lock roadius2/ultra_laya; ultra_laya REVIEW defects; default branch claude/laya-jev-review-gg5ppo; XNLI EN 88.3% ECE 0.032 → RU 77.3% ECE 0.096; Δ −11.0 pp [−14.2,−7.8]; ECE +0.063; MASSIVE no detectable difference at n=600; confidence is function of p_max (r=1.000); pointer-not-generator 400 human-authored responses; proposed ≠ authorized; FewRel 160: Jev 85.0% vs lexical 13.125%; gated 100% (95/95) coverage 59.375%; J++ composable semantic computation language; judge-jev 0.5 still soft; 947 repos scored; A 273 / B 302 / C 372; LLM rubric ≠ benches; No benchmark winner is claimed; phishing: naive 62.6% vs regex 91.8%; 5-atomic + LR 95.0% *theirs*; AITuber tension ±15; README npm global; repo is Rust; git-confess code owns counting/blame/ratio; httpx exhibit 11% (13/119) *theirs*; 90d trend +12.40% vs random +12.75% vs BH +41.71%; 5m win rate 25%; Awesomejev 656 entries / 38,160 stars; tracker likes 64 (+4) lastModified UNCHANGED; Laya present; Blackwood ABSENT; Archer still promised_not_landed; do not reopen or amend PR #23/#24/#25/#26. Apply 0541 (`notes.md` §110): Blackwood tracker ABSENT; likes 2 gated manual; r = c - p_a; ECE 0.021; acc 0.807 vs warmup 0.746; calibration beyond ~500 tokens unmeasured; Independent primitive; 11.57s vs 54.10s · 4.67× · 120/128 *theirs*; default path is pretrained Gemma probs not trained RLCD head; GH Meanblock 404; lock leesk212/JEV-CPU; softmax over letter slots ≠ Noul; WANLI 0.741 vs openjev v2 0.77 *theirs*; 3-way NLI ≠ Noul; priority 0.464 = majority floor; banking77 contaminated; raw margins not probabilities; GH jev-haiku-benchmarking 404; do not distill Jev as teacher of record (they distilled Haiku); “0.9 is not one number”; ranking ≠ calibration; banking77 0.8–0.9 stated 0.86 actual 0.73 over-confident *theirs*; ≠ Praveenrajus/jev-bench ≠ fstandhartinger/jevbench; $0.0000153–$0.0000226 vs circulating $0.0004 (~20×); Score is 0..n-1 expectation not 0–1; Noul has no confidence field; TCP floor 198.8 ms; type reliability is not a reason to choose Jev (json_schema 5/5); gateway tax not one number; Function-only 5/8 vs hybrid 8/8; 4/8 without Jev; 8 designed cases not conversion lift; ≠ RadRebelSam/awesome-jev; 200-row pilot Jev 86.5% 173/200 vs Gemini Flash-Lite 86.0% 172/200 vs Pro 87.0% 174/200 *theirs*; not a ranking; NLI Tetris argmax P(entail)−P(contradict); 情緒測謊器; 1q 396ms / 30q 567ms; ±0.03; 33q $0.000045 vs Gemini ~5× slower ~60× cost *theirs*; ≠ realZachi/jevtest; 8-example Jev vs GPT-5.6 Sol ~64× cost 5.4× latency *theirs*; synthetic; no inference; ≠ JevBench v1.2 §78; Judged 3317 / listed 2560; Jev judges, code applies policy; catalog ≠ endorsement; APA “microsecond policy / zero hallucination” overclaim; Client-side quiz; pointer from held docs; scanned-PDF warn; CSP only api.typesafe.ai; Jev judges / agent reasons / user decides; selecting an option is not permission to implement; degraded fallback; pattern exact, judgement must clear floor; no matching pattern → no model call; not a correctness oracle; $0.00022 vs chat $0.00306 *theirs*; Spec vs artifact remainder; treating 0.85 as 85% / minProbability hard-gate as Harbor; VERIFY acquires discriminating evidence, never same-pool confidence-only rescoring; fast/full/max are ceilings not sizes; Solar writes, Jev chooses NEXT ACTION; SemIf 2186★ (+20 vs §109 2166); jevlike 1038★ (+7 vs 1031); TypeAR 14★ flat; AnotiaWang 96★ (+1 vs 95); yibie/awesome-jev 490★; Laya likes 802 (was 783); tracker likes 64 flat, lastModified UNCHANGED; do not reopen or amend PR #23/#24/#25/#26/#27. Apply 0743 (`notes.md` §113): Hourly 0743 uniqueness lock: Heman10x-NGU/Verdict-open-jev ≠ Heman10x-NGU/openJev-verdict-2.0; TF-IDF + LogReg ECE 0.0207 vs Jev 0.1440; Verdict-open-jev 48.07% vs Jev 90.80%; abstention combined recall 10.00%; p50 35.58 ms; K=25 (maximum capacity) 72.00%; 0.85 coverage 84.60% selective risk 1.18%; 26.1× faster than standard Qwen JSON generation; Jevify 90.0% / 167 ms CUDA graphs disabled; Finding 1: Brier on stated confidence alone is a trap; grpo_rlcr 0.78 / ECE 0.084; reliability 0.007 but resolution 0.000; 27 900 schema-driven decisions; 13 600 / 13 600 questions; candidate mass min 0.99999624; 22 configs · 166,054 rows · 4 calibration-gold; sha a39eba3f; Student B MAE 0.148 / Pearson 0.836 / 86.0%; pngwn/open-jev-laya-bench README 404; sha 9f69c742 likes 2; HDFS 0.9933 (745/750) / retain 0.0084; BGL ERROR/FATAL protection 1.0000; 2,479 / 2,500 HDFS uncertain; cache hit 0.9648 (2412/2500); $0.153936 estimated; E2 recomputes from saved probabilities; Space sha eda59e0a; MASSIVE English 0.783 / Khmer 0.033 / Hindi 0.133; 40–48 rows too small to ship T; T never changes argmax; siren2345/jev-single-decode-transformers ≠ siren2345/jev-single-decode; Split Transformers experiment from llama.cpp runtime; tanayvasishtha/jev-lab ≠ dairui1/jev-lab ≠ BrendanH18/jev-lab ≠ yibie/laya-jev-lab; Four experiments stress-testing TypeSafe's Jev: calibration, bundle bias, label bias, and ensembling; second pass must be $0.00 from cache; The pages never call Jev; Gemma 4 31B 77.0% / Jev 1.13.0 61.4% / Laya 322M 0.0%; restriction state 95.0% against 84.4%; None of the systems are particularly good at knowing when to stop and ask; They skip the question and call a tool directly; 100% schema pass; six-field joint 48.8% vs 72.8%; ywchiu/jev_benchmark ≠ Running-Dolphins/jev-bench ≠ Praveenrajus/jev-bench; ACT / REVIEW / FALLBACK; A provider failure, timeout, malformed output, or missing answer is **not** a policy outcome; confidence is descriptive provider output, not a substitute for probability; Quality denominators include only valid scored answers; an exact halfway tie chooses the lower level; aiwithenoch/Jev-Skill ≠ simplosophy/jev-skill ≠ laguagu/jev-skills; The local path does not claim to turn a smaller checkpoint into Jev; Low support becomes decision: "review"; MIT-0 SPDX NOASSERTION; current-llm; 结构兼容,不是 Jev 模型能力; altryne/jevify ≠ Mintzs/jevify ≠ gulagala001/jevify ≠ uspraveen/Jevify; Find where Jev belongs. Design the questions. Measure the difference; TypeAR-AI/TypeAR 301 → TypeLLM/TypeLLM; TypeLLM/TypeLLM 16★; SemIf 2241★ (+34 vs §111 2207); jevlike 1051★ (+8 vs 1043); AnotiaWang 98★ (+1 vs 97); yibie/awesome-jev 525★ (+19 vs 506); Laya likes 864 (was 822); tracker likes 67 (+3 vs 64); lastModified UNCHANGED `2026-09-20T04:29:16.000Z`; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#32/#33; notes.md §113 Apply 0646 (`notes.md` §111): Calibration is not alpha; NO CURRENT ALPHA CANDIDATE; ΔR² approximately +0.00084; Brier 0.2131387; ECE 0.0421875; Adding Jev probability to deterministic volatility improved Brier by only 1.4058e-05; default 0.5 keeps zero non pinned; keepResult median 0.14 to 0.17; keepCall median 0.28 to 0.35; usable range is about 0.10 to 0.25; 7.8% to 57.9%; judges results it never sees; task-finish eval not built yet; $0.002 per compaction; slavadubrov/sgr-judge-bench ≠ slavadubrov/jev-judge-bench; Jev 108/120 $0.083 0.34 s; Luna SGR 114/120; paired Jev accuracy-difference intervals include zero; not evidence of equivalence; GLM SGR 26/120 93 format failures; Terra-planned Jev hybrid 55/120; rule-based by default, optionally Jev-backed; empty README; missing key cannot break the experience; prefill plus exactly one decode; softmax over A/B/C ≠ Noul; BBQ 9,053/10,000 (90.53%); ECE 0.0890; Mean confidence 0.9943; overconfident; score and noul not implemented; DGUI 12 rows (was 6); INSTRUCT 119 rows likes 2; encode the state once, decide everything in parallel; 0.740 accuracy against a 0.508 majority; ECE 0.047; fine-tune's advantage ends where its 384-token training data does; jasonkneen/open-jev ≠ pngwn/open-jev; same sha d41dc3cd; Space does not call Jev; recomputes routing from saved probabilities; 200-case Jev 97.0% / 100.0% / 95.0% / MAE 9.22; synthetic repository benchmark; Jev evaluations are advisory; YehuiTang0316/jev-nlgrep ≠ Bentlybro/jevgrep ≠ can1357/jegrep ≠ uehaj/jev-semgrep; default threshold 0.8 still soft; 40-line windows cannot prove whole function; token-native sequential start/end Choice; Gemini/Haiku stubs not configured yet; handful of hand-written examples, not a benchmark; Jev judged exactly what it was given; laguagu/jev-skills ≠ laguagu/jev-evidence-lab ≠ Pleo2/awesome-jev-agent-skills; contract_passed is not a claim of guaranteed factual truth; Wilson lower bound 0.85 floor; fixture mode no savings claim; SemIf 2207★ (+21 vs §110 2186); jevlike 1043★ (+5 vs 1038); TypeAR 15★ (+1 vs 14); AnotiaWang 97★ (+1 vs 96); yibie/awesome-jev 506★ (+16 vs 490); Laya likes 822 (was 802); tracker likes 64 flat, lastModified UNCHANGED; do not reopen or amend PR #23/#24/#25/#26/#27/#28.) | -| Context sieve | Relevance Noul per block; always-keep set in code; stub + recall key. Encoder cousin: GLiNER2.5 retention Choice + char-offset spans (gliner25-compaction); fail-closed keep_full; shadowMode default. Stdout cousin: jev-pruner (Jev Noul after hard ≤10k/JSON-diff envelope; fail-safe original; archive). Session-ledger cousin: carryforward (verbatim facts; Jev scores recall; rules never judged; fail-open dump). Classify-first MCP cousin: jev-sift (batch path/url/text → Jev without entering main agent context; uncertain/errors/truncation ≠ irrelevant). **Framework-agnostic compact+gate:** Jev judges relevance, code decides structure; never rewrite; regex floor in code; compaction fail-open if Jev down, safety gate fail-closed (jev-compactor later bench **73%** / 350 ms / 4 of 4 *theirs*; OpenCode port fast-jev-opencode already §62). **Pre-send views:** dizk/jev-lens (79% fewer tokens on 500 SWE-rebench trajectories; code full unless confident). **Observational memory:** pi-om keep/kind verbatim. **tools≠use:** carryforward SessionStart hook > MCP sitting there. **Hermes WHETHER/HOW/WHAT:** compact original chunks; skip next main-model when evidence is enough (hermes-jev-router; needs core patch; fail-open). **Empty compaction-proxy skip:** IPECTER/jev-context-pruner **and** IPECTER/jev-runway LICENSE-only. **Pi verbatim summarizer replacement:** pi-jev-compact (keep/drop tool calls; fail-open to LLM summary; ≠ vava-nessa/pi-jev-compaction). **Dedicated Pi port of fast-jev-compaction:** zaycruz/fast-jev-compaction-pi (verbatim keep/drop; fail-open to built-in LLM summary; ~50× *theirs*; **≠** pi-jev-compact **≠** pi-jev-compaction). **Atom then sense MCP** (enzo-mcp independently falsifiable claims; enzo-mcp UNKNOWN useful; **≠** jev-sift). **Hourly 1347:** 0thernet/system-one-skills deterministic verify; nanoprune 2.8MB ECE 2.58%. **OpenCode stdout-prune host:** indiejoseph/opencode-jev-pruner (jev-zen / jev-1.13-free; zen-chat ≠ Noul; fail-open original; keepScore >0.1 floor; ≠ tamaratran ≠ fast-jev-opencode; `notes.md` §96). **Hourly 1740 DGP ≠ this job** (protocol envelope around assessor; commit is the guard; `notes.md` §98). **Hourly 1843 gut auto-batch** is same-object fan-out, not a sieve. **Hourly 1943:** retrieve by relevance not resemblance; one calibrated yes/no per memory in one request; pointer mode 17/18 19/20 *theirs*; embedding resemblance misses the allergy. **Hourly 2041:** resume-screening bias audit methodology; Plan/PRD panel → code-owned pass|review|block; cost-aware multi-model routing/escalation; frozen-protocol zero-shot bench; context-window admission control; typed decision control plane; receipt ≠ authorization; live 15-dim typed rubric re-score per pause; adversarial pre-registered Jev eval; confidence does not track ignorance; provider-neutral Elixir/BEAM Noul/Choice/Score SDK **Hourly 2145:** question-linting of Jev questions themselves; nine jaggedness rules, no API key, no labelled data; static lint ≠ measured separation; yodablocks/jevq ≠ tenbin ≠ JevLint ≠ commitjev; open-weights Laya as class exemplar (binding); Nx/Bumblebee runtime; host chooses backend; ChristianAlexander/laya_ex ≠ system_one_sdk ≠ dannote/jev ≠ NandhaKishorM/laya; on-chain/edge Laya deploy; parity_verified stays false; model output never grants Tx; humandebri/IC-Laya ≠ laya_ex; auditable weekend replica; Jev outputs never used for training; unpaired 0.577 vs 0.727; agilabs-ai/jev48 ≠ JevBench ≠ Mapika/decider; adversarial dual-judge / framing attack surface; comparative framing is the usable judgment; prior injection crowds out evidence; copyleftdev/ember ≠ ember.js; Laya specialist fine-tune pipeline; training still GPU-pending; PIXELZX0/XERON ≠ convaiinnovations/laya; Hub Laya replica drop; daliborsb/laya ≠ convaiinnovations/laya ≠ NandhaKishorM/laya; System One student distillation corpus; gold is programmatic; teacher is closed-API clone; MagaBitmex/jev-4b-distill-data ≠ missing student checkpoint; non-LLM VIN System One; planning depth not chat; lewislululu/jevon ≠ douglance/jevon; source-bound evidence checks; local quote mismatch needs no API; exit 0 ≠ claim truth; WaynezProg/jev-kit ≠ jonathanavis96/jev-kit (Airlock) ≠ jev-use ≠ jev-mcp **Hourly 2246:** independent System One evidence catalog; 19 reviewed records; scores not one leaderboard; no external record currently reproduced; TokenTrim no-Jev matched hybrid 62.4%; reachjalil/system-one-bench ≠ mallahyari/system-one-benchmark; 21 tasks · 134 items · 208 questions; scenes from public GitHub contracts, not production logs; SivletLabs/jev-eval ≠ willkelly/jev-evaluation ≠ 4esv/jev-eval ≠ xxkuboxx/jev-eval ≠ onlyoneaman/jev-eval ≠ dayhaysoos/jevals; option isolation (sibling-blind); permutation-equivariant; Hub OWNER not published; nafisazizir/hev ≠ jaredpalmer/kev; frozen local LLM logits, no trained decision head; residual-head 9,222-param decreased 73/96→67/96; confidence = 1−normalized entropy, not P(correct); yuki-oshio/mini-jev ≠ r-ms/mini-jev; Jev classifier as autoregressive next-token predictor; ChatJev-style soundness theater; erik-dunteman/ChatJev ≠ dannote/jev ≠ jev-gpt; calibrated decision head × AlphaProof value head; implementation-layer isomorphism, semantic difference; timeout = censoring; do not launder Noul as proof; parallel rank-prediction vs serial selection; independent questions can conflict; zzzzzec/jevsort ≠ keltokhy/jsort; curated open System One ecosystem catalog; rupeshpoojary9/awesome-open-system-one ≠ AnotiaWang/awesome-jev; arXiv paper radar with Jev relevance scoring; ranking ≠ calibration / 0.5 still soft; fail-open failed evals not marked seen **Hourly 2340:** train calibrated ~27M from scratch; typed Q→prob dist / one forward pass / no LLM decode; hyusi2003/MiniSystemOne ≠ Colvin0315/MiniSystemOne; description-only stub / size 5; ESCI hard probe fails four of six; jev_bool ECE 0.242 inversion 0.255; do not re-fold §60 six-gates as new; jobbyjev one-request-per-company from batch-size result; find/design/evaluate TypeSafe Jev decision loops; karanb192/jev-architect ≠ samtay32/jev-system-architect; Jairik/jev-distiller size 1; distill-Jev UI stub / do not distill Jev as teacher of record; post-launch scored use-case map / Jev self-scores then human curation; licensedsaucer9-web/jev-opportunities; Jev-inize a use case into classifier/router; gavinHuang/jevinize → simple-jev not TypeSafe; featherless-ai/simple-jev; compare saved decisions / same label can still change the branch; VihaanAgarwal/jev-diff ≠ Saik0s/diffusiongemma-jev-macos; not tested with a live Jev API key; constrained logprob + temp/Platt ≠ Noul; OpenJevPro pastes openjev-sglang JevBench as own; zhangcy122/OpenJevPro ≠ IamBusy/OpenJev ≠ ekzhang/openjev-sglang; PolyForm Noncommercial; SmolLM-135M / sub-70ms / 0 output tokens; demo P(True) 0.5052 / Choice conf 0.2872 / Score conf 0.0055; README claims MIT / GitHub license null / no LICENSE file; patelvishwa112/jev-system-one-rlcd ≠ arnabgho/rlcd-lite ≠ blackwood-rlcd; source-backed Awesome Jev radar / 306+ commit-pinned; logicrw/awesome-jev-projects ≠ AnotiaWang/awesome-jev ≠ yibie/awesome-jev ≠ cobanov/awesome-jev ≠ rupeshpoojary9/awesome-open-system-one; auto GitHub sync / Issue-only submissions; hashed n-gram encoder / rival-aware attention; olanotolu/jevbetter vs jevlike starter; synthetic hard menus top-1 0.916 vs 0.873 / ECE 0.0182 vs 0.0367 / 40 vs 4608 menus/sec; shuffled-context control 0.335 | `references/applied-mappings.md#1-context-sieve` Apply 0042 (`notes.md` §105): structured probability readouts; distribution > argmax; Noul 0.5 midpoint; score is expectation not integer; bare HTTP not SDK; Arohtea/jev-readout; Jev-style Choice/Score/Noul from ordinary models; optional DSH plugin; schema-valid ≠ calibrated; gulagala001/jevify ≠ Mintzs/jevify; Laya RLCD benchmark; 40.3% below constant-answer; open-weight measurement; mourad-ghafiri/laya-rlcd-benchmark ≠ yibie/laya-jev-lab; cheap fail-open semantic edge; second signal not sole; FastLoopError catch; SupremeDreamZ/jev-fastloop ≠ jev-ultrafast; asking more questions in one call; 0.980 at every N; nearly not fully deterministic; TheWebDevel/jev-fanout; Qwen3-VL perception + Jev decisions train RL; 0 model calls at deployment; VLM alone 1.7 vs +Jev 4.4; harneet2512/reflexrl ≠ khordoo/jev-reflex-autonomy-lab; independent Jev API vs Laya; cascade 0.60 matches 78% at 1.8×; noul facts not judgements; yibie/laya-jev-lab ≠ dairui1/jev-lab ≠ BrendanH18/jev-lab; GLiNER vs GLiFormer vs Laya vs Jev; extractors ≠ decision engines; Laya dict-instructions collapse 58.3%; umstek/zero-shot-ie-bench; decisions-per-minute & cost; 204 moves vs 73; throughput not intelligence; angelgalvisc/snake-arena-jev-vs-llms ≠ vtrivedy/jev-plays-games; behavioral contracts; pin expectations eval upgrades; raw 0.94 is not a release; sathariels/jevcheck ≠ dayhaysoos/jevals ≠ SivletLabs/jev-eval; evidence-linked dependency upgrade; Jev never generates filenames; no_direct_evidence ≠ safe to merge; GaneshVG18/upgrade-radar ≠ LYchoon/paper-radar-jev; discography theme/mood/complexity; five atomic questions one call; lirantal/discoprint | +| Context sieve | Relevance Noul per block; always-keep set in code; stub + recall key. Encoder cousin: GLiNER2.5 retention Choice + char-offset spans (gliner25-compaction); fail-closed keep_full; shadowMode default. Stdout cousin: jev-pruner (Jev Noul after hard ≤10k/JSON-diff envelope; fail-safe original; archive). Session-ledger cousin: carryforward (verbatim facts; Jev scores recall; rules never judged; fail-open dump). Classify-first MCP cousin: jev-sift (batch path/url/text → Jev without entering main agent context; uncertain/errors/truncation ≠ irrelevant). **Capability-tree lookup:** NiazMorshed2007/jcr one tool; returns context; **does not execute** (`notes.md` §116). **Framework-agnostic compact+gate:** Jev judges relevance, code decides structure; never rewrite; regex floor in code; compaction fail-open if Jev down, safety gate fail-closed (jev-compactor later bench **73%** / 350 ms / 4 of 4 *theirs*; OpenCode port fast-jev-opencode already §62). **Pre-send views:** dizk/jev-lens (79% fewer tokens on 500 SWE-rebench trajectories; code full unless confident). **Observational memory:** pi-om keep/kind verbatim. **tools≠use:** carryforward SessionStart hook > MCP sitting there. **Hermes WHETHER/HOW/WHAT:** compact original chunks; skip next main-model when evidence is enough (hermes-jev-router; needs core patch; fail-open). **Empty compaction-proxy skip:** IPECTER/jev-context-pruner **and** IPECTER/jev-runway LICENSE-only. **Pi verbatim summarizer replacement:** pi-jev-compact (keep/drop tool calls; fail-open to LLM summary; ≠ vava-nessa/pi-jev-compaction). **Dedicated Pi port of fast-jev-compaction:** zaycruz/fast-jev-compaction-pi (verbatim keep/drop; fail-open to built-in LLM summary; ~50× *theirs*; **≠** pi-jev-compact **≠** pi-jev-compaction). **Atom then sense MCP** (enzo-mcp independently falsifiable claims; enzo-mcp UNKNOWN useful; **≠** jev-sift). **Hourly 1347:** 0thernet/system-one-skills deterministic verify; nanoprune 2.8MB ECE 2.58%. **OpenCode stdout-prune host:** indiejoseph/opencode-jev-pruner (jev-zen / jev-1.13-free; zen-chat ≠ Noul; fail-open original; keepScore >0.1 floor; ≠ tamaratran ≠ fast-jev-opencode; `notes.md` §96). **Hourly 1740 DGP ≠ this job** (protocol envelope around assessor; commit is the guard; `notes.md` §98). **Hourly 1843 gut auto-batch** is same-object fan-out, not a sieve. **Hourly 1943:** retrieve by relevance not resemblance; one calibrated yes/no per memory in one request; pointer mode 17/18 19/20 *theirs*; embedding resemblance misses the allergy. **Hourly 2041:** resume-screening bias audit methodology; Plan/PRD panel → code-owned pass|review|block; cost-aware multi-model routing/escalation; frozen-protocol zero-shot bench; context-window admission control; typed decision control plane; receipt ≠ authorization; live 15-dim typed rubric re-score per pause; adversarial pre-registered Jev eval; confidence does not track ignorance; provider-neutral Elixir/BEAM Noul/Choice/Score SDK **Hourly 2145:** question-linting of Jev questions themselves; nine jaggedness rules, no API key, no labelled data; static lint ≠ measured separation; yodablocks/jevq ≠ tenbin ≠ JevLint ≠ commitjev; open-weights Laya as class exemplar (binding); Nx/Bumblebee runtime; host chooses backend; ChristianAlexander/laya_ex ≠ system_one_sdk ≠ dannote/jev ≠ NandhaKishorM/laya; on-chain/edge Laya deploy; parity_verified stays false; model output never grants Tx; humandebri/IC-Laya ≠ laya_ex; auditable weekend replica; Jev outputs never used for training; unpaired 0.577 vs 0.727; agilabs-ai/jev48 ≠ JevBench ≠ Mapika/decider; adversarial dual-judge / framing attack surface; comparative framing is the usable judgment; prior injection crowds out evidence; copyleftdev/ember ≠ ember.js; Laya specialist fine-tune pipeline; training still GPU-pending; PIXELZX0/XERON ≠ convaiinnovations/laya; Hub Laya replica drop; daliborsb/laya ≠ convaiinnovations/laya ≠ NandhaKishorM/laya; System One student distillation corpus; gold is programmatic; teacher is closed-API clone; MagaBitmex/jev-4b-distill-data ≠ missing student checkpoint; non-LLM VIN System One; planning depth not chat; lewislululu/jevon ≠ douglance/jevon; source-bound evidence checks; local quote mismatch needs no API; exit 0 ≠ claim truth; WaynezProg/jev-kit ≠ jonathanavis96/jev-kit (Airlock) ≠ jev-use ≠ jev-mcp **Hourly 2246:** independent System One evidence catalog; 19 reviewed records; scores not one leaderboard; no external record currently reproduced; TokenTrim no-Jev matched hybrid 62.4%; reachjalil/system-one-bench ≠ mallahyari/system-one-benchmark; 21 tasks · 134 items · 208 questions; scenes from public GitHub contracts, not production logs; SivletLabs/jev-eval ≠ willkelly/jev-evaluation ≠ 4esv/jev-eval ≠ xxkuboxx/jev-eval ≠ onlyoneaman/jev-eval ≠ dayhaysoos/jevals; option isolation (sibling-blind); permutation-equivariant; Hub OWNER not published; nafisazizir/hev ≠ jaredpalmer/kev; frozen local LLM logits, no trained decision head; residual-head 9,222-param decreased 73/96→67/96; confidence = 1−normalized entropy, not P(correct); yuki-oshio/mini-jev ≠ r-ms/mini-jev; Jev classifier as autoregressive next-token predictor; ChatJev-style soundness theater; erik-dunteman/ChatJev ≠ dannote/jev ≠ jev-gpt; calibrated decision head × AlphaProof value head; implementation-layer isomorphism, semantic difference; timeout = censoring; do not launder Noul as proof; parallel rank-prediction vs serial selection; independent questions can conflict; zzzzzec/jevsort ≠ keltokhy/jsort; curated open System One ecosystem catalog; rupeshpoojary9/awesome-open-system-one ≠ AnotiaWang/awesome-jev; arXiv paper radar with Jev relevance scoring; ranking ≠ calibration / 0.5 still soft; fail-open failed evals not marked seen **Hourly 2340:** train calibrated ~27M from scratch; typed Q→prob dist / one forward pass / no LLM decode; hyusi2003/MiniSystemOne ≠ Colvin0315/MiniSystemOne; description-only stub / size 5; ESCI hard probe fails four of six; jev_bool ECE 0.242 inversion 0.255; do not re-fold §60 six-gates as new; jobbyjev one-request-per-company from batch-size result; find/design/evaluate TypeSafe Jev decision loops; karanb192/jev-architect ≠ samtay32/jev-system-architect; Jairik/jev-distiller size 1; distill-Jev UI stub / do not distill Jev as teacher of record; post-launch scored use-case map / Jev self-scores then human curation; licensedsaucer9-web/jev-opportunities; Jev-inize a use case into classifier/router; gavinHuang/jevinize → simple-jev not TypeSafe; featherless-ai/simple-jev; compare saved decisions / same label can still change the branch; VihaanAgarwal/jev-diff ≠ Saik0s/diffusiongemma-jev-macos; not tested with a live Jev API key; constrained logprob + temp/Platt ≠ Noul; OpenJevPro pastes openjev-sglang JevBench as own; zhangcy122/OpenJevPro ≠ IamBusy/OpenJev ≠ ekzhang/openjev-sglang; PolyForm Noncommercial; SmolLM-135M / sub-70ms / 0 output tokens; demo P(True) 0.5052 / Choice conf 0.2872 / Score conf 0.0055; README claims MIT / GitHub license null / no LICENSE file; patelvishwa112/jev-system-one-rlcd ≠ arnabgho/rlcd-lite ≠ blackwood-rlcd; source-backed Awesome Jev radar / 306+ commit-pinned; logicrw/awesome-jev-projects ≠ AnotiaWang/awesome-jev ≠ yibie/awesome-jev ≠ cobanov/awesome-jev ≠ rupeshpoojary9/awesome-open-system-one; auto GitHub sync / Issue-only submissions; hashed n-gram encoder / rival-aware attention; olanotolu/jevbetter vs jevlike starter; synthetic hard menus top-1 0.916 vs 0.873 / ECE 0.0182 vs 0.0367 / 40 vs 4608 menus/sec; shuffled-context control 0.335 | `references/applied-mappings.md#1-context-sieve` Apply 0042 (`notes.md` §105): structured probability readouts; distribution > argmax; Noul 0.5 midpoint; score is expectation not integer; bare HTTP not SDK; Arohtea/jev-readout; Jev-style Choice/Score/Noul from ordinary models; optional DSH plugin; schema-valid ≠ calibrated; gulagala001/jevify ≠ Mintzs/jevify; Laya RLCD benchmark; 40.3% below constant-answer; open-weight measurement; mourad-ghafiri/laya-rlcd-benchmark ≠ yibie/laya-jev-lab; cheap fail-open semantic edge; second signal not sole; FastLoopError catch; SupremeDreamZ/jev-fastloop ≠ jev-ultrafast; asking more questions in one call; 0.980 at every N; nearly not fully deterministic; TheWebDevel/jev-fanout; Qwen3-VL perception + Jev decisions train RL; 0 model calls at deployment; VLM alone 1.7 vs +Jev 4.4; harneet2512/reflexrl ≠ khordoo/jev-reflex-autonomy-lab; independent Jev API vs Laya; cascade 0.60 matches 78% at 1.8×; noul facts not judgements; yibie/laya-jev-lab ≠ dairui1/jev-lab ≠ BrendanH18/jev-lab; GLiNER vs GLiFormer vs Laya vs Jev; extractors ≠ decision engines; Laya dict-instructions collapse 58.3%; umstek/zero-shot-ie-bench; decisions-per-minute & cost; 204 moves vs 73; throughput not intelligence; angelgalvisc/snake-arena-jev-vs-llms ≠ vtrivedy/jev-plays-games; behavioral contracts; pin expectations eval upgrades; raw 0.94 is not a release; sathariels/jevcheck ≠ dayhaysoos/jevals ≠ SivletLabs/jev-eval; evidence-linked dependency upgrade; Jev never generates filenames; no_direct_evidence ≠ safe to merge; GaneshVG18/upgrade-radar ≠ LYchoon/paper-radar-jev; discography theme/mood/complexity; five atomic questions one call; lirantal/discoprint | | Exact-text keep / drop | Choice include/exclude/mixed over candidates code already holds. Extractive quotes / pointer-not-generator (model never writes the excerpt; char-offset compaction same species). Observed a11y/DOM controls: score among them; code clicks (Jev or GLiNER2 or Cua-S1 option-attention). Harness productization: Stagehand extract pick-and-copy (Jev picks; code copies; schema/gate else LLM). **Closed-vote CU:** code builds options, Jev only picks, no planner LLM (JevOnly); host-owned handlers/permissions (waymode). **Hot-click CU:** indexed element table → operation+target in one request; code owns observe/execute/verify; text model only for type (ego-jev; ~2× vs per-step LLM, n=3, not a bench). **OCR+AX desktop CU:** numbered OCR+AX items; hosted Jev; writer only for free text (typesafe-computer-use; exclusive actions; split kind/item/site; **≠** jev-ultrafast). **ASR voice-browser CU:** regex spans; Jev picks; code copies (jev-voice-browser; numbered overlay, no second model). Meaning-as-spec: Cucumber .feature only; Jev picks among observed controls (jevcumber). **Evidence-synthesis pointer (choxos/jev-reviewer, ≠ egma-ai):** two-pass Choice (which line) + Noul (does this line itself answer); *Not found* is an answer; human check never overwritten. **Adversarial browser:** Playwright executes, Jev chooses next act; sample from the distribution not argmax; fail only high conf **and** high severity (browser-jev). **Decision-validated UI:** derive candidates, Jev selects, compiler emits; Jev never authors text (gram-render Telegram GramSpec; jev2ui A2UI + leftover Gemini). **Pointer search:** path Noul then window; Jev picks files/line ranges; code copies (JevFind; 0.25/0.55 still soft). **NL memory → beam-search FS:** findme; **≠** JevFind path-then-window. **Decision-as-filing:** pigeonhole OTHER skip. **Observe→score→act namesake:** ZHUBoer/ego-jev reserved `__none__`; runWorkflow completed ≠ success. **Downloads fail-open:** tidy none-of-folders stay. **Hourly 1541:** quarry evidence projection. **Hourly 1639:** OpenCode jev-pruner context sieve (verbatim chunks + markers). **SIGNAL §97 ≠ this job** (schema→spans locate, not keep/drop of held candidates). **Hourly 1740 jegrep** returns path+range pointers (ranking/gather, not keep/drop of held candidates). **Hourly 1843 jev-forge** scores *supplied* candidate bags (selector, not keep/drop of held bytes). **Hourly 1943:** name↔body / comment truth / test-claims; mizchi/jev-lint is mizchi/jevlint rename; no shipped rule has severity error. **Hourly 2041:** resume-screening bias audit methodology; Plan/PRD panel → code-owned pass|review|block; cost-aware multi-model routing/escalation; frozen-protocol zero-shot bench; context-window admission control; typed decision control plane; receipt ≠ authorization; live 15-dim typed rubric re-score per pause; adversarial pre-registered Jev eval; confidence does not track ignorance; provider-neutral Elixir/BEAM Noul/Choice/Score SDK **Hourly 2145:** question-linting of Jev questions themselves; nine jaggedness rules, no API key, no labelled data; static lint ≠ measured separation; yodablocks/jevq ≠ tenbin ≠ JevLint ≠ commitjev; open-weights Laya as class exemplar (binding); Nx/Bumblebee runtime; host chooses backend; ChristianAlexander/laya_ex ≠ system_one_sdk ≠ dannote/jev ≠ NandhaKishorM/laya; on-chain/edge Laya deploy; parity_verified stays false; model output never grants Tx; humandebri/IC-Laya ≠ laya_ex; auditable weekend replica; Jev outputs never used for training; unpaired 0.577 vs 0.727; agilabs-ai/jev48 ≠ JevBench ≠ Mapika/decider; adversarial dual-judge / framing attack surface; comparative framing is the usable judgment; prior injection crowds out evidence; copyleftdev/ember ≠ ember.js; Laya specialist fine-tune pipeline; training still GPU-pending; PIXELZX0/XERON ≠ convaiinnovations/laya; Hub Laya replica drop; daliborsb/laya ≠ convaiinnovations/laya ≠ NandhaKishorM/laya; System One student distillation corpus; gold is programmatic; teacher is closed-API clone; MagaBitmex/jev-4b-distill-data ≠ missing student checkpoint; non-LLM VIN System One; planning depth not chat; lewislululu/jevon ≠ douglance/jevon; source-bound evidence checks; local quote mismatch needs no API; exit 0 ≠ claim truth; WaynezProg/jev-kit ≠ jonathanavis96/jev-kit (Airlock) ≠ jev-use ≠ jev-mcp **Hourly 2246:** independent System One evidence catalog; 19 reviewed records; scores not one leaderboard; no external record currently reproduced; TokenTrim no-Jev matched hybrid 62.4%; reachjalil/system-one-bench ≠ mallahyari/system-one-benchmark; 21 tasks · 134 items · 208 questions; scenes from public GitHub contracts, not production logs; SivletLabs/jev-eval ≠ willkelly/jev-evaluation ≠ 4esv/jev-eval ≠ xxkuboxx/jev-eval ≠ onlyoneaman/jev-eval ≠ dayhaysoos/jevals; option isolation (sibling-blind); permutation-equivariant; Hub OWNER not published; nafisazizir/hev ≠ jaredpalmer/kev; frozen local LLM logits, no trained decision head; residual-head 9,222-param decreased 73/96→67/96; confidence = 1−normalized entropy, not P(correct); yuki-oshio/mini-jev ≠ r-ms/mini-jev; Jev classifier as autoregressive next-token predictor; ChatJev-style soundness theater; erik-dunteman/ChatJev ≠ dannote/jev ≠ jev-gpt; calibrated decision head × AlphaProof value head; implementation-layer isomorphism, semantic difference; timeout = censoring; do not launder Noul as proof; parallel rank-prediction vs serial selection; independent questions can conflict; zzzzzec/jevsort ≠ keltokhy/jsort; curated open System One ecosystem catalog; rupeshpoojary9/awesome-open-system-one ≠ AnotiaWang/awesome-jev; arXiv paper radar with Jev relevance scoring; ranking ≠ calibration / 0.5 still soft; fail-open failed evals not marked seen **Hourly 2340:** train calibrated ~27M from scratch; typed Q→prob dist / one forward pass / no LLM decode; hyusi2003/MiniSystemOne ≠ Colvin0315/MiniSystemOne; description-only stub / size 5; ESCI hard probe fails four of six; jev_bool ECE 0.242 inversion 0.255; do not re-fold §60 six-gates as new; jobbyjev one-request-per-company from batch-size result; find/design/evaluate TypeSafe Jev decision loops; karanb192/jev-architect ≠ samtay32/jev-system-architect; Jairik/jev-distiller size 1; distill-Jev UI stub / do not distill Jev as teacher of record; post-launch scored use-case map / Jev self-scores then human curation; licensedsaucer9-web/jev-opportunities; Jev-inize a use case into classifier/router; gavinHuang/jevinize → simple-jev not TypeSafe; featherless-ai/simple-jev; compare saved decisions / same label can still change the branch; VihaanAgarwal/jev-diff ≠ Saik0s/diffusiongemma-jev-macos; not tested with a live Jev API key; constrained logprob + temp/Platt ≠ Noul; OpenJevPro pastes openjev-sglang JevBench as own; zhangcy122/OpenJevPro ≠ IamBusy/OpenJev ≠ ekzhang/openjev-sglang; PolyForm Noncommercial; SmolLM-135M / sub-70ms / 0 output tokens; demo P(True) 0.5052 / Choice conf 0.2872 / Score conf 0.0055; README claims MIT / GitHub license null / no LICENSE file; patelvishwa112/jev-system-one-rlcd ≠ arnabgho/rlcd-lite ≠ blackwood-rlcd; source-backed Awesome Jev radar / 306+ commit-pinned; logicrw/awesome-jev-projects ≠ AnotiaWang/awesome-jev ≠ yibie/awesome-jev ≠ cobanov/awesome-jev ≠ rupeshpoojary9/awesome-open-system-one; auto GitHub sync / Issue-only submissions; hashed n-gram encoder / rival-aware attention; olanotolu/jevbetter vs jevlike starter; synthetic hard menus top-1 0.916 vs 0.873 / ECE 0.0182 vs 0.0367 / 40 vs 4608 menus/sec; shuffled-context control 0.335 | `references/applied-mappings.md#2-exact-text-keep--drop` Apply 0042 (`notes.md` §105): structured probability readouts; distribution > argmax; Noul 0.5 midpoint; score is expectation not integer; bare HTTP not SDK; Arohtea/jev-readout; Jev-style Choice/Score/Noul from ordinary models; optional DSH plugin; schema-valid ≠ calibrated; gulagala001/jevify ≠ Mintzs/jevify; Laya RLCD benchmark; 40.3% below constant-answer; open-weight measurement; mourad-ghafiri/laya-rlcd-benchmark ≠ yibie/laya-jev-lab; cheap fail-open semantic edge; second signal not sole; FastLoopError catch; SupremeDreamZ/jev-fastloop ≠ jev-ultrafast; asking more questions in one call; 0.980 at every N; nearly not fully deterministic; TheWebDevel/jev-fanout; Qwen3-VL perception + Jev decisions train RL; 0 model calls at deployment; VLM alone 1.7 vs +Jev 4.4; harneet2512/reflexrl ≠ khordoo/jev-reflex-autonomy-lab; independent Jev API vs Laya; cascade 0.60 matches 78% at 1.8×; noul facts not judgements; yibie/laya-jev-lab ≠ dairui1/jev-lab ≠ BrendanH18/jev-lab; GLiNER vs GLiFormer vs Laya vs Jev; extractors ≠ decision engines; Laya dict-instructions collapse 58.3%; umstek/zero-shot-ie-bench; decisions-per-minute & cost; 204 moves vs 73; throughput not intelligence; angelgalvisc/snake-arena-jev-vs-llms ≠ vtrivedy/jev-plays-games; behavioral contracts; pin expectations eval upgrades; raw 0.94 is not a release; sathariels/jevcheck ≠ dayhaysoos/jevals ≠ SivletLabs/jev-eval; evidence-linked dependency upgrade; Jev never generates filenames; no_direct_evidence ≠ safe to merge; GaneshVG18/upgrade-radar ≠ LYchoon/paper-radar-jev; discography theme/mood/complexity; five atomic questions one call; lirantal/discoprint | | Environment / harness triage | Scan every step for env failure; LLM autopsy only on flags. Merge-gate cousin: cluster in code, judge labels cause, policy owns PASS/BLOCK (latch; judge never says ignore alone). **Pre-review typed gate:** should_review/risk/route/touches_secrets → auto-approve|human-review|block (ci-gatekeeper; cheap before expensive LLM/human; operator-owned thresholds). **Stop-hook attention redirect:** eight risk axes, assist=one reinspect, fail-open, uncalibrated 0.85 (jev-preflight; not a merge blocker). **VOI hunk prune:** Jev scores PR hunks before expensive generative review (prune-review; 22-run 1.18% with 305% outlier *theirs*; ~20% target). **Whole-repo intent:** VERIFIED/VIOLATION/UNKNOWN beyond the diff (jev-intent-review). **Commit pre-review:** message vs diff / cohesion / omitted change; middle band is review not a verdict; regex proves literals (commitjev). **Hourly 1241:** jev-crawlers risk bands never raw boolean. **Hourly 1347:** alsoleg89/decide packing VOI; 0.8 ≠ 80% accuracy; 0thernet/system-one-skills deterministic verify | `references/applied-mappings.md#3-environment--harness-triage` | | Moderation and ranking | Hold-before-publish vs graded rerank; fail policy per action. Meaning-search without embeddings (jevgrep packed parallel; 79% top-5 on stripped repos; keyword still wins exact strings). **Meaning-grep** AND/OR/NOT over *thresholded* line Nouls; proposition≠embedding; Semgrep.dev collision; not a gate (jev-semgrep §86). **Evidence-packet explorer:** index-once ask-many, citable source_of_truth/tests/callers (jevex). Measured RAG rerank vs generative rerank (Jev-RAG one-run ≥70% cost / 72% latency vs Spark rerank; full-context Spark still faster). **Local rules first then remainder Nouls:** never auto-train on the model's own hides (x-reply-filter). Minimal consumer labels: bohutang/sift ~$0.00003/post. **Worth-your-attention VOI:** ThinkyMiner/Winnow read/skim/save/skip from typed answers (80%/90% *theirs*; always qualify vs kevinpita/winnow). **Personal-history ranking without a social graph:** jevfeed (one Jev request per batch of ten; distribution *is* ranking; history never uploaded). **OpenRouter recipe atlas:** jev-cookbook (triage/PII/rerank/browser; samples not benches). **Decision-native inbox:** mailordinal (nine typed signals → 100-point policy; humans own ambiguity). **Evidence-packet explorer delta:** jevex n=16 SWE 160s→69s / $8.74→$3.13 *theirs* (rename of jev-semantic-explorer). **Productized classification API:** classifier.dev (label+calibrated p; spam/inbox/feedback; **185★**). **Open NAR packaging:** NandhaKishorM/laya PyPI+Router (**710★**; Hub weights; not HTTP Jev). **Pointer search:** JevFind path then window (0.25/0.55 still soft). **Authorship named escape:** `human`/`ai_generated`/`uncertain` (not evidence). **n8n classify/route/score:** Route by Choice + Low Confidence (unofficial; 0.5 still soft). **Four-engine Harbor:** job-posting-triage majority floor 0.947; fitted tfidf wins; calibration ≠ discrimination. **Read-only Gmail trays:** jevmail gmail.readonly; **macOS inbox writes after review:** mailjay. **Hourly 1241:** jsort scores are relative; Noul not Choice for scale; muhammedilyasy/jev-mail metadata only; lkclean Show fail-open; jev-yt-time-saver Show anyway. **Hourly 1347:** nanoprune 2.8MB ECE 2.58%; alsoleg89/decide packing VOI. **Hourly 1740:** jegrep calibrated path+range Nouls; no embeddings/index/daemon; ~$0.01–0.03 typical; agent --json; can1357/jegrep ≠ Bentlybro/jevgrep ≠ uehaj/jev-semgrep. OpenRouter/TypeSafe auto-failover is silent FALLBACK. **Hourly 1843:** typed judgments vs chat judges on guardrailing; ishaannk/llm-vs-jev cross-note only; deeper integrity fold is rh-guard; nothing wins outright. **Hourly 1943:** ~1 in 5 findings wrong *theirs*; mizchi/jev-lint ≠ huntedman/JevLint ≠ MichitoSugawara/jev-lint **Hourly 2041:** resume-screening bias audit methodology; Plan/PRD panel → code-owned pass|review|block; cost-aware multi-model routing/escalation; frozen-protocol zero-shot bench; context-window admission control; typed decision control plane; receipt ≠ authorization; live 15-dim typed rubric re-score per pause; adversarial pre-registered Jev eval; confidence does not track ignorance; provider-neutral Elixir/BEAM Noul/Choice/Score SDK **Hourly 2145:** question-linting of Jev questions themselves; nine jaggedness rules, no API key, no labelled data; static lint ≠ measured separation; yodablocks/jevq ≠ tenbin ≠ JevLint ≠ commitjev; open-weights Laya as class exemplar (binding); Nx/Bumblebee runtime; host chooses backend; ChristianAlexander/laya_ex ≠ system_one_sdk ≠ dannote/jev ≠ NandhaKishorM/laya; on-chain/edge Laya deploy; parity_verified stays false; model output never grants Tx; humandebri/IC-Laya ≠ laya_ex; auditable weekend replica; Jev outputs never used for training; unpaired 0.577 vs 0.727; agilabs-ai/jev48 ≠ JevBench ≠ Mapika/decider; adversarial dual-judge / framing attack surface; comparative framing is the usable judgment; prior injection crowds out evidence; copyleftdev/ember ≠ ember.js; Laya specialist fine-tune pipeline; training still GPU-pending; PIXELZX0/XERON ≠ convaiinnovations/laya; Hub Laya replica drop; daliborsb/laya ≠ convaiinnovations/laya ≠ NandhaKishorM/laya; System One student distillation corpus; gold is programmatic; teacher is closed-API clone; MagaBitmex/jev-4b-distill-data ≠ missing student checkpoint; non-LLM VIN System One; planning depth not chat; lewislululu/jevon ≠ douglance/jevon; source-bound evidence checks; local quote mismatch needs no API; exit 0 ≠ claim truth; WaynezProg/jev-kit ≠ jonathanavis96/jev-kit (Airlock) ≠ jev-use ≠ jev-mcp **Hourly 2246:** independent System One evidence catalog; 19 reviewed records; scores not one leaderboard; no external record currently reproduced; TokenTrim no-Jev matched hybrid 62.4%; reachjalil/system-one-bench ≠ mallahyari/system-one-benchmark; 21 tasks · 134 items · 208 questions; scenes from public GitHub contracts, not production logs; SivletLabs/jev-eval ≠ willkelly/jev-evaluation ≠ 4esv/jev-eval ≠ xxkuboxx/jev-eval ≠ onlyoneaman/jev-eval ≠ dayhaysoos/jevals; option isolation (sibling-blind); permutation-equivariant; Hub OWNER not published; nafisazizir/hev ≠ jaredpalmer/kev; frozen local LLM logits, no trained decision head; residual-head 9,222-param decreased 73/96→67/96; confidence = 1−normalized entropy, not P(correct); yuki-oshio/mini-jev ≠ r-ms/mini-jev; Jev classifier as autoregressive next-token predictor; ChatJev-style soundness theater; erik-dunteman/ChatJev ≠ dannote/jev ≠ jev-gpt; calibrated decision head × AlphaProof value head; implementation-layer isomorphism, semantic difference; timeout = censoring; do not launder Noul as proof; parallel rank-prediction vs serial selection; independent questions can conflict; zzzzzec/jevsort ≠ keltokhy/jsort; curated open System One ecosystem catalog; rupeshpoojary9/awesome-open-system-one ≠ AnotiaWang/awesome-jev; arXiv paper radar with Jev relevance scoring; ranking ≠ calibration / 0.5 still soft; fail-open failed evals not marked seen **Hourly 2340:** train calibrated ~27M from scratch; typed Q→prob dist / one forward pass / no LLM decode; hyusi2003/MiniSystemOne ≠ Colvin0315/MiniSystemOne; description-only stub / size 5; ESCI hard probe fails four of six; jev_bool ECE 0.242 inversion 0.255; do not re-fold §60 six-gates as new; jobbyjev one-request-per-company from batch-size result; find/design/evaluate TypeSafe Jev decision loops; karanb192/jev-architect ≠ samtay32/jev-system-architect; Jairik/jev-distiller size 1; distill-Jev UI stub / do not distill Jev as teacher of record; post-launch scored use-case map / Jev self-scores then human curation; licensedsaucer9-web/jev-opportunities; Jev-inize a use case into classifier/router; gavinHuang/jevinize → simple-jev not TypeSafe; featherless-ai/simple-jev; compare saved decisions / same label can still change the branch; VihaanAgarwal/jev-diff ≠ Saik0s/diffusiongemma-jev-macos; not tested with a live Jev API key; constrained logprob + temp/Platt ≠ Noul; OpenJevPro pastes openjev-sglang JevBench as own; zhangcy122/OpenJevPro ≠ IamBusy/OpenJev ≠ ekzhang/openjev-sglang; PolyForm Noncommercial; SmolLM-135M / sub-70ms / 0 output tokens; demo P(True) 0.5052 / Choice conf 0.2872 / Score conf 0.0055; README claims MIT / GitHub license null / no LICENSE file; patelvishwa112/jev-system-one-rlcd ≠ arnabgho/rlcd-lite ≠ blackwood-rlcd; source-backed Awesome Jev radar / 306+ commit-pinned; logicrw/awesome-jev-projects ≠ AnotiaWang/awesome-jev ≠ yibie/awesome-jev ≠ cobanov/awesome-jev ≠ rupeshpoojary9/awesome-open-system-one; auto GitHub sync / Issue-only submissions; hashed n-gram encoder / rival-aware attention; olanotolu/jevbetter vs jevlike starter; synthetic hard menus top-1 0.916 vs 0.873 / ECE 0.0182 vs 0.0367 / 40 vs 4608 menus/sec; shuffled-context control 0.335 | `references/applied-mappings.md#4-moderation-and-ranking` Apply 0042 (`notes.md` §105): structured probability readouts; distribution > argmax; Noul 0.5 midpoint; score is expectation not integer; bare HTTP not SDK; Arohtea/jev-readout; Jev-style Choice/Score/Noul from ordinary models; optional DSH plugin; schema-valid ≠ calibrated; gulagala001/jevify ≠ Mintzs/jevify; Laya RLCD benchmark; 40.3% below constant-answer; open-weight measurement; mourad-ghafiri/laya-rlcd-benchmark ≠ yibie/laya-jev-lab; cheap fail-open semantic edge; second signal not sole; FastLoopError catch; SupremeDreamZ/jev-fastloop ≠ jev-ultrafast; asking more questions in one call; 0.980 at every N; nearly not fully deterministic; TheWebDevel/jev-fanout; Qwen3-VL perception + Jev decisions train RL; 0 model calls at deployment; VLM alone 1.7 vs +Jev 4.4; harneet2512/reflexrl ≠ khordoo/jev-reflex-autonomy-lab; independent Jev API vs Laya; cascade 0.60 matches 78% at 1.8×; noul facts not judgements; yibie/laya-jev-lab ≠ dairui1/jev-lab ≠ BrendanH18/jev-lab; GLiNER vs GLiFormer vs Laya vs Jev; extractors ≠ decision engines; Laya dict-instructions collapse 58.3%; umstek/zero-shot-ie-bench; decisions-per-minute & cost; 204 moves vs 73; throughput not intelligence; angelgalvisc/snake-arena-jev-vs-llms ≠ vtrivedy/jev-plays-games; behavioral contracts; pin expectations eval upgrades; raw 0.94 is not a release; sathariels/jevcheck ≠ dayhaysoos/jevals ≠ SivletLabs/jev-eval; evidence-linked dependency upgrade; Jev never generates filenames; no_direct_evidence ≠ safe to merge; GaneshVG18/upgrade-radar ≠ LYchoon/paper-radar-jev; discography theme/mood/complexity; five atomic questions one call; lirantal/discoprint | -| Skill / tool routing | Choice over a closed catalog + whether-anything-fits; code dispatches. Route ≠ memory: cheap intent gate skips memory tours on easy routes. Session-sticky first-prompt classification (lock for the session; fail-closed to a declared fallback). OMP/pi: `jev_route` topology/tier + `jev_acceptance_gate` before done (**fail-open**; contrast pi-jev-approver fail-closed). OMP prompt suppression: Jev grades gated calls; operator owns the bar; plugin never self-tunes (omp-greenlight; not a sandbox). **Outline only:** Hermes pre-agent skill broker — code owns grants; Jev never grants access (skill-broker; not a production recipe). **Constrained optimizer:** Jev supplies task/exactness/evidence features; controller owns SLO/quality floors; fail-open local features (slo-router; measured p95 77.93→490.38 same routes). **VOI over skill library:** two-pass + none-of-these; hook never blocks (skillranker 52★). **Sibling contrast:** skill-broker grants in code (fail-closed foundation-only) vs skillranker advisory vs turnstile runtime authorize. **Codex MCP adapter:** jev_select_capability / jev_search / jev_triage; caller supplies catalog; lexical fallback (jev-in-codex). **Harbor roster-size harness:** BM25 vs Jev at 50–500 (pi-jev-skill-bench; 43 gold; no live numbers this pass). **Pi strip-roster:** two-stage gate 0.30 / fits 0.40; no key → no-op; tool mode is tools≠use cousin (pi-jev-skill-suggestion). **Pi System-One control plane:** pi-jev-control (task/model/skill/memory/review/GUI; no live quality numbers; license null). **Host-adapter surface delta:** jev-routing adds Cursor Agent CLI / Devin CLI (still not MCP). **Hermes plugin ≠ TypeSafe:** hermes-plugin-jev is Agnes 3.0 Flash chat-completions branded as Jev. **Competing NAR agent-routing:** Cerebellum-2B pointer over candidates on `/v1/decide` (**not** TypeSafe `/v1/systemone`; wire-compat vs agent-routing as separate Harbor axes; do not endorse vs-Jev 94.92%/81.1%). **Jev-first bounded agent:** stanley-code; Empty findings ≠ approval. **Price workers, not the conversation:** jevsubrouter; Stats are counts, not dollars. **Cheap decision layer / skill honor:** grok-bot-jev; **APA harness cousin:** apa-agent-harness **≠** AntonioCoppe/jev-harness. **Hourly 1241:** yuyang2230/jev-agent-skill jev-1.13-free; jev-techstack-classifier stack_config.json; tpellet/hunch exit 3. **Hourly 1347:** hermes-switchyard ≠ hermes-jev-router ≠ hermes-plugin-jev; ctmx/openrouter-jev-mcp Decision-as-Plugin; claudecode-jev-marketplace fail-open not hot path; FrancoisChastel/jev-code ≠ npm jev-code. **Hourly 1541:** difficulty + policy thresholds + JSONL trace; jev-codex-pilot model + reasoning depth | `references/applied-mappings.md#5-skill--tool-routing` | +| Skill / tool routing | Choice over a closed catalog + whether-anything-fits; code dispatches. Route ≠ memory: cheap intent gate skips memory tours on easy routes. Session-sticky first-prompt classification (lock for the session; fail-closed to a declared fallback). OMP/pi: `jev_route` topology/tier + `jev_acceptance_gate` before done (**fail-open**; contrast pi-jev-approver fail-closed). OMP prompt suppression: Jev grades gated calls; operator owns the bar; plugin never self-tunes (omp-greenlight; not a sandbox). **Outline only:** Hermes pre-agent skill broker — code owns grants; Jev never grants access (skill-broker; not a production recipe). **Skills vs capabilities:** JCR looks up operation docs; it does not grant and does not execute (`notes.md` §116). **Constrained optimizer:** Jev supplies task/exactness/evidence features; controller owns SLO/quality floors; fail-open local features (slo-router; measured p95 77.93→490.38 same routes). **VOI over skill library:** two-pass + none-of-these; hook never blocks (skillranker 52★). **Sibling contrast:** skill-broker grants in code (fail-closed foundation-only) vs skillranker advisory vs turnstile runtime authorize. **Codex MCP adapter:** jev_select_capability / jev_search / jev_triage; caller supplies catalog; lexical fallback (jev-in-codex). **Harbor roster-size harness:** BM25 vs Jev at 50–500 (pi-jev-skill-bench; 43 gold; no live numbers this pass). **Pi strip-roster:** two-stage gate 0.30 / fits 0.40; no key → no-op; tool mode is tools≠use cousin (pi-jev-skill-suggestion). **Pi System-One control plane:** pi-jev-control (task/model/skill/memory/review/GUI; no live quality numbers; license null). **Host-adapter surface delta:** jev-routing adds Cursor Agent CLI / Devin CLI (still not MCP). **Hermes plugin ≠ TypeSafe:** hermes-plugin-jev is Agnes 3.0 Flash chat-completions branded as Jev. **Competing NAR agent-routing:** Cerebellum-2B pointer over candidates on `/v1/decide` (**not** TypeSafe `/v1/systemone`; wire-compat vs agent-routing as separate Harbor axes; do not endorse vs-Jev 94.92%/81.1%). **Jev-first bounded agent:** stanley-code; Empty findings ≠ approval. **Price workers, not the conversation:** jevsubrouter; Stats are counts, not dollars. **Cheap decision layer / skill honor:** grok-bot-jev; **APA harness cousin:** apa-agent-harness **≠** AntonioCoppe/jev-harness. **Hourly 1241:** yuyang2230/jev-agent-skill jev-1.13-free; jev-techstack-classifier stack_config.json; tpellet/hunch exit 3. **Hourly 1347:** hermes-switchyard ≠ hermes-jev-router ≠ hermes-plugin-jev; ctmx/openrouter-jev-mcp Decision-as-Plugin; claudecode-jev-marketplace fail-open not hot path; FrancoisChastel/jev-code ≠ npm jev-code. **Hourly 1541:** difficulty + policy thresholds + JSONL trace; jev-codex-pilot model + reasoning depth | `references/applied-mappings.md#5-skill--tool-routing` | | Expensive observation router | Structural prove (text layer) ∩ remainder Noul (needs OCR?) | `references/applied-mappings.md#6-expensive-observation-router` | | Capability kernel / human-confirmed gate | Secrets never in the agent; closed action space; Jev SENSOR; policy BLOCK/ASK/ALLOW. Distinct from pre-exec toolgate. Human is the only kill trigger; identity re-check; shields override; mapped explanations not raw model prose. **Permission vs probability:** operator owns auto-approve thresholds; plugin never self-tunes the safety bar; host deny stays above (omp-greenlight; not a sandbox). **Spoken confirm ≠ auth:** jev-voice-browser `destructive` spoken "confirm" is convenience not a guarantee (control-port reach is the grant; rh-guard owns the gate cousin). **Privilege ≠ verdict:** effect-based shell gate; fast structural rules then Jev Choice + independent risk Nouls; fail-closed (construct-auto-classifier; 0 dangerous / 975 *theirs*). **Jev supplies evidence, code owns authority** (actiongate-jev; a positive score never overrides a deterministic security failure). **Turnstile:** policy first, Jev after permit, missing Jev → Review, replay thresholds. **SEAL:** no seal, no advance; coverage.path visible; mint ≠ product brain. **Never confidently wrong:** TLA+ kernel may escalate, must not return a confident wrong (jev-labs). **Landed-script trust / headless≠auto-approve:** construct (byte-identical to remote default branch; headless escalation is deny-and-report, not auto-approve). **Typed baton:** escalate/continue/abort; gate `allow` never grants (jev-handoff). **Persist constraints across compaction:** conversational policy as structured state; Jev never writes policy; fail-open (pi-heed 98.5%/0 false block *theirs*). **Pi control-plane tool gate / GUI:** pi-jev-control (deterministic fast-path then Jev; GUI < threshold → unknown, never force-click). **jev-use PreToolUse gate:** deny/ask, fail-open, 12/12 *theirs*; Vercel margin fallback. **Commit-msg hook:** commitjev blocks only on a warning; instrument failure is not a refuse. Toolbelt notes: jev-security-scan / jev-decisions / TeoMastro; unofficial jev-cli not ready (≠ jevql); actiongate slogan already §64; **rh-guard owns reward-hack**. **Wrap-as-execution ALLOW/ASK/DENY:** the wrap *is* the tool function; rules first; ASK throws; fail-closed (AgentGhost; **≠** actiongate **≠** toolgate **≠** jev-use; rh-guard owns the gate cousin). **Tool-risk as placement (not a wrap card):** Akshay pedagogy cites LangChain middleware; AgentGhost owns wrap-as-execution. **Human every action:** Essentiel-Jev never authority. **Hourly 1241:** never-execute list; tab-bouncer pinned/audio/current never closed; ORIGIN pause-if-no-Jev. **Hourly 1347:** typed-gate band [0.40,0.60] is refusal; pi-jev-gate fail-closed; choice is the verdict; jev-calibration-arena never acts. **SIGNAL §93:** Cache hit ≠ correctness; score never auto-accepts (rh-guard thin). **SIGNAL §94:** guidance ≠ hook; unofficial ≠ TypeSafe; hosted bootstrap ≠ silent TypeSafe; Nemotron ≠ TypeSafe Jev; not a calibrated replacement (rh-guard thin). **Hourly 1541:** jevguard calibrator/cache/escape; jev-ci-selector CI shadow mode; one-dollar-tahoe TypeSafe Jev defense eval. **Hourly 1740:** Decision Graph Protocol frame→assess→commit; app retains permissions/effects; Jev-first assessor-neutral; guarded commit / receipt/next frame; assessment batching; hard-gating DGP as safety theater; numerous-com/dgp ≠ TypeSafe official. **Hourly 1843:** YES / NO / UNSURE from cost_false_yes / cost_false_no / cost_human; default on_unsure=raise is app policy not a System One hard gate; cost_human is VOI not a grant. **Hourly 1943:** memory leases ended by new evidence; six Nouls then fixed rules in code; 0 of 157 false invalidations; unsure → review queue; host keeps the store **Hourly 2041:** resume-screening bias audit methodology; Plan/PRD panel → code-owned pass|review|block; cost-aware multi-model routing/escalation; frozen-protocol zero-shot bench; context-window admission control; typed decision control plane; receipt ≠ authorization; live 15-dim typed rubric re-score per pause; adversarial pre-registered Jev eval; confidence does not track ignorance; provider-neutral Elixir/BEAM Noul/Choice/Score SDK **Hourly 2145:** question-linting of Jev questions themselves; nine jaggedness rules, no API key, no labelled data; static lint ≠ measured separation; yodablocks/jevq ≠ tenbin ≠ JevLint ≠ commitjev; open-weights Laya as class exemplar (binding); Nx/Bumblebee runtime; host chooses backend; ChristianAlexander/laya_ex ≠ system_one_sdk ≠ dannote/jev ≠ NandhaKishorM/laya; on-chain/edge Laya deploy; parity_verified stays false; model output never grants Tx; humandebri/IC-Laya ≠ laya_ex; auditable weekend replica; Jev outputs never used for training; unpaired 0.577 vs 0.727; agilabs-ai/jev48 ≠ JevBench ≠ Mapika/decider; adversarial dual-judge / framing attack surface; comparative framing is the usable judgment; prior injection crowds out evidence; copyleftdev/ember ≠ ember.js; Laya specialist fine-tune pipeline; training still GPU-pending; PIXELZX0/XERON ≠ convaiinnovations/laya; Hub Laya replica drop; daliborsb/laya ≠ convaiinnovations/laya ≠ NandhaKishorM/laya; System One student distillation corpus; gold is programmatic; teacher is closed-API clone; MagaBitmex/jev-4b-distill-data ≠ missing student checkpoint; non-LLM VIN System One; planning depth not chat; lewislululu/jevon ≠ douglance/jevon; source-bound evidence checks; local quote mismatch needs no API; exit 0 ≠ claim truth; WaynezProg/jev-kit ≠ jonathanavis96/jev-kit (Airlock) ≠ jev-use ≠ jev-mcp **Hourly 2246:** independent System One evidence catalog; 19 reviewed records; scores not one leaderboard; no external record currently reproduced; TokenTrim no-Jev matched hybrid 62.4%; reachjalil/system-one-bench ≠ mallahyari/system-one-benchmark; 21 tasks · 134 items · 208 questions; scenes from public GitHub contracts, not production logs; SivletLabs/jev-eval ≠ willkelly/jev-evaluation ≠ 4esv/jev-eval ≠ xxkuboxx/jev-eval ≠ onlyoneaman/jev-eval ≠ dayhaysoos/jevals; option isolation (sibling-blind); permutation-equivariant; Hub OWNER not published; nafisazizir/hev ≠ jaredpalmer/kev; frozen local LLM logits, no trained decision head; residual-head 9,222-param decreased 73/96→67/96; confidence = 1−normalized entropy, not P(correct); yuki-oshio/mini-jev ≠ r-ms/mini-jev; Jev classifier as autoregressive next-token predictor; ChatJev-style soundness theater; erik-dunteman/ChatJev ≠ dannote/jev ≠ jev-gpt; calibrated decision head × AlphaProof value head; implementation-layer isomorphism, semantic difference; timeout = censoring; do not launder Noul as proof; parallel rank-prediction vs serial selection; independent questions can conflict; zzzzzec/jevsort ≠ keltokhy/jsort; curated open System One ecosystem catalog; rupeshpoojary9/awesome-open-system-one ≠ AnotiaWang/awesome-jev; arXiv paper radar with Jev relevance scoring; ranking ≠ calibration / 0.5 still soft; fail-open failed evals not marked seen **Hourly 2340:** train calibrated ~27M from scratch; typed Q→prob dist / one forward pass / no LLM decode; hyusi2003/MiniSystemOne ≠ Colvin0315/MiniSystemOne; description-only stub / size 5; ESCI hard probe fails four of six; jev_bool ECE 0.242 inversion 0.255; do not re-fold §60 six-gates as new; jobbyjev one-request-per-company from batch-size result; find/design/evaluate TypeSafe Jev decision loops; karanb192/jev-architect ≠ samtay32/jev-system-architect; Jairik/jev-distiller size 1; distill-Jev UI stub / do not distill Jev as teacher of record; post-launch scored use-case map / Jev self-scores then human curation; licensedsaucer9-web/jev-opportunities; Jev-inize a use case into classifier/router; gavinHuang/jevinize → simple-jev not TypeSafe; featherless-ai/simple-jev; compare saved decisions / same label can still change the branch; VihaanAgarwal/jev-diff ≠ Saik0s/diffusiongemma-jev-macos; not tested with a live Jev API key; constrained logprob + temp/Platt ≠ Noul; OpenJevPro pastes openjev-sglang JevBench as own; zhangcy122/OpenJevPro ≠ IamBusy/OpenJev ≠ ekzhang/openjev-sglang; PolyForm Noncommercial; SmolLM-135M / sub-70ms / 0 output tokens; demo P(True) 0.5052 / Choice conf 0.2872 / Score conf 0.0055; README claims MIT / GitHub license null / no LICENSE file; patelvishwa112/jev-system-one-rlcd ≠ arnabgho/rlcd-lite ≠ blackwood-rlcd; source-backed Awesome Jev radar / 306+ commit-pinned; logicrw/awesome-jev-projects ≠ AnotiaWang/awesome-jev ≠ yibie/awesome-jev ≠ cobanov/awesome-jev ≠ rupeshpoojary9/awesome-open-system-one; auto GitHub sync / Issue-only submissions; hashed n-gram encoder / rival-aware attention; olanotolu/jevbetter vs jevlike starter; synthetic hard menus top-1 0.916 vs 0.873 / ECE 0.0182 vs 0.0367 / 40 vs 4608 menus/sec; shuffled-context control 0.335 | `references/applied-mappings.md#7-capability-kernel--human-confirmed-gate` Apply 0042 (`notes.md` §105): structured probability readouts; distribution > argmax; Noul 0.5 midpoint; score is expectation not integer; bare HTTP not SDK; Arohtea/jev-readout; Jev-style Choice/Score/Noul from ordinary models; optional DSH plugin; schema-valid ≠ calibrated; gulagala001/jevify ≠ Mintzs/jevify; Laya RLCD benchmark; 40.3% below constant-answer; open-weight measurement; mourad-ghafiri/laya-rlcd-benchmark ≠ yibie/laya-jev-lab; cheap fail-open semantic edge; second signal not sole; FastLoopError catch; SupremeDreamZ/jev-fastloop ≠ jev-ultrafast; asking more questions in one call; 0.980 at every N; nearly not fully deterministic; TheWebDevel/jev-fanout; Qwen3-VL perception + Jev decisions train RL; 0 model calls at deployment; VLM alone 1.7 vs +Jev 4.4; harneet2512/reflexrl ≠ khordoo/jev-reflex-autonomy-lab; independent Jev API vs Laya; cascade 0.60 matches 78% at 1.8×; noul facts not judgements; yibie/laya-jev-lab ≠ dairui1/jev-lab ≠ BrendanH18/jev-lab; GLiNER vs GLiFormer vs Laya vs Jev; extractors ≠ decision engines; Laya dict-instructions collapse 58.3%; umstek/zero-shot-ie-bench; decisions-per-minute & cost; 204 moves vs 73; throughput not intelligence; angelgalvisc/snake-arena-jev-vs-llms ≠ vtrivedy/jev-plays-games; behavioral contracts; pin expectations eval upgrades; raw 0.94 is not a release; sathariels/jevcheck ≠ dayhaysoos/jevals ≠ SivletLabs/jev-eval; evidence-linked dependency upgrade; Jev never generates filenames; no_direct_evidence ≠ safe to merge; GaneshVG18/upgrade-radar ≠ LYchoon/paper-radar-jev; discography theme/mood/complexity; five atomic questions one call; lirantal/discoprint | | Decide → policy → LLM leftover | Typed decide backends share one Answer schema; policy in code routes auto/review/llm; generator writes leftover text only. Noul 0.5 = cannot-tell, never rounded. Score conf 0.0 = flat, never acted on. Hard flags always review. **Inbox cousin:** mailordinal (typed signals → deterministic priority; no leftover LLM required). **Public decide-backend:** classifier.dev (HTTP classification API; leftover LLM is fallback when Jev is down). **Open NAR cousin:** NandhaKishorM/laya (self-hosted Choice/Score/Noul; Router picks ckpt; policy still in the caller). **Compile leftover:** byenzyme/enzyme (Decisions uses Jev; catalyst generation uses a separate LLM; guidance ≠ hook; hosted bootstrap ≠ silent TypeSafe). **Read-only Gmail trays** (jevmail leftover); **macOS proposed writes** (mailjay). **Hourly 1241:** ORIGIN pause-if-no-Jev; validResponse sums-to-1. **Hourly 1347:** jev-calibration-arena never acts; codaaiteam/jev-skill jevtypesafeai.com ≠ TypeSafe. **Hourly 1441:** ha-conversation-jev Jev→Grok; dsh-jev can only gate. **Hourly 1541:** keep/shadow/hybrid/reject; difficulty + policy thresholds + JSONL trace; jev-codex-pilot model + reasoning depth | `references/applied-mappings.md#8-decide--policy--llm-leftover-cascade` | @@ -135,7 +135,7 @@ classical method you already trust, substitute it, classify the win | Feature engineering / multi-criteria analysis | Nouls + Score distributions as named features, weights in code | `references/mappings.md#1-semantic-judgments--features-and-explicit-utility` | | Selective classification / decision theory | Thresholds from action costs, abstention paths. Train a domain specialist when downstream code **reads the probability**; few-shot hosted API when only **argmax** matters (calibration/VOI, not an accuracy bake-off). **Active-learning triage:** high conf accept / middling expensive teacher / low-or-boundary human; log full distributions. **Do not distill Jev as teacher of record** (~68% ceiling compounds errors; real outcomes stay the targets). **SIGNAL §93:** human labels only; production capture flywheel; score never auto-accepts | `references/mappings.md#2-probabilistic-judgments--cost-sensitive-decisions` | | Decision tables / circuits / state machines | Judgment predicates, code owns transitions. Language primitive: Ruby `chance`/`pick`/`rate` as control flow (hunch; English-as-config; fail polarity per action). feelings `.feels()` default 0.5 is Noul-0.5-never-rounded; exhaustive BAML `match`; **≠** hunch **≠** southpolesteve/probably (language primitive, not a `.feels()` overlay). apa-persona-engine SM then leftover LLM. **Hourly 1241:** s1_ruby collapse late; `undecided?` abstain; tpellet/hunch exit 3. **Hourly 1347:** typed-judge-kit verdict-in-code; judgekit YAML classify/score/route/verify | `references/mappings.md#3-semantic-predicates--decision-circuits` | -| Retrieve + expensive relevance fn | Bounded rerank of a retrieved shortlist. Decision-native RAG: retrieve wide → decide explicitly → evidence set → conflict resolve → reason only over kept evidence (embeddings stay candidate generators; no universal benchmark). Classify-first MCP: same sandwich on agent I/O (path/url/text → judge; main LLM opens survivors). Meaning-search without embeddings (packed parallel relevance; two-stage outline→zoom). **Meaning-grep** AND/OR/NOT over *thresholded* line Nouls; proposition≠embedding; not a gate (jev-semgrep §86). **Evidence-packet explorer:** index-once, BM25 shortlist, Jev ranks, citable source_of_truth (jevex). Measured pointwise rerank vs a generative reranker (one-run; name the no-RAG arm). **RAG rerank harness:** jev-rag-benchmark; “Jev wins” is not an assumption. **Hourly 1241:** jsort scores are relative; Noul not Choice for scale; groundedness-judge-bench native vs schema-guided; implicit_true included in yes. **Hourly 1347:** nanoprune 2.8MB ECE 2.58%; alsoleg89/decide packing VOI. **Hourly 1541:** quarry evidence projection. **Hourly 1740:** jegrep calibrated path+range Nouls; no embeddings/index/daemon; ~$0.01–0.03 typical; agent --json; can1357/jegrep ≠ Bentlybro/jevgrep ≠ uehaj/jev-semgrep. OpenRouter/TypeSafe auto-failover is silent FALLBACK. **Hourly 1843:** variable-N option scoring as the trainable object (dynamic candidate bags not fixed label sets). **Hourly 1943:** retrieve by relevance not resemblance; samdotmak/jev-recall ≠ jev-search ≠ jev-sift ≠ carryforward ≠ chopratejas/invalidate **Hourly 2041:** resume-screening bias audit methodology; Plan/PRD panel → code-owned pass|review|block; cost-aware multi-model routing/escalation; frozen-protocol zero-shot bench; context-window admission control; typed decision control plane; receipt ≠ authorization; live 15-dim typed rubric re-score per pause; adversarial pre-registered Jev eval; confidence does not track ignorance; provider-neutral Elixir/BEAM Noul/Choice/Score SDK **Hourly 2145:** question-linting of Jev questions themselves; nine jaggedness rules, no API key, no labelled data; static lint ≠ measured separation; yodablocks/jevq ≠ tenbin ≠ JevLint ≠ commitjev; open-weights Laya as class exemplar (binding); Nx/Bumblebee runtime; host chooses backend; ChristianAlexander/laya_ex ≠ system_one_sdk ≠ dannote/jev ≠ NandhaKishorM/laya; on-chain/edge Laya deploy; parity_verified stays false; model output never grants Tx; humandebri/IC-Laya ≠ laya_ex; auditable weekend replica; Jev outputs never used for training; unpaired 0.577 vs 0.727; agilabs-ai/jev48 ≠ JevBench ≠ Mapika/decider; adversarial dual-judge / framing attack surface; comparative framing is the usable judgment; prior injection crowds out evidence; copyleftdev/ember ≠ ember.js; Laya specialist fine-tune pipeline; training still GPU-pending; PIXELZX0/XERON ≠ convaiinnovations/laya; Hub Laya replica drop; daliborsb/laya ≠ convaiinnovations/laya ≠ NandhaKishorM/laya; System One student distillation corpus; gold is programmatic; teacher is closed-API clone; MagaBitmex/jev-4b-distill-data ≠ missing student checkpoint; non-LLM VIN System One; planning depth not chat; lewislululu/jevon ≠ douglance/jevon; source-bound evidence checks; local quote mismatch needs no API; exit 0 ≠ claim truth; WaynezProg/jev-kit ≠ jonathanavis96/jev-kit (Airlock) ≠ jev-use ≠ jev-mcp **Hourly 2246:** independent System One evidence catalog; 19 reviewed records; scores not one leaderboard; no external record currently reproduced; TokenTrim no-Jev matched hybrid 62.4%; reachjalil/system-one-bench ≠ mallahyari/system-one-benchmark; 21 tasks · 134 items · 208 questions; scenes from public GitHub contracts, not production logs; SivletLabs/jev-eval ≠ willkelly/jev-evaluation ≠ 4esv/jev-eval ≠ xxkuboxx/jev-eval ≠ onlyoneaman/jev-eval ≠ dayhaysoos/jevals; option isolation (sibling-blind); permutation-equivariant; Hub OWNER not published; nafisazizir/hev ≠ jaredpalmer/kev; frozen local LLM logits, no trained decision head; residual-head 9,222-param decreased 73/96→67/96; confidence = 1−normalized entropy, not P(correct); yuki-oshio/mini-jev ≠ r-ms/mini-jev; Jev classifier as autoregressive next-token predictor; ChatJev-style soundness theater; erik-dunteman/ChatJev ≠ dannote/jev ≠ jev-gpt; calibrated decision head × AlphaProof value head; implementation-layer isomorphism, semantic difference; timeout = censoring; do not launder Noul as proof; parallel rank-prediction vs serial selection; independent questions can conflict; zzzzzec/jevsort ≠ keltokhy/jsort; curated open System One ecosystem catalog; rupeshpoojary9/awesome-open-system-one ≠ AnotiaWang/awesome-jev; arXiv paper radar with Jev relevance scoring; ranking ≠ calibration / 0.5 still soft; fail-open failed evals not marked seen **Hourly 2340:** train calibrated ~27M from scratch; typed Q→prob dist / one forward pass / no LLM decode; hyusi2003/MiniSystemOne ≠ Colvin0315/MiniSystemOne; description-only stub / size 5; ESCI hard probe fails four of six; jev_bool ECE 0.242 inversion 0.255; do not re-fold §60 six-gates as new; jobbyjev one-request-per-company from batch-size result; find/design/evaluate TypeSafe Jev decision loops; karanb192/jev-architect ≠ samtay32/jev-system-architect; Jairik/jev-distiller size 1; distill-Jev UI stub / do not distill Jev as teacher of record; post-launch scored use-case map / Jev self-scores then human curation; licensedsaucer9-web/jev-opportunities; Jev-inize a use case into classifier/router; gavinHuang/jevinize → simple-jev not TypeSafe; featherless-ai/simple-jev; compare saved decisions / same label can still change the branch; VihaanAgarwal/jev-diff ≠ Saik0s/diffusiongemma-jev-macos; not tested with a live Jev API key; constrained logprob + temp/Platt ≠ Noul; OpenJevPro pastes openjev-sglang JevBench as own; zhangcy122/OpenJevPro ≠ IamBusy/OpenJev ≠ ekzhang/openjev-sglang; PolyForm Noncommercial; SmolLM-135M / sub-70ms / 0 output tokens; demo P(True) 0.5052 / Choice conf 0.2872 / Score conf 0.0055; README claims MIT / GitHub license null / no LICENSE file; patelvishwa112/jev-system-one-rlcd ≠ arnabgho/rlcd-lite ≠ blackwood-rlcd; source-backed Awesome Jev radar / 306+ commit-pinned; logicrw/awesome-jev-projects ≠ AnotiaWang/awesome-jev ≠ yibie/awesome-jev ≠ cobanov/awesome-jev ≠ rupeshpoojary9/awesome-open-system-one; auto GitHub sync / Issue-only submissions; hashed n-gram encoder / rival-aware attention; olanotolu/jevbetter vs jevlike starter; synthetic hard menus top-1 0.916 vs 0.873 / ECE 0.0182 vs 0.0367 / 40 vs 4608 menus/sec; shuffled-context control 0.335 | `references/mappings.md#4-retrieval--bounded-semantic-reranking` (independent TREC DL2019 benchmark: Jev zero-shot best MAP 0.4748, nDCG@10 0.683 vs tuned monoBERT 0.718 — competitive, not dominant) Apply 0042 (`notes.md` §105): structured probability readouts; distribution > argmax; Noul 0.5 midpoint; score is expectation not integer; bare HTTP not SDK; Arohtea/jev-readout; Jev-style Choice/Score/Noul from ordinary models; optional DSH plugin; schema-valid ≠ calibrated; gulagala001/jevify ≠ Mintzs/jevify; Laya RLCD benchmark; 40.3% below constant-answer; open-weight measurement; mourad-ghafiri/laya-rlcd-benchmark ≠ yibie/laya-jev-lab; cheap fail-open semantic edge; second signal not sole; FastLoopError catch; SupremeDreamZ/jev-fastloop ≠ jev-ultrafast; asking more questions in one call; 0.980 at every N; nearly not fully deterministic; TheWebDevel/jev-fanout; Qwen3-VL perception + Jev decisions train RL; 0 model calls at deployment; VLM alone 1.7 vs +Jev 4.4; harneet2512/reflexrl ≠ khordoo/jev-reflex-autonomy-lab; independent Jev API vs Laya; cascade 0.60 matches 78% at 1.8×; noul facts not judgements; yibie/laya-jev-lab ≠ dairui1/jev-lab ≠ BrendanH18/jev-lab; GLiNER vs GLiFormer vs Laya vs Jev; extractors ≠ decision engines; Laya dict-instructions collapse 58.3%; umstek/zero-shot-ie-bench; decisions-per-minute & cost; 204 moves vs 73; throughput not intelligence; angelgalvisc/snake-arena-jev-vs-llms ≠ vtrivedy/jev-plays-games; behavioral contracts; pin expectations eval upgrades; raw 0.94 is not a release; sathariels/jevcheck ≠ dayhaysoos/jevals ≠ SivletLabs/jev-eval; evidence-linked dependency upgrade; Jev never generates filenames; no_direct_evidence ≠ safe to merge; GaneshVG18/upgrade-radar ≠ LYchoon/paper-radar-jev; discography theme/mood/complexity; five atomic questions one call; lirantal/discoprint | +| Retrieve + expensive relevance fn | Bounded rerank of a retrieved shortlist. Decision-native RAG: retrieve wide → decide explicitly → evidence set → conflict resolve → reason only over kept evidence (embeddings stay candidate generators; no universal benchmark). **Capability tree:** NiazMorshed2007/jcr applies the same sandwich to documented commands (`notes.md` §116). Classify-first MCP: same sandwich on agent I/O (path/url/text → judge; main LLM opens survivors). Meaning-search without embeddings (packed parallel relevance; two-stage outline→zoom). **Meaning-grep** AND/OR/NOT over *thresholded* line Nouls; proposition≠embedding; not a gate (jev-semgrep §86). **Evidence-packet explorer:** index-once, BM25 shortlist, Jev ranks, citable source_of_truth (jevex). Measured pointwise rerank vs a generative reranker (one-run; name the no-RAG arm). **RAG rerank harness:** jev-rag-benchmark; “Jev wins” is not an assumption. **Hourly 1241:** jsort scores are relative; Noul not Choice for scale; groundedness-judge-bench native vs schema-guided; implicit_true included in yes. **Hourly 1347:** nanoprune 2.8MB ECE 2.58%; alsoleg89/decide packing VOI. **Hourly 1541:** quarry evidence projection. **Hourly 1740:** jegrep calibrated path+range Nouls; no embeddings/index/daemon; ~$0.01–0.03 typical; agent --json; can1357/jegrep ≠ Bentlybro/jevgrep ≠ uehaj/jev-semgrep. OpenRouter/TypeSafe auto-failover is silent FALLBACK. **Hourly 1843:** variable-N option scoring as the trainable object (dynamic candidate bags not fixed label sets). **Hourly 1943:** retrieve by relevance not resemblance; samdotmak/jev-recall ≠ jev-search ≠ jev-sift ≠ carryforward ≠ chopratejas/invalidate **Hourly 2041:** resume-screening bias audit methodology; Plan/PRD panel → code-owned pass|review|block; cost-aware multi-model routing/escalation; frozen-protocol zero-shot bench; context-window admission control; typed decision control plane; receipt ≠ authorization; live 15-dim typed rubric re-score per pause; adversarial pre-registered Jev eval; confidence does not track ignorance; provider-neutral Elixir/BEAM Noul/Choice/Score SDK **Hourly 2145:** question-linting of Jev questions themselves; nine jaggedness rules, no API key, no labelled data; static lint ≠ measured separation; yodablocks/jevq ≠ tenbin ≠ JevLint ≠ commitjev; open-weights Laya as class exemplar (binding); Nx/Bumblebee runtime; host chooses backend; ChristianAlexander/laya_ex ≠ system_one_sdk ≠ dannote/jev ≠ NandhaKishorM/laya; on-chain/edge Laya deploy; parity_verified stays false; model output never grants Tx; humandebri/IC-Laya ≠ laya_ex; auditable weekend replica; Jev outputs never used for training; unpaired 0.577 vs 0.727; agilabs-ai/jev48 ≠ JevBench ≠ Mapika/decider; adversarial dual-judge / framing attack surface; comparative framing is the usable judgment; prior injection crowds out evidence; copyleftdev/ember ≠ ember.js; Laya specialist fine-tune pipeline; training still GPU-pending; PIXELZX0/XERON ≠ convaiinnovations/laya; Hub Laya replica drop; daliborsb/laya ≠ convaiinnovations/laya ≠ NandhaKishorM/laya; System One student distillation corpus; gold is programmatic; teacher is closed-API clone; MagaBitmex/jev-4b-distill-data ≠ missing student checkpoint; non-LLM VIN System One; planning depth not chat; lewislululu/jevon ≠ douglance/jevon; source-bound evidence checks; local quote mismatch needs no API; exit 0 ≠ claim truth; WaynezProg/jev-kit ≠ jonathanavis96/jev-kit (Airlock) ≠ jev-use ≠ jev-mcp **Hourly 2246:** independent System One evidence catalog; 19 reviewed records; scores not one leaderboard; no external record currently reproduced; TokenTrim no-Jev matched hybrid 62.4%; reachjalil/system-one-bench ≠ mallahyari/system-one-benchmark; 21 tasks · 134 items · 208 questions; scenes from public GitHub contracts, not production logs; SivletLabs/jev-eval ≠ willkelly/jev-evaluation ≠ 4esv/jev-eval ≠ xxkuboxx/jev-eval ≠ onlyoneaman/jev-eval ≠ dayhaysoos/jevals; option isolation (sibling-blind); permutation-equivariant; Hub OWNER not published; nafisazizir/hev ≠ jaredpalmer/kev; frozen local LLM logits, no trained decision head; residual-head 9,222-param decreased 73/96→67/96; confidence = 1−normalized entropy, not P(correct); yuki-oshio/mini-jev ≠ r-ms/mini-jev; Jev classifier as autoregressive next-token predictor; ChatJev-style soundness theater; erik-dunteman/ChatJev ≠ dannote/jev ≠ jev-gpt; calibrated decision head × AlphaProof value head; implementation-layer isomorphism, semantic difference; timeout = censoring; do not launder Noul as proof; parallel rank-prediction vs serial selection; independent questions can conflict; zzzzzec/jevsort ≠ keltokhy/jsort; curated open System One ecosystem catalog; rupeshpoojary9/awesome-open-system-one ≠ AnotiaWang/awesome-jev; arXiv paper radar with Jev relevance scoring; ranking ≠ calibration / 0.5 still soft; fail-open failed evals not marked seen **Hourly 2340:** train calibrated ~27M from scratch; typed Q→prob dist / one forward pass / no LLM decode; hyusi2003/MiniSystemOne ≠ Colvin0315/MiniSystemOne; description-only stub / size 5; ESCI hard probe fails four of six; jev_bool ECE 0.242 inversion 0.255; do not re-fold §60 six-gates as new; jobbyjev one-request-per-company from batch-size result; find/design/evaluate TypeSafe Jev decision loops; karanb192/jev-architect ≠ samtay32/jev-system-architect; Jairik/jev-distiller size 1; distill-Jev UI stub / do not distill Jev as teacher of record; post-launch scored use-case map / Jev self-scores then human curation; licensedsaucer9-web/jev-opportunities; Jev-inize a use case into classifier/router; gavinHuang/jevinize → simple-jev not TypeSafe; featherless-ai/simple-jev; compare saved decisions / same label can still change the branch; VihaanAgarwal/jev-diff ≠ Saik0s/diffusiongemma-jev-macos; not tested with a live Jev API key; constrained logprob + temp/Platt ≠ Noul; OpenJevPro pastes openjev-sglang JevBench as own; zhangcy122/OpenJevPro ≠ IamBusy/OpenJev ≠ ekzhang/openjev-sglang; PolyForm Noncommercial; SmolLM-135M / sub-70ms / 0 output tokens; demo P(True) 0.5052 / Choice conf 0.2872 / Score conf 0.0055; README claims MIT / GitHub license null / no LICENSE file; patelvishwa112/jev-system-one-rlcd ≠ arnabgho/rlcd-lite ≠ blackwood-rlcd; source-backed Awesome Jev radar / 306+ commit-pinned; logicrw/awesome-jev-projects ≠ AnotiaWang/awesome-jev ≠ yibie/awesome-jev ≠ cobanov/awesome-jev ≠ rupeshpoojary9/awesome-open-system-one; auto GitHub sync / Issue-only submissions; hashed n-gram encoder / rival-aware attention; olanotolu/jevbetter vs jevlike starter; synthetic hard menus top-1 0.916 vs 0.873 / ECE 0.0182 vs 0.0367 / 40 vs 4608 menus/sec; shuffled-context control 0.335 | `references/mappings.md#4-retrieval--bounded-semantic-reranking` (independent TREC DL2019 benchmark: Jev zero-shot best MAP 0.4748, nDCG@10 0.683 vs tuned monoBERT 0.718 — competitive, not dominant) Apply 0042 (`notes.md` §105): structured probability readouts; distribution > argmax; Noul 0.5 midpoint; score is expectation not integer; bare HTTP not SDK; Arohtea/jev-readout; Jev-style Choice/Score/Noul from ordinary models; optional DSH plugin; schema-valid ≠ calibrated; gulagala001/jevify ≠ Mintzs/jevify; Laya RLCD benchmark; 40.3% below constant-answer; open-weight measurement; mourad-ghafiri/laya-rlcd-benchmark ≠ yibie/laya-jev-lab; cheap fail-open semantic edge; second signal not sole; FastLoopError catch; SupremeDreamZ/jev-fastloop ≠ jev-ultrafast; asking more questions in one call; 0.980 at every N; nearly not fully deterministic; TheWebDevel/jev-fanout; Qwen3-VL perception + Jev decisions train RL; 0 model calls at deployment; VLM alone 1.7 vs +Jev 4.4; harneet2512/reflexrl ≠ khordoo/jev-reflex-autonomy-lab; independent Jev API vs Laya; cascade 0.60 matches 78% at 1.8×; noul facts not judgements; yibie/laya-jev-lab ≠ dairui1/jev-lab ≠ BrendanH18/jev-lab; GLiNER vs GLiFormer vs Laya vs Jev; extractors ≠ decision engines; Laya dict-instructions collapse 58.3%; umstek/zero-shot-ie-bench; decisions-per-minute & cost; 204 moves vs 73; throughput not intelligence; angelgalvisc/snake-arena-jev-vs-llms ≠ vtrivedy/jev-plays-games; behavioral contracts; pin expectations eval upgrades; raw 0.94 is not a release; sathariels/jevcheck ≠ dayhaysoos/jevals ≠ SivletLabs/jev-eval; evidence-linked dependency upgrade; Jev never generates filenames; no_direct_evidence ≠ safe to merge; GaneshVG18/upgrade-radar ≠ LYchoon/paper-radar-jev; discography theme/mood/complexity; five atomic questions one call; lirantal/discoprint | | Store as semantic index (SQL / SQLite / zoxide / dataframe) | Cheap exact predicates first; typed questions on the remainder. In-engine extension (sqlite-jev / pg-jev) vs **judgment outside the store** (jevql CLI; vanilla Postgres never sees `jev()`) vs path index (joxide) vs dataframe columns (jevpandas / jevframe). **ORDER BY over probs is a ranking job:** calibration ≠ sortable; measure pairwise inversion / Score ordinality / two-decimal ties (jev-orderby-bench). **2340 upgrade:** ESCI hard probe fails four of six; jev_bool ECE 0.242 inversion 0.255; do not re-fold §60 six-gates as new (`notes.md` §104) | `references/mappings.md#4-retrieval--bounded-semantic-reranking` | | Soft judgment inside a hard envelope | Model may only match the deterministic policy or be more conservative (bitrate ABR; query-planner override-when-confident; compaction mutations/shell operators → keep_full; stdout prune: ≤10k/JSON-diff-whole-doc untouched, then Noul; Cua-S1: plan≠execute, dry-run, fail-closed checkbox/fill; Stagehand extract: schema/completion-gate/screenshot-always-LLM then pick, else LLM; skills→oxlint: AST/precheck prove, guidance whole-file in state, remainder Noul — not a hard gate; pre-exec toolgate: allow/block/review — Jev is not authorization; guard error/timeout stops; distinct from capability kernel interlock — secrets never in agent, closed action space. OMP prompt suppression: operator owns thresholds; plugin never self-tunes; host `bash.patterns: deny` stays the floor (omp-greenlight — not a sandbox). Human-confirmed kill: Jev recommends, human is the only trigger, identity re-check, shields override, mapped explanations (port-cleanup). Effect-based shell: fast-allow/deny prove, then Jev remainder; privilege ≠ verdict (construct-auto-classifier). Wrap-as-execution: rules prove allow/deny, Jev remainder, ASK throws, fail-closed (AgentGhost). SLO router: exactness raises quality floor, never overrides capability; Jev features fail-open to local (slo-router). **Hourly 1740 DGP:** JSON Schema proves record structure; app owns freshness/authorization/idempotency; assessment is the remainder sensor; commit fail-closed (numerous-com/dgp; hard-gating DGP as safety theater). **Hourly 1843:** thresholds derived from costs not hard-coded; the cost table is policy; the Noul is a SENSOR; hard-gating 0.038 is theater). **Hourly 1943:** noul_threshold 0.5 decoder not a proof; softmax over allowed tokens ≠ Noul; no shipped rule has severity error **Hourly 2041:** resume-screening bias audit methodology; Plan/PRD panel → code-owned pass|review|block; cost-aware multi-model routing/escalation; frozen-protocol zero-shot bench; context-window admission control; typed decision control plane; receipt ≠ authorization; live 15-dim typed rubric re-score per pause; adversarial pre-registered Jev eval; confidence does not track ignorance; provider-neutral Elixir/BEAM Noul/Choice/Score SDK **Hourly 2145:** question-linting of Jev questions themselves; nine jaggedness rules, no API key, no labelled data; static lint ≠ measured separation; yodablocks/jevq ≠ tenbin ≠ JevLint ≠ commitjev; open-weights Laya as class exemplar (binding); Nx/Bumblebee runtime; host chooses backend; ChristianAlexander/laya_ex ≠ system_one_sdk ≠ dannote/jev ≠ NandhaKishorM/laya; on-chain/edge Laya deploy; parity_verified stays false; model output never grants Tx; humandebri/IC-Laya ≠ laya_ex; auditable weekend replica; Jev outputs never used for training; unpaired 0.577 vs 0.727; agilabs-ai/jev48 ≠ JevBench ≠ Mapika/decider; adversarial dual-judge / framing attack surface; comparative framing is the usable judgment; prior injection crowds out evidence; copyleftdev/ember ≠ ember.js; Laya specialist fine-tune pipeline; training still GPU-pending; PIXELZX0/XERON ≠ convaiinnovations/laya; Hub Laya replica drop; daliborsb/laya ≠ convaiinnovations/laya ≠ NandhaKishorM/laya; System One student distillation corpus; gold is programmatic; teacher is closed-API clone; MagaBitmex/jev-4b-distill-data ≠ missing student checkpoint; non-LLM VIN System One; planning depth not chat; lewislululu/jevon ≠ douglance/jevon; source-bound evidence checks; local quote mismatch needs no API; exit 0 ≠ claim truth; WaynezProg/jev-kit ≠ jonathanavis96/jev-kit (Airlock) ≠ jev-use ≠ jev-mcp **Hourly 2246:** independent System One evidence catalog; 19 reviewed records; scores not one leaderboard; no external record currently reproduced; TokenTrim no-Jev matched hybrid 62.4%; reachjalil/system-one-bench ≠ mallahyari/system-one-benchmark; 21 tasks · 134 items · 208 questions; scenes from public GitHub contracts, not production logs; SivletLabs/jev-eval ≠ willkelly/jev-evaluation ≠ 4esv/jev-eval ≠ xxkuboxx/jev-eval ≠ onlyoneaman/jev-eval ≠ dayhaysoos/jevals; option isolation (sibling-blind); permutation-equivariant; Hub OWNER not published; nafisazizir/hev ≠ jaredpalmer/kev; frozen local LLM logits, no trained decision head; residual-head 9,222-param decreased 73/96→67/96; confidence = 1−normalized entropy, not P(correct); yuki-oshio/mini-jev ≠ r-ms/mini-jev; Jev classifier as autoregressive next-token predictor; ChatJev-style soundness theater; erik-dunteman/ChatJev ≠ dannote/jev ≠ jev-gpt; calibrated decision head × AlphaProof value head; implementation-layer isomorphism, semantic difference; timeout = censoring; do not launder Noul as proof; parallel rank-prediction vs serial selection; independent questions can conflict; zzzzzec/jevsort ≠ keltokhy/jsort; curated open System One ecosystem catalog; rupeshpoojary9/awesome-open-system-one ≠ AnotiaWang/awesome-jev; arXiv paper radar with Jev relevance scoring; ranking ≠ calibration / 0.5 still soft; fail-open failed evals not marked seen **Hourly 2340:** train calibrated ~27M from scratch; typed Q→prob dist / one forward pass / no LLM decode; hyusi2003/MiniSystemOne ≠ Colvin0315/MiniSystemOne; description-only stub / size 5; ESCI hard probe fails four of six; jev_bool ECE 0.242 inversion 0.255; do not re-fold §60 six-gates as new; jobbyjev one-request-per-company from batch-size result; find/design/evaluate TypeSafe Jev decision loops; karanb192/jev-architect ≠ samtay32/jev-system-architect; Jairik/jev-distiller size 1; distill-Jev UI stub / do not distill Jev as teacher of record; post-launch scored use-case map / Jev self-scores then human curation; licensedsaucer9-web/jev-opportunities; Jev-inize a use case into classifier/router; gavinHuang/jevinize → simple-jev not TypeSafe; featherless-ai/simple-jev; compare saved decisions / same label can still change the branch; VihaanAgarwal/jev-diff ≠ Saik0s/diffusiongemma-jev-macos; not tested with a live Jev API key; constrained logprob + temp/Platt ≠ Noul; OpenJevPro pastes openjev-sglang JevBench as own; zhangcy122/OpenJevPro ≠ IamBusy/OpenJev ≠ ekzhang/openjev-sglang; PolyForm Noncommercial; SmolLM-135M / sub-70ms / 0 output tokens; demo P(True) 0.5052 / Choice conf 0.2872 / Score conf 0.0055; README claims MIT / GitHub license null / no LICENSE file; patelvishwa112/jev-system-one-rlcd ≠ arnabgho/rlcd-lite ≠ blackwood-rlcd; source-backed Awesome Jev radar / 306+ commit-pinned; logicrw/awesome-jev-projects ≠ AnotiaWang/awesome-jev ≠ yibie/awesome-jev ≠ cobanov/awesome-jev ≠ rupeshpoojary9/awesome-open-system-one; auto GitHub sync / Issue-only submissions; hashed n-gram encoder / rival-aware attention; olanotolu/jevbetter vs jevlike starter; synthetic hard menus top-1 0.916 vs 0.873 / ECE 0.0182 vs 0.0367 / 40 vs 4608 menus/sec; shuffled-context control 0.335 | `references/mappings.md#12-runtime-assurance-sandwich-hypothesis`; `references/mappings.md#18-structural-prove--soft-remainder-hypothesis-as-domain-general-empirical-as-named-shapes` Apply 0042 (`notes.md` §105): structured probability readouts; distribution > argmax; Noul 0.5 midpoint; score is expectation not integer; bare HTTP not SDK; Arohtea/jev-readout; Jev-style Choice/Score/Noul from ordinary models; optional DSH plugin; schema-valid ≠ calibrated; gulagala001/jevify ≠ Mintzs/jevify; Laya RLCD benchmark; 40.3% below constant-answer; open-weight measurement; mourad-ghafiri/laya-rlcd-benchmark ≠ yibie/laya-jev-lab; cheap fail-open semantic edge; second signal not sole; FastLoopError catch; SupremeDreamZ/jev-fastloop ≠ jev-ultrafast; asking more questions in one call; 0.980 at every N; nearly not fully deterministic; TheWebDevel/jev-fanout; Qwen3-VL perception + Jev decisions train RL; 0 model calls at deployment; VLM alone 1.7 vs +Jev 4.4; harneet2512/reflexrl ≠ khordoo/jev-reflex-autonomy-lab; independent Jev API vs Laya; cascade 0.60 matches 78% at 1.8×; noul facts not judgements; yibie/laya-jev-lab ≠ dairui1/jev-lab ≠ BrendanH18/jev-lab; GLiNER vs GLiFormer vs Laya vs Jev; extractors ≠ decision engines; Laya dict-instructions collapse 58.3%; umstek/zero-shot-ie-bench; decisions-per-minute & cost; 204 moves vs 73; throughput not intelligence; angelgalvisc/snake-arena-jev-vs-llms ≠ vtrivedy/jev-plays-games; behavioral contracts; pin expectations eval upgrades; raw 0.94 is not a release; sathariels/jevcheck ≠ dayhaysoos/jevals ≠ SivletLabs/jev-eval; evidence-linked dependency upgrade; Jev never generates filenames; no_direct_evidence ≠ safe to merge; GaneshVG18/upgrade-radar ≠ LYchoon/paper-radar-jev; discography theme/mood/complexity; five atomic questions one call; lirantal/discoprint | | Value of information / gather as an act | Pay for another observation only if EV(decision) improves more than cost; abstain from calling *any* model when a regex already answers (meta-VOI). Fail-open wake/resume: skip the LLM turn only if the judge answers and p(wake) is low (Horvitz). Selective memory: verbatim ledger + scored recall (carryforward; rules never judged; 9×3 is a hint). Classify-first read: pay for a full agent open iff relevance (or typed question) says it might change the act (jev-sift; errors/truncation ≠ irrelevant). **Training-data VOI:** spend expensive teacher/human labels only where confidence says they change the outcome (jev-triage); do not distill Jev as teacher of record. **Decision-model latency cost:** measure p95 of sync Jev on the hot path before claiming routing (slo-router 77.93→490.38 same routes). **Human-review VOI:** attention filter that never blocks the agent and never says green unless sure (jev-lens). **Skill-library VOI:** load a skill iff it changes the next step; abstention first-class (skillranker). **tools≠use:** an MCP memory tool the agent never calls is not VOI (carryforward 0/4; SessionStart hook). **Pre-send token-econ:** compress before first send (dizk/jev-lens; post-send prune cost 17% more *theirs*). **Same-intent cache admit:** skip the LLM iff Jev says same intent (jevcache 0 FP/100 *theirs*; fail-open). **Human-feed VOI:** worth-your-attention before click (ThinkyMiner/Winnow). **Second-call VOI:** skip the narrating main-model turn (hermes-jev-router WHETHER/HOW/WHAT). **Hunk-review VOI:** score PR hunks before the generative reviewer (prune-review). **Intent-search VOI:** pay to judge places the diff did not touch (jev-intent-review). **Empty compact-proxy skip:** IPECTER context-pruner **and** jev-runway slogan only. **Batch ranking VOI:** one request per ten (jevfeed). **No-text-step VOI:** hand the judgment, keep writing on the LLM (jev-use 186 vs 2,672 ms *theirs*). **Files-to-read VOI:** jevex n=16 160s→69s *theirs*. **Commit-attention VOI:** commitjev middle band never rounded. **Pi compact VOI:** pi-jev-compact keep/drop vs LLM summary. **Escalate-under-threshold VOI:** classifier.dev smart re-asks single-label <0.7; multi-label ignores (re-judge worse). **Escalate-without-stall VOI:** khordoo S2 is async one-use; S1 never waits; log consumed not arrived. **Split-question CU VOI:** kind/item/site/offscreen in one request; exclusive actions; perception rebuild is the expensive gather (typesafe-computer-use; 155× *theirs* one screenshot). **Partial-speech VOI:** complete Noul; closed-set may fire; free-text waits (jev-voice-browser). **Evidence-synthesis two-pass VOI:** choxos/jev-reviewer fan-out then absolute Noul; *Not found* cheaper than paraphrase (≠ egma-ai). **Script-before-p VOI:** NandhaKishorM/laya Router routes before the forward pass because Khmer 0.000@0.952 gating cannot catch. **Path-then-window VOI:** JevFind scores paths first (0.25) then windows (0.55); `--file-threshold 0` when recall matters. **Frontier cascade VOI:** jev-frontier-bench pay Fable only when Jev top p < 0.9 — cut is **in-sample**. **Compaction-pi VOI:** fail-open to built-in LLM summary. **Beam-search FS VOI** (findme). **Cache-vs-worker VOI** (jevsubrouter). **Cheap decision-layer VOI** (grok-bot-jev). **Question-preflight VOI** (jev-reliability). **Hourly 1241:** jev-file-search scores not calibrated accuracy; jev-linkmap Jev never sees S2 prose. **Hourly 1347:** alsoleg89/decide packing VOI; 0.8 ≠ 80% accuracy; Jev-Calibration Platt ECE 0.117→0.052. **SIGNAL §93:** fingerprint after redact; recall vs decide; publish fingerprints+answers; VOI of cache hit; hyperspaceai/jevcache ≠ kushals256/jevcache; human labels only (uncertainty acquisition). **Compile-time index VOI:** pay for Jev at compile to shape catalysts, then retrieve through questions (byenzyme/enzyme; catalysts ≠ summaries; guidance ≠ hook; hosted bootstrap ≠ silent TypeSafe). **Hourly 1541:** quarry evidence projection. **Hourly 1740:** assessment batching; ~$0.01–0.03 typical; no embeddings/index/daemon; agent --json . **Hourly 1843:** cost_human is a gather act; auto-batching same-object questions is measurement economics; YES / NO / UNSURE from cost_false_yes / cost_false_no / cost_human. **Hourly 1943:** 133★ / forks 10 live; build calibrated classifiers from human feedback; one calibrated yes/no per memory in one request **Hourly 2041:** resume-screening bias audit methodology; Plan/PRD panel → code-owned pass|review|block; cost-aware multi-model routing/escalation; frozen-protocol zero-shot bench; context-window admission control; typed decision control plane; receipt ≠ authorization; live 15-dim typed rubric re-score per pause; adversarial pre-registered Jev eval; confidence does not track ignorance; provider-neutral Elixir/BEAM Noul/Choice/Score SDK **Hourly 2145:** question-linting of Jev questions themselves; nine jaggedness rules, no API key, no labelled data; static lint ≠ measured separation; yodablocks/jevq ≠ tenbin ≠ JevLint ≠ commitjev; open-weights Laya as class exemplar (binding); Nx/Bumblebee runtime; host chooses backend; ChristianAlexander/laya_ex ≠ system_one_sdk ≠ dannote/jev ≠ NandhaKishorM/laya; on-chain/edge Laya deploy; parity_verified stays false; model output never grants Tx; humandebri/IC-Laya ≠ laya_ex; auditable weekend replica; Jev outputs never used for training; unpaired 0.577 vs 0.727; agilabs-ai/jev48 ≠ JevBench ≠ Mapika/decider; adversarial dual-judge / framing attack surface; comparative framing is the usable judgment; prior injection crowds out evidence; copyleftdev/ember ≠ ember.js; Laya specialist fine-tune pipeline; training still GPU-pending; PIXELZX0/XERON ≠ convaiinnovations/laya; Hub Laya replica drop; daliborsb/laya ≠ convaiinnovations/laya ≠ NandhaKishorM/laya; System One student distillation corpus; gold is programmatic; teacher is closed-API clone; MagaBitmex/jev-4b-distill-data ≠ missing student checkpoint; non-LLM VIN System One; planning depth not chat; lewislululu/jevon ≠ douglance/jevon; source-bound evidence checks; local quote mismatch needs no API; exit 0 ≠ claim truth; WaynezProg/jev-kit ≠ jonathanavis96/jev-kit (Airlock) ≠ jev-use ≠ jev-mcp **Hourly 2246:** independent System One evidence catalog; 19 reviewed records; scores not one leaderboard; no external record currently reproduced; TokenTrim no-Jev matched hybrid 62.4%; reachjalil/system-one-bench ≠ mallahyari/system-one-benchmark; 21 tasks · 134 items · 208 questions; scenes from public GitHub contracts, not production logs; SivletLabs/jev-eval ≠ willkelly/jev-evaluation ≠ 4esv/jev-eval ≠ xxkuboxx/jev-eval ≠ onlyoneaman/jev-eval ≠ dayhaysoos/jevals; option isolation (sibling-blind); permutation-equivariant; Hub OWNER not published; nafisazizir/hev ≠ jaredpalmer/kev; frozen local LLM logits, no trained decision head; residual-head 9,222-param decreased 73/96→67/96; confidence = 1−normalized entropy, not P(correct); yuki-oshio/mini-jev ≠ r-ms/mini-jev; Jev classifier as autoregressive next-token predictor; ChatJev-style soundness theater; erik-dunteman/ChatJev ≠ dannote/jev ≠ jev-gpt; calibrated decision head × AlphaProof value head; implementation-layer isomorphism, semantic difference; timeout = censoring; do not launder Noul as proof; parallel rank-prediction vs serial selection; independent questions can conflict; zzzzzec/jevsort ≠ keltokhy/jsort; curated open System One ecosystem catalog; rupeshpoojary9/awesome-open-system-one ≠ AnotiaWang/awesome-jev; arXiv paper radar with Jev relevance scoring; ranking ≠ calibration / 0.5 still soft; fail-open failed evals not marked seen **Hourly 2340:** train calibrated ~27M from scratch; typed Q→prob dist / one forward pass / no LLM decode; hyusi2003/MiniSystemOne ≠ Colvin0315/MiniSystemOne; description-only stub / size 5; ESCI hard probe fails four of six; jev_bool ECE 0.242 inversion 0.255; do not re-fold §60 six-gates as new; jobbyjev one-request-per-company from batch-size result; find/design/evaluate TypeSafe Jev decision loops; karanb192/jev-architect ≠ samtay32/jev-system-architect; Jairik/jev-distiller size 1; distill-Jev UI stub / do not distill Jev as teacher of record; post-launch scored use-case map / Jev self-scores then human curation; licensedsaucer9-web/jev-opportunities; Jev-inize a use case into classifier/router; gavinHuang/jevinize → simple-jev not TypeSafe; featherless-ai/simple-jev; compare saved decisions / same label can still change the branch; VihaanAgarwal/jev-diff ≠ Saik0s/diffusiongemma-jev-macos; not tested with a live Jev API key; constrained logprob + temp/Platt ≠ Noul; OpenJevPro pastes openjev-sglang JevBench as own; zhangcy122/OpenJevPro ≠ IamBusy/OpenJev ≠ ekzhang/openjev-sglang; PolyForm Noncommercial; SmolLM-135M / sub-70ms / 0 output tokens; demo P(True) 0.5052 / Choice conf 0.2872 / Score conf 0.0055; README claims MIT / GitHub license null / no LICENSE file; patelvishwa112/jev-system-one-rlcd ≠ arnabgho/rlcd-lite ≠ blackwood-rlcd; source-backed Awesome Jev radar / 306+ commit-pinned; logicrw/awesome-jev-projects ≠ AnotiaWang/awesome-jev ≠ yibie/awesome-jev ≠ cobanov/awesome-jev ≠ rupeshpoojary9/awesome-open-system-one; auto GitHub sync / Issue-only submissions; hashed n-gram encoder / rival-aware attention; olanotolu/jevbetter vs jevlike starter; synthetic hard menus top-1 0.916 vs 0.873 / ECE 0.0182 vs 0.0367 / 40 vs 4608 menus/sec; shuffled-context control 0.335 | `references/mappings.md#6-value-of-information--gather-as-an-enumerated-act` (**Hypothesis** until a labeled act/outcome log; 149-row receipt is Empirical as a shape; wakegate 21/21 is smoke; jev-triage is Empirical as README architecture; slo-router live analysis is Empirical as a *negative* on sync Jev; jev-lens is Empirical as README architecture) Apply 0042 (`notes.md` §105): structured probability readouts; distribution > argmax; Noul 0.5 midpoint; score is expectation not integer; bare HTTP not SDK; Arohtea/jev-readout; Jev-style Choice/Score/Noul from ordinary models; optional DSH plugin; schema-valid ≠ calibrated; gulagala001/jevify ≠ Mintzs/jevify; Laya RLCD benchmark; 40.3% below constant-answer; open-weight measurement; mourad-ghafiri/laya-rlcd-benchmark ≠ yibie/laya-jev-lab; cheap fail-open semantic edge; second signal not sole; FastLoopError catch; SupremeDreamZ/jev-fastloop ≠ jev-ultrafast; asking more questions in one call; 0.980 at every N; nearly not fully deterministic; TheWebDevel/jev-fanout; Qwen3-VL perception + Jev decisions train RL; 0 model calls at deployment; VLM alone 1.7 vs +Jev 4.4; harneet2512/reflexrl ≠ khordoo/jev-reflex-autonomy-lab; independent Jev API vs Laya; cascade 0.60 matches 78% at 1.8×; noul facts not judgements; yibie/laya-jev-lab ≠ dairui1/jev-lab ≠ BrendanH18/jev-lab; GLiNER vs GLiFormer vs Laya vs Jev; extractors ≠ decision engines; Laya dict-instructions collapse 58.3%; umstek/zero-shot-ie-bench; decisions-per-minute & cost; 204 moves vs 73; throughput not intelligence; angelgalvisc/snake-arena-jev-vs-llms ≠ vtrivedy/jev-plays-games; behavioral contracts; pin expectations eval upgrades; raw 0.94 is not a release; sathariels/jevcheck ≠ dayhaysoos/jevals ≠ SivletLabs/jev-eval; evidence-linked dependency upgrade; Jev never generates filenames; no_direct_evidence ≠ safe to merge; GaneshVG18/upgrade-radar ≠ LYchoon/paper-radar-jev; discography theme/mood/complexity; five atomic questions one call; lirantal/discoprint | @@ -260,3 +260,6 @@ Protocol: quote *theirs*; mixed-architecture fail polarity per act; judgment-class replica honesty. User-provided 0915 uniqueness lock: TianyuCodings/NanoJev unified-games-v1 densify; A 0.6B parallel decision model: states and questions in, complete probability distributions out. Zero output-token decoding.; One model, four games; ViZDoom Basic 128/128 vs Jev 56/128; Predict Position 27/128 vs Jev 11/128; Maze 225 attempts vs Jev 2738; Snake 30 food / 256 steps; held-out Maze 4/10 Snake 8/8 Basic 128/128 Predict 27/128; Untuned Qwen3-0.6B baseline; 18,760 questions per variant; 16,333 ViZDoom; 896 Predict Position expert episodes; hard_lr1e5; mix weights 1/3, 1/3, 1/6, 1/6; Hub C-Tianyu/NanoJev revision unified-games-v1 likes 58; dataset C-Tianyu/NanoJev-Data likes 5; HEAD 618cea6d906d54e128360786d12f703fff2b1245; 1289★ / 158 forks / size 64035; README SHA 4190093c64ee75b26e9726daa3b00cbcf6d3157a; MIT; caijinchun/nanojev-arena ≠ liao96312/jev-arena-nanojev ≠ zwliJay/jev-forge ≠ NanoJev; not TypeSafe Jev; open replica / specialist gameplay S1; soft scores ≠ hard gates; Game success ≠ calibrated Noul; local type boolean ≠ TypeSafe noul; invented_signal false; do not reopen or amend PR #31/#32/#33/#35; notes.md §115 + + +User-provided 0920 jcr uniqueness lock: NiazMorshed2007/jcr MIT; site https://jcr.niazmorshed.dev; topics ai-agents,jev,mcp; **4★**; HEAD `138b3832`; README SHA `2a49dbc1`; LICENSE SHA `46231303`; size **14850**; Jev Capability Resolver; one tool to find documented deterministic commands in a nested capability tree; returns context; **does not execute**; skills = workflow+judgment; capabilities = individual operations; format independent of Jev; proposed open standard exploration; classify (Jev) → optional OpenAI decompose compound → beam search geometric mean of routing probs; keep up to 3 paths ≥60% of best (JCR_BAND_RATIO 0.6); ambiguity / no-match / depth-limit explicit; soft scores ≠ hard gates; 0.6 band is application policy; routing ≠ permission; docs ≠ authority to run; sol-vs-opus5-20 *theirs*: 20 scenarios × 4 variants = 80 runs; lookup+explain only, no execution; Claude Opus 5: agent input 108,585→15,819 (−85%), cost $0.3700→$0.1222 (−67%), wall 105.5s→77.7s; Codex GPT-5.6-Sol: 61,952→47,669 (−23%), $0.1377→$0.1151 (−16%), wall 25.3s→62.4s (Sol slower with JCR in 19/20); One Sol outlier 372.6s / 193 Jev calls; n=1 per cell; Not Harbor task-execution; Claude/Codex harnesses; compare mode; 50 scenarios bundled; 11 groups, 960 nodes, 11,360 items; 16 routing rounds per step; NiazMorshed2007/jcr ≠ skill-broker ≠ skillranker ≠ jev-sift ≠ jev-lens ≠ jevusher ≠ jev_select_capability; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33/#34; notes.md §116 diff --git a/.agents/skills/augustus/references/agent-self-assessment.md b/.agents/skills/augustus/references/agent-self-assessment.md index e3eab6ee..87ce5c54 100644 --- a/.agents/skills/augustus/references/agent-self-assessment.md +++ b/.agents/skills/augustus/references/agent-self-assessment.md @@ -931,3 +931,8 @@ User-provided 0806 uniqueness lock: people who compare Jev against GPT-5.6 has n Hourly 0843 uniqueness lock: A hunch is a probability with a policy attached; { enter: 0.8, exit: 0.6 } is hysteresis; replay a policy change without inference; Decision models are providers, not the product; huncho ≠ Kungie/gut ≠ carldaws/hunch ≠ tpellet/hunch; pretrained Qwen2.5 base ECE 0.030 (0.5B) / 0.040 (7B); instruct 0.302 / 0.269; 70.9% → 70.0% mean conf 74.1% → 96.7%; temperature scaling still matches it in-distribution; No Jev API was called; Qwen2.5 ≠ Archer; Qwen/Qwen3.8-27B ≠ Archer; 学習済みモデル v0.1 は準備中です; bool AUROC 0.523; 先頭だと0件、末尾だと250件; 温度を渡さない場合、確率は較正されていません; このリポジトリには Jev を呼ぶコードが存在しません; g0runmezadam/what-is-jev IS tunahansahin897/what-is-jev (same GitHub id 1378007307); 947 repos scored; A 273 · B 302 · C 372; LLM rubric ≠ benches; Data as of 2026-09-20; HEAD 895b9498; README SHA 3ae98c56; 13 focused checks and one mutually exclusive outcome; Probabilities are advisory, not calibrated guarantees; omni-/ask-jev ≠ pedroknigge/mcp_jev; pd.cut bins by equal width while jeval bins by quantile; ECE 0.113 and ECE 0.076; jeval drift is not implemented yet; rlaope/jeval ≠ dayhaysoos/jevals; calibration does not compose; ECE has exactly zero statistical power to detect the failure mode that kills trajectories; 25–60× headline withdrawn; P(all-correct): 0.0071 vs 0.0001; TCE / AMS; Qwen 3.8 sparring ≠ Archer; Deferred Crispification; light_cutoff_applied_to_combination 0; recorded run, kinematic animation; BANKING77 Accuracy BERT-Base 93.02 Jev 79.90; Analyse jev calibration (NLL, ECE) backlog; BERT figures are published supervised references, not zero-shot; 档位措辞效应 分数极差中位 0.50、最大 1.32; 修好后对照组是 0.01; 不是 benchmark; 概率没做 calibration; ~1,430 API calls, about $0.15; xiaohuaxi/jev-study ≠ wjdjdakf17/jev-study ≠ baekenough/jev-study; AND: product (independence assumed and recorded in the trace); chat model's stated confidence is not calibrated; circuit-vl-4b ≠ Archer; Bring your own API key; vamsikrishna2421/jev-usecases ≠ whyashthakker/awesome-jev-use-cases; catalog ≠ endorsement; SemIf 2237★ (+30 vs §111 2207); jevlike 1049★ (+6 vs 1043); TypeLLM/TypeLLM 16★; AnotiaWang 98★; yibie/awesome-jev 520★ (+14 vs 506); Laya likes 861 (was 822); tracker likes 66 (+2 vs 64) lastModified UNCHANGED; Blackwood likes 2 gated manual; Archer still promised_not_landed; Hub archerhume/4rcherhume HTTP 401; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33; notes.md §114 User-provided 0915 uniqueness lock: TianyuCodings/NanoJev unified-games-v1 densify; A 0.6B parallel decision model: states and questions in, complete probability distributions out. Zero output-token decoding.; One model, four games; ViZDoom Basic 128/128 vs Jev 56/128; Predict Position 27/128 vs Jev 11/128; Maze 225 attempts vs Jev 2738; Snake 30 food / 256 steps; held-out Maze 4/10 Snake 8/8 Basic 128/128 Predict 27/128; Untuned Qwen3-0.6B baseline; 18,760 questions per variant; 16,333 ViZDoom; 896 Predict Position expert episodes; hard_lr1e5; mix weights 1/3, 1/3, 1/6, 1/6; Hub C-Tianyu/NanoJev revision unified-games-v1 likes 58; dataset C-Tianyu/NanoJev-Data likes 5; HEAD 618cea6d906d54e128360786d12f703fff2b1245; 1289★ / 158 forks / size 64035; README SHA 4190093c64ee75b26e9726daa3b00cbcf6d3157a; MIT; caijinchun/nanojev-arena ≠ liao96312/jev-arena-nanojev ≠ zwliJay/jev-forge ≠ NanoJev; not TypeSafe Jev; open replica / specialist gameplay S1; soft scores ≠ hard gates; Game success ≠ calibrated Noul; local type boolean ≠ TypeSafe noul; invented_signal false; do not reopen or amend PR #31/#32/#33/#35; notes.md §115 + + +**0920 jcr:** pre-action lookup is not a done-check and not a permission gate. Ambiguity / no-match / depth-limit are explicit abstention paths. NiazMorshed2007/jcr. `notes.md` §116. + +User-provided 0920 jcr uniqueness lock: NiazMorshed2007/jcr MIT; site https://jcr.niazmorshed.dev; topics ai-agents,jev,mcp; **4★**; HEAD `138b3832`; README SHA `2a49dbc1`; LICENSE SHA `46231303`; size **14850**; Jev Capability Resolver; one tool to find documented deterministic commands in a nested capability tree; returns context; **does not execute**; skills = workflow+judgment; capabilities = individual operations; format independent of Jev; proposed open standard exploration; classify (Jev) → optional OpenAI decompose compound → beam search geometric mean of routing probs; keep up to 3 paths ≥60% of best (JCR_BAND_RATIO 0.6); ambiguity / no-match / depth-limit explicit; soft scores ≠ hard gates; 0.6 band is application policy; routing ≠ permission; docs ≠ authority to run; sol-vs-opus5-20 *theirs*: 20 scenarios × 4 variants = 80 runs; lookup+explain only, no execution; Claude Opus 5: agent input 108,585→15,819 (−85%), cost $0.3700→$0.1222 (−67%), wall 105.5s→77.7s; Codex GPT-5.6-Sol: 61,952→47,669 (−23%), $0.1377→$0.1151 (−16%), wall 25.3s→62.4s (Sol slower with JCR in 19/20); One Sol outlier 372.6s / 193 Jev calls; n=1 per cell; Not Harbor task-execution; Claude/Codex harnesses; compare mode; 50 scenarios bundled; 11 groups, 960 nodes, 11,360 items; 16 routing rounds per step; NiazMorshed2007/jcr ≠ skill-broker ≠ skillranker ≠ jev-sift ≠ jev-lens ≠ jevusher ≠ jev_select_capability; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33/#34; notes.md §116 diff --git a/.agents/skills/augustus/references/applied-mappings.md b/.agents/skills/augustus/references/applied-mappings.md index 95f83616..70688caf 100644 --- a/.agents/skills/augustus/references/applied-mappings.md +++ b/.agents/skills/augustus/references/applied-mappings.md @@ -2421,6 +2421,11 @@ lastModified UNCHANGED `2026-09-20T04:29:16.000Z`; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#32/#33. Do not copy keys / `npx` / `pip` / `uv` / `.env` / `TYPESAFE_API_KEY`. Soft Noul ≠ hard safety. + +### 0920 jcr capability resolver (`notes.md` §116) + +Place JCR as retrieve-wide → decide → evidence-set on a nested capability tree. Skills stay workflow+judgment; capabilities are individual operations. The resolver returns documentation; the harness owns execution, credentials, and checks. Soft scores ≠ hard gates. routing ≠ permission. docs ≠ authority to run. NiazMorshed2007/jcr ≠ skill-broker ≠ skillranker ≠ jev-sift ≠ jev-lens ≠ jevusher ≠ jev_select_capability. + Hourly 0743 uniqueness lock: Heman10x-NGU/Verdict-open-jev ≠ Heman10x-NGU/openJev-verdict-2.0; TF-IDF + LogReg ECE 0.0207 vs Jev 0.1440; Verdict-open-jev 48.07% vs Jev 90.80%; abstention combined recall 10.00%; p50 35.58 ms; K=25 (maximum capacity) 72.00%; 0.85 coverage 84.60% selective risk 1.18%; 26.1× faster than standard Qwen JSON generation; Jevify 90.0% / 167 ms CUDA graphs disabled; Finding 1: Brier on stated confidence alone is a trap; grpo_rlcr 0.78 / ECE 0.084; reliability 0.007 but resolution 0.000; 27 900 schema-driven decisions; 13 600 / 13 600 questions; candidate mass min 0.99999624; 22 configs · 166,054 rows · 4 calibration-gold; sha a39eba3f; Student B MAE 0.148 / Pearson 0.836 / 86.0%; pngwn/open-jev-laya-bench README 404; sha 9f69c742 likes 2; HDFS 0.9933 (745/750) / retain 0.0084; BGL ERROR/FATAL protection 1.0000; 2,479 / 2,500 HDFS uncertain; cache hit 0.9648 (2412/2500); $0.153936 estimated; E2 recomputes from saved probabilities; Space sha eda59e0a; MASSIVE English 0.783 / Khmer 0.033 / Hindi 0.133; 40–48 rows too small to ship T; T never changes argmax; siren2345/jev-single-decode-transformers ≠ siren2345/jev-single-decode; Split Transformers experiment from llama.cpp runtime; tanayvasishtha/jev-lab ≠ dairui1/jev-lab ≠ BrendanH18/jev-lab ≠ yibie/laya-jev-lab; Four experiments stress-testing TypeSafe's Jev: calibration, bundle bias, label bias, and ensembling; second pass must be $0.00 from cache; The pages never call Jev; Gemma 4 31B 77.0% / Jev 1.13.0 61.4% / Laya 322M 0.0%; restriction state 95.0% against 84.4%; None of the systems are particularly good at knowing when to stop and ask; They skip the question and call a tool directly; 100% schema pass; six-field joint 48.8% vs 72.8%; ywchiu/jev_benchmark ≠ Running-Dolphins/jev-bench ≠ Praveenrajus/jev-bench; ACT / REVIEW / FALLBACK; A provider failure, timeout, malformed output, or missing answer is **not** a policy outcome; confidence is descriptive provider output, not a substitute for probability; Quality denominators include only valid scored answers; an exact halfway tie chooses the lower level; aiwithenoch/Jev-Skill ≠ simplosophy/jev-skill ≠ laguagu/jev-skills; The local path does not claim to turn a smaller checkpoint into Jev; Low support becomes decision: "review"; MIT-0 SPDX NOASSERTION; current-llm; 结构兼容,不是 Jev 模型能力; altryne/jevify ≠ Mintzs/jevify ≠ gulagala001/jevify ≠ uspraveen/Jevify; Find where Jev belongs. Design the questions. Measure the difference; TypeAR-AI/TypeAR 301 → TypeLLM/TypeLLM; TypeLLM/TypeLLM 16★; SemIf 2241★ (+34 vs §111 2207); jevlike 1051★ (+8 vs 1043); AnotiaWang 98★ (+1 vs 97); yibie/awesome-jev 525★ (+19 vs 506); Laya likes 864 (was 822); tracker likes 67 (+3 vs 64); lastModified UNCHANGED `2026-09-20T04:29:16.000Z`; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#32/#33; notes.md §113 **Hourly 0646 HIGH (`notes.md` §111).** Measurement PRIMARY + datasets/Spaces + applied/skills/economics: @@ -2497,3 +2502,6 @@ not TypeSafe Jev; open replica / specialist gameplay S1. Game success ≠ calibr caijinchun/nanojev-arena ≠ liao96312/jev-arena-nanojev ≠ zwliJay/jev-forge ≠ NanoJev. Do not copy `pip` / `snapshot_download` / `serve_decisions`. User-provided 0915 uniqueness lock: TianyuCodings/NanoJev unified-games-v1 densify; A 0.6B parallel decision model: states and questions in, complete probability distributions out. Zero output-token decoding.; One model, four games; ViZDoom Basic 128/128 vs Jev 56/128; Predict Position 27/128 vs Jev 11/128; Maze 225 attempts vs Jev 2738; Snake 30 food / 256 steps; held-out Maze 4/10 Snake 8/8 Basic 128/128 Predict 27/128; Untuned Qwen3-0.6B baseline; 18,760 questions per variant; 16,333 ViZDoom; 896 Predict Position expert episodes; hard_lr1e5; mix weights 1/3, 1/3, 1/6, 1/6; Hub C-Tianyu/NanoJev revision unified-games-v1 likes 58; dataset C-Tianyu/NanoJev-Data likes 5; HEAD 618cea6d906d54e128360786d12f703fff2b1245; 1289★ / 158 forks / size 64035; README SHA 4190093c64ee75b26e9726daa3b00cbcf6d3157a; MIT; caijinchun/nanojev-arena ≠ liao96312/jev-arena-nanojev ≠ zwliJay/jev-forge ≠ NanoJev; not TypeSafe Jev; open replica / specialist gameplay S1; soft scores ≠ hard gates; Game success ≠ calibrated Noul; local type boolean ≠ TypeSafe noul; invented_signal false; do not reopen or amend PR #31/#32/#33/#35; notes.md §115 + + +User-provided 0920 jcr uniqueness lock: NiazMorshed2007/jcr MIT; site https://jcr.niazmorshed.dev; topics ai-agents,jev,mcp; **4★**; HEAD `138b3832`; README SHA `2a49dbc1`; LICENSE SHA `46231303`; size **14850**; Jev Capability Resolver; one tool to find documented deterministic commands in a nested capability tree; returns context; **does not execute**; skills = workflow+judgment; capabilities = individual operations; format independent of Jev; proposed open standard exploration; classify (Jev) → optional OpenAI decompose compound → beam search geometric mean of routing probs; keep up to 3 paths ≥60% of best (JCR_BAND_RATIO 0.6); ambiguity / no-match / depth-limit explicit; soft scores ≠ hard gates; 0.6 band is application policy; routing ≠ permission; docs ≠ authority to run; sol-vs-opus5-20 *theirs*: 20 scenarios × 4 variants = 80 runs; lookup+explain only, no execution; Claude Opus 5: agent input 108,585→15,819 (−85%), cost $0.3700→$0.1222 (−67%), wall 105.5s→77.7s; Codex GPT-5.6-Sol: 61,952→47,669 (−23%), $0.1377→$0.1151 (−16%), wall 25.3s→62.4s (Sol slower with JCR in 19/20); One Sol outlier 372.6s / 193 Jev calls; n=1 per cell; Not Harbor task-execution; Claude/Codex harnesses; compare mode; 50 scenarios bundled; 11 groups, 960 nodes, 11,360 items; 16 routing rounds per step; NiazMorshed2007/jcr ≠ skill-broker ≠ skillranker ≠ jev-sift ≠ jev-lens ≠ jevusher ≠ jev_select_capability; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33/#34; notes.md §116 diff --git a/.agents/skills/augustus/references/composition-algebra.md b/.agents/skills/augustus/references/composition-algebra.md index 7f6c8211..e12327b2 100644 --- a/.agents/skills/augustus/references/composition-algebra.md +++ b/.agents/skills/augustus/references/composition-algebra.md @@ -1926,6 +1926,7 @@ User-provided 0806 uniqueness lock: people who compare Jev against GPT-5.6 has n position 1 (Operand). current-llm. 结构兼容,不是 Jev 模型能力. Full cards: `applied-mappings.md`, `faq.md`. + 289. **Hysteresis as policy** (edgardcham/huncho): positions 3 (Gate) × 7 (Policy). A hunch is a probability with a policy attached. { enter: 0.8, exit: 0.6 } is hysteresis. replay a policy change without inference. @@ -2047,6 +2048,55 @@ Soft Noul ≠ hard safety. User-provided 0915 uniqueness lock: TianyuCodings/NanoJev unified-games-v1 densify; A 0.6B parallel decision model: states and questions in, complete probability distributions out. Zero output-token decoding.; One model, four games; ViZDoom Basic 128/128 vs Jev 56/128; Predict Position 27/128 vs Jev 11/128; Maze 225 attempts vs Jev 2738; Snake 30 food / 256 steps; held-out Maze 4/10 Snake 8/8 Basic 128/128 Predict 27/128; Untuned Qwen3-0.6B baseline; 18,760 questions per variant; 16,333 ViZDoom; 896 Predict Position expert episodes; hard_lr1e5; mix weights 1/3, 1/3, 1/6, 1/6; Hub C-Tianyu/NanoJev revision unified-games-v1 likes 58; dataset C-Tianyu/NanoJev-Data likes 5; HEAD 618cea6d906d54e128360786d12f703fff2b1245; 1289★ / 158 forks / size 64035; README SHA 4190093c64ee75b26e9726daa3b00cbcf6d3157a; MIT; caijinchun/nanojev-arena ≠ liao96312/jev-arena-nanojev ≠ zwliJay/jev-forge ≠ NanoJev; not TypeSafe Jev; open replica / specialist gameplay S1; soft scores ≠ hard gates; Game success ≠ calibrated Noul; local type boolean ≠ TypeSafe noul; invented_signal false; do not reopen or amend PR #31/#32/#33/#35; notes.md §115 +309. **Capability-tree retrieve-wide→decide→evidence-set** (NiazMorshed2007/jcr): + positions 2 (Post-judge) × 4 (Selector) × 5 (Prior). + NiazMorshed2007/jcr ≠ skill-broker ≠ skillranker ≠ jev-sift ≠ jev-lens ≠ jevusher ≠ jev_select_capability. + one tool to find documented deterministic commands in a nested capability tree. + returns context. **does not execute**. + Full cards: `mappings.md` §4, `faq.md`, `mental-models.md`. +310. **Skills vs capability catalogs** (NiazMorshed2007/jcr): + position 2 (Post-judge). skills = workflow+judgment; capabilities = individual operations. + format independent of Jev. proposed open standard exploration. + Full cards: `applied-mappings.md`, `faq.md`. +311. **0.6 band is application policy** (NiazMorshed2007/jcr): + positions 3 (Gate) × 7 (Policy). + keep up to 3 paths ≥60% of best (JCR_BAND_RATIO 0.6). + soft scores ≠ hard gates. 0.6 band is application policy. + Full cards: `mixed-architecture.md`, `validation.md`. +312. **Routing ≠ permission / docs ≠ authority to run** (NiazMorshed2007/jcr): + positions 7 (Policy) × 11 (Bounds). + routing ≠ permission. docs ≠ authority to run. + JCR returns documentation. It does not execute commands. + Full cards: `formal-methods.md`, `faq.md`. +313. **Beam as control (geometric mean)** (NiazMorshed2007/jcr): + position 4 (Selector). Cookbook cousin (`notes.md` §2 K=3), not a new species. + classify (Jev) → optional OpenAI decompose compound → beam search geometric mean of routing probs. + ambiguity / no-match / depth-limit explicit. 16 routing rounds per step. + Full cards: `mental-models.md`, `question-design.md`. +314. **VOI of context admission** (NiazMorshed2007/jcr): + position 5 (Prior). Search stays outside the main agent; selected `context` is the evidence set. + 11 groups, 960 nodes, 11,360 items. + Full cards: `mixed-architecture.md`, `agent-self-assessment.md`. +315. **Measurement honesty / wall-time mixed** (NiazMorshed2007/jcr): + position 8 (Metric). sol-vs-opus5-20 *theirs*: 20 scenarios × 4 variants = 80 runs. + lookup+explain only, no execution. n=1 per cell. Not Harbor task-execution. + Claude Opus 5: agent input 108,585→15,819 (−85%), cost $0.3700→$0.1222 (−67%), wall 105.5s→77.7s. + Codex GPT-5.6-Sol: 61,952→47,669 (−23%), $0.1377→$0.1151 (−16%), wall 25.3s→62.4s (Sol slower with JCR in 19/20). + One Sol outlier 372.6s / 193 Jev calls. + Full cards: `validation.md`, `faq.md`. +316. **Lookup+explain only / Claude and Codex harnesses** (NiazMorshed2007/jcr): + positions 8 (Metric) × 11 (Bounds). Claude/Codex harnesses. compare mode. 50 scenarios bundled. + Full cards: `validation.md`. + +User-provided 0920 jcr items 309–316 (`notes.md` §116). Do **not** +re-fold 0743 items 273–288 / merged #30 items 268–272 / 0646 items 248–267. +Merged #35 owns §114 / items 289–302 / batch #97. Merged #36 owns §115 / 303–308 / #98. Open #37 owns §117 / 315–321 / #100 (item overlap 315–316 is #37's remap). +NiazMorshed2007/jcr ≠ skill-broker ≠ skillranker ≠ jev-sift ≠ jev-lens ≠ jevusher ≠ jev_select_capability. +soft scores ≠ hard gates; 0.6 band is application policy; +routing ≠ permission; docs ≠ authority to run; +n=1 per cell; Not Harbor task-execution; +do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33/#34. +Soft Noul ≠ hard safety. Hourly 0743 items 273–288 (`notes.md` §113). Do **not** re-fold merged #30 items 268–272 / 0646 items 248–267 / 0541 items 226–247 / 0439 items 202–225 / 0345 items 186–201 / 0243 items 178–185 / 0145 items 161–177 / 0042 items 149–160 / 2340 items 138–148 / 2246 items 129–137 / 2145 items 120–128 / 2041 items 111–119 / 1943 items 102–110 / @@ -2140,3 +2190,5 @@ mechanism / §60 six-gates / §78 v1.2 board. Soft Noul ≠ hard safety. Hourly 0843 uniqueness lock: A hunch is a probability with a policy attached; { enter: 0.8, exit: 0.6 } is hysteresis; replay a policy change without inference; Decision models are providers, not the product; huncho ≠ Kungie/gut ≠ carldaws/hunch ≠ tpellet/hunch; pretrained Qwen2.5 base ECE 0.030 (0.5B) / 0.040 (7B); instruct 0.302 / 0.269; 70.9% → 70.0% mean conf 74.1% → 96.7%; temperature scaling still matches it in-distribution; No Jev API was called; Qwen2.5 ≠ Archer; Qwen/Qwen3.8-27B ≠ Archer; 学習済みモデル v0.1 は準備中です; bool AUROC 0.523; 先頭だと0件、末尾だと250件; 温度を渡さない場合、確率は較正されていません; このリポジトリには Jev を呼ぶコードが存在しません; g0runmezadam/what-is-jev IS tunahansahin897/what-is-jev (same GitHub id 1378007307); 947 repos scored; A 273 · B 302 · C 372; LLM rubric ≠ benches; Data as of 2026-09-20; HEAD 895b9498; README SHA 3ae98c56; 13 focused checks and one mutually exclusive outcome; Probabilities are advisory, not calibrated guarantees; omni-/ask-jev ≠ pedroknigge/mcp_jev; pd.cut bins by equal width while jeval bins by quantile; ECE 0.113 and ECE 0.076; jeval drift is not implemented yet; rlaope/jeval ≠ dayhaysoos/jevals; calibration does not compose; ECE has exactly zero statistical power to detect the failure mode that kills trajectories; 25–60× headline withdrawn; P(all-correct): 0.0071 vs 0.0001; TCE / AMS; Qwen 3.8 sparring ≠ Archer; Deferred Crispification; light_cutoff_applied_to_combination 0; recorded run, kinematic animation; BANKING77 Accuracy BERT-Base 93.02 Jev 79.90; Analyse jev calibration (NLL, ECE) backlog; BERT figures are published supervised references, not zero-shot; 档位措辞效应 分数极差中位 0.50、最大 1.32; 修好后对照组是 0.01; 不是 benchmark; 概率没做 calibration; ~1,430 API calls, about $0.15; xiaohuaxi/jev-study ≠ wjdjdakf17/jev-study ≠ baekenough/jev-study; AND: product (independence assumed and recorded in the trace); chat model's stated confidence is not calibrated; circuit-vl-4b ≠ Archer; Bring your own API key; vamsikrishna2421/jev-usecases ≠ whyashthakker/awesome-jev-use-cases; catalog ≠ endorsement; SemIf 2237★ (+30 vs §111 2207); jevlike 1049★ (+6 vs 1043); TypeLLM/TypeLLM 16★; AnotiaWang 98★; yibie/awesome-jev 520★ (+14 vs 506); Laya likes 861 (was 822); tracker likes 66 (+2 vs 64) lastModified UNCHANGED; Blackwood likes 2 gated manual; Archer still promised_not_landed; Hub archerhume/4rcherhume HTTP 401; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33; notes.md §114 + +User-provided 0920 jcr uniqueness lock: NiazMorshed2007/jcr MIT; site https://jcr.niazmorshed.dev; topics ai-agents,jev,mcp; **4★**; HEAD `138b3832`; README SHA `2a49dbc1`; LICENSE SHA `46231303`; size **14850**; Jev Capability Resolver; one tool to find documented deterministic commands in a nested capability tree; returns context; **does not execute**; skills = workflow+judgment; capabilities = individual operations; format independent of Jev; proposed open standard exploration; classify (Jev) → optional OpenAI decompose compound → beam search geometric mean of routing probs; keep up to 3 paths ≥60% of best (JCR_BAND_RATIO 0.6); ambiguity / no-match / depth-limit explicit; soft scores ≠ hard gates; 0.6 band is application policy; routing ≠ permission; docs ≠ authority to run; sol-vs-opus5-20 *theirs*: 20 scenarios × 4 variants = 80 runs; lookup+explain only, no execution; Claude Opus 5: agent input 108,585→15,819 (−85%), cost $0.3700→$0.1222 (−67%), wall 105.5s→77.7s; Codex GPT-5.6-Sol: 61,952→47,669 (−23%), $0.1377→$0.1151 (−16%), wall 25.3s→62.4s (Sol slower with JCR in 19/20); One Sol outlier 372.6s / 193 Jev calls; n=1 per cell; Not Harbor task-execution; Claude/Codex harnesses; compare mode; 50 scenarios bundled; 11 groups, 960 nodes, 11,360 items; 16 routing rounds per step; NiazMorshed2007/jcr ≠ skill-broker ≠ skillranker ≠ jev-sift ≠ jev-lens ≠ jevusher ≠ jev_select_capability; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33/#34; notes.md §116 diff --git a/.agents/skills/augustus/references/faq.md b/.agents/skills/augustus/references/faq.md index f0e4c5f9..7e2f4b85 100644 --- a/.agents/skills/augustus/references/faq.md +++ b/.agents/skills/augustus/references/faq.md @@ -539,6 +539,38 @@ without entering main agent context first (paths/URLs). Uncertain / errors / truncation ≠ irrelevant. Transport tests ≠ accuracy. `mappings.md` §4; `notes.md` §55, §56. + +## Is JCR a skill loader? + +No. [NiazMorshed2007/jcr](https://github.com/NiazMorshed2007/jcr) +is a **capability resolver**. Skills = workflow+judgment; +capabilities = individual operations; format independent of +Jev (*theirs*). One tool finds documented deterministic +commands in a nested capability tree and **returns context**. +It **does not execute**. Distinct from skill-broker (grants), +skillranker (advisory VOI over skills), jev-sift (classify +path/url/text), jev-lens (pre-send views), jevusher +(token admission), and Codex `jev_select_capability`. +`notes.md` §116. + +## Does a 0.6 band authorize the returned command? + +No. Soft scores ≠ hard gates. `JCR_BAND_RATIO` 0.6 is +application policy: keep up to 3 paths ≥60% of the best +geometric-mean routing score. Routing ≠ permission. Docs ≠ +authority to run. A match can still be wrong (*theirs*). +The harness still supplies values, credentials, order, and +checks. `notes.md` §116. + +## Did JCR win a Harbor task-execution bench? + +No. `sol-vs-opus5-20` *theirs* is lookup+explain only, 20 +scenarios × 4 variants = 80 runs, n=1 per cell. Token and +cost fell in both harnesses; wall time mixed (Sol slower +with JCR in 19/20; one 372.6s / 193 Jev-call outlier). +Not Harbor task-execution. Do not steal −85% as a class +constant. `validation.md`; `notes.md` §116. + ## Dump files into context, or classify first? Classify first when the items are **not** already in the main agent @@ -3589,3 +3621,5 @@ Hourly 0646 uniqueness lock: Calibration is not alpha; NO CURRENT ALPHA CANDIDAT Hourly 0843 uniqueness lock: A hunch is a probability with a policy attached; { enter: 0.8, exit: 0.6 } is hysteresis; replay a policy change without inference; Decision models are providers, not the product; huncho ≠ Kungie/gut ≠ carldaws/hunch ≠ tpellet/hunch; pretrained Qwen2.5 base ECE 0.030 (0.5B) / 0.040 (7B); instruct 0.302 / 0.269; 70.9% → 70.0% mean conf 74.1% → 96.7%; temperature scaling still matches it in-distribution; No Jev API was called; Qwen2.5 ≠ Archer; Qwen/Qwen3.8-27B ≠ Archer; 学習済みモデル v0.1 は準備中です; bool AUROC 0.523; 先頭だと0件、末尾だと250件; 温度を渡さない場合、確率は較正されていません; このリポジトリには Jev を呼ぶコードが存在しません; g0runmezadam/what-is-jev IS tunahansahin897/what-is-jev (same GitHub id 1378007307); 947 repos scored; A 273 · B 302 · C 372; LLM rubric ≠ benches; Data as of 2026-09-20; HEAD 895b9498; README SHA 3ae98c56; 13 focused checks and one mutually exclusive outcome; Probabilities are advisory, not calibrated guarantees; omni-/ask-jev ≠ pedroknigge/mcp_jev; pd.cut bins by equal width while jeval bins by quantile; ECE 0.113 and ECE 0.076; jeval drift is not implemented yet; rlaope/jeval ≠ dayhaysoos/jevals; calibration does not compose; ECE has exactly zero statistical power to detect the failure mode that kills trajectories; 25–60× headline withdrawn; P(all-correct): 0.0071 vs 0.0001; TCE / AMS; Qwen 3.8 sparring ≠ Archer; Deferred Crispification; light_cutoff_applied_to_combination 0; recorded run, kinematic animation; BANKING77 Accuracy BERT-Base 93.02 Jev 79.90; Analyse jev calibration (NLL, ECE) backlog; BERT figures are published supervised references, not zero-shot; 档位措辞效应 分数极差中位 0.50、最大 1.32; 修好后对照组是 0.01; 不是 benchmark; 概率没做 calibration; ~1,430 API calls, about $0.15; xiaohuaxi/jev-study ≠ wjdjdakf17/jev-study ≠ baekenough/jev-study; AND: product (independence assumed and recorded in the trace); chat model's stated confidence is not calibrated; circuit-vl-4b ≠ Archer; Bring your own API key; vamsikrishna2421/jev-usecases ≠ whyashthakker/awesome-jev-use-cases; catalog ≠ endorsement; SemIf 2237★ (+30 vs §111 2207); jevlike 1049★ (+6 vs 1043); TypeLLM/TypeLLM 16★; AnotiaWang 98★; yibie/awesome-jev 520★ (+14 vs 506); Laya likes 861 (was 822); tracker likes 66 (+2 vs 64) lastModified UNCHANGED; Blackwood likes 2 gated manual; Archer still promised_not_landed; Hub archerhume/4rcherhume HTTP 401; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33; notes.md §114 + +User-provided 0920 jcr uniqueness lock: NiazMorshed2007/jcr MIT; site https://jcr.niazmorshed.dev; topics ai-agents,jev,mcp; **4★**; HEAD `138b3832`; README SHA `2a49dbc1`; LICENSE SHA `46231303`; size **14850**; Jev Capability Resolver; one tool to find documented deterministic commands in a nested capability tree; returns context; **does not execute**; skills = workflow+judgment; capabilities = individual operations; format independent of Jev; proposed open standard exploration; classify (Jev) → optional OpenAI decompose compound → beam search geometric mean of routing probs; keep up to 3 paths ≥60% of best (JCR_BAND_RATIO 0.6); ambiguity / no-match / depth-limit explicit; soft scores ≠ hard gates; 0.6 band is application policy; routing ≠ permission; docs ≠ authority to run; sol-vs-opus5-20 *theirs*: 20 scenarios × 4 variants = 80 runs; lookup+explain only, no execution; Claude Opus 5: agent input 108,585→15,819 (−85%), cost $0.3700→$0.1222 (−67%), wall 105.5s→77.7s; Codex GPT-5.6-Sol: 61,952→47,669 (−23%), $0.1377→$0.1151 (−16%), wall 25.3s→62.4s (Sol slower with JCR in 19/20); One Sol outlier 372.6s / 193 Jev calls; n=1 per cell; Not Harbor task-execution; Claude/Codex harnesses; compare mode; 50 scenarios bundled; 11 groups, 960 nodes, 11,360 items; 16 routing rounds per step; NiazMorshed2007/jcr ≠ skill-broker ≠ skillranker ≠ jev-sift ≠ jev-lens ≠ jevusher ≠ jev_select_capability; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33/#34; notes.md §116 diff --git a/.agents/skills/augustus/references/formal-methods.md b/.agents/skills/augustus/references/formal-methods.md index 7ba393e8..75e795f9 100644 --- a/.agents/skills/augustus/references/formal-methods.md +++ b/.agents/skills/augustus/references/formal-methods.md @@ -1351,6 +1351,8 @@ Noul is the same theater. User-provided 0915 uniqueness lock: TianyuCodings/NanoJev unified-games-v1 densify; A 0.6B parallel decision model: states and questions in, complete probability distributions out. Zero output-token decoding.; One model, four games; ViZDoom Basic 128/128 vs Jev 56/128; Predict Position 27/128 vs Jev 11/128; Maze 225 attempts vs Jev 2738; Snake 30 food / 256 steps; held-out Maze 4/10 Snake 8/8 Basic 128/128 Predict 27/128; Untuned Qwen3-0.6B baseline; 18,760 questions per variant; 16,333 ViZDoom; 896 Predict Position expert episodes; hard_lr1e5; mix weights 1/3, 1/3, 1/6, 1/6; Hub C-Tianyu/NanoJev revision unified-games-v1 likes 58; dataset C-Tianyu/NanoJev-Data likes 5; HEAD 618cea6d906d54e128360786d12f703fff2b1245; 1289★ / 158 forks / size 64035; README SHA 4190093c64ee75b26e9726daa3b00cbcf6d3157a; MIT; caijinchun/nanojev-arena ≠ liao96312/jev-arena-nanojev ≠ zwliJay/jev-forge ≠ NanoJev; not TypeSafe Jev; open replica / specialist gameplay S1; soft scores ≠ hard gates; Game success ≠ calibrated Noul; local type boolean ≠ TypeSafe noul; invented_signal false; do not reopen or amend PR #31/#32/#33/#35; notes.md §115 +JCR compose (`notes.md` §116): routing ≠ permission; docs ≠ authority to run. Returned `context` is evidence. Credentials, side effects, and verification stay exact. Soft scores ≠ hard gates. 0.6 band is application policy, not a discharged proof. + Propose two placements if the hole is mixed (e.g. TLA+ on the protocol + DST on the SDK + judgment triaging failing seeds). Do not invent a hybrid "verified by Noul" API. @@ -1361,3 +1363,5 @@ Related: `formal-semi-formal.md` (one screen); `mental-models.md`; preference lint; `faq.md`; `boundary-audit.md`. Hourly 0843 uniqueness lock: A hunch is a probability with a policy attached; { enter: 0.8, exit: 0.6 } is hysteresis; replay a policy change without inference; Decision models are providers, not the product; huncho ≠ Kungie/gut ≠ carldaws/hunch ≠ tpellet/hunch; pretrained Qwen2.5 base ECE 0.030 (0.5B) / 0.040 (7B); instruct 0.302 / 0.269; 70.9% → 70.0% mean conf 74.1% → 96.7%; temperature scaling still matches it in-distribution; No Jev API was called; Qwen2.5 ≠ Archer; Qwen/Qwen3.8-27B ≠ Archer; 学習済みモデル v0.1 は準備中です; bool AUROC 0.523; 先頭だと0件、末尾だと250件; 温度を渡さない場合、確率は較正されていません; このリポジトリには Jev を呼ぶコードが存在しません; g0runmezadam/what-is-jev IS tunahansahin897/what-is-jev (same GitHub id 1378007307); 947 repos scored; A 273 · B 302 · C 372; LLM rubric ≠ benches; Data as of 2026-09-20; HEAD 895b9498; README SHA 3ae98c56; 13 focused checks and one mutually exclusive outcome; Probabilities are advisory, not calibrated guarantees; omni-/ask-jev ≠ pedroknigge/mcp_jev; pd.cut bins by equal width while jeval bins by quantile; ECE 0.113 and ECE 0.076; jeval drift is not implemented yet; rlaope/jeval ≠ dayhaysoos/jevals; calibration does not compose; ECE has exactly zero statistical power to detect the failure mode that kills trajectories; 25–60× headline withdrawn; P(all-correct): 0.0071 vs 0.0001; TCE / AMS; Qwen 3.8 sparring ≠ Archer; Deferred Crispification; light_cutoff_applied_to_combination 0; recorded run, kinematic animation; BANKING77 Accuracy BERT-Base 93.02 Jev 79.90; Analyse jev calibration (NLL, ECE) backlog; BERT figures are published supervised references, not zero-shot; 档位措辞效应 分数极差中位 0.50、最大 1.32; 修好后对照组是 0.01; 不是 benchmark; 概率没做 calibration; ~1,430 API calls, about $0.15; xiaohuaxi/jev-study ≠ wjdjdakf17/jev-study ≠ baekenough/jev-study; AND: product (independence assumed and recorded in the trace); chat model's stated confidence is not calibrated; circuit-vl-4b ≠ Archer; Bring your own API key; vamsikrishna2421/jev-usecases ≠ whyashthakker/awesome-jev-use-cases; catalog ≠ endorsement; SemIf 2237★ (+30 vs §111 2207); jevlike 1049★ (+6 vs 1043); TypeLLM/TypeLLM 16★; AnotiaWang 98★; yibie/awesome-jev 520★ (+14 vs 506); Laya likes 861 (was 822); tracker likes 66 (+2 vs 64) lastModified UNCHANGED; Blackwood likes 2 gated manual; Archer still promised_not_landed; Hub archerhume/4rcherhume HTTP 401; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33; notes.md §114 + +User-provided 0920 jcr uniqueness lock: NiazMorshed2007/jcr MIT; site https://jcr.niazmorshed.dev; topics ai-agents,jev,mcp; **4★**; HEAD `138b3832`; README SHA `2a49dbc1`; LICENSE SHA `46231303`; size **14850**; Jev Capability Resolver; one tool to find documented deterministic commands in a nested capability tree; returns context; **does not execute**; skills = workflow+judgment; capabilities = individual operations; format independent of Jev; proposed open standard exploration; classify (Jev) → optional OpenAI decompose compound → beam search geometric mean of routing probs; keep up to 3 paths ≥60% of best (JCR_BAND_RATIO 0.6); ambiguity / no-match / depth-limit explicit; soft scores ≠ hard gates; 0.6 band is application policy; routing ≠ permission; docs ≠ authority to run; sol-vs-opus5-20 *theirs*: 20 scenarios × 4 variants = 80 runs; lookup+explain only, no execution; Claude Opus 5: agent input 108,585→15,819 (−85%), cost $0.3700→$0.1222 (−67%), wall 105.5s→77.7s; Codex GPT-5.6-Sol: 61,952→47,669 (−23%), $0.1377→$0.1151 (−16%), wall 25.3s→62.4s (Sol slower with JCR in 19/20); One Sol outlier 372.6s / 193 Jev calls; n=1 per cell; Not Harbor task-execution; Claude/Codex harnesses; compare mode; 50 scenarios bundled; 11 groups, 960 nodes, 11,360 items; 16 routing rounds per step; NiazMorshed2007/jcr ≠ skill-broker ≠ skillranker ≠ jev-sift ≠ jev-lens ≠ jevusher ≠ jev_select_capability; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33/#34; notes.md §116 diff --git a/.agents/skills/augustus/references/formal-semi-formal.md b/.agents/skills/augustus/references/formal-semi-formal.md index 5f8272d2..413bf649 100644 --- a/.agents/skills/augustus/references/formal-semi-formal.md +++ b/.agents/skills/augustus/references/formal-semi-formal.md @@ -73,3 +73,8 @@ establish empirical probability calibration. Do not launder 128/128 as Harbor or as a Noul. User-provided 0915 uniqueness lock: TianyuCodings/NanoJev unified-games-v1 densify; A 0.6B parallel decision model: states and questions in, complete probability distributions out. Zero output-token decoding.; One model, four games; ViZDoom Basic 128/128 vs Jev 56/128; Predict Position 27/128 vs Jev 11/128; Maze 225 attempts vs Jev 2738; Snake 30 food / 256 steps; held-out Maze 4/10 Snake 8/8 Basic 128/128 Predict 27/128; Untuned Qwen3-0.6B baseline; 18,760 questions per variant; 16,333 ViZDoom; 896 Predict Position expert episodes; hard_lr1e5; mix weights 1/3, 1/3, 1/6, 1/6; Hub C-Tianyu/NanoJev revision unified-games-v1 likes 58; dataset C-Tianyu/NanoJev-Data likes 5; HEAD 618cea6d906d54e128360786d12f703fff2b1245; 1289★ / 158 forks / size 64035; README SHA 4190093c64ee75b26e9726daa3b00cbcf6d3157a; MIT; caijinchun/nanojev-arena ≠ liao96312/jev-arena-nanojev ≠ zwliJay/jev-forge ≠ NanoJev; not TypeSafe Jev; open replica / specialist gameplay S1; soft scores ≠ hard gates; Game success ≠ calibrated Noul; local type boolean ≠ TypeSafe noul; invented_signal false; do not reopen or amend PR #31/#32/#33/#35; notes.md §115 + + +JCR (`notes.md` §116): docs ≠ authority to run. routing ≠ permission. A Noul / routing probability is a SENSOR, not a grant. + +User-provided 0920 jcr uniqueness lock: NiazMorshed2007/jcr MIT; site https://jcr.niazmorshed.dev; topics ai-agents,jev,mcp; **4★**; HEAD `138b3832`; README SHA `2a49dbc1`; LICENSE SHA `46231303`; size **14850**; Jev Capability Resolver; one tool to find documented deterministic commands in a nested capability tree; returns context; **does not execute**; skills = workflow+judgment; capabilities = individual operations; format independent of Jev; proposed open standard exploration; classify (Jev) → optional OpenAI decompose compound → beam search geometric mean of routing probs; keep up to 3 paths ≥60% of best (JCR_BAND_RATIO 0.6); ambiguity / no-match / depth-limit explicit; soft scores ≠ hard gates; 0.6 band is application policy; routing ≠ permission; docs ≠ authority to run; sol-vs-opus5-20 *theirs*: 20 scenarios × 4 variants = 80 runs; lookup+explain only, no execution; Claude Opus 5: agent input 108,585→15,819 (−85%), cost $0.3700→$0.1222 (−67%), wall 105.5s→77.7s; Codex GPT-5.6-Sol: 61,952→47,669 (−23%), $0.1377→$0.1151 (−16%), wall 25.3s→62.4s (Sol slower with JCR in 19/20); One Sol outlier 372.6s / 193 Jev calls; n=1 per cell; Not Harbor task-execution; Claude/Codex harnesses; compare mode; 50 scenarios bundled; 11 groups, 960 nodes, 11,360 items; 16 routing rounds per step; NiazMorshed2007/jcr ≠ skill-broker ≠ skillranker ≠ jev-sift ≠ jev-lens ≠ jevusher ≠ jev_select_capability; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33/#34; notes.md §116 diff --git a/.agents/skills/augustus/references/judgment-class.md b/.agents/skills/augustus/references/judgment-class.md index 2777042b..5f716e1f 100644 --- a/.agents/skills/augustus/references/judgment-class.md +++ b/.agents/skills/augustus/references/judgment-class.md @@ -1439,6 +1439,8 @@ decision gate). Do not invent a hybrid API. +**0920 jcr class port (`notes.md` §116):** decide/rank over a nested catalog of deterministic operations; Jev is the routing sensor, not the executor. Beam geometric mean is control, not a new head. NiazMorshed2007/jcr ≠ skill-broker ≠ skillranker ≠ jev-sift ≠ jev-lens ≠ jevusher ≠ jev_select_capability. + **0743 HIGH class ports (`notes.md` §113):** siren2345/jev-single-decode-transformers ≠ siren2345/jev-single-decode; Split Transformers experiment from llama.cpp runtime; softmax over A/B/C ≠ Noul; Heman10x-NGU/Verdict-open-jev ≠ Heman10x-NGU/openJev-verdict-2.0; TF-IDF + LogReg ECE 0.0207 vs Jev 0.1440; current-llm; 结构兼容,不是 Jev 模型能力; The local path does not claim to turn a smaller checkpoint into Jev; TypeAR-AI/TypeAR 301 → TypeLLM/TypeLLM; TypeLLM/TypeLLM 16★. Heman10x-NGU/Verdict-open-jev ≠ Heman10x-NGU/openJev-verdict-2.0; TF-IDF + LogReg ECE 0.0207 vs Jev 0.1440; Verdict-open-jev 48.07% vs Jev 90.80%; abstention combined recall 10.00%; p50 35.58 ms; K=25 (maximum capacity) 72.00%; 0.85 coverage 84.60% selective risk 1.18%; 26.1× faster than standard Qwen JSON generation; Jevify 90.0% / 167 ms CUDA graphs disabled; Finding 1: Brier on stated confidence alone is a trap; grpo_rlcr 0.78 / ECE 0.084; reliability 0.007 but resolution 0.000; 27 900 schema-driven decisions; 13 600 / 13 600 questions; candidate mass min 0.99999624; 22 configs · 166,054 rows · 4 calibration-gold; sha a39eba3f; Student B MAE 0.148 / Pearson 0.836 / 86.0%; pngwn/open-jev-laya-bench README 404; sha 9f69c742 likes 2; HDFS 0.9933 (745/750) / retain 0.0084; BGL ERROR/FATAL protection 1.0000; 2,479 / 2,500 HDFS uncertain; cache hit 0.9648 (2412/2500); $0.153936 estimated; E2 recomputes from saved probabilities; Space sha eda59e0a; MASSIVE English 0.783 / Khmer 0.033 / Hindi 0.133; 40–48 rows too small to ship T; T never changes argmax; siren2345/jev-single-decode-transformers ≠ siren2345/jev-single-decode; Split Transformers experiment from llama.cpp runtime; tanayvasishtha/jev-lab ≠ dairui1/jev-lab ≠ BrendanH18/jev-lab ≠ yibie/laya-jev-lab; Four experiments stress-testing TypeSafe's Jev: calibration, bundle bias, label bias, and ensembling; second pass must be $0.00 from cache; The pages never call Jev; Gemma 4 31B 77.0% / Jev 1.13.0 61.4% / Laya 322M 0.0%; restriction state 95.0% against 84.4%; None of the systems are particularly good at knowing when to stop and ask; They skip the question and call a tool directly; 100% schema pass; six-field joint 48.8% vs 72.8%; ywchiu/jev_benchmark ≠ Running-Dolphins/jev-bench ≠ Praveenrajus/jev-bench; ACT / REVIEW / FALLBACK; A provider failure, timeout, malformed output, or missing answer is **not** a policy outcome; confidence is descriptive provider output, not a substitute for probability; Quality denominators include only valid scored answers; an exact halfway tie chooses the lower level; aiwithenoch/Jev-Skill ≠ simplosophy/jev-skill ≠ laguagu/jev-skills; The local path does not claim to turn a smaller checkpoint into Jev; Low support becomes decision: "review"; MIT-0 SPDX NOASSERTION; current-llm; 结构兼容,不是 Jev 模型能力; altryne/jevify ≠ Mintzs/jevify ≠ gulagala001/jevify ≠ uspraveen/Jevify; Find where Jev belongs. Design the questions. Measure the difference; TypeAR-AI/TypeAR 301 → TypeLLM/TypeLLM; TypeLLM/TypeLLM 16★; SemIf 2241★ (+34 vs §111 2207); jevlike 1051★ (+8 vs 1043); AnotiaWang 98★ (+1 vs 97); yibie/awesome-jev 525★ (+19 vs 506); Laya likes 864 (was 822); tracker likes 67 (+3 vs 64); lastModified UNCHANGED `2026-09-20T04:29:16.000Z`; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#32/#33; notes.md §113 Hourly 0743 uniqueness lock: Heman10x-NGU/Verdict-open-jev ≠ Heman10x-NGU/openJev-verdict-2.0; TF-IDF + LogReg ECE 0.0207 vs Jev 0.1440; Verdict-open-jev 48.07% vs Jev 90.80%; abstention combined recall 10.00%; p50 35.58 ms; K=25 (maximum capacity) 72.00%; 0.85 coverage 84.60% selective risk 1.18%; 26.1× faster than standard Qwen JSON generation; Jevify 90.0% / 167 ms CUDA graphs disabled; Finding 1: Brier on stated confidence alone is a trap; grpo_rlcr 0.78 / ECE 0.084; reliability 0.007 but resolution 0.000; 27 900 schema-driven decisions; 13 600 / 13 600 questions; candidate mass min 0.99999624; 22 configs · 166,054 rows · 4 calibration-gold; sha a39eba3f; Student B MAE 0.148 / Pearson 0.836 / 86.0%; pngwn/open-jev-laya-bench README 404; sha 9f69c742 likes 2; HDFS 0.9933 (745/750) / retain 0.0084; BGL ERROR/FATAL protection 1.0000; 2,479 / 2,500 HDFS uncertain; cache hit 0.9648 (2412/2500); $0.153936 estimated; E2 recomputes from saved probabilities; Space sha eda59e0a; MASSIVE English 0.783 / Khmer 0.033 / Hindi 0.133; 40–48 rows too small to ship T; T never changes argmax; siren2345/jev-single-decode-transformers ≠ siren2345/jev-single-decode; Split Transformers experiment from llama.cpp runtime; tanayvasishtha/jev-lab ≠ dairui1/jev-lab ≠ BrendanH18/jev-lab ≠ yibie/laya-jev-lab; Four experiments stress-testing TypeSafe's Jev: calibration, bundle bias, label bias, and ensembling; second pass must be $0.00 from cache; The pages never call Jev; Gemma 4 31B 77.0% / Jev 1.13.0 61.4% / Laya 322M 0.0%; restriction state 95.0% against 84.4%; None of the systems are particularly good at knowing when to stop and ask; They skip the question and call a tool directly; 100% schema pass; six-field joint 48.8% vs 72.8%; ywchiu/jev_benchmark ≠ Running-Dolphins/jev-bench ≠ Praveenrajus/jev-bench; ACT / REVIEW / FALLBACK; A provider failure, timeout, malformed output, or missing answer is **not** a policy outcome; confidence is descriptive provider output, not a substitute for probability; Quality denominators include only valid scored answers; an exact halfway tie chooses the lower level; aiwithenoch/Jev-Skill ≠ simplosophy/jev-skill ≠ laguagu/jev-skills; The local path does not claim to turn a smaller checkpoint into Jev; Low support becomes decision: "review"; MIT-0 SPDX NOASSERTION; current-llm; 结构兼容,不是 Jev 模型能力; altryne/jevify ≠ Mintzs/jevify ≠ gulagala001/jevify ≠ uspraveen/Jevify; Find where Jev belongs. Design the questions. Measure the difference; TypeAR-AI/TypeAR 301 → TypeLLM/TypeLLM; TypeLLM/TypeLLM 16★; SemIf 2241★ (+34 vs §111 2207); jevlike 1051★ (+8 vs 1043); AnotiaWang 98★ (+1 vs 97); yibie/awesome-jev 525★ (+19 vs 506); Laya likes 864 (was 822); tracker likes 67 (+3 vs 64); lastModified UNCHANGED `2026-09-20T04:29:16.000Z`; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#32/#33; notes.md §113 @@ -1453,3 +1455,6 @@ User-provided 0806 uniqueness lock: people who compare Jev against GPT-5.6 has n Hourly 0843 uniqueness lock: A hunch is a probability with a policy attached; { enter: 0.8, exit: 0.6 } is hysteresis; replay a policy change without inference; Decision models are providers, not the product; huncho ≠ Kungie/gut ≠ carldaws/hunch ≠ tpellet/hunch; pretrained Qwen2.5 base ECE 0.030 (0.5B) / 0.040 (7B); instruct 0.302 / 0.269; 70.9% → 70.0% mean conf 74.1% → 96.7%; temperature scaling still matches it in-distribution; No Jev API was called; Qwen2.5 ≠ Archer; Qwen/Qwen3.8-27B ≠ Archer; 学習済みモデル v0.1 は準備中です; bool AUROC 0.523; 先頭だと0件、末尾だと250件; 温度を渡さない場合、確率は較正されていません; このリポジトリには Jev を呼ぶコードが存在しません; g0runmezadam/what-is-jev IS tunahansahin897/what-is-jev (same GitHub id 1378007307); 947 repos scored; A 273 · B 302 · C 372; LLM rubric ≠ benches; Data as of 2026-09-20; HEAD 895b9498; README SHA 3ae98c56; 13 focused checks and one mutually exclusive outcome; Probabilities are advisory, not calibrated guarantees; omni-/ask-jev ≠ pedroknigge/mcp_jev; pd.cut bins by equal width while jeval bins by quantile; ECE 0.113 and ECE 0.076; jeval drift is not implemented yet; rlaope/jeval ≠ dayhaysoos/jevals; calibration does not compose; ECE has exactly zero statistical power to detect the failure mode that kills trajectories; 25–60× headline withdrawn; P(all-correct): 0.0071 vs 0.0001; TCE / AMS; Qwen 3.8 sparring ≠ Archer; Deferred Crispification; light_cutoff_applied_to_combination 0; recorded run, kinematic animation; BANKING77 Accuracy BERT-Base 93.02 Jev 79.90; Analyse jev calibration (NLL, ECE) backlog; BERT figures are published supervised references, not zero-shot; 档位措辞效应 分数极差中位 0.50、最大 1.32; 修好后对照组是 0.01; 不是 benchmark; 概率没做 calibration; ~1,430 API calls, about $0.15; xiaohuaxi/jev-study ≠ wjdjdakf17/jev-study ≠ baekenough/jev-study; AND: product (independence assumed and recorded in the trace); chat model's stated confidence is not calibrated; circuit-vl-4b ≠ Archer; Bring your own API key; vamsikrishna2421/jev-usecases ≠ whyashthakker/awesome-jev-use-cases; catalog ≠ endorsement; SemIf 2237★ (+30 vs §111 2207); jevlike 1049★ (+6 vs 1043); TypeLLM/TypeLLM 16★; AnotiaWang 98★; yibie/awesome-jev 520★ (+14 vs 506); Laya likes 861 (was 822); tracker likes 66 (+2 vs 64) lastModified UNCHANGED; Blackwood likes 2 gated manual; Archer still promised_not_landed; Hub archerhume/4rcherhume HTTP 401; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33; notes.md §114 User-provided 0915 uniqueness lock: TianyuCodings/NanoJev unified-games-v1 densify; A 0.6B parallel decision model: states and questions in, complete probability distributions out. Zero output-token decoding.; One model, four games; ViZDoom Basic 128/128 vs Jev 56/128; Predict Position 27/128 vs Jev 11/128; Maze 225 attempts vs Jev 2738; Snake 30 food / 256 steps; held-out Maze 4/10 Snake 8/8 Basic 128/128 Predict 27/128; Untuned Qwen3-0.6B baseline; 18,760 questions per variant; 16,333 ViZDoom; 896 Predict Position expert episodes; hard_lr1e5; mix weights 1/3, 1/3, 1/6, 1/6; Hub C-Tianyu/NanoJev revision unified-games-v1 likes 58; dataset C-Tianyu/NanoJev-Data likes 5; HEAD 618cea6d906d54e128360786d12f703fff2b1245; 1289★ / 158 forks / size 64035; README SHA 4190093c64ee75b26e9726daa3b00cbcf6d3157a; MIT; caijinchun/nanojev-arena ≠ liao96312/jev-arena-nanojev ≠ zwliJay/jev-forge ≠ NanoJev; not TypeSafe Jev; open replica / specialist gameplay S1; soft scores ≠ hard gates; Game success ≠ calibrated Noul; local type boolean ≠ TypeSafe noul; invented_signal false; do not reopen or amend PR #31/#32/#33/#35; notes.md §115 + + +User-provided 0920 jcr uniqueness lock: NiazMorshed2007/jcr MIT; site https://jcr.niazmorshed.dev; topics ai-agents,jev,mcp; **4★**; HEAD `138b3832`; README SHA `2a49dbc1`; LICENSE SHA `46231303`; size **14850**; Jev Capability Resolver; one tool to find documented deterministic commands in a nested capability tree; returns context; **does not execute**; skills = workflow+judgment; capabilities = individual operations; format independent of Jev; proposed open standard exploration; classify (Jev) → optional OpenAI decompose compound → beam search geometric mean of routing probs; keep up to 3 paths ≥60% of best (JCR_BAND_RATIO 0.6); ambiguity / no-match / depth-limit explicit; soft scores ≠ hard gates; 0.6 band is application policy; routing ≠ permission; docs ≠ authority to run; sol-vs-opus5-20 *theirs*: 20 scenarios × 4 variants = 80 runs; lookup+explain only, no execution; Claude Opus 5: agent input 108,585→15,819 (−85%), cost $0.3700→$0.1222 (−67%), wall 105.5s→77.7s; Codex GPT-5.6-Sol: 61,952→47,669 (−23%), $0.1377→$0.1151 (−16%), wall 25.3s→62.4s (Sol slower with JCR in 19/20); One Sol outlier 372.6s / 193 Jev calls; n=1 per cell; Not Harbor task-execution; Claude/Codex harnesses; compare mode; 50 scenarios bundled; 11 groups, 960 nodes, 11,360 items; 16 routing rounds per step; NiazMorshed2007/jcr ≠ skill-broker ≠ skillranker ≠ jev-sift ≠ jev-lens ≠ jevusher ≠ jev_select_capability; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33/#34; notes.md §116 diff --git a/.agents/skills/augustus/references/mappings.md b/.agents/skills/augustus/references/mappings.md index 0c517f5b..ad7db704 100644 --- a/.agents/skills/augustus/references/mappings.md +++ b/.agents/skills/augustus/references/mappings.md @@ -2450,6 +2450,8 @@ orchestration; `notes.md` §28. +**0920 jcr HIGH (`notes.md` §116):** retrieve-wide → decide → evidence-set on a capability tree. one tool; returns context; **does not execute**. 0.6 band is application policy. routing ≠ permission. docs ≠ authority to run. n=1 lookup+explain; Not Harbor task-execution. NiazMorshed2007/jcr ≠ skill-broker ≠ skillranker ≠ jev-sift ≠ jev-lens ≠ jevusher ≠ jev_select_capability. + **0743 HIGH (`notes.md` §113):** Heman10x-NGU/Verdict-open-jev ≠ Heman10x-NGU/openJev-verdict-2.0; TF-IDF + LogReg ECE 0.0207 vs Jev 0.1440; Verdict-open-jev 48.07% vs Jev 90.80%; abstention combined recall 10.00%; p50 35.58 ms; K=25 (maximum capacity) 72.00%; 0.85 coverage 84.60% selective risk 1.18%; 26.1× faster than standard Qwen JSON generation; Jevify 90.0% / 167 ms CUDA graphs disabled; Finding 1: Brier on stated confidence alone is a trap; grpo_rlcr 0.78 / ECE 0.084; reliability 0.007 but resolution 0.000; 27 900 schema-driven decisions; 13 600 / 13 600 questions; candidate mass min 0.99999624; 22 configs · 166,054 rows · 4 calibration-gold; sha a39eba3f; Student B MAE 0.148 / Pearson 0.836 / 86.0%; pngwn/open-jev-laya-bench README 404; sha 9f69c742 likes 2; HDFS 0.9933 (745/750) / retain 0.0084; BGL ERROR/FATAL protection 1.0000; 2,479 / 2,500 HDFS uncertain; cache hit 0.9648 (2412/2500); $0.153936 estimated; E2 recomputes from saved probabilities; Space sha eda59e0a; MASSIVE English 0.783 / Khmer 0.033 / Hindi 0.133; 40–48 rows too small to ship T; T never changes argmax; siren2345/jev-single-decode-transformers ≠ siren2345/jev-single-decode; Split Transformers experiment from llama.cpp runtime; tanayvasishtha/jev-lab ≠ dairui1/jev-lab ≠ BrendanH18/jev-lab ≠ yibie/laya-jev-lab; Four experiments stress-testing TypeSafe's Jev: calibration, bundle bias, label bias, and ensembling; second pass must be $0.00 from cache; The pages never call Jev; Gemma 4 31B 77.0% / Jev 1.13.0 61.4% / Laya 322M 0.0%; restriction state 95.0% against 84.4%; None of the systems are particularly good at knowing when to stop and ask; They skip the question and call a tool directly; 100% schema pass; six-field joint 48.8% vs 72.8%; ywchiu/jev_benchmark ≠ Running-Dolphins/jev-bench ≠ Praveenrajus/jev-bench; ACT / REVIEW / FALLBACK; A provider failure, timeout, malformed output, or missing answer is **not** a policy outcome; confidence is descriptive provider output, not a substitute for probability; Quality denominators include only valid scored answers; an exact halfway tie chooses the lower level; aiwithenoch/Jev-Skill ≠ simplosophy/jev-skill ≠ laguagu/jev-skills; The local path does not claim to turn a smaller checkpoint into Jev; Low support becomes decision: "review"; MIT-0 SPDX NOASSERTION; current-llm; 结构兼容,不是 Jev 模型能力; altryne/jevify ≠ Mintzs/jevify ≠ gulagala001/jevify ≠ uspraveen/Jevify; Find where Jev belongs. Design the questions. Measure the difference; TypeAR-AI/TypeAR 301 → TypeLLM/TypeLLM; TypeLLM/TypeLLM 16★; SemIf 2241★ (+34 vs §111 2207); jevlike 1051★ (+8 vs 1043); AnotiaWang 98★ (+1 vs 97); yibie/awesome-jev 525★ (+19 vs 506); Laya likes 864 (was 822); tracker likes 67 (+3 vs 64); lastModified UNCHANGED `2026-09-20T04:29:16.000Z`; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#32/#33; notes.md §113 Hourly 0743 uniqueness lock: Heman10x-NGU/Verdict-open-jev ≠ Heman10x-NGU/openJev-verdict-2.0; TF-IDF + LogReg ECE 0.0207 vs Jev 0.1440; Verdict-open-jev 48.07% vs Jev 90.80%; abstention combined recall 10.00%; p50 35.58 ms; K=25 (maximum capacity) 72.00%; 0.85 coverage 84.60% selective risk 1.18%; 26.1× faster than standard Qwen JSON generation; Jevify 90.0% / 167 ms CUDA graphs disabled; Finding 1: Brier on stated confidence alone is a trap; grpo_rlcr 0.78 / ECE 0.084; reliability 0.007 but resolution 0.000; 27 900 schema-driven decisions; 13 600 / 13 600 questions; candidate mass min 0.99999624; 22 configs · 166,054 rows · 4 calibration-gold; sha a39eba3f; Student B MAE 0.148 / Pearson 0.836 / 86.0%; pngwn/open-jev-laya-bench README 404; sha 9f69c742 likes 2; HDFS 0.9933 (745/750) / retain 0.0084; BGL ERROR/FATAL protection 1.0000; 2,479 / 2,500 HDFS uncertain; cache hit 0.9648 (2412/2500); $0.153936 estimated; E2 recomputes from saved probabilities; Space sha eda59e0a; MASSIVE English 0.783 / Khmer 0.033 / Hindi 0.133; 40–48 rows too small to ship T; T never changes argmax; siren2345/jev-single-decode-transformers ≠ siren2345/jev-single-decode; Split Transformers experiment from llama.cpp runtime; tanayvasishtha/jev-lab ≠ dairui1/jev-lab ≠ BrendanH18/jev-lab ≠ yibie/laya-jev-lab; Four experiments stress-testing TypeSafe's Jev: calibration, bundle bias, label bias, and ensembling; second pass must be $0.00 from cache; The pages never call Jev; Gemma 4 31B 77.0% / Jev 1.13.0 61.4% / Laya 322M 0.0%; restriction state 95.0% against 84.4%; None of the systems are particularly good at knowing when to stop and ask; They skip the question and call a tool directly; 100% schema pass; six-field joint 48.8% vs 72.8%; ywchiu/jev_benchmark ≠ Running-Dolphins/jev-bench ≠ Praveenrajus/jev-bench; ACT / REVIEW / FALLBACK; A provider failure, timeout, malformed output, or missing answer is **not** a policy outcome; confidence is descriptive provider output, not a substitute for probability; Quality denominators include only valid scored answers; an exact halfway tie chooses the lower level; aiwithenoch/Jev-Skill ≠ simplosophy/jev-skill ≠ laguagu/jev-skills; The local path does not claim to turn a smaller checkpoint into Jev; Low support becomes decision: "review"; MIT-0 SPDX NOASSERTION; current-llm; 结构兼容,不是 Jev 模型能力; altryne/jevify ≠ Mintzs/jevify ≠ gulagala001/jevify ≠ uspraveen/Jevify; Find where Jev belongs. Design the questions. Measure the difference; TypeAR-AI/TypeAR 301 → TypeLLM/TypeLLM; TypeLLM/TypeLLM 16★; SemIf 2241★ (+34 vs §111 2207); jevlike 1051★ (+8 vs 1043); AnotiaWang 98★ (+1 vs 97); yibie/awesome-jev 525★ (+19 vs 506); Laya likes 864 (was 822); tracker likes 67 (+3 vs 64); lastModified UNCHANGED `2026-09-20T04:29:16.000Z`; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#32/#33; notes.md §113 @@ -2465,3 +2467,6 @@ User-provided 0806 uniqueness lock: people who compare Jev against GPT-5.6 has n Hourly 0843 uniqueness lock: A hunch is a probability with a policy attached; { enter: 0.8, exit: 0.6 } is hysteresis; replay a policy change without inference; Decision models are providers, not the product; huncho ≠ Kungie/gut ≠ carldaws/hunch ≠ tpellet/hunch; pretrained Qwen2.5 base ECE 0.030 (0.5B) / 0.040 (7B); instruct 0.302 / 0.269; 70.9% → 70.0% mean conf 74.1% → 96.7%; temperature scaling still matches it in-distribution; No Jev API was called; Qwen2.5 ≠ Archer; Qwen/Qwen3.8-27B ≠ Archer; 学習済みモデル v0.1 は準備中です; bool AUROC 0.523; 先頭だと0件、末尾だと250件; 温度を渡さない場合、確率は較正されていません; このリポジトリには Jev を呼ぶコードが存在しません; g0runmezadam/what-is-jev IS tunahansahin897/what-is-jev (same GitHub id 1378007307); 947 repos scored; A 273 · B 302 · C 372; LLM rubric ≠ benches; Data as of 2026-09-20; HEAD 895b9498; README SHA 3ae98c56; 13 focused checks and one mutually exclusive outcome; Probabilities are advisory, not calibrated guarantees; omni-/ask-jev ≠ pedroknigge/mcp_jev; pd.cut bins by equal width while jeval bins by quantile; ECE 0.113 and ECE 0.076; jeval drift is not implemented yet; rlaope/jeval ≠ dayhaysoos/jevals; calibration does not compose; ECE has exactly zero statistical power to detect the failure mode that kills trajectories; 25–60× headline withdrawn; P(all-correct): 0.0071 vs 0.0001; TCE / AMS; Qwen 3.8 sparring ≠ Archer; Deferred Crispification; light_cutoff_applied_to_combination 0; recorded run, kinematic animation; BANKING77 Accuracy BERT-Base 93.02 Jev 79.90; Analyse jev calibration (NLL, ECE) backlog; BERT figures are published supervised references, not zero-shot; 档位措辞效应 分数极差中位 0.50、最大 1.32; 修好后对照组是 0.01; 不是 benchmark; 概率没做 calibration; ~1,430 API calls, about $0.15; xiaohuaxi/jev-study ≠ wjdjdakf17/jev-study ≠ baekenough/jev-study; AND: product (independence assumed and recorded in the trace); chat model's stated confidence is not calibrated; circuit-vl-4b ≠ Archer; Bring your own API key; vamsikrishna2421/jev-usecases ≠ whyashthakker/awesome-jev-use-cases; catalog ≠ endorsement; SemIf 2237★ (+30 vs §111 2207); jevlike 1049★ (+6 vs 1043); TypeLLM/TypeLLM 16★; AnotiaWang 98★; yibie/awesome-jev 520★ (+14 vs 506); Laya likes 861 (was 822); tracker likes 66 (+2 vs 64) lastModified UNCHANGED; Blackwood likes 2 gated manual; Archer still promised_not_landed; Hub archerhume/4rcherhume HTTP 401; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33; notes.md §114 **0915 HIGH (`notes.md` §115):** TianyuCodings/NanoJev unified-games-v1 densify; A 0.6B parallel decision model: states and questions in, complete probability distributions out. Zero output-token decoding.; One model, four games; ViZDoom Basic 128/128 vs Jev 56/128; Predict Position 27/128 vs Jev 11/128; Maze 225 attempts vs Jev 2738; Snake 30 food / 256 steps; held-out Maze 4/10 Snake 8/8 Basic 128/128 Predict 27/128; Untuned Qwen3-0.6B baseline; not TypeSafe Jev; open replica / specialist gameplay S1; Game success ≠ calibrated Noul; local type boolean ≠ TypeSafe noul; caijinchun/nanojev-arena ≠ liao96312/jev-arena-nanojev ≠ zwliJay/jev-forge ≠ NanoJev; do not reopen or amend PR #31/#32/#33/#35; notes.md §115 User-provided 0915 uniqueness lock: TianyuCodings/NanoJev unified-games-v1 densify; A 0.6B parallel decision model: states and questions in, complete probability distributions out. Zero output-token decoding.; One model, four games; ViZDoom Basic 128/128 vs Jev 56/128; Predict Position 27/128 vs Jev 11/128; Maze 225 attempts vs Jev 2738; Snake 30 food / 256 steps; held-out Maze 4/10 Snake 8/8 Basic 128/128 Predict 27/128; Untuned Qwen3-0.6B baseline; 18,760 questions per variant; 16,333 ViZDoom; 896 Predict Position expert episodes; hard_lr1e5; mix weights 1/3, 1/3, 1/6, 1/6; Hub C-Tianyu/NanoJev revision unified-games-v1 likes 58; dataset C-Tianyu/NanoJev-Data likes 5; HEAD 618cea6d906d54e128360786d12f703fff2b1245; 1289★ / 158 forks / size 64035; README SHA 4190093c64ee75b26e9726daa3b00cbcf6d3157a; MIT; caijinchun/nanojev-arena ≠ liao96312/jev-arena-nanojev ≠ zwliJay/jev-forge ≠ NanoJev; not TypeSafe Jev; open replica / specialist gameplay S1; soft scores ≠ hard gates; Game success ≠ calibrated Noul; local type boolean ≠ TypeSafe noul; invented_signal false; do not reopen or amend PR #31/#32/#33/#35; notes.md §115 + + +User-provided 0920 jcr uniqueness lock: NiazMorshed2007/jcr MIT; site https://jcr.niazmorshed.dev; topics ai-agents,jev,mcp; **4★**; HEAD `138b3832`; README SHA `2a49dbc1`; LICENSE SHA `46231303`; size **14850**; Jev Capability Resolver; one tool to find documented deterministic commands in a nested capability tree; returns context; **does not execute**; skills = workflow+judgment; capabilities = individual operations; format independent of Jev; proposed open standard exploration; classify (Jev) → optional OpenAI decompose compound → beam search geometric mean of routing probs; keep up to 3 paths ≥60% of best (JCR_BAND_RATIO 0.6); ambiguity / no-match / depth-limit explicit; soft scores ≠ hard gates; 0.6 band is application policy; routing ≠ permission; docs ≠ authority to run; sol-vs-opus5-20 *theirs*: 20 scenarios × 4 variants = 80 runs; lookup+explain only, no execution; Claude Opus 5: agent input 108,585→15,819 (−85%), cost $0.3700→$0.1222 (−67%), wall 105.5s→77.7s; Codex GPT-5.6-Sol: 61,952→47,669 (−23%), $0.1377→$0.1151 (−16%), wall 25.3s→62.4s (Sol slower with JCR in 19/20); One Sol outlier 372.6s / 193 Jev calls; n=1 per cell; Not Harbor task-execution; Claude/Codex harnesses; compare mode; 50 scenarios bundled; 11 groups, 960 nodes, 11,360 items; 16 routing rounds per step; NiazMorshed2007/jcr ≠ skill-broker ≠ skillranker ≠ jev-sift ≠ jev-lens ≠ jevusher ≠ jev_select_capability; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33/#34; notes.md §116 diff --git a/.agents/skills/augustus/references/mental-models.md b/.agents/skills/augustus/references/mental-models.md index 723619a4..f4ec7163 100644 --- a/.agents/skills/augustus/references/mental-models.md +++ b/.agents/skills/augustus/references/mental-models.md @@ -2484,6 +2484,58 @@ do not reopen or amend PR #23/#24/#25/#26/#27. + +## Apply 0920 jcr (`notes.md` §116) + +Same pillars, one productized HIGH (capability-tree +lookup as retrieve-wide → decide → evidence-set; +skills vs capability catalogs; VOI of context +admission; measurement honesty on mixed wall-time) +across AI, SWE, business, knowledge work — not +SWE-only. Do **not** re-fold 0743 / §113 / 0806 / +§112 / 0646 / §111 or cookbook beam §2 as if they +were this product. Fresh PR; never reopen merged +#7–**#33**. do not reopen or amend PR #23 or #24 +or #25 or #26 or #27 or #28 or #29 or #30 or #31 +or #32 or #33 or #34 or #35 or #36. Do **not** push onto open +#37 (SemIf). Skip Archer rewrite (still +**promised_not_landed**). Soft Noul ≠ hard safety. +Formal methods **compose** with scoring; a Noul +is a SENSOR. Jev is the hot exemplar, not the +whole mandate. Treating 0.6 as a proof, −85% as +Harbor, or returned docs as permission to run is +the anti-pattern. + +1. **NiazMorshed2007/jcr** (PRIMARY) — + NiazMorshed2007/jcr ≠ skill-broker ≠ skillranker ≠ jev-sift ≠ jev-lens ≠ jevusher ≠ jev_select_capability. + Jev Capability Resolver. one tool to find documented deterministic commands in a nested capability tree. + returns context. **does not execute**. + skills = workflow+judgment. capabilities = individual operations. + format independent of Jev. proposed open standard exploration. + classify (Jev) → optional OpenAI decompose compound → beam search geometric mean of routing probs. + keep up to 3 paths ≥60% of best (JCR_BAND_RATIO 0.6). + ambiguity / no-match / depth-limit explicit. + soft scores ≠ hard gates. 0.6 band is application policy. + routing ≠ permission. docs ≠ authority to run. + sol-vs-opus5-20 *theirs*: 20 scenarios × 4 variants = 80 runs. + lookup+explain only, no execution. n=1 per cell. Not Harbor task-execution. + Claude Opus 5: agent input 108,585→15,819 (−85%), cost $0.3700→$0.1222 (−67%), wall 105.5s→77.7s. + Codex GPT-5.6-Sol: 61,952→47,669 (−23%), $0.1377→$0.1151 (−16%), wall 25.3s→62.4s (Sol slower with JCR in 19/20). + One Sol outlier 372.6s / 193 Jev calls. + 11 groups, 960 nodes, 11,360 items. 16 routing rounds per step. + **4★**. HEAD `138b3832`. README SHA `2a49dbc1`. site https://jcr.niazmorshed.dev. + +Soft Noul ≠ hard safety: 0.6 / −85% / 372.6s are +**sensors**. Treating 0.6 as a fail-closed grant, +−85% as Harbor, or docs as authority to run is +the same theater as jev-gate §79. + +Formal methods **compose** with scoring. A Noul +is a SENSOR. Catalog audit / beam width / band +ratio / depth budget are exact work. Ranking ≠ +calibration theater. + + ## Apply 0743 (`notes.md` §113) Same pillars, sixteen HIGH clusters / three themes @@ -2825,3 +2877,5 @@ Related: `mappings.md` §1–§18, `methods-catalog.md`, `toolbox-mapping.md`, `faq.md`. Hourly 0843 uniqueness lock: A hunch is a probability with a policy attached; { enter: 0.8, exit: 0.6 } is hysteresis; replay a policy change without inference; Decision models are providers, not the product; huncho ≠ Kungie/gut ≠ carldaws/hunch ≠ tpellet/hunch; pretrained Qwen2.5 base ECE 0.030 (0.5B) / 0.040 (7B); instruct 0.302 / 0.269; 70.9% → 70.0% mean conf 74.1% → 96.7%; temperature scaling still matches it in-distribution; No Jev API was called; Qwen2.5 ≠ Archer; Qwen/Qwen3.8-27B ≠ Archer; 学習済みモデル v0.1 は準備中です; bool AUROC 0.523; 先頭だと0件、末尾だと250件; 温度を渡さない場合、確率は較正されていません; このリポジトリには Jev を呼ぶコードが存在しません; g0runmezadam/what-is-jev IS tunahansahin897/what-is-jev (same GitHub id 1378007307); 947 repos scored; A 273 · B 302 · C 372; LLM rubric ≠ benches; Data as of 2026-09-20; HEAD 895b9498; README SHA 3ae98c56; 13 focused checks and one mutually exclusive outcome; Probabilities are advisory, not calibrated guarantees; omni-/ask-jev ≠ pedroknigge/mcp_jev; pd.cut bins by equal width while jeval bins by quantile; ECE 0.113 and ECE 0.076; jeval drift is not implemented yet; rlaope/jeval ≠ dayhaysoos/jevals; calibration does not compose; ECE has exactly zero statistical power to detect the failure mode that kills trajectories; 25–60× headline withdrawn; P(all-correct): 0.0071 vs 0.0001; TCE / AMS; Qwen 3.8 sparring ≠ Archer; Deferred Crispification; light_cutoff_applied_to_combination 0; recorded run, kinematic animation; BANKING77 Accuracy BERT-Base 93.02 Jev 79.90; Analyse jev calibration (NLL, ECE) backlog; BERT figures are published supervised references, not zero-shot; 档位措辞效应 分数极差中位 0.50、最大 1.32; 修好后对照组是 0.01; 不是 benchmark; 概率没做 calibration; ~1,430 API calls, about $0.15; xiaohuaxi/jev-study ≠ wjdjdakf17/jev-study ≠ baekenough/jev-study; AND: product (independence assumed and recorded in the trace); chat model's stated confidence is not calibrated; circuit-vl-4b ≠ Archer; Bring your own API key; vamsikrishna2421/jev-usecases ≠ whyashthakker/awesome-jev-use-cases; catalog ≠ endorsement; SemIf 2237★ (+30 vs §111 2207); jevlike 1049★ (+6 vs 1043); TypeLLM/TypeLLM 16★; AnotiaWang 98★; yibie/awesome-jev 520★ (+14 vs 506); Laya likes 861 (was 822); tracker likes 66 (+2 vs 64) lastModified UNCHANGED; Blackwood likes 2 gated manual; Archer still promised_not_landed; Hub archerhume/4rcherhume HTTP 401; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33; notes.md §114 + +User-provided 0920 jcr uniqueness lock: NiazMorshed2007/jcr MIT; site https://jcr.niazmorshed.dev; topics ai-agents,jev,mcp; **4★**; HEAD `138b3832`; README SHA `2a49dbc1`; LICENSE SHA `46231303`; size **14850**; Jev Capability Resolver; one tool to find documented deterministic commands in a nested capability tree; returns context; **does not execute**; skills = workflow+judgment; capabilities = individual operations; format independent of Jev; proposed open standard exploration; classify (Jev) → optional OpenAI decompose compound → beam search geometric mean of routing probs; keep up to 3 paths ≥60% of best (JCR_BAND_RATIO 0.6); ambiguity / no-match / depth-limit explicit; soft scores ≠ hard gates; 0.6 band is application policy; routing ≠ permission; docs ≠ authority to run; sol-vs-opus5-20 *theirs*: 20 scenarios × 4 variants = 80 runs; lookup+explain only, no execution; Claude Opus 5: agent input 108,585→15,819 (−85%), cost $0.3700→$0.1222 (−67%), wall 105.5s→77.7s; Codex GPT-5.6-Sol: 61,952→47,669 (−23%), $0.1377→$0.1151 (−16%), wall 25.3s→62.4s (Sol slower with JCR in 19/20); One Sol outlier 372.6s / 193 Jev calls; n=1 per cell; Not Harbor task-execution; Claude/Codex harnesses; compare mode; 50 scenarios bundled; 11 groups, 960 nodes, 11,360 items; 16 routing rounds per step; NiazMorshed2007/jcr ≠ skill-broker ≠ skillranker ≠ jev-sift ≠ jev-lens ≠ jevusher ≠ jev_select_capability; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33/#34; notes.md §116 diff --git a/.agents/skills/augustus/references/methods-catalog.md b/.agents/skills/augustus/references/methods-catalog.md index 0e30b508..bf792235 100644 --- a/.agents/skills/augustus/references/methods-catalog.md +++ b/.agents/skills/augustus/references/methods-catalog.md @@ -268,3 +268,7 @@ calibration), you don't yet have a mapping — you have a metaphor. to shortlists (heuristics→rerank shape). **Hypothesis.** Hourly 0843 uniqueness lock: A hunch is a probability with a policy attached; { enter: 0.8, exit: 0.6 } is hysteresis; replay a policy change without inference; Decision models are providers, not the product; huncho ≠ Kungie/gut ≠ carldaws/hunch ≠ tpellet/hunch; pretrained Qwen2.5 base ECE 0.030 (0.5B) / 0.040 (7B); instruct 0.302 / 0.269; 70.9% → 70.0% mean conf 74.1% → 96.7%; temperature scaling still matches it in-distribution; No Jev API was called; Qwen2.5 ≠ Archer; Qwen/Qwen3.8-27B ≠ Archer; 学習済みモデル v0.1 は準備中です; bool AUROC 0.523; 先頭だと0件、末尾だと250件; 温度を渡さない場合、確率は較正されていません; このリポジトリには Jev を呼ぶコードが存在しません; g0runmezadam/what-is-jev IS tunahansahin897/what-is-jev (same GitHub id 1378007307); 947 repos scored; A 273 · B 302 · C 372; LLM rubric ≠ benches; Data as of 2026-09-20; HEAD 895b9498; README SHA 3ae98c56; 13 focused checks and one mutually exclusive outcome; Probabilities are advisory, not calibrated guarantees; omni-/ask-jev ≠ pedroknigge/mcp_jev; pd.cut bins by equal width while jeval bins by quantile; ECE 0.113 and ECE 0.076; jeval drift is not implemented yet; rlaope/jeval ≠ dayhaysoos/jevals; calibration does not compose; ECE has exactly zero statistical power to detect the failure mode that kills trajectories; 25–60× headline withdrawn; P(all-correct): 0.0071 vs 0.0001; TCE / AMS; Qwen 3.8 sparring ≠ Archer; Deferred Crispification; light_cutoff_applied_to_combination 0; recorded run, kinematic animation; BANKING77 Accuracy BERT-Base 93.02 Jev 79.90; Analyse jev calibration (NLL, ECE) backlog; BERT figures are published supervised references, not zero-shot; 档位措辞效应 分数极差中位 0.50、最大 1.32; 修好后对照组是 0.01; 不是 benchmark; 概率没做 calibration; ~1,430 API calls, about $0.15; xiaohuaxi/jev-study ≠ wjdjdakf17/jev-study ≠ baekenough/jev-study; AND: product (independence assumed and recorded in the trace); chat model's stated confidence is not calibrated; circuit-vl-4b ≠ Archer; Bring your own API key; vamsikrishna2421/jev-usecases ≠ whyashthakker/awesome-jev-use-cases; catalog ≠ endorsement; SemIf 2237★ (+30 vs §111 2207); jevlike 1049★ (+6 vs 1043); TypeLLM/TypeLLM 16★; AnotiaWang 98★; yibie/awesome-jev 520★ (+14 vs 506); Laya likes 861 (was 822); tracker likes 66 (+2 vs 64) lastModified UNCHANGED; Blackwood likes 2 gated manual; Archer still promised_not_landed; Hub archerhume/4rcherhume HTTP 401; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33; notes.md §114 + +**0920 jcr:** geometric-mean path score is a control statistic (cookbook beam cousin), not a calibrated Noul and not a permission bit. `notes.md` §116. + +User-provided 0920 jcr uniqueness lock: NiazMorshed2007/jcr MIT; site https://jcr.niazmorshed.dev; topics ai-agents,jev,mcp; **4★**; HEAD `138b3832`; README SHA `2a49dbc1`; LICENSE SHA `46231303`; size **14850**; Jev Capability Resolver; one tool to find documented deterministic commands in a nested capability tree; returns context; **does not execute**; skills = workflow+judgment; capabilities = individual operations; format independent of Jev; proposed open standard exploration; classify (Jev) → optional OpenAI decompose compound → beam search geometric mean of routing probs; keep up to 3 paths ≥60% of best (JCR_BAND_RATIO 0.6); ambiguity / no-match / depth-limit explicit; soft scores ≠ hard gates; 0.6 band is application policy; routing ≠ permission; docs ≠ authority to run; sol-vs-opus5-20 *theirs*: 20 scenarios × 4 variants = 80 runs; lookup+explain only, no execution; Claude Opus 5: agent input 108,585→15,819 (−85%), cost $0.3700→$0.1222 (−67%), wall 105.5s→77.7s; Codex GPT-5.6-Sol: 61,952→47,669 (−23%), $0.1377→$0.1151 (−16%), wall 25.3s→62.4s (Sol slower with JCR in 19/20); One Sol outlier 372.6s / 193 Jev calls; n=1 per cell; Not Harbor task-execution; Claude/Codex harnesses; compare mode; 50 scenarios bundled; 11 groups, 960 nodes, 11,360 items; 16 routing rounds per step; NiazMorshed2007/jcr ≠ skill-broker ≠ skillranker ≠ jev-sift ≠ jev-lens ≠ jevusher ≠ jev_select_capability; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33/#34; notes.md §116 diff --git a/.agents/skills/augustus/references/mixed-architecture.md b/.agents/skills/augustus/references/mixed-architecture.md index 7dd88094..44cb020b 100644 --- a/.agents/skills/augustus/references/mixed-architecture.md +++ b/.agents/skills/augustus/references/mixed-architecture.md @@ -466,6 +466,13 @@ not a global virtue: | Treat Brier-on-stated-conf as RLCD / distill Jev as teacher of record | **Fail closed** (trap; teacher-copy) | Finding 1: Brier on stated confidence alone is a trap. grpo_rlcr 0.78 / ECE 0.084. reliability 0.007 but resolution 0.000. Student B MAE 0.148 / Pearson 0.836 / 86.0%. `notes.md` §113 | | Treat HDFS 0.9933 as Harbor / ship T from 40–48 rows / re-fold pngwn RESULTS | **Fail closed** (re-policy / n too small / densify) | HDFS 0.9933 (745/750) / retain 0.0084. BGL ERROR/FATAL protection 1.0000. 2,479 / 2,500 HDFS uncertain. cache hit 0.9648 (2412/2500). $0.153936 estimated. E2 recomputes from saved probabilities. Space sha eda59e0a. MASSIVE English 0.783 / Khmer 0.033 / Hindi 0.133. 40–48 rows too small to ship T. T never changes argmax. pngwn/open-jev-laya-bench README 404. sha 9f69c742 likes 2. 22 configs · 166,054 rows · 4 calibration-gold. sha a39eba3f. 27 900 schema-driven decisions. 13 600 / 13 600 questions. candidate mass min 0.99999624. `notes.md` §113 | | Treat ACT as a provider proof / collapse Jev-Skill into simplosophy / treat current-llm as a replica | **Fail closed** (policy ≠ provider; structure-compat) | ACT / REVIEW / FALLBACK. A provider failure, timeout, malformed output, or missing answer is **not** a policy outcome. confidence is descriptive provider output, not a substitute for probability. Quality denominators include only valid scored answers. an exact halfway tie chooses the lower level. aiwithenoch/Jev-Skill ≠ simplosophy/jev-skill ≠ laguagu/jev-skills. The local path does not claim to turn a smaller checkpoint into Jev. Low support becomes decision: "review". MIT-0 SPDX NOASSERTION. current-llm. 结构兼容,不是 Jev 模型能力. `notes.md` §113 | + +| Treat JCR as a skill loader / collapse into skill-broker grants | **Fail closed** (lookup ≠ grant) | NiazMorshed2007/jcr ≠ skill-broker ≠ skillranker ≠ jev-sift ≠ jev-lens ≠ jevusher ≠ jev_select_capability. skills = workflow+judgment. capabilities = individual operations. format independent of Jev. `notes.md` §116 | +| Treat 0.6 band as a proof / hard-gate geometric-mean routing | **Fail closed** (application policy; soft scores ≠ hard gates) | keep up to 3 paths ≥60% of best (JCR_BAND_RATIO 0.6). 0.6 band is application policy. A match can still be wrong. `notes.md` §116 | +| Execute returned docs / treat routing as permission | **Fail closed** (docs ≠ authority to run) | returns context. **does not execute**. routing ≠ permission. docs ≠ authority to run. JCR returns documentation. It does not execute commands. `notes.md` §116 | +| Quote −85% / 372.6s as Harbor task-execution | **Fail closed** (lookup+explain; n=1; wall-time mixed) | sol-vs-opus5-20 *theirs*: 20 scenarios × 4 variants = 80 runs. lookup+explain only, no execution. n=1 per cell. Not Harbor task-execution. Claude Opus 5: agent input 108,585→15,819 (−85%), cost $0.3700→$0.1222 (−67%), wall 105.5s→77.7s. Codex GPT-5.6-Sol: 61,952→47,669 (−23%), $0.1377→$0.1151 (−16%), wall 25.3s→62.4s (Sol slower with JCR in 19/20). One Sol outlier 372.6s / 193 Jev calls. `notes.md` §116 | +| Copy npm / .env / keys / reopen or amend PR #23–#36 / push onto #37 | **Fail closed** (docs/skill only; ID collision reserved) | do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33/#34. notes.md §116. Merged #35 owns §114 / 289–302 / #97. Merged #36 owns §115 / 303–308 / #98. Open #37 §117. Fresh PR off main. `notes.md` §116 | + | Treat TypeAR-AI/TypeAR as the live name / tracker likes 66 as Archer landing / reopen or amend PR #23–#30 | **Fail closed** (301 rename; likes ephemeral; prior PRs MERGED) | TypeAR-AI/TypeAR 301 → TypeLLM/TypeLLM. TypeLLM/TypeLLM 16★. SemIf 2241★ (+34 vs §111 2207). jevlike 1051★ (+8 vs 1043). AnotiaWang 98★ (+1 vs 97). yibie/awesome-jev 525★ (+19 vs 506). Laya likes 864 (was 822). tracker likes 67 (+3 vs 64). lastModified UNCHANGED `2026-09-20T04:29:16.000Z`. Archer still promised_not_landed. do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#32/#33. Fresh PR off main. `notes.md` §113 | @@ -1426,3 +1433,5 @@ Propose three *placements* (prefilter vs post-verify vs replace-this-one- classifier-step), recommend one, and name the experiment that kills it. Hourly 0843 uniqueness lock: A hunch is a probability with a policy attached; { enter: 0.8, exit: 0.6 } is hysteresis; replay a policy change without inference; Decision models are providers, not the product; huncho ≠ Kungie/gut ≠ carldaws/hunch ≠ tpellet/hunch; pretrained Qwen2.5 base ECE 0.030 (0.5B) / 0.040 (7B); instruct 0.302 / 0.269; 70.9% → 70.0% mean conf 74.1% → 96.7%; temperature scaling still matches it in-distribution; No Jev API was called; Qwen2.5 ≠ Archer; Qwen/Qwen3.8-27B ≠ Archer; 学習済みモデル v0.1 は準備中です; bool AUROC 0.523; 先頭だと0件、末尾だと250件; 温度を渡さない場合、確率は較正されていません; このリポジトリには Jev を呼ぶコードが存在しません; g0runmezadam/what-is-jev IS tunahansahin897/what-is-jev (same GitHub id 1378007307); 947 repos scored; A 273 · B 302 · C 372; LLM rubric ≠ benches; Data as of 2026-09-20; HEAD 895b9498; README SHA 3ae98c56; 13 focused checks and one mutually exclusive outcome; Probabilities are advisory, not calibrated guarantees; omni-/ask-jev ≠ pedroknigge/mcp_jev; pd.cut bins by equal width while jeval bins by quantile; ECE 0.113 and ECE 0.076; jeval drift is not implemented yet; rlaope/jeval ≠ dayhaysoos/jevals; calibration does not compose; ECE has exactly zero statistical power to detect the failure mode that kills trajectories; 25–60× headline withdrawn; P(all-correct): 0.0071 vs 0.0001; TCE / AMS; Qwen 3.8 sparring ≠ Archer; Deferred Crispification; light_cutoff_applied_to_combination 0; recorded run, kinematic animation; BANKING77 Accuracy BERT-Base 93.02 Jev 79.90; Analyse jev calibration (NLL, ECE) backlog; BERT figures are published supervised references, not zero-shot; 档位措辞效应 分数极差中位 0.50、最大 1.32; 修好后对照组是 0.01; 不是 benchmark; 概率没做 calibration; ~1,430 API calls, about $0.15; xiaohuaxi/jev-study ≠ wjdjdakf17/jev-study ≠ baekenough/jev-study; AND: product (independence assumed and recorded in the trace); chat model's stated confidence is not calibrated; circuit-vl-4b ≠ Archer; Bring your own API key; vamsikrishna2421/jev-usecases ≠ whyashthakker/awesome-jev-use-cases; catalog ≠ endorsement; SemIf 2237★ (+30 vs §111 2207); jevlike 1049★ (+6 vs 1043); TypeLLM/TypeLLM 16★; AnotiaWang 98★; yibie/awesome-jev 520★ (+14 vs 506); Laya likes 861 (was 822); tracker likes 66 (+2 vs 64) lastModified UNCHANGED; Blackwood likes 2 gated manual; Archer still promised_not_landed; Hub archerhume/4rcherhume HTTP 401; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33; notes.md §114 + +User-provided 0920 jcr uniqueness lock: NiazMorshed2007/jcr MIT; site https://jcr.niazmorshed.dev; topics ai-agents,jev,mcp; **4★**; HEAD `138b3832`; README SHA `2a49dbc1`; LICENSE SHA `46231303`; size **14850**; Jev Capability Resolver; one tool to find documented deterministic commands in a nested capability tree; returns context; **does not execute**; skills = workflow+judgment; capabilities = individual operations; format independent of Jev; proposed open standard exploration; classify (Jev) → optional OpenAI decompose compound → beam search geometric mean of routing probs; keep up to 3 paths ≥60% of best (JCR_BAND_RATIO 0.6); ambiguity / no-match / depth-limit explicit; soft scores ≠ hard gates; 0.6 band is application policy; routing ≠ permission; docs ≠ authority to run; sol-vs-opus5-20 *theirs*: 20 scenarios × 4 variants = 80 runs; lookup+explain only, no execution; Claude Opus 5: agent input 108,585→15,819 (−85%), cost $0.3700→$0.1222 (−67%), wall 105.5s→77.7s; Codex GPT-5.6-Sol: 61,952→47,669 (−23%), $0.1377→$0.1151 (−16%), wall 25.3s→62.4s (Sol slower with JCR in 19/20); One Sol outlier 372.6s / 193 Jev calls; n=1 per cell; Not Harbor task-execution; Claude/Codex harnesses; compare mode; 50 scenarios bundled; 11 groups, 960 nodes, 11,360 items; 16 routing rounds per step; NiazMorshed2007/jcr ≠ skill-broker ≠ skillranker ≠ jev-sift ≠ jev-lens ≠ jevusher ≠ jev_select_capability; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33/#34; notes.md §116 diff --git a/.agents/skills/augustus/references/question-design.md b/.agents/skills/augustus/references/question-design.md index d9fb197c..c9e734b7 100644 --- a/.agents/skills/augustus/references/question-design.md +++ b/.agents/skills/augustus/references/question-design.md @@ -396,6 +396,8 @@ Hourly 0646 uniqueness lock: Calibration is not alpha; NO CURRENT ALPHA CANDIDAT Hourly 0843 question-design: 档位措辞效应 分数极差中位 0.50、最大 1.32 — option-label wording moves Score more than language. Position bias 先頭だと0件、末尾だと250件. Screening cutoff 0.49 / light_cutoff_applied_to_combination 0. Independence for AND-product must be recorded. `notes.md` §114. User-provided 0806 uniqueness lock: people who compare Jev against GPT-5.6 has never fine-tuned BERTForXYZ for living and it shows; zero shot classifiers; scale them as much as decoder only models; many problems solved with LLMs could have been solved with them, it was a skill issue; opt for DeBERTa and ModernBERT ones; BERTForXYZ → DeBERTa → ModernBERT; Jev vs GPT-5.6 bakeoffs are a category error; encoder / ZS classifiers; institutional HF voice; quote *theirs*; do not invent accuracy numbers; softmax/ZS scores still ≠ calibrated Noul; soft scores ≠ hard gates; @mervenoyann; likes 421 / 189; impressions 35498 / 9613; multimodal image<>text ZS as perception front-end; hf:MoritzLaurer/deberta-v3-large-zeroshot-v2.0 likes 139; hf:MoritzLaurer/ModernBERT-large-zeroshot-v2.0 likes 72; Bart, bert, deberta, modernbert, these are all LLMs; Maziyar quoted; Jev is exemplar not the mandate; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29. +JCR routing questions (`notes.md` §116): each round is a Choice over **direct children** plus no-match. Compound vs single is a prior Choice. Do not stuff the command `context` into routing questions. Band 0.6 is policy in code after the distribution. + Hourly 0541 uniqueness lock: calibration beyond ~500 tokens unmeasured; 11.57s vs 54.10s · 4.67× · 120/128 *theirs*; GH jev-haiku-benchmarking 404; TCP floor 198.8 ms; gateway tax not one number; ≠ RadRebelSam/awesome-jev; NLI Tetris argmax P(entail)−P(contradict); 1q 396ms / 30q 567ms; ±0.03; 33q $0.000045 vs Gemini ~5× slower ~60× cost *theirs*; catalog ≠ endorsement; Client-side quiz; pointer from held docs; scanned-PDF warn; CSP only api.typesafe.ai; $0.00022 vs chat $0.00306 *theirs*; SemIf 2186★ (+20 vs §109 2166); jevlike 1038★ (+7 vs 1031); TypeAR 14★ flat; AnotiaWang 96★ (+1 vs 95); yibie/awesome-jev 490★; Laya likes 802 (was 783); tracker likes 64 flat, lastModified UNCHANGED; do not reopen or amend PR #23/#24/#25/#26/#27. | Each answer right, decision wrong | Policy wrong | Change weights/thresholds in code, leave questions alone | @@ -410,3 +412,5 @@ Hourly 0541 uniqueness lock: calibration beyond ~500 tokens unmeasured; 11.57s v Confidence-gated routing (doc defaults): act / confirm-or-flag / hand off, floors 0.5–0.6, 0.85–0.9 for high-stakes — starting points, not constants (calibration must be re-measured per dataset; see validation.md). Hourly 0843 uniqueness lock: A hunch is a probability with a policy attached; { enter: 0.8, exit: 0.6 } is hysteresis; replay a policy change without inference; Decision models are providers, not the product; huncho ≠ Kungie/gut ≠ carldaws/hunch ≠ tpellet/hunch; pretrained Qwen2.5 base ECE 0.030 (0.5B) / 0.040 (7B); instruct 0.302 / 0.269; 70.9% → 70.0% mean conf 74.1% → 96.7%; temperature scaling still matches it in-distribution; No Jev API was called; Qwen2.5 ≠ Archer; Qwen/Qwen3.8-27B ≠ Archer; 学習済みモデル v0.1 は準備中です; bool AUROC 0.523; 先頭だと0件、末尾だと250件; 温度を渡さない場合、確率は較正されていません; このリポジトリには Jev を呼ぶコードが存在しません; g0runmezadam/what-is-jev IS tunahansahin897/what-is-jev (same GitHub id 1378007307); 947 repos scored; A 273 · B 302 · C 372; LLM rubric ≠ benches; Data as of 2026-09-20; HEAD 895b9498; README SHA 3ae98c56; 13 focused checks and one mutually exclusive outcome; Probabilities are advisory, not calibrated guarantees; omni-/ask-jev ≠ pedroknigge/mcp_jev; pd.cut bins by equal width while jeval bins by quantile; ECE 0.113 and ECE 0.076; jeval drift is not implemented yet; rlaope/jeval ≠ dayhaysoos/jevals; calibration does not compose; ECE has exactly zero statistical power to detect the failure mode that kills trajectories; 25–60× headline withdrawn; P(all-correct): 0.0071 vs 0.0001; TCE / AMS; Qwen 3.8 sparring ≠ Archer; Deferred Crispification; light_cutoff_applied_to_combination 0; recorded run, kinematic animation; BANKING77 Accuracy BERT-Base 93.02 Jev 79.90; Analyse jev calibration (NLL, ECE) backlog; BERT figures are published supervised references, not zero-shot; 档位措辞效应 分数极差中位 0.50、最大 1.32; 修好后对照组是 0.01; 不是 benchmark; 概率没做 calibration; ~1,430 API calls, about $0.15; xiaohuaxi/jev-study ≠ wjdjdakf17/jev-study ≠ baekenough/jev-study; AND: product (independence assumed and recorded in the trace); chat model's stated confidence is not calibrated; circuit-vl-4b ≠ Archer; Bring your own API key; vamsikrishna2421/jev-usecases ≠ whyashthakker/awesome-jev-use-cases; catalog ≠ endorsement; SemIf 2237★ (+30 vs §111 2207); jevlike 1049★ (+6 vs 1043); TypeLLM/TypeLLM 16★; AnotiaWang 98★; yibie/awesome-jev 520★ (+14 vs 506); Laya likes 861 (was 822); tracker likes 66 (+2 vs 64) lastModified UNCHANGED; Blackwood likes 2 gated manual; Archer still promised_not_landed; Hub archerhume/4rcherhume HTTP 401; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33; notes.md §114 + +User-provided 0920 jcr uniqueness lock: NiazMorshed2007/jcr MIT; site https://jcr.niazmorshed.dev; topics ai-agents,jev,mcp; **4★**; HEAD `138b3832`; README SHA `2a49dbc1`; LICENSE SHA `46231303`; size **14850**; Jev Capability Resolver; one tool to find documented deterministic commands in a nested capability tree; returns context; **does not execute**; skills = workflow+judgment; capabilities = individual operations; format independent of Jev; proposed open standard exploration; classify (Jev) → optional OpenAI decompose compound → beam search geometric mean of routing probs; keep up to 3 paths ≥60% of best (JCR_BAND_RATIO 0.6); ambiguity / no-match / depth-limit explicit; soft scores ≠ hard gates; 0.6 band is application policy; routing ≠ permission; docs ≠ authority to run; sol-vs-opus5-20 *theirs*: 20 scenarios × 4 variants = 80 runs; lookup+explain only, no execution; Claude Opus 5: agent input 108,585→15,819 (−85%), cost $0.3700→$0.1222 (−67%), wall 105.5s→77.7s; Codex GPT-5.6-Sol: 61,952→47,669 (−23%), $0.1377→$0.1151 (−16%), wall 25.3s→62.4s (Sol slower with JCR in 19/20); One Sol outlier 372.6s / 193 Jev calls; n=1 per cell; Not Harbor task-execution; Claude/Codex harnesses; compare mode; 50 scenarios bundled; 11 groups, 960 nodes, 11,360 items; 16 routing rounds per step; NiazMorshed2007/jcr ≠ skill-broker ≠ skillranker ≠ jev-sift ≠ jev-lens ≠ jevusher ≠ jev_select_capability; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33/#34; notes.md §116 diff --git a/.agents/skills/augustus/references/toolbox-mapping.md b/.agents/skills/augustus/references/toolbox-mapping.md index 77c9574b..223f42ec 100644 --- a/.agents/skills/augustus/references/toolbox-mapping.md +++ b/.agents/skills/augustus/references/toolbox-mapping.md @@ -343,3 +343,8 @@ with an acceptance test that ran. Hourly 0843 uniqueness lock: A hunch is a probability with a policy attached; { enter: 0.8, exit: 0.6 } is hysteresis; replay a policy change without inference; Decision models are providers, not the product; huncho ≠ Kungie/gut ≠ carldaws/hunch ≠ tpellet/hunch; pretrained Qwen2.5 base ECE 0.030 (0.5B) / 0.040 (7B); instruct 0.302 / 0.269; 70.9% → 70.0% mean conf 74.1% → 96.7%; temperature scaling still matches it in-distribution; No Jev API was called; Qwen2.5 ≠ Archer; Qwen/Qwen3.8-27B ≠ Archer; 学習済みモデル v0.1 は準備中です; bool AUROC 0.523; 先頭だと0件、末尾だと250件; 温度を渡さない場合、確率は較正されていません; このリポジトリには Jev を呼ぶコードが存在しません; g0runmezadam/what-is-jev IS tunahansahin897/what-is-jev (same GitHub id 1378007307); 947 repos scored; A 273 · B 302 · C 372; LLM rubric ≠ benches; Data as of 2026-09-20; HEAD 895b9498; README SHA 3ae98c56; 13 focused checks and one mutually exclusive outcome; Probabilities are advisory, not calibrated guarantees; omni-/ask-jev ≠ pedroknigge/mcp_jev; pd.cut bins by equal width while jeval bins by quantile; ECE 0.113 and ECE 0.076; jeval drift is not implemented yet; rlaope/jeval ≠ dayhaysoos/jevals; calibration does not compose; ECE has exactly zero statistical power to detect the failure mode that kills trajectories; 25–60× headline withdrawn; P(all-correct): 0.0071 vs 0.0001; TCE / AMS; Qwen 3.8 sparring ≠ Archer; Deferred Crispification; light_cutoff_applied_to_combination 0; recorded run, kinematic animation; BANKING77 Accuracy BERT-Base 93.02 Jev 79.90; Analyse jev calibration (NLL, ECE) backlog; BERT figures are published supervised references, not zero-shot; 档位措辞效应 分数极差中位 0.50、最大 1.32; 修好后对照组是 0.01; 不是 benchmark; 概率没做 calibration; ~1,430 API calls, about $0.15; xiaohuaxi/jev-study ≠ wjdjdakf17/jev-study ≠ baekenough/jev-study; AND: product (independence assumed and recorded in the trace); chat model's stated confidence is not calibrated; circuit-vl-4b ≠ Archer; Bring your own API key; vamsikrishna2421/jev-usecases ≠ whyashthakker/awesome-jev-use-cases; catalog ≠ endorsement; SemIf 2237★ (+30 vs §111 2207); jevlike 1049★ (+6 vs 1043); TypeLLM/TypeLLM 16★; AnotiaWang 98★; yibie/awesome-jev 520★ (+14 vs 506); Laya likes 861 (was 822); tracker likes 66 (+2 vs 64) lastModified UNCHANGED; Blackwood likes 2 gated manual; Archer still promised_not_landed; Hub archerhume/4rcherhume HTTP 401; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33; notes.md §114 User-provided 0915 uniqueness lock: TianyuCodings/NanoJev unified-games-v1 densify; A 0.6B parallel decision model: states and questions in, complete probability distributions out. Zero output-token decoding.; One model, four games; ViZDoom Basic 128/128 vs Jev 56/128; Predict Position 27/128 vs Jev 11/128; Maze 225 attempts vs Jev 2738; Snake 30 food / 256 steps; held-out Maze 4/10 Snake 8/8 Basic 128/128 Predict 27/128; Untuned Qwen3-0.6B baseline; 18,760 questions per variant; 16,333 ViZDoom; 896 Predict Position expert episodes; hard_lr1e5; mix weights 1/3, 1/3, 1/6, 1/6; Hub C-Tianyu/NanoJev revision unified-games-v1 likes 58; dataset C-Tianyu/NanoJev-Data likes 5; HEAD 618cea6d906d54e128360786d12f703fff2b1245; 1289★ / 158 forks / size 64035; README SHA 4190093c64ee75b26e9726daa3b00cbcf6d3157a; MIT; caijinchun/nanojev-arena ≠ liao96312/jev-arena-nanojev ≠ zwliJay/jev-forge ≠ NanoJev; not TypeSafe Jev; open replica / specialist gameplay S1; soft scores ≠ hard gates; Game success ≠ calibrated Noul; local type boolean ≠ TypeSafe noul; invented_signal false; do not reopen or amend PR #31/#32/#33/#35; notes.md §115 + + +**0920 jcr sweep:** judgment-shaped hole is *which documented operation fits this request* on a tree the code already holds. Substitute hierarchical Choice + geometric-mean beam; keep execution exact. NiazMorshed2007/jcr. `notes.md` §116. + +User-provided 0920 jcr uniqueness lock: NiazMorshed2007/jcr MIT; site https://jcr.niazmorshed.dev; topics ai-agents,jev,mcp; **4★**; HEAD `138b3832`; README SHA `2a49dbc1`; LICENSE SHA `46231303`; size **14850**; Jev Capability Resolver; one tool to find documented deterministic commands in a nested capability tree; returns context; **does not execute**; skills = workflow+judgment; capabilities = individual operations; format independent of Jev; proposed open standard exploration; classify (Jev) → optional OpenAI decompose compound → beam search geometric mean of routing probs; keep up to 3 paths ≥60% of best (JCR_BAND_RATIO 0.6); ambiguity / no-match / depth-limit explicit; soft scores ≠ hard gates; 0.6 band is application policy; routing ≠ permission; docs ≠ authority to run; sol-vs-opus5-20 *theirs*: 20 scenarios × 4 variants = 80 runs; lookup+explain only, no execution; Claude Opus 5: agent input 108,585→15,819 (−85%), cost $0.3700→$0.1222 (−67%), wall 105.5s→77.7s; Codex GPT-5.6-Sol: 61,952→47,669 (−23%), $0.1377→$0.1151 (−16%), wall 25.3s→62.4s (Sol slower with JCR in 19/20); One Sol outlier 372.6s / 193 Jev calls; n=1 per cell; Not Harbor task-execution; Claude/Codex harnesses; compare mode; 50 scenarios bundled; 11 groups, 960 nodes, 11,360 items; 16 routing rounds per step; NiazMorshed2007/jcr ≠ skill-broker ≠ skillranker ≠ jev-sift ≠ jev-lens ≠ jevusher ≠ jev_select_capability; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33/#34; notes.md §116 diff --git a/.agents/skills/augustus/references/validation.md b/.agents/skills/augustus/references/validation.md index 7a0ff703..6fa65543 100644 --- a/.agents/skills/augustus/references/validation.md +++ b/.agents/skills/augustus/references/validation.md @@ -1087,6 +1087,16 @@ HDFS 0.9933 (745/750) / retain 0.0084. BGL ERROR/FATAL protection 1.0000. E2 recomputes from saved probabilities. `notes.md` §113. + +**JCR lookup+explain (Empirical as sol-vs-opus5-20 *theirs*, not Harbor).** +[`NiazMorshed2007/jcr`](https://github.com/NiazMorshed2007/jcr) +20 scenarios × 4 variants = 80 runs. lookup+explain only, no execution. +n=1 per cell. Not Harbor task-execution. Wall-time mixed: Opus mean +105.5s→77.7s; Sol 25.3s→62.4s (slower with JCR in 19/20). One Sol +outlier 372.6s / 193 Jev calls. Token/cost *theirs* are instruction-lookup +deltas, not task success. Soft scores ≠ hard gates. `notes.md` §116. + + Hourly 0743 uniqueness lock: Heman10x-NGU/Verdict-open-jev ≠ Heman10x-NGU/openJev-verdict-2.0; TF-IDF + LogReg ECE 0.0207 vs Jev 0.1440; Verdict-open-jev 48.07% vs Jev 90.80%; abstention combined recall 10.00%; p50 35.58 ms; K=25 (maximum capacity) 72.00%; 0.85 coverage 84.60% selective risk 1.18%; 26.1× faster than standard Qwen JSON generation; Jevify 90.0% / 167 ms CUDA graphs disabled; Finding 1: Brier on stated confidence alone is a trap; grpo_rlcr 0.78 / ECE 0.084; reliability 0.007 but resolution 0.000; 27 900 schema-driven decisions; 13 600 / 13 600 questions; candidate mass min 0.99999624; 22 configs · 166,054 rows · 4 calibration-gold; sha a39eba3f; Student B MAE 0.148 / Pearson 0.836 / 86.0%; pngwn/open-jev-laya-bench README 404; sha 9f69c742 likes 2; HDFS 0.9933 (745/750) / retain 0.0084; BGL ERROR/FATAL protection 1.0000; 2,479 / 2,500 HDFS uncertain; cache hit 0.9648 (2412/2500); $0.153936 estimated; E2 recomputes from saved probabilities; Space sha eda59e0a; MASSIVE English 0.783 / Khmer 0.033 / Hindi 0.133; 40–48 rows too small to ship T; T never changes argmax; siren2345/jev-single-decode-transformers ≠ siren2345/jev-single-decode; Split Transformers experiment from llama.cpp runtime; tanayvasishtha/jev-lab ≠ dairui1/jev-lab ≠ BrendanH18/jev-lab ≠ yibie/laya-jev-lab; Four experiments stress-testing TypeSafe's Jev: calibration, bundle bias, label bias, and ensembling; second pass must be $0.00 from cache; The pages never call Jev; Gemma 4 31B 77.0% / Jev 1.13.0 61.4% / Laya 322M 0.0%; restriction state 95.0% against 84.4%; None of the systems are particularly good at knowing when to stop and ask; They skip the question and call a tool directly; 100% schema pass; six-field joint 48.8% vs 72.8%; ywchiu/jev_benchmark ≠ Running-Dolphins/jev-bench ≠ Praveenrajus/jev-bench; ACT / REVIEW / FALLBACK; A provider failure, timeout, malformed output, or missing answer is **not** a policy outcome; confidence is descriptive provider output, not a substitute for probability; Quality denominators include only valid scored answers; an exact halfway tie chooses the lower level; aiwithenoch/Jev-Skill ≠ simplosophy/jev-skill ≠ laguagu/jev-skills; The local path does not claim to turn a smaller checkpoint into Jev; Low support becomes decision: "review"; MIT-0 SPDX NOASSERTION; current-llm; 结构兼容,不是 Jev 模型能力; altryne/jevify ≠ Mintzs/jevify ≠ gulagala001/jevify ≠ uspraveen/Jevify; Find where Jev belongs. Design the questions. Measure the difference; TypeAR-AI/TypeAR 301 → TypeLLM/TypeLLM; TypeLLM/TypeLLM 16★; SemIf 2241★ (+34 vs §111 2207); jevlike 1051★ (+8 vs 1043); AnotiaWang 98★ (+1 vs 97); yibie/awesome-jev 525★ (+19 vs 506); Laya likes 864 (was 822); tracker likes 67 (+3 vs 64); lastModified UNCHANGED `2026-09-20T04:29:16.000Z`; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#32/#33; notes.md §113 **Instruct-tuning honesty collapse (Empirical as 15-bin ECE *theirs*, PRIMARY this hour, not a Jev measurement).** @@ -1137,3 +1147,5 @@ Hourly 0646 uniqueness lock: Calibration is not alpha; NO CURRENT ALPHA CANDIDAT User-provided 0806 uniqueness lock: people who compare Jev against GPT-5.6 has never fine-tuned BERTForXYZ for living and it shows; zero shot classifiers; scale them as much as decoder only models; many problems solved with LLMs could have been solved with them, it was a skill issue; opt for DeBERTa and ModernBERT ones; BERTForXYZ → DeBERTa → ModernBERT; Jev vs GPT-5.6 bakeoffs are a category error; encoder / ZS classifiers; institutional HF voice; quote *theirs*; do not invent accuracy numbers; softmax/ZS scores still ≠ calibrated Noul; soft scores ≠ hard gates; @mervenoyann; likes 421 / 189; impressions 35498 / 9613; multimodal image<>text ZS as perception front-end; hf:MoritzLaurer/deberta-v3-large-zeroshot-v2.0 likes 139; hf:MoritzLaurer/ModernBERT-large-zeroshot-v2.0 likes 72; Bart, bert, deberta, modernbert, these are all LLMs; Maziyar quoted; Jev is exemplar not the mandate; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29. Hourly 0843 uniqueness lock: A hunch is a probability with a policy attached; { enter: 0.8, exit: 0.6 } is hysteresis; replay a policy change without inference; Decision models are providers, not the product; huncho ≠ Kungie/gut ≠ carldaws/hunch ≠ tpellet/hunch; pretrained Qwen2.5 base ECE 0.030 (0.5B) / 0.040 (7B); instruct 0.302 / 0.269; 70.9% → 70.0% mean conf 74.1% → 96.7%; temperature scaling still matches it in-distribution; No Jev API was called; Qwen2.5 ≠ Archer; Qwen/Qwen3.8-27B ≠ Archer; 学習済みモデル v0.1 は準備中です; bool AUROC 0.523; 先頭だと0件、末尾だと250件; 温度を渡さない場合、確率は較正されていません; このリポジトリには Jev を呼ぶコードが存在しません; g0runmezadam/what-is-jev IS tunahansahin897/what-is-jev (same GitHub id 1378007307); 947 repos scored; A 273 · B 302 · C 372; LLM rubric ≠ benches; Data as of 2026-09-20; HEAD 895b9498; README SHA 3ae98c56; 13 focused checks and one mutually exclusive outcome; Probabilities are advisory, not calibrated guarantees; omni-/ask-jev ≠ pedroknigge/mcp_jev; pd.cut bins by equal width while jeval bins by quantile; ECE 0.113 and ECE 0.076; jeval drift is not implemented yet; rlaope/jeval ≠ dayhaysoos/jevals; calibration does not compose; ECE has exactly zero statistical power to detect the failure mode that kills trajectories; 25–60× headline withdrawn; P(all-correct): 0.0071 vs 0.0001; TCE / AMS; Qwen 3.8 sparring ≠ Archer; Deferred Crispification; light_cutoff_applied_to_combination 0; recorded run, kinematic animation; BANKING77 Accuracy BERT-Base 93.02 Jev 79.90; Analyse jev calibration (NLL, ECE) backlog; BERT figures are published supervised references, not zero-shot; 档位措辞效应 分数极差中位 0.50、最大 1.32; 修好后对照组是 0.01; 不是 benchmark; 概率没做 calibration; ~1,430 API calls, about $0.15; xiaohuaxi/jev-study ≠ wjdjdakf17/jev-study ≠ baekenough/jev-study; AND: product (independence assumed and recorded in the trace); chat model's stated confidence is not calibrated; circuit-vl-4b ≠ Archer; Bring your own API key; vamsikrishna2421/jev-usecases ≠ whyashthakker/awesome-jev-use-cases; catalog ≠ endorsement; SemIf 2237★ (+30 vs §111 2207); jevlike 1049★ (+6 vs 1043); TypeLLM/TypeLLM 16★; AnotiaWang 98★; yibie/awesome-jev 520★ (+14 vs 506); Laya likes 861 (was 822); tracker likes 66 (+2 vs 64) lastModified UNCHANGED; Blackwood likes 2 gated manual; Archer still promised_not_landed; Hub archerhume/4rcherhume HTTP 401; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33; notes.md §114 + +User-provided 0920 jcr uniqueness lock: NiazMorshed2007/jcr MIT; site https://jcr.niazmorshed.dev; topics ai-agents,jev,mcp; **4★**; HEAD `138b3832`; README SHA `2a49dbc1`; LICENSE SHA `46231303`; size **14850**; Jev Capability Resolver; one tool to find documented deterministic commands in a nested capability tree; returns context; **does not execute**; skills = workflow+judgment; capabilities = individual operations; format independent of Jev; proposed open standard exploration; classify (Jev) → optional OpenAI decompose compound → beam search geometric mean of routing probs; keep up to 3 paths ≥60% of best (JCR_BAND_RATIO 0.6); ambiguity / no-match / depth-limit explicit; soft scores ≠ hard gates; 0.6 band is application policy; routing ≠ permission; docs ≠ authority to run; sol-vs-opus5-20 *theirs*: 20 scenarios × 4 variants = 80 runs; lookup+explain only, no execution; Claude Opus 5: agent input 108,585→15,819 (−85%), cost $0.3700→$0.1222 (−67%), wall 105.5s→77.7s; Codex GPT-5.6-Sol: 61,952→47,669 (−23%), $0.1377→$0.1151 (−16%), wall 25.3s→62.4s (Sol slower with JCR in 19/20); One Sol outlier 372.6s / 193 Jev calls; n=1 per cell; Not Harbor task-execution; Claude/Codex harnesses; compare mode; 50 scenarios bundled; 11 groups, 960 nodes, 11,360 items; 16 routing rounds per step; NiazMorshed2007/jcr ≠ skill-broker ≠ skillranker ≠ jev-sift ≠ jev-lens ≠ jevusher ≠ jev_select_capability; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33/#34; notes.md §116 diff --git a/.agents/skills/augustus/scripts/uniqueness_gate.py b/.agents/skills/augustus/scripts/uniqueness_gate.py index dc43b420..6a158975 100644 --- a/.agents/skills/augustus/scripts/uniqueness_gate.py +++ b/.agents/skills/augustus/scripts/uniqueness_gate.py @@ -1,50 +1,33 @@ #!/usr/bin/env python3 -"""Hourly fold uniqueness gate (0843). +"""Uniqueness gate for merged 0843 (§114), merged 0915 NanoJev (§115), +and user-provided 0920 jcr (§116). -The uniqueness lock must appear as one consecutive substring in every -listed overlay. Fragments scattered across files do not count. +Each lock must appear as one consecutive substring in every listed overlay. +Fragments scattered across files do not count. -Also: YAML-parse SKILL.md frontmatter; notes.md owns §114; composition -items 289–302 exist. Does not fetch the network. Does not treat a lock +Also: YAML-parse SKILL.md frontmatter; notes.md owns §114, §115, and §116; +composition items 289–302, 303–308, and 309–316 exist; findings batches +#97, #98, and #99 exist. Does not fetch the network. Does not treat a lock as a Harbor score. """ from pathlib import Path import sys +import yaml + ROOT = Path(__file__).resolve().parents[4] -UNIQ = ( - "Hourly 0843 uniqueness lock: A hunch is a probability with a policy attached; " - "{ enter: 0.8, exit: 0.6 } is hysteresis; replay a policy change without inference; " - "Decision models are providers, not the product; huncho ≠ Kungie/gut ≠ carldaws/hunch ≠ tpellet/hunch; " - "pretrained Qwen2.5 base ECE 0.030 (0.5B) / 0.040 (7B); instruct 0.302 / 0.269; " - "70.9% → 70.0% mean conf 74.1% → 96.7%; temperature scaling still matches it in-distribution; " - "No Jev API was called; Qwen2.5 ≠ Archer; Qwen/Qwen3.8-27B ≠ Archer; " - "学習済みモデル v0.1 は準備中です; bool AUROC 0.523; 先頭だと0件、末尾だと250件; " - "温度を渡さない場合、確率は較正されていません; このリポジトリには Jev を呼ぶコードが存在しません; " - "g0runmezadam/what-is-jev IS tunahansahin897/what-is-jev (same GitHub id 1378007307); " - "947 repos scored; A 273 · B 302 · C 372; LLM rubric ≠ benches; Data as of 2026-09-20; " - "HEAD 895b9498; README SHA 3ae98c56; 13 focused checks and one mutually exclusive outcome; " - "Probabilities are advisory, not calibrated guarantees; omni-/ask-jev ≠ pedroknigge/mcp_jev; " - "pd.cut bins by equal width while jeval bins by quantile; ECE 0.113 and ECE 0.076; " - "jeval drift is not implemented yet; rlaope/jeval ≠ dayhaysoos/jevals; " - "calibration does not compose; ECE has exactly zero statistical power to detect the failure mode that kills trajectories; " - "25–60× headline withdrawn; P(all-correct): 0.0071 vs 0.0001; TCE / AMS; " - "Qwen 3.8 sparring ≠ Archer; Deferred Crispification; light_cutoff_applied_to_combination 0; " - "recorded run, kinematic animation; BANKING77 Accuracy BERT-Base 93.02 Jev 79.90; " - "Analyse jev calibration (NLL, ECE) backlog; BERT figures are published supervised references, not zero-shot; " - "档位措辞效应 分数极差中位 0.50、最大 1.32; 修好后对照组是 0.01; 不是 benchmark; 概率没做 calibration; " - "~1,430 API calls, about $0.15; xiaohuaxi/jev-study ≠ wjdjdakf17/jev-study ≠ baekenough/jev-study; " - "AND: product (independence assumed and recorded in the trace); " - "chat model's stated confidence is not calibrated; circuit-vl-4b ≠ Archer; " - "Bring your own API key; vamsikrishna2421/jev-usecases ≠ whyashthakker/awesome-jev-use-cases; " - "catalog ≠ endorsement; SemIf 2237★ (+30 vs §111 2207); jevlike 1049★ (+6 vs 1043); " - "TypeLLM/TypeLLM 16★; AnotiaWang 98★; yibie/awesome-jev 520★ (+14 vs 506); " - "Laya likes 861 (was 822); tracker likes 66 (+2 vs 64) lastModified UNCHANGED; " - "Blackwood likes 2 gated manual; Archer still promised_not_landed; " - "Hub archerhume/4rcherhume HTTP 401; " - "do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33; notes.md §114" +UNIQ_0843 = ( + "Hourly 0843 uniqueness lock: A hunch is a probability with a policy attached; { enter: 0.8, exit: 0.6 } is hysteresis; replay a policy change without inference; Decision models are providers, not the product; huncho ≠ Kungie/gut ≠ carldaws/hunch ≠ tpellet/hunch; pretrained Qwen2.5 base ECE 0.030 (0.5B) / 0.040 (7B); instruct 0.302 / 0.269; 70.9% → 70.0% mean conf 74.1% → 96.7%; temperature scaling still matches it in-distribution; No Jev API was called; Qwen2.5 ≠ Archer; Qwen/Qwen3.8-27B ≠ Archer; 学習済みモデル v0.1 は準備中です; bool AUROC 0.523; 先頭だと0件、末尾だと250件; 温度を渡さない場合、確率は較正されていません; このリポジトリには Jev を呼ぶコードが存在しません; g0runmezadam/what-is-jev IS tunahansahin897/what-is-jev (same GitHub id 1378007307); 947 repos scored; A 273 · B 302 · C 372; LLM rubric ≠ benches; Data as of 2026-09-20; HEAD 895b9498; README SHA 3ae98c56; 13 focused checks and one mutually exclusive outcome; Probabilities are advisory, not calibrated guarantees; omni-/ask-jev ≠ pedroknigge/mcp_jev; pd.cut bins by equal width while jeval bins by quantile; ECE 0.113 and ECE 0.076; jeval drift is not implemented yet; rlaope/jeval ≠ dayhaysoos/jevals; calibration does not compose; ECE has exactly zero statistical power to detect the failure mode that kills trajectories; 25–60× headline withdrawn; P(all-correct): 0.0071 vs 0.0001; TCE / AMS; Qwen 3.8 sparring ≠ Archer; Deferred Crispification; light_cutoff_applied_to_combination 0; recorded run, kinematic animation; BANKING77 Accuracy BERT-Base 93.02 Jev 79.90; Analyse jev calibration (NLL, ECE) backlog; BERT figures are published supervised references, not zero-shot; 档位措辞效应 分数极差中位 0.50、最大 1.32; 修好后对照组是 0.01; 不是 benchmark; 概率没做 calibration; ~1,430 API calls, about $0.15; xiaohuaxi/jev-study ≠ wjdjdakf17/jev-study ≠ baekenough/jev-study; AND: product (independence assumed and recorded in the trace); chat model's stated confidence is not calibrated; circuit-vl-4b ≠ Archer; Bring your own API key; vamsikrishna2421/jev-usecases ≠ whyashthakker/awesome-jev-use-cases; catalog ≠ endorsement; SemIf 2237★ (+30 vs §111 2207); jevlike 1049★ (+6 vs 1043); TypeLLM/TypeLLM 16★; AnotiaWang 98★; yibie/awesome-jev 520★ (+14 vs 506); Laya likes 861 (was 822); tracker likes 66 (+2 vs 64) lastModified UNCHANGED; Blackwood likes 2 gated manual; Archer still promised_not_landed; Hub archerhume/4rcherhume HTTP 401; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33; notes.md §114" +) + +UNIQ_0915 = ( + 'User-provided 0915 uniqueness lock: TianyuCodings/NanoJev unified-games-v1 densify; A 0.6B parallel decision model: states and questions in, complete probability distributions out. Zero output-token decoding.; One model, four games; ViZDoom Basic 128/128 vs Jev 56/128; Predict Position 27/128 vs Jev 11/128; Maze 225 attempts vs Jev 2738; Snake 30 food / 256 steps; held-out Maze 4/10 Snake 8/8 Basic 128/128 Predict 27/128; Untuned Qwen3-0.6B baseline; 18,760 questions per variant; 16,333 ViZDoom; 896 Predict Position expert episodes; hard_lr1e5; mix weights 1/3, 1/3, 1/6, 1/6; Hub C-Tianyu/NanoJev revision unified-games-v1 likes 58; dataset C-Tianyu/NanoJev-Data likes 5; HEAD 618cea6d906d54e128360786d12f703fff2b1245; 1289★ / 158 forks / size 64035; README SHA 4190093c64ee75b26e9726daa3b00cbcf6d3157a; MIT; caijinchun/nanojev-arena ≠ liao96312/jev-arena-nanojev ≠ zwliJay/jev-forge ≠ NanoJev; not TypeSafe Jev; open replica / specialist gameplay S1; soft scores ≠ hard gates; Game success ≠ calibrated Noul; local type boolean ≠ TypeSafe noul; invented_signal false; do not reopen or amend PR #31/#32/#33/#35; notes.md §115' +) + +UNIQ_JCR = ( + 'User-provided 0920 jcr uniqueness lock: NiazMorshed2007/jcr MIT; site https://jcr.niazmorshed.dev; topics ai-agents,jev,mcp; **4★**; HEAD `138b3832`; README SHA `2a49dbc1`; LICENSE SHA `46231303`; size **14850**; Jev Capability Resolver; one tool to find documented deterministic commands in a nested capability tree; returns context; **does not execute**; skills = workflow+judgment; capabilities = individual operations; format independent of Jev; proposed open standard exploration; classify (Jev) → optional OpenAI decompose compound → beam search geometric mean of routing probs; keep up to 3 paths ≥60% of best (JCR_BAND_RATIO 0.6); ambiguity / no-match / depth-limit explicit; soft scores ≠ hard gates; 0.6 band is application policy; routing ≠ permission; docs ≠ authority to run; sol-vs-opus5-20 *theirs*: 20 scenarios × 4 variants = 80 runs; lookup+explain only, no execution; Claude Opus 5: agent input 108,585→15,819 (−85%), cost $0.3700→$0.1222 (−67%), wall 105.5s→77.7s; Codex GPT-5.6-Sol: 61,952→47,669 (−23%), $0.1377→$0.1151 (−16%), wall 25.3s→62.4s (Sol slower with JCR in 19/20); One Sol outlier 372.6s / 193 Jev calls; n=1 per cell; Not Harbor task-execution; Claude/Codex harnesses; compare mode; 50 scenarios bundled; 11 groups, 960 nodes, 11,360 items; 16 routing rounds per step; NiazMorshed2007/jcr ≠ skill-broker ≠ skillranker ≠ jev-sift ≠ jev-lens ≠ jevusher ≠ jev_select_capability; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33/#34; notes.md §116' ) OVERLAYS = [ @@ -68,13 +51,11 @@ ".agents/skills/augustus/references/toolbox-mapping.md", ".agents/skills/augustus/references/mappings.md", ".agents/skills/augustus/references/agent-self-assessment.md", - ".agents/skills/augustus/references/question-design.md", + ".agents/skills/augustus/references/question-design.md" ] def load_skill_frontmatter(text: str) -> dict: - import yaml - if not text.startswith("---"): raise AssertionError("SKILL.md missing YAML frontmatter") end = text.find("\n---\n", 3) @@ -94,18 +75,30 @@ def main() -> int: failed.append(f"missing {rel}") continue body = path.read_text(encoding="utf-8") - if UNIQ not in body: - failed.append(f"lock missing as one substring: {rel}") + if UNIQ_0843 not in body: + failed.append(f"0843 lock missing as one substring: {rel}") + if UNIQ_0915 not in body: + failed.append(f"0915 lock missing as one substring: {rel}") + if UNIQ_JCR not in body: + failed.append(f"jcr lock missing as one substring: {rel}") notes = (ROOT / "research/notes.md").read_text(encoding="utf-8") if "## 114. Hourly 0843 HIGH" not in notes: failed.append("notes.md missing §114 heading") + if "## 115. User-provided HIGH — TianyuCodings/NanoJev" not in notes: + failed.append("notes.md missing §115 heading") + if "## 116. User-provided HIGH — NiazMorshed2007/jcr" not in notes: + failed.append("notes.md missing §116 heading") algebra = (ROOT / ".agents/skills/augustus/references/composition-algebra.md").read_text( encoding="utf-8" ) - for n in range(289, 303): + for n in list(range(289, 303)) + list(range(303, 309)) + list(range(309, 317)): needle = f"{n}. **" if needle not in algebra: failed.append(f"composition-algebra missing item {n}") + findings = (ROOT / "research/archive/findings.md").read_text(encoding="utf-8") + for batch in ("## Batch #97", "## Batch #98", "## Batch #99"): + if batch not in findings: + failed.append(f"findings.md missing {batch}") skill = (ROOT / ".agents/skills/augustus/SKILL.md").read_text(encoding="utf-8") try: fm = load_skill_frontmatter(skill) @@ -116,16 +109,19 @@ def main() -> int: if fm.get("name") != "augustus": failed.append("SKILL.md name != augustus") desc = fm.get("description") or "" + combined = desc + "\n" + skill for frag in ( "A hunch is a probability with a policy attached", "calibration does not compose", - "Qwen2.5 ≠ Archer", - "Qwen/Qwen3.8-27B ≠ Archer", - "Deferred Crispification", - "ranking ≠ calibration", - "g0runmezadam/what-is-jev IS tunahansahin897/what-is-jev", + "NiazMorshed2007/jcr", + "does not execute", + "JCR_BAND_RATIO 0.6", + "routing ≠ permission", + "docs ≠ authority to run", + "Not Harbor task-execution", + "TianyuCodings/NanoJev", ): - if frag not in desc and frag not in skill: + if frag not in combined: failed.append(f"SKILL.md missing fragment {frag!r}") if failed: print("uniqueness-gate FAIL") @@ -133,7 +129,10 @@ def main() -> int: print(" -", line) return 1 print("uniqueness-gate ok") - print(f"lock chars={len(UNIQ)} overlays={len(OVERLAYS)}") + print( + f"0843 chars={len(UNIQ_0843)} 0915 chars={len(UNIQ_0915)} " + f"jcr chars={len(UNIQ_JCR)} overlays={len(OVERLAYS)}" + ) return 0 diff --git a/CHANGELOG.md b/CHANGELOG.md index 7be93474..588622d7 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -16,12 +16,14 @@ folds: `research/notes.md`. ## [Unreleased] -User-provided 0915 HIGH (`research/notes.md` §115 / -composition items 303–308 / findings batch #98). Does -**not** bump the 0.4.0 pin. Do not reopen or amend PR -#31/#32/#33/#35. Merged #35 owns §114 / -items 289–302 / batch #97. This fold stays §115 / -303–308 / #98. +Hourly 0843 HIGH (`research/notes.md` §114 / composition +items 289–302 / findings batch #97) plus merged #36 NanoJev +(`research/notes.md` §115 / items 303–308 / batch #98) plus +user-provided HIGH NiazMorshed2007/jcr (`research/notes.md` +§116 / composition items 309–316 / findings batch #99). Does +**not** bump the 0.4.0 pin. Do not reopen or amend PR #23–#36. +Open #37 owns §117 / 315–321 / #100 (item overlap 315–316 is +#37's remap). ### Added @@ -50,10 +52,6 @@ items 289–302 / batch #97. This fold stays §115 / [`research/changelog-hourly.md`](research/changelog-hourly.md). User-provided 0915 uniqueness lock: TianyuCodings/NanoJev unified-games-v1 densify; A 0.6B parallel decision model: states and questions in, complete probability distributions out. Zero output-token decoding.; One model, four games; ViZDoom Basic 128/128 vs Jev 56/128; Predict Position 27/128 vs Jev 11/128; Maze 225 attempts vs Jev 2738; Snake 30 food / 256 steps; held-out Maze 4/10 Snake 8/8 Basic 128/128 Predict 27/128; Untuned Qwen3-0.6B baseline; 18,760 questions per variant; 16,333 ViZDoom; 896 Predict Position expert episodes; hard_lr1e5; mix weights 1/3, 1/3, 1/6, 1/6; Hub C-Tianyu/NanoJev revision unified-games-v1 likes 58; dataset C-Tianyu/NanoJev-Data likes 5; HEAD 618cea6d906d54e128360786d12f703fff2b1245; 1289★ / 158 forks / size 64035; README SHA 4190093c64ee75b26e9726daa3b00cbcf6d3157a; MIT; caijinchun/nanojev-arena ≠ liao96312/jev-arena-nanojev ≠ zwliJay/jev-forge ≠ NanoJev; not TypeSafe Jev; open replica / specialist gameplay S1; soft scores ≠ hard gates; Game success ≠ calibrated Noul; local type boolean ≠ TypeSafe noul; invented_signal false; do not reopen or amend PR #31/#32/#33/#35; notes.md §115 -Hourly 0743 HIGH (`research/notes.md` §113 / composition -items 273–288 / findings batch #96). Does **not** bump -the 0.4.0 pin. Do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#32/#33. - ### Added - **Hourly 0843 HIGH (`notes.md` §114).** Measurement / judgment fold @@ -75,6 +73,23 @@ the 0.4.0 pin. Do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#32/#33 - Uniqueness lock archive: `research/changelog-hourly.md` (this hour's lock is there; this file stays scannable). +- User-provided HIGH NiazMorshed2007/jcr + (`research/notes.md` §116): **Skip Archer rewrite.** + Docs-only, rebased onto merged #35 (`0189825`). + **HARD RULE:** do not reopen or amend PR #23–#35. + Merged #36 owns §115 / 303–308 / #98; open #37 owns + §117 / 315–321 / #100 (item overlap 315–316 is + #37's remap). This fold keeps §116 / items + 309–316 / batch #99. Quote *theirs*. Jev Capability + Resolver: one tool, nested capability tree, returns + context, **does not execute**. skills vs capabilities. + 0.6 band is application policy. routing ≠ permission. + docs ≠ authority to run. sol-vs-opus5-20 lookup+explain + only; n=1; wall-time mixed; Not Harbor task-execution. + Live REST: **4★**; HEAD `138b3832`; README SHA + `2a49dbc1`; size **14850**. `invented_signal: false`. + Uniqueness dump in `research/changelog-hourly.md`. + ### Changed - `evaluate_decisions.py` reports equal-width and quantile ECE, AUC, @@ -467,3 +482,4 @@ The dated passes below are how 0.1.0 was assembled. relations, logical-operator combination rules, and the position×construct traversal as the systematic application generator; wired into SKILL.md index + toolbox sweep. + diff --git a/README.md b/README.md index 7295b895..0479a65a 100644 --- a/README.md +++ b/README.md @@ -63,6 +63,8 @@ One line per file. The living catalog is in the reference cards and - `research/README.md`: evidence archive index (sources, refresh log, hourly dumps) - `research/changelog-hourly.md`: hourly uniqueness dumps after v0.3.0 - User-provided 0915 uniqueness lock: TianyuCodings/NanoJev unified-games-v1 densify; A 0.6B parallel decision model: states and questions in, complete probability distributions out. Zero output-token decoding.; One model, four games; ViZDoom Basic 128/128 vs Jev 56/128; Predict Position 27/128 vs Jev 11/128; Maze 225 attempts vs Jev 2738; Snake 30 food / 256 steps; held-out Maze 4/10 Snake 8/8 Basic 128/128 Predict 27/128; Untuned Qwen3-0.6B baseline; 18,760 questions per variant; 16,333 ViZDoom; 896 Predict Position expert episodes; hard_lr1e5; mix weights 1/3, 1/3, 1/6, 1/6; Hub C-Tianyu/NanoJev revision unified-games-v1 likes 58; dataset C-Tianyu/NanoJev-Data likes 5; HEAD 618cea6d906d54e128360786d12f703fff2b1245; 1289★ / 158 forks / size 64035; README SHA 4190093c64ee75b26e9726daa3b00cbcf6d3157a; MIT; caijinchun/nanojev-arena ≠ liao96312/jev-arena-nanojev ≠ zwliJay/jev-forge ≠ NanoJev; not TypeSafe Jev; open replica / specialist gameplay S1; soft scores ≠ hard gates; Game success ≠ calibrated Noul; local type boolean ≠ TypeSafe noul; invented_signal false; do not reopen or amend PR #31/#32/#33/#35; notes.md §115 + +- User-provided HIGH NiazMorshed2007/jcr (`research/notes.md` §116 / items 309–316 / batch #99). Capability-tree lookup; **does not execute**. Merged #35 owns §114. Merged #36 owns §115. Open #37 §117. - Hourly 0743 uniqueness lock: Heman10x-NGU/Verdict-open-jev ≠ Heman10x-NGU/openJev-verdict-2.0; TF-IDF + LogReg ECE 0.0207 vs Jev 0.1440; Verdict-open-jev 48.07% vs Jev 90.80%; abstention combined recall 10.00%; p50 35.58 ms; K=25 (maximum capacity) 72.00%; 0.85 coverage 84.60% selective risk 1.18%; 26.1× faster than standard Qwen JSON generation; Jevify 90.0% / 167 ms CUDA graphs disabled; Finding 1: Brier on stated confidence alone is a trap; grpo_rlcr 0.78 / ECE 0.084; reliability 0.007 but resolution 0.000; 27 900 schema-driven decisions; 13 600 / 13 600 questions; candidate mass min 0.99999624; 22 configs · 166,054 rows · 4 calibration-gold; sha a39eba3f; Student B MAE 0.148 / Pearson 0.836 / 86.0%; pngwn/open-jev-laya-bench README 404; sha 9f69c742 likes 2; HDFS 0.9933 (745/750) / retain 0.0084; BGL ERROR/FATAL protection 1.0000; 2,479 / 2,500 HDFS uncertain; cache hit 0.9648 (2412/2500); $0.153936 estimated; E2 recomputes from saved probabilities; Space sha eda59e0a; MASSIVE English 0.783 / Khmer 0.033 / Hindi 0.133; 40–48 rows too small to ship T; T never changes argmax; siren2345/jev-single-decode-transformers ≠ siren2345/jev-single-decode; Split Transformers experiment from llama.cpp runtime; tanayvasishtha/jev-lab ≠ dairui1/jev-lab ≠ BrendanH18/jev-lab ≠ yibie/laya-jev-lab; Four experiments stress-testing TypeSafe's Jev: calibration, bundle bias, label bias, and ensembling; second pass must be $0.00 from cache; The pages never call Jev; Gemma 4 31B 77.0% / Jev 1.13.0 61.4% / Laya 322M 0.0%; restriction state 95.0% against 84.4%; None of the systems are particularly good at knowing when to stop and ask; They skip the question and call a tool directly; 100% schema pass; six-field joint 48.8% vs 72.8%; ywchiu/jev_benchmark ≠ Running-Dolphins/jev-bench ≠ Praveenrajus/jev-bench; ACT / REVIEW / FALLBACK; A provider failure, timeout, malformed output, or missing answer is **not** a policy outcome; confidence is descriptive provider output, not a substitute for probability; Quality denominators include only valid scored answers; an exact halfway tie chooses the lower level; aiwithenoch/Jev-Skill ≠ simplosophy/jev-skill ≠ laguagu/jev-skills; The local path does not claim to turn a smaller checkpoint into Jev; Low support becomes decision: "review"; MIT-0 SPDX NOASSERTION; current-llm; 结构兼容,不是 Jev 模型能力; altryne/jevify ≠ Mintzs/jevify ≠ gulagala001/jevify ≠ uspraveen/Jevify; Find where Jev belongs. Design the questions. Measure the difference; TypeAR-AI/TypeAR 301 → TypeLLM/TypeLLM; TypeLLM/TypeLLM 16★; SemIf 2241★ (+34 vs §111 2207); jevlike 1051★ (+8 vs 1043); AnotiaWang 98★ (+1 vs 97); yibie/awesome-jev 525★ (+19 vs 506); Laya likes 864 (was 822); tracker likes 67 (+3 vs 64); lastModified UNCHANGED `2026-09-20T04:29:16.000Z`; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#32/#33; notes.md §113 - Hourly 0843 uniqueness lock: A hunch is a probability with a policy attached; { enter: 0.8, exit: 0.6 } is hysteresis; replay a policy change without inference; Decision models are providers, not the product; huncho ≠ Kungie/gut ≠ carldaws/hunch ≠ tpellet/hunch; pretrained Qwen2.5 base ECE 0.030 (0.5B) / 0.040 (7B); instruct 0.302 / 0.269; 70.9% → 70.0% mean conf 74.1% → 96.7%; temperature scaling still matches it in-distribution; No Jev API was called; Qwen2.5 ≠ Archer; Qwen/Qwen3.8-27B ≠ Archer; 学習済みモデル v0.1 は準備中です; bool AUROC 0.523; 先頭だと0件、末尾だと250件; 温度を渡さない場合、確率は較正されていません; このリポジトリには Jev を呼ぶコードが存在しません; g0runmezadam/what-is-jev IS tunahansahin897/what-is-jev (same GitHub id 1378007307); 947 repos scored; A 273 · B 302 · C 372; LLM rubric ≠ benches; Data as of 2026-09-20; HEAD 895b9498; README SHA 3ae98c56; 13 focused checks and one mutually exclusive outcome; Probabilities are advisory, not calibrated guarantees; omni-/ask-jev ≠ pedroknigge/mcp_jev; pd.cut bins by equal width while jeval bins by quantile; ECE 0.113 and ECE 0.076; jeval drift is not implemented yet; rlaope/jeval ≠ dayhaysoos/jevals; calibration does not compose; ECE has exactly zero statistical power to detect the failure mode that kills trajectories; 25–60× headline withdrawn; P(all-correct): 0.0071 vs 0.0001; TCE / AMS; Qwen 3.8 sparring ≠ Archer; Deferred Crispification; light_cutoff_applied_to_combination 0; recorded run, kinematic animation; BANKING77 Accuracy BERT-Base 93.02 Jev 79.90; Analyse jev calibration (NLL, ECE) backlog; BERT figures are published supervised references, not zero-shot; 档位措辞效应 分数极差中位 0.50、最大 1.32; 修好后对照组是 0.01; 不是 benchmark; 概率没做 calibration; ~1,430 API calls, about $0.15; xiaohuaxi/jev-study ≠ wjdjdakf17/jev-study ≠ baekenough/jev-study; AND: product (independence assumed and recorded in the trace); chat model's stated confidence is not calibrated; circuit-vl-4b ≠ Archer; Bring your own API key; vamsikrishna2421/jev-usecases ≠ whyashthakker/awesome-jev-use-cases; catalog ≠ endorsement; SemIf 2237★ (+30 vs §111 2207); jevlike 1049★ (+6 vs 1043); TypeLLM/TypeLLM 16★; AnotiaWang 98★; yibie/awesome-jev 520★ (+14 vs 506); Laya likes 861 (was 822); tracker likes 66 (+2 vs 64) lastModified UNCHANGED; Blackwood likes 2 gated manual; Archer still promised_not_landed; Hub archerhume/4rcherhume HTTP 401; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33; notes.md §114 @@ -121,3 +123,5 @@ before treating that pin as current API behavior. MIT. See [LICENSE](LICENSE). Security reports: [SECURITY.md](SECURITY.md). Contributions: [CONTRIBUTING.md](CONTRIBUTING.md). + +User-provided 0920 jcr uniqueness lock: NiazMorshed2007/jcr MIT; site https://jcr.niazmorshed.dev; topics ai-agents,jev,mcp; **4★**; HEAD `138b3832`; README SHA `2a49dbc1`; LICENSE SHA `46231303`; size **14850**; Jev Capability Resolver; one tool to find documented deterministic commands in a nested capability tree; returns context; **does not execute**; skills = workflow+judgment; capabilities = individual operations; format independent of Jev; proposed open standard exploration; classify (Jev) → optional OpenAI decompose compound → beam search geometric mean of routing probs; keep up to 3 paths ≥60% of best (JCR_BAND_RATIO 0.6); ambiguity / no-match / depth-limit explicit; soft scores ≠ hard gates; 0.6 band is application policy; routing ≠ permission; docs ≠ authority to run; sol-vs-opus5-20 *theirs*: 20 scenarios × 4 variants = 80 runs; lookup+explain only, no execution; Claude Opus 5: agent input 108,585→15,819 (−85%), cost $0.3700→$0.1222 (−67%), wall 105.5s→77.7s; Codex GPT-5.6-Sol: 61,952→47,669 (−23%), $0.1377→$0.1151 (−16%), wall 25.3s→62.4s (Sol slower with JCR in 19/20); One Sol outlier 372.6s / 193 Jev calls; n=1 per cell; Not Harbor task-execution; Claude/Codex harnesses; compare mode; 50 scenarios bundled; 11 groups, 960 nodes, 11,360 items; 16 routing rounds per step; NiazMorshed2007/jcr ≠ skill-broker ≠ skillranker ≠ jev-sift ≠ jev-lens ≠ jevusher ≠ jev_select_capability; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33/#34; notes.md §116 diff --git a/docs/ecosystem.md b/docs/ecosystem.md index 113c1561..c6d48f94 100644 --- a/docs/ecosystem.md +++ b/docs/ecosystem.md @@ -35,6 +35,7 @@ weekdays. Jev is the densest public corpus, not the class monopoly. - **jamesward/zio-typesafe-ai** — Effect-oriented (ZIO) client: Jev is the outer Choice; the handler runs the effect. Not Effect.ts and not the Jev HTTP contract. Hypothesis: `mappings.md` §19 (`notes.md` §28). ### Agent harnesses & self-supervision +- **NiazMorshed2007/jcr** — JavaScript MIT; **4★**; HEAD `138b3832`; README SHA `2a49dbc1`; LICENSE SHA `46231303`; size **14850**; topics ai-agents,jev,mcp; site https://jcr.niazmorshed.dev. Jev Capability Resolver: one tool to find documented deterministic commands in a nested capability tree; returns context; **does not execute**. skills = workflow+judgment; capabilities = individual operations; format independent of Jev. Beam geometric mean; JCR_BAND_RATIO 0.6 is application policy. routing ≠ permission; docs ≠ authority to run. sol-vs-opus5-20 *theirs* lookup+explain; n=1; Not Harbor task-execution. NiazMorshed2007/jcr ≠ skill-broker ≠ skillranker ≠ jev-sift ≠ jev-lens ≠ jevusher ≠ jev_select_capability. `notes.md` §116. - **Kevthetech143/super-jev** — domain-independent loop: observe → questions → decide → **permit (independent of confidence)** → execute (idempotency key) → verify → JSONL replay. - **AntonioCoppe/jev-harness** — policy + confidence gate + shadow mode + offline eval CLI asserting on the **action**; 24-row filter 48.9s (Claude CLI) vs 1.3s Jev. Harbor/jevals-adjacent practice. `notes.md` §33, §44. - **khordoo/jev-reflex-autonomy-lab** — S1 Jev reflex keeps control; optional S2 planner is one-use advice on low confidence. **Delta:** escalate without stalling; purple telemetry = consumed not arrived; Local controller (rule-based) ≠ githubnext/localjev; 20% starting gate still soft; seed = geometry not replay; no pixels; experimental viz, not a flight controller. `notes.md` §46, §80. @@ -788,6 +789,8 @@ Pulse (do not invent): Demo `nanojev-dev.tianyuchen99.chatgpt.site` HTTP **401** Hourly 0743 uniqueness lock: Heman10x-NGU/Verdict-open-jev ≠ Heman10x-NGU/openJev-verdict-2.0; TF-IDF + LogReg ECE 0.0207 vs Jev 0.1440; Verdict-open-jev 48.07% vs Jev 90.80%; abstention combined recall 10.00%; p50 35.58 ms; K=25 (maximum capacity) 72.00%; 0.85 coverage 84.60% selective risk 1.18%; 26.1× faster than standard Qwen JSON generation; Jevify 90.0% / 167 ms CUDA graphs disabled; Finding 1: Brier on stated confidence alone is a trap; grpo_rlcr 0.78 / ECE 0.084; reliability 0.007 but resolution 0.000; 27 900 schema-driven decisions; 13 600 / 13 600 questions; candidate mass min 0.99999624; 22 configs · 166,054 rows · 4 calibration-gold; sha a39eba3f; Student B MAE 0.148 / Pearson 0.836 / 86.0%; pngwn/open-jev-laya-bench README 404; sha 9f69c742 likes 2; HDFS 0.9933 (745/750) / retain 0.0084; BGL ERROR/FATAL protection 1.0000; 2,479 / 2,500 HDFS uncertain; cache hit 0.9648 (2412/2500); $0.153936 estimated; E2 recomputes from saved probabilities; Space sha eda59e0a; MASSIVE English 0.783 / Khmer 0.033 / Hindi 0.133; 40–48 rows too small to ship T; T never changes argmax; siren2345/jev-single-decode-transformers ≠ siren2345/jev-single-decode; Split Transformers experiment from llama.cpp runtime; tanayvasishtha/jev-lab ≠ dairui1/jev-lab ≠ BrendanH18/jev-lab ≠ yibie/laya-jev-lab; Four experiments stress-testing TypeSafe's Jev: calibration, bundle bias, label bias, and ensembling; second pass must be $0.00 from cache; The pages never call Jev; Gemma 4 31B 77.0% / Jev 1.13.0 61.4% / Laya 322M 0.0%; restriction state 95.0% against 84.4%; None of the systems are particularly good at knowing when to stop and ask; They skip the question and call a tool directly; 100% schema pass; six-field joint 48.8% vs 72.8%; ywchiu/jev_benchmark ≠ Running-Dolphins/jev-bench ≠ Praveenrajus/jev-bench; ACT / REVIEW / FALLBACK; A provider failure, timeout, malformed output, or missing answer is **not** a policy outcome; confidence is descriptive provider output, not a substitute for probability; Quality denominators include only valid scored answers; an exact halfway tie chooses the lower level; aiwithenoch/Jev-Skill ≠ simplosophy/jev-skill ≠ laguagu/jev-skills; The local path does not claim to turn a smaller checkpoint into Jev; Low support becomes decision: "review"; MIT-0 SPDX NOASSERTION; current-llm; 结构兼容,不是 Jev 模型能力; altryne/jevify ≠ Mintzs/jevify ≠ gulagala001/jevify ≠ uspraveen/Jevify; Find where Jev belongs. Design the questions. Measure the difference; TypeAR-AI/TypeAR 301 → TypeLLM/TypeLLM; TypeLLM/TypeLLM 16★; SemIf 2241★ (+34 vs §111 2207); jevlike 1051★ (+8 vs 1043); AnotiaWang 98★ (+1 vs 97); yibie/awesome-jev 525★ (+19 vs 506); Laya likes 864 (was 822); tracker likes 67 (+3 vs 64); lastModified UNCHANGED `2026-09-20T04:29:16.000Z`; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#32/#33; notes.md §113 +Architecture / mental models / VOI admission / measurement honesty, not a JCR / npm / skill-broker tutorial. `notes.md` §116. Skip Archer rewrite. Quote README. Soft Noul ≠ hard safety. Do **not** re-fold 0843 / §114 / 0915 / §115 / 0743 / §113 / 0806 / §112. Fresh PR; never reopen merged #7–**#36**. **HARD RULE:** do not reopen or amend PR #23–#36. Do not push onto open #37. + Architecture / mental models / Harbor-jevals / toolbelt, not a Verdict-open-jev / Mintzs / rlcd-lite / ywchiu / DecisionOps / Jev-Skill tutorial. `notes.md` §113. Skip Archer rewrite. Quote READMEs. Soft Noul ≠ hard safety. Do **not** re-fold 0646 / §111 / 0541 / §110 / 0439 / §109 / 0345 / §108 / 0243 / §107 / 0145 / §106 / 0042 / §105 / 2340 / §104 / 2246 / §103 / 2145 / §102 / 2041 / §101 / 1943 / §100 / 1843 / §99 / 1740 / §98 / 1639 / §96 / gliner-native-runtime / §97 / 1541 / §95 / jev-align *mechanism* / §93 / jev-orderby-bench *six-gates* / §60 / JevBench v1.2 *board* / §78 / openJev-verdict *claim-audit* / §71 / pngwn RESULTS / §46 / yuki-oshio/mini-jev *93.25%* / §103. Fresh PR; never reopen merged #7–**#29** / merged **#30** / merged **#32**. **HARD RULE:** do not reopen or amend PR #23 or #24 or #25 or #26 or #27 or #28 or #29 or #30 or #32. 0★ HIGH still gets a real card. Size **0** WITH CONTENTS still gets a real card. Measurement densifies PRIMARY (ywchiu routing-across-a-conversation) plus Transformers one-decode sibling, tanay scaffold, dataset/Space densifies plus open-weight/RLCD plus skills/DecisionOps are the *class* exemplars this hour, not TypeSafe drop-ins. Quote live REST over watch claims. `invented_signal: false`. - **ywchiu/jev_benchmark** — GitHub license **null**; **1★**; HEAD `4322c350`; README SHA `75e6a338`; size **173**. PRIMARY. ywchiu/jev_benchmark ≠ Running-Dolphins/jev-bench ≠ Praveenrajus/jev-bench. Gemma 4 31B 77.0% / Jev 1.13.0 61.4% / Laya 322M 0.0%. restriction state 95.0% against 84.4%. None of the systems are particularly good at knowing when to stop and ask. They skip the question and call a tool directly. 100% schema pass. six-field joint 48.8% vs 72.8%. @@ -1112,3 +1115,5 @@ placement, applied placements, method substitution, composition algebra, question-design, validation, and optimizer coupling. Hourly 0843 uniqueness lock: A hunch is a probability with a policy attached; { enter: 0.8, exit: 0.6 } is hysteresis; replay a policy change without inference; Decision models are providers, not the product; huncho ≠ Kungie/gut ≠ carldaws/hunch ≠ tpellet/hunch; pretrained Qwen2.5 base ECE 0.030 (0.5B) / 0.040 (7B); instruct 0.302 / 0.269; 70.9% → 70.0% mean conf 74.1% → 96.7%; temperature scaling still matches it in-distribution; No Jev API was called; Qwen2.5 ≠ Archer; Qwen/Qwen3.8-27B ≠ Archer; 学習済みモデル v0.1 は準備中です; bool AUROC 0.523; 先頭だと0件、末尾だと250件; 温度を渡さない場合、確率は較正されていません; このリポジトリには Jev を呼ぶコードが存在しません; g0runmezadam/what-is-jev IS tunahansahin897/what-is-jev (same GitHub id 1378007307); 947 repos scored; A 273 · B 302 · C 372; LLM rubric ≠ benches; Data as of 2026-09-20; HEAD 895b9498; README SHA 3ae98c56; 13 focused checks and one mutually exclusive outcome; Probabilities are advisory, not calibrated guarantees; omni-/ask-jev ≠ pedroknigge/mcp_jev; pd.cut bins by equal width while jeval bins by quantile; ECE 0.113 and ECE 0.076; jeval drift is not implemented yet; rlaope/jeval ≠ dayhaysoos/jevals; calibration does not compose; ECE has exactly zero statistical power to detect the failure mode that kills trajectories; 25–60× headline withdrawn; P(all-correct): 0.0071 vs 0.0001; TCE / AMS; Qwen 3.8 sparring ≠ Archer; Deferred Crispification; light_cutoff_applied_to_combination 0; recorded run, kinematic animation; BANKING77 Accuracy BERT-Base 93.02 Jev 79.90; Analyse jev calibration (NLL, ECE) backlog; BERT figures are published supervised references, not zero-shot; 档位措辞效应 分数极差中位 0.50、最大 1.32; 修好后对照组是 0.01; 不是 benchmark; 概率没做 calibration; ~1,430 API calls, about $0.15; xiaohuaxi/jev-study ≠ wjdjdakf17/jev-study ≠ baekenough/jev-study; AND: product (independence assumed and recorded in the trace); chat model's stated confidence is not calibrated; circuit-vl-4b ≠ Archer; Bring your own API key; vamsikrishna2421/jev-usecases ≠ whyashthakker/awesome-jev-use-cases; catalog ≠ endorsement; SemIf 2237★ (+30 vs §111 2207); jevlike 1049★ (+6 vs 1043); TypeLLM/TypeLLM 16★; AnotiaWang 98★; yibie/awesome-jev 520★ (+14 vs 506); Laya likes 861 (was 822); tracker likes 66 (+2 vs 64) lastModified UNCHANGED; Blackwood likes 2 gated manual; Archer still promised_not_landed; Hub archerhume/4rcherhume HTTP 401; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33; notes.md §114 + +User-provided 0920 jcr uniqueness lock: NiazMorshed2007/jcr MIT; site https://jcr.niazmorshed.dev; topics ai-agents,jev,mcp; **4★**; HEAD `138b3832`; README SHA `2a49dbc1`; LICENSE SHA `46231303`; size **14850**; Jev Capability Resolver; one tool to find documented deterministic commands in a nested capability tree; returns context; **does not execute**; skills = workflow+judgment; capabilities = individual operations; format independent of Jev; proposed open standard exploration; classify (Jev) → optional OpenAI decompose compound → beam search geometric mean of routing probs; keep up to 3 paths ≥60% of best (JCR_BAND_RATIO 0.6); ambiguity / no-match / depth-limit explicit; soft scores ≠ hard gates; 0.6 band is application policy; routing ≠ permission; docs ≠ authority to run; sol-vs-opus5-20 *theirs*: 20 scenarios × 4 variants = 80 runs; lookup+explain only, no execution; Claude Opus 5: agent input 108,585→15,819 (−85%), cost $0.3700→$0.1222 (−67%), wall 105.5s→77.7s; Codex GPT-5.6-Sol: 61,952→47,669 (−23%), $0.1377→$0.1151 (−16%), wall 25.3s→62.4s (Sol slower with JCR in 19/20); One Sol outlier 372.6s / 193 Jev calls; n=1 per cell; Not Harbor task-execution; Claude/Codex harnesses; compare mode; 50 scenarios bundled; 11 groups, 960 nodes, 11,360 items; 16 routing rounds per step; NiazMorshed2007/jcr ≠ skill-broker ≠ skillranker ≠ jev-sift ≠ jev-lens ≠ jevusher ≠ jev_select_capability; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33/#34; notes.md §116 diff --git a/research/archive/findings.md b/research/archive/findings.md index ab5e917a..f9eccfd0 100644 --- a/research/archive/findings.md +++ b/research/archive/findings.md @@ -1,5 +1,63 @@ # Deep-read findings (evidence for research/notes.md) + +## Batch #99 (2026-09-20 ~15:20 UTC / ~09:20 Boise) — user-provided HIGH NiazMorshed2007/jcr + +User-provided 0920 jcr uniqueness lock: NiazMorshed2007/jcr MIT; site https://jcr.niazmorshed.dev; topics ai-agents,jev,mcp; **4★**; HEAD `138b3832`; README SHA `2a49dbc1`; LICENSE SHA `46231303`; size **14850**; Jev Capability Resolver; one tool to find documented deterministic commands in a nested capability tree; returns context; **does not execute**; skills = workflow+judgment; capabilities = individual operations; format independent of Jev; proposed open standard exploration; classify (Jev) → optional OpenAI decompose compound → beam search geometric mean of routing probs; keep up to 3 paths ≥60% of best (JCR_BAND_RATIO 0.6); ambiguity / no-match / depth-limit explicit; soft scores ≠ hard gates; 0.6 band is application policy; routing ≠ permission; docs ≠ authority to run; sol-vs-opus5-20 *theirs*: 20 scenarios × 4 variants = 80 runs; lookup+explain only, no execution; Claude Opus 5: agent input 108,585→15,819 (−85%), cost $0.3700→$0.1222 (−67%), wall 105.5s→77.7s; Codex GPT-5.6-Sol: 61,952→47,669 (−23%), $0.1377→$0.1151 (−16%), wall 25.3s→62.4s (Sol slower with JCR in 19/20); One Sol outlier 372.6s / 193 Jev calls; n=1 per cell; Not Harbor task-execution; Claude/Codex harnesses; compare mode; 50 scenarios bundled; 11 groups, 960 nodes, 11,360 items; 16 routing rounds per step; NiazMorshed2007/jcr ≠ skill-broker ≠ skillranker ≠ jev-sift ≠ jev-lens ≠ jevusher ≠ jev_select_capability; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33/#34; notes.md §116 + +Note: `research/notes.md` §116. Docs-only, rebased onto +merged #35 (`0189825` / §114 / 289–302 / #97) after +merged #34 Pages (`712223b`) and merged #31 (`35bec95`, +§113 / 273–288 / #96). Merged #36 owns §115 / 303–308 / +#98. Open #37 claims §117 / 315–321 / #100 (item overlap +315–316 is #37's remap). This fold keeps §116 / items +309–316 / batch #99. +**HARD RULE:** do not reopen or amend PR #23 or #24 or #25 or #26 or #27 or #28 or #29 or #30 or #31 or #32 or #33 or #34 or #35 or #36. +Never reopen merged #7–**#36**. Do **not** push onto +open #37. Do **not** re-fold §115 / §114 / §113 / §112 / §111 +/ cookbook beam §2 as this product. Skip Archer rewrite. +No invented metrics. Hunches labeled. Quote README + +site. Soft Noul ≠ hard safety. Augustus owns placement. +Quote live REST over watch. +`invented_signal: false`. + +- **NiazMorshed2007/jcr (PRIMARY).** NiazMorshed2007/jcr + (JavaScript MIT; **4★**; HEAD `138b3832`; README SHA + `2a49dbc1`; LICENSE SHA `46231303`; size **14850**; + topics ai-agents,jev,mcp; site + https://jcr.niazmorshed.dev). Jev Capability Resolver. + one tool to find documented deterministic commands in a nested capability tree. + returns context. **does not execute**. + skills = workflow+judgment; capabilities = individual operations. + format independent of Jev; proposed open standard exploration. + classify (Jev) → optional OpenAI decompose compound → beam search geometric mean of routing probs. + keep up to 3 paths ≥60% of best (JCR_BAND_RATIO 0.6). + ambiguity / no-match / depth-limit explicit. + soft scores ≠ hard gates. 0.6 band is application policy. + routing ≠ permission. docs ≠ authority to run. + sol-vs-opus5-20 *theirs*: 20 scenarios × 4 variants = 80 runs. + lookup+explain only, no execution. n=1 per cell. Not Harbor task-execution. + Claude Opus 5: agent input 108,585→15,819 (−85%), cost $0.3700→$0.1222 (−67%), wall 105.5s→77.7s. + Codex GPT-5.6-Sol: 61,952→47,669 (−23%), $0.1377→$0.1151 (−16%), wall 25.3s→62.4s (Sol slower with JCR in 19/20). + One Sol outlier 372.6s / 193 Jev calls. + Claude/Codex harnesses. compare mode. 50 scenarios bundled. + 11 groups, 960 nodes, 11,360 items. 16 routing rounds per step. + NiazMorshed2007/jcr ≠ skill-broker ≠ skillranker ≠ jev-sift ≠ jev-lens ≠ jevusher ≠ jev_select_capability. + +Pulse: jcr **4★** this pass (+1 vs fold-time 3★; stars ephemeral). +HEAD `138b3832`. README SHA `2a49dbc1`. GitHub search +`jcr jev` = one repo. Archer still promised_not_landed. +X MCP not used. `invented_signal: false`. + +Cross-repo addition: (jcr-a) retrieve-wide→decide→evidence-set +on a capability tree; (jcr-b) skills vs capability catalogs; +(jcr-c) 0.6 band as application policy; (jcr-d) routing ≠ +permission / docs ≠ authority; (jcr-e) beam geometric mean +as control; (jcr-f) VOI of context admission; (jcr-g) +measurement honesty / wall-time mixed / n=1; (jcr-h) +lookup+explain only. + + Derived from `analysis/evidence.csv` (169 repos, README-lines / langs / noul / choice / score / api / tests / threshold histogram) plus full deep reads. Status labels: **Contract** = verified from source/docs, **Empirical** = repo-published measurement, **Hypothesis**. ## dabit3/jev-experiments — 21 latency-first applications (Empirical, self-reported) diff --git a/research/changelog-hourly.md b/research/changelog-hourly.md index e1029bfc..36c54a0b 100644 --- a/research/changelog-hourly.md +++ b/research/changelog-hourly.md @@ -1,4 +1,4 @@ -# Hourly uniqueness dump (pre-0.4.0 + 0743 + 0843 + 0915) +# Hourly uniqueness dump (pre-0.4.0 + 0743 + 0843 + 0915 + 0920 jcr) This is the pre-0.4.0 `CHANGELOG.md` after hourly folds (#2–#30 / notes §44–§112) stuffed uniqueness locks into Keep-a-Changelog sections, plus @@ -15,6 +15,16 @@ the merged **0743 HIGH** dump (PR #31 / notes.md §113 / items 273–288 --- +## User-provided 0920 jcr HIGH (notes.md §116) + +- User-provided HIGH NiazMorshed2007/jcr (`research/notes.md` §116): + **Skip Archer rewrite.** Docs-only on a **fresh PR off latest main**. + **HARD RULE:** do not reopen or amend PR #23 or #24 or #25 or #26 or + #27 or #28 or #29 or #30 or #31 or #32 or #33 or #34 or #35 or #36. Do + **not** push onto open #37. This fold keeps `notes.md` §116 / items + 309–316 / batch #99. Merged #35 owns §114 / 289–302 / #97. + Merged #36 owns §115 / 303–308 / #98. + ## [Unreleased] hourly 0843 (not a SemVer bump) @@ -3722,3 +3732,5 @@ The dated passes below are how 0.1.0 was assembled. index + toolbox sweep. Hourly 0843 uniqueness lock: A hunch is a probability with a policy attached; { enter: 0.8, exit: 0.6 } is hysteresis; replay a policy change without inference; Decision models are providers, not the product; huncho ≠ Kungie/gut ≠ carldaws/hunch ≠ tpellet/hunch; pretrained Qwen2.5 base ECE 0.030 (0.5B) / 0.040 (7B); instruct 0.302 / 0.269; 70.9% → 70.0% mean conf 74.1% → 96.7%; temperature scaling still matches it in-distribution; No Jev API was called; Qwen2.5 ≠ Archer; Qwen/Qwen3.8-27B ≠ Archer; 学習済みモデル v0.1 は準備中です; bool AUROC 0.523; 先頭だと0件、末尾だと250件; 温度を渡さない場合、確率は較正されていません; このリポジトリには Jev を呼ぶコードが存在しません; g0runmezadam/what-is-jev IS tunahansahin897/what-is-jev (same GitHub id 1378007307); 947 repos scored; A 273 · B 302 · C 372; LLM rubric ≠ benches; Data as of 2026-09-20; HEAD 895b9498; README SHA 3ae98c56; 13 focused checks and one mutually exclusive outcome; Probabilities are advisory, not calibrated guarantees; omni-/ask-jev ≠ pedroknigge/mcp_jev; pd.cut bins by equal width while jeval bins by quantile; ECE 0.113 and ECE 0.076; jeval drift is not implemented yet; rlaope/jeval ≠ dayhaysoos/jevals; calibration does not compose; ECE has exactly zero statistical power to detect the failure mode that kills trajectories; 25–60× headline withdrawn; P(all-correct): 0.0071 vs 0.0001; TCE / AMS; Qwen 3.8 sparring ≠ Archer; Deferred Crispification; light_cutoff_applied_to_combination 0; recorded run, kinematic animation; BANKING77 Accuracy BERT-Base 93.02 Jev 79.90; Analyse jev calibration (NLL, ECE) backlog; BERT figures are published supervised references, not zero-shot; 档位措辞效应 分数极差中位 0.50、最大 1.32; 修好后对照组是 0.01; 不是 benchmark; 概率没做 calibration; ~1,430 API calls, about $0.15; xiaohuaxi/jev-study ≠ wjdjdakf17/jev-study ≠ baekenough/jev-study; AND: product (independence assumed and recorded in the trace); chat model's stated confidence is not calibrated; circuit-vl-4b ≠ Archer; Bring your own API key; vamsikrishna2421/jev-usecases ≠ whyashthakker/awesome-jev-use-cases; catalog ≠ endorsement; SemIf 2237★ (+30 vs §111 2207); jevlike 1049★ (+6 vs 1043); TypeLLM/TypeLLM 16★; AnotiaWang 98★; yibie/awesome-jev 520★ (+14 vs 506); Laya likes 861 (was 822); tracker likes 66 (+2 vs 64) lastModified UNCHANGED; Blackwood likes 2 gated manual; Archer still promised_not_landed; Hub archerhume/4rcherhume HTTP 401; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33; notes.md §114 + +User-provided 0920 jcr uniqueness lock: NiazMorshed2007/jcr MIT; site https://jcr.niazmorshed.dev; topics ai-agents,jev,mcp; **4★**; HEAD `138b3832`; README SHA `2a49dbc1`; LICENSE SHA `46231303`; size **14850**; Jev Capability Resolver; one tool to find documented deterministic commands in a nested capability tree; returns context; **does not execute**; skills = workflow+judgment; capabilities = individual operations; format independent of Jev; proposed open standard exploration; classify (Jev) → optional OpenAI decompose compound → beam search geometric mean of routing probs; keep up to 3 paths ≥60% of best (JCR_BAND_RATIO 0.6); ambiguity / no-match / depth-limit explicit; soft scores ≠ hard gates; 0.6 band is application policy; routing ≠ permission; docs ≠ authority to run; sol-vs-opus5-20 *theirs*: 20 scenarios × 4 variants = 80 runs; lookup+explain only, no execution; Claude Opus 5: agent input 108,585→15,819 (−85%), cost $0.3700→$0.1222 (−67%), wall 105.5s→77.7s; Codex GPT-5.6-Sol: 61,952→47,669 (−23%), $0.1377→$0.1151 (−16%), wall 25.3s→62.4s (Sol slower with JCR in 19/20); One Sol outlier 372.6s / 193 Jev calls; n=1 per cell; Not Harbor task-execution; Claude/Codex harnesses; compare mode; 50 scenarios bundled; 11 groups, 960 nodes, 11,360 items; 16 routing rounds per step; NiazMorshed2007/jcr ≠ skill-broker ≠ skillranker ≠ jev-sift ≠ jev-lens ≠ jevusher ≠ jev_select_capability; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33/#34; notes.md §116 diff --git a/research/notes.md b/research/notes.md index 45f601cc..98780a70 100644 --- a/research/notes.md +++ b/research/notes.md @@ -28650,3 +28650,259 @@ Parent merge only after **CLEAN** adversarial review do not extend. User-provided 0915 uniqueness lock: TianyuCodings/NanoJev unified-games-v1 densify; A 0.6B parallel decision model: states and questions in, complete probability distributions out. Zero output-token decoding.; One model, four games; ViZDoom Basic 128/128 vs Jev 56/128; Predict Position 27/128 vs Jev 11/128; Maze 225 attempts vs Jev 2738; Snake 30 food / 256 steps; held-out Maze 4/10 Snake 8/8 Basic 128/128 Predict 27/128; Untuned Qwen3-0.6B baseline; 18,760 questions per variant; 16,333 ViZDoom; 896 Predict Position expert episodes; hard_lr1e5; mix weights 1/3, 1/3, 1/6, 1/6; Hub C-Tianyu/NanoJev revision unified-games-v1 likes 58; dataset C-Tianyu/NanoJev-Data likes 5; HEAD 618cea6d906d54e128360786d12f703fff2b1245; 1289★ / 158 forks / size 64035; README SHA 4190093c64ee75b26e9726daa3b00cbcf6d3157a; MIT; caijinchun/nanojev-arena ≠ liao96312/jev-arena-nanojev ≠ zwliJay/jev-forge ≠ NanoJev; not TypeSafe Jev; open replica / specialist gameplay S1; soft scores ≠ hard gates; Game success ≠ calibrated Noul; local type boolean ≠ TypeSafe noul; invented_signal false; do not reopen or amend PR #31/#32/#33/#35; notes.md §115 + + +## 116. User-provided HIGH — NiazMorshed2007/jcr (2026-09-20 ~09:20 Boise / ~15:20 UTC) + +Docs-only on a **fresh PR off latest main** +(`cursor/fold-jcr-capability-resolver-19c4`), rebased +onto merged **#35** (`0189825`, hourly 0843 / +`notes.md` §114 / items 289–302 / batch #97) after +merged #34 Pages (`712223b`) and merged #31 (`35bec95`, +hourly 0743 / `notes.md` §113 / items 273–288 / +batch #96). Merged **#36** (NanoJev unified-games-v1) owns +**§115 / items 303–308 / batch #98**. Open **#37** +(SemIf densify) claims **§117 / items 315–321 / +batch #100** (item overlap 315–316 is #37's remap). +This fold keeps **`notes.md` §116 / items 309–316 / +batch #99**. Do **not** collide with merged §113–§115. + +**Never reopen merged** Augustus PR #7 / **#8** / +**#9** / **#10** / **#12** / **#13** / **#14** / +**#15** / **#16** / **#17** / **#18** / **#19** / +**#20** / **#21** / **#22** / **#23** / **#24** / +**#25** / **#26** / **#27** / **#28** / **#29** / +**#30** / **#31** / **#32** / **#33** / **#34**. +**HARD RULE:** do not reopen or amend PR #23 or #24 +or #25 or #26 or #27 or #28 or #29 or #30 or #31 or +#32 or #33 or #34 or #35 or #36. Do **not** push onto +open #37. Do **not** bump the 0.4.0 pin. Do +**not** merge from this agent. + +Do **not** re-fold §113 0743, §112 Merve, §111 +0646, cookbook hierarchy/beam (`notes.md` §2 K=3 +geometric mean) as if they were this product, or +skill-broker / skillranker / jev-sift / jev-lens / +jevusher / Codex `jev_select_capability` as a +second JCR. Quote **their** README and site. +Mark *theirs*. No invented metrics. Hunches labeled. +No wrappers, `npm ci` / `npm start` / `.env` / +`TYPESAFE_API_KEY` / `ANTHROPIC_API_KEY` / +`OPENAI_API_KEY` as recipes. `invented_signal: +false`. Skip Archer rewrite. X MCP not used this +pass. Do not invent tweets. + +Lane is Augustus: **mental models / architecture / +VOI of context admission / measurement honesty**. +Backend-agnostic categorization/scoring/decision +class; Jev is the hot **exemplar**, not the mandate. +Soft scores ≠ hard gates. Formal methods +**compose**: routing ≠ permission; docs ≠ authority +to run. 0.6 band is application policy, not a +proof. Ranking ≠ calibration theater. + +Live REST this pass (quote over watch; review +relock ~2026-09-20T15:51Z): +[NiazMorshed2007/jcr](https://github.com/NiazMorshed2007/jcr) +JavaScript MIT; **4★**; forks **0**; open issues **0**; +topics `ai-agents`, `jev`, `mcp`; homepage +https://jcr.niazmorshed.dev; created +2026-09-20T12:55:26Z; pushed 2026-09-20T15:00:32Z; +updated 2026-09-20T15:51:43Z; GitHub size +**14850**. HEAD `138b3832` ("Date the article and +link the author handle."). README SHA `2a49dbc1`. +LICENSE SHA `46231303` (Copyright (c) 2026 Niaz +Morshed). GitHub search `jcr jev` returned **one** +repo. Star counts **ephemeral** (fold-time 3★). + +Unique consecutive fragments (this fold): +NiazMorshed2007/jcr MIT; +site https://jcr.niazmorshed.dev; +topics ai-agents,jev,mcp; +**4★**; +HEAD `138b3832`; +README SHA `2a49dbc1`; +LICENSE SHA `46231303`; +size **14850**; +Jev Capability Resolver; +one tool to find documented deterministic commands in a nested capability tree; +returns context; +**does not execute**; +skills = workflow+judgment; +capabilities = individual operations; +format independent of Jev; +proposed open standard exploration; +classify (Jev) → optional OpenAI decompose compound → beam search geometric mean of routing probs; +keep up to 3 paths ≥60% of best (JCR_BAND_RATIO 0.6); +ambiguity / no-match / depth-limit explicit; +soft scores ≠ hard gates; +0.6 band is application policy; +routing ≠ permission; +docs ≠ authority to run; +sol-vs-opus5-20 *theirs*: 20 scenarios × 4 variants = 80 runs; +lookup+explain only, no execution; +Claude Opus 5: agent input 108,585→15,819 (−85%), cost $0.3700→$0.1222 (−67%), wall 105.5s→77.7s; +Codex GPT-5.6-Sol: 61,952→47,669 (−23%), $0.1377→$0.1151 (−16%), wall 25.3s→62.4s (Sol slower with JCR in 19/20); +One Sol outlier 372.6s / 193 Jev calls; +n=1 per cell; +Not Harbor task-execution; +Claude/Codex harnesses; +compare mode; +50 scenarios bundled; +11 groups, 960 nodes, 11,360 items; +16 routing rounds per step; +NiazMorshed2007/jcr ≠ skill-broker ≠ skillranker ≠ jev-sift ≠ jev-lens ≠ jevusher ≠ jev_select_capability; +do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33/#34; +notes.md §116. + +### How-to-apply (four placements / one product) + +When someone pastes JCR, extract the *placement +theses* and refuse execution, Harbor-task, and +hard-gating the 0.6 band. Same retrieve-wide → +decide → evidence-set sandwich as decision-native +RAG / jev-sift, on a **capability tree** rather +than files or RAG chunks. Cookbook beam (K=3, +geometric mean, `notes.md` §2) is the control +cousin, **not** a new species. + +1. **Retrieve-wide → decide → evidence-set** — + README *theirs*: "JCR gives an agent one tool to + find the documented commands it needs for a + task." "The main agent gets the selected + instructions while the search stays outside its + context." Site *theirs*: "I kept thinking about + how much reading agents do before finding the + command they need." Placement: VOI of context + admission. Cousin of jev-sift / jev-lens / + jevusher. NiazMorshed2007/jcr ≠ those. +2. **Skills vs capability catalogs** — README + *theirs*: "Skills can describe a workflow and + the judgment it needs. Capabilities can document + the individual operations used along the way. + The format is independent of Jev, and I would + like to explore whether it should become an + open standard with the community." Do not + collapse a skill pack into a command leaf. + skill-broker still owns **grants**; JCR owns + **lookup**. +3. **Soft scores ≠ hard gates; routing ≠ + permission; docs ≠ authority to run** — README + *theirs*: "JCR returns documentation. It does + not execute commands." "The commands being + documented can be deterministic. The + model-based choice of which command fits a + request is probabilistic." "The probabilities + help guide the search, though a match can still + be wrong." `JCR_BAND_RATIO` default 0.6 is + application policy, not a Harbor bar and not + permission to run the returned `PUT`/`POST`. +4. **Measurement honesty (wall-time mixed)** — + `sol-vs-opus5-20` *theirs*: 20 scenarios × 4 + variants = 80 runs; lookup+explain only; n=1 + per cell; not a Harbor task-execution bench. + Quote the mixed wall-time; do not steal −85% + tokens as a class constant. One Sol outlier + 372.6s / 193 Jev calls does not cancel + "Sol was still slower with JCR in 19 of 20." + +Receipts: user-linked 2026-09-20 ~09:20 Boise. +GitHub REST + README SHA + site fetch this pass +~2026-09-20T15:20Z; review relock **4★** +~2026-09-20T15:51Z. X MCP not used. + +### HIGH + +1. **[`NiazMorshed2007/jcr`](https://github.com/NiazMorshed2007/jcr)** + — NEW HIGH (productized capability-tree + resolver). MIT; **4★**. Site + https://jcr.niazmorshed.dev. + + **Quote README (*theirs*).** "JCR gives an + agent one tool to find the documented commands + it needs for a task. It uses Jev to search a + nested capability tree and return the context + attached to selected operations." "JCR returns + documentation. It does not execute commands. + The included harnesses also stop at explaining + the steps needed to carry out a task." "Skills + can describe a workflow and the judgment it + needs. Capabilities can document the individual + operations used along the way. The format is + independent of Jev." + + **Flow (*theirs*).** Classify (Jev single vs + compound) → optional OpenAI decompose → beam + search; path score = geometric mean of routing + probabilities; keep up to three paths whose + score is at least 60% of the best + (`JCR_BEAM_WIDTH` 3 / `JCR_BAND_RATIO` 0.6). + Ambiguity / no-match / depth-limit (16 routing + rounds per step) are explicit. Catalog *theirs*: + 11 groups, 960 nodes, 11,360 items; max node + depth 6; largest direct child set 175. + + **Bench sol-vs-opus5-20 (*theirs*; averages per + run; skills first, JCR second).** Claude Opus 5: + agent input 108,585→15,819 (−85%), cost + $0.3700→$0.1222 (−67%), wall 105.5s→77.7s + (median 86.6s→57.5s). Codex GPT-5.6-Sol: + 61,952→47,669 (−23%), $0.1377→$0.1151 (−16%), + wall 25.3s→62.4s (median 23.2s→45.4s). Sol + slower with JCR in 19/20. One Sol JCR run + 372.6s with two resolver calls and 193 Jev + calls. n=1 per cell. All 80 completed without + runtime errors. "Read these measurements as a + comparison of the included instruction-lookup + setups, rather than a task-execution + benchmark." Not Harbor. + + **How to apply:** one `resolve_capabilities` + tool; search stays outside the main agent; + returned `context` is evidence, not a grant. + Formal compose: credentials, side effects, + order, and verification stay in the harness. + Do **not** copy keys. Do **not** hard-gate 0.6. + Do **not** quote −85% as Harbor. + +### Skip Archer + +Text-only routing over names/descriptions of +direct children. Do not wait for Archer. Do not +send pixels. Compound split is a small language +model, not omni System One. + +### Not + +Not a TypeSafe how-to. Not an npm library +(root package is private *theirs*). Not a skill +loader. Not skill-broker grants. Not jev-sift +file/URL classify-first. Not jev-lens pre-send +views. Not jevusher token-admission. Not Codex +`jev_select_capability`. Not Harbor +task-execution. Do not reopen or amend PR +#23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33. +Do not re-fold cookbook beam as a new species. + +### Curated status + +Productized HIGH **folded** (capability-tree +lookup; docs-only; does not execute). Archer +still **promised_not_landed**. Cookbook beam / +retrieve-then-judge already in `notes.md` §2 / +§55 **not re-derived** as new products. +`invented_signal: false`. + +### Cross-links + +Cards: `mixed-architecture.md` (fail table); +`faq.md`; `mental-models.md` Apply 0920 jcr; +`validation.md`; `applied-mappings.md`; +`mappings.md` §4; `toolbox-mapping.md`; +`composition-algebra.md` items 309–316; +`question-design.md`; `methods-catalog.md`; +`formal-methods.md`; `agent-self-assessment.md`; +`judgment-class.md`. Hunches labeled. No wrapper. + + +User-provided 0920 jcr uniqueness lock: NiazMorshed2007/jcr MIT; site https://jcr.niazmorshed.dev; topics ai-agents,jev,mcp; **4★**; HEAD `138b3832`; README SHA `2a49dbc1`; LICENSE SHA `46231303`; size **14850**; Jev Capability Resolver; one tool to find documented deterministic commands in a nested capability tree; returns context; **does not execute**; skills = workflow+judgment; capabilities = individual operations; format independent of Jev; proposed open standard exploration; classify (Jev) → optional OpenAI decompose compound → beam search geometric mean of routing probs; keep up to 3 paths ≥60% of best (JCR_BAND_RATIO 0.6); ambiguity / no-match / depth-limit explicit; soft scores ≠ hard gates; 0.6 band is application policy; routing ≠ permission; docs ≠ authority to run; sol-vs-opus5-20 *theirs*: 20 scenarios × 4 variants = 80 runs; lookup+explain only, no execution; Claude Opus 5: agent input 108,585→15,819 (−85%), cost $0.3700→$0.1222 (−67%), wall 105.5s→77.7s; Codex GPT-5.6-Sol: 61,952→47,669 (−23%), $0.1377→$0.1151 (−16%), wall 25.3s→62.4s (Sol slower with JCR in 19/20); One Sol outlier 372.6s / 193 Jev calls; n=1 per cell; Not Harbor task-execution; Claude/Codex harnesses; compare mode; 50 scenarios bundled; 11 groups, 960 nodes, 11,360 items; 16 routing rounds per step; NiazMorshed2007/jcr ≠ skill-broker ≠ skillranker ≠ jev-sift ≠ jev-lens ≠ jevusher ≠ jev_select_capability; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33/#34; notes.md §116 diff --git a/research/refresh-log.md b/research/refresh-log.md index 035ea1d9..d7769380 100644 --- a/research/refresh-log.md +++ b/research/refresh-log.md @@ -3060,6 +3060,38 @@ likes 421 / 189; impressions 35498 / 9613. - User-provided 0806 uniqueness lock: people who compare Jev against GPT-5.6 has never fine-tuned BERTForXYZ for living and it shows; zero shot classifiers; scale them as much as decoder only models; many problems solved with LLMs could have been solved with them, it was a skill issue; opt for DeBERTa and ModernBERT ones; BERTForXYZ → DeBERTa → ModernBERT; Jev vs GPT-5.6 bakeoffs are a category error; encoder / ZS classifiers; institutional HF voice; quote *theirs*; do not invent accuracy numbers; softmax/ZS scores still ≠ calibrated Noul; soft scores ≠ hard gates; @mervenoyann; likes 421 / 189; impressions 35498 / 9613; multimodal image<>text ZS as perception front-end; hf:MoritzLaurer/deberta-v3-large-zeroshot-v2.0 likes 139; hf:MoritzLaurer/ModernBERT-large-zeroshot-v2.0 likes 72; Bart, bert, deberta, modernbert, these are all LLMs; Maziyar quoted; Jev is exemplar not the mandate; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29. + + +## 2026-09-20 ~15:20 UTC — user-provided HIGH NiazMorshed2007/jcr +- Docs-only on a **fresh PR off latest main** + (`cursor/fold-jcr-capability-resolver-19c4`) rebased onto + merged #35 (`0189825`, hourly 0843 / `notes.md` §114 / + items 289–302 / batch #97) after merged #34 Pages and + merged #31 (`35bec95`, §113 / 273–288 / #96). Merged #36 + owns §115 / 303–308 / #98; open #37 owns §117 / 315–321 / + #100. This fold keeps §116 / items 309–316 / batch #99. + **HARD RULE:** do not reopen or amend PR #23–#36. Do not + push onto #37. Do not bump 0.4.0. Do not merge. +- Live REST: NiazMorshed2007/jcr JavaScript MIT; **4★**; + forks 0; topics ai-agents,jev,mcp; site + https://jcr.niazmorshed.dev; HEAD `138b3832`; README SHA + `2a49dbc1`; LICENSE SHA `46231303`; size **14850**; + created 2026-09-20T12:55:26Z; pushed 2026-09-20T15:00:32Z; + updated 2026-09-20T15:20:06Z. GitHub search `jcr jev` = 1 repo. +- Quote README + site. X MCP not used. No wrappers / keys. + `invented_signal: false`. +- Folded into `notes.md` §116, SKILL.md (description uniqueness + + protocol triggers + mapping index), mental-models Apply 0920 jcr, + faq, mixed-architecture fail table, applied-mappings, validation, + judgment-class, composition-algebra items 309–316, toolbox-mapping, + methods-catalog, formal-methods, question-design, + agent-self-assessment, mappings, CHANGELOG, README, + docs/ecosystem, findings batch #99, sources.json, + changelog-hourly.md. Hunches labeled. No wrapper. + + +User-provided 0920 jcr uniqueness lock: NiazMorshed2007/jcr MIT; site https://jcr.niazmorshed.dev; topics ai-agents,jev,mcp; **4★**; HEAD `138b3832`; README SHA `2a49dbc1`; LICENSE SHA `46231303`; size **14850**; Jev Capability Resolver; one tool to find documented deterministic commands in a nested capability tree; returns context; **does not execute**; skills = workflow+judgment; capabilities = individual operations; format independent of Jev; proposed open standard exploration; classify (Jev) → optional OpenAI decompose compound → beam search geometric mean of routing probs; keep up to 3 paths ≥60% of best (JCR_BAND_RATIO 0.6); ambiguity / no-match / depth-limit explicit; soft scores ≠ hard gates; 0.6 band is application policy; routing ≠ permission; docs ≠ authority to run; sol-vs-opus5-20 *theirs*: 20 scenarios × 4 variants = 80 runs; lookup+explain only, no execution; Claude Opus 5: agent input 108,585→15,819 (−85%), cost $0.3700→$0.1222 (−67%), wall 105.5s→77.7s; Codex GPT-5.6-Sol: 61,952→47,669 (−23%), $0.1377→$0.1151 (−16%), wall 25.3s→62.4s (Sol slower with JCR in 19/20); One Sol outlier 372.6s / 193 Jev calls; n=1 per cell; Not Harbor task-execution; Claude/Codex harnesses; compare mode; 50 scenarios bundled; 11 groups, 960 nodes, 11,360 items; 16 routing rounds per step; NiazMorshed2007/jcr ≠ skill-broker ≠ skillranker ≠ jev-sift ≠ jev-lens ≠ jevusher ≠ jev_select_capability; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33/#34; notes.md §116 + ## 2026-09-20 ~14:43 UTC — hourly 0843 HIGH measurement / judgment - Docs + evaluator on a **fresh PR off main** (`cursor/hourly-0843-augustus-fold-220d`). Rebased onto latest `main` after merged #31 (0743, `notes.md` §113 / diff --git a/research/sources.json b/research/sources.json index bbe30793..2f974eb2 100644 --- a/research/sources.json +++ b/research/sources.json @@ -4933,6 +4933,18 @@ "title": "NanoJev demo (gated)", "url": "https://nanojev-dev.tianyuchen99.chatgpt.site/?autoplay=1", "note": "HTTP 401 gated this pass. Recordings local *theirs*. notes.md §115." + }, + { + "kind": "github", + "title": "NiazMorshed2007/jcr — Jev Capability Resolver", + "url": "https://github.com/NiazMorshed2007/jcr", + "note": "MIT; 4★ this pass; HEAD 138b3832; README SHA 2a49dbc1; LICENSE SHA 46231303; size 14850; topics ai-agents,jev,mcp. Docs-only fold notes.md §116. Does not execute. invented_signal false." + }, + { + "kind": "site", + "title": "How I’m using Jev in an agent harness — JCR", + "url": "https://jcr.niazmorshed.dev", + "note": "Author article 2026-09-20. Quote *theirs*. Complement to README. notes.md §116." } ] }