From f1e82778614168ac48ee5400849565da74fb56b4 Mon Sep 17 00:00:00 2001 From: Cursor Agent Date: Sun, 20 Sep 2026 17:26:09 +0000 Subject: [PATCH 1/2] Fold hourly 1049 HIGH: ggmlc GGUF serving, option-order, catalogs MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Census serving substrates of Laya (ggmlc GGUF is not llama.cpp), option-order vs exact-p Minesweeper measurements, adapter heads, and catalog/life placements. Allocates notes.md §120 / composition 353-368 / findings batch #103 past merged #41. Adds pick_by_id vs pick_second evaluator trap. uniqueness_gate now also requires the 1049 wall. Does not bump 0.5.0. Co-authored-by: Basit Mustafa <24601@users.noreply.github.com> --- .agents/skills/augustus/SKILL.md | 13 +- .../references/agent-self-assessment.md | 4 + .../augustus/references/applied-mappings.md | 4 + .../references/composition-algebra.md | 73 + .agents/skills/augustus/references/faq.md | 23 + .../augustus/references/formal-methods.md | 4 + .../augustus/references/formal-semi-formal.md | 4 + .../augustus/references/judgment-class.md | 4 + .../skills/augustus/references/mappings.md | 4 + .../augustus/references/mental-models.md | 14 + .../augustus/references/methods-catalog.md | 4 + .../augustus/references/mixed-architecture.md | 4 + .../augustus/references/question-design.md | 4 + .../augustus/references/toolbox-mapping.md | 4 + .../skills/augustus/references/validation.md | 4 + .../augustus/scripts/evaluate_decisions.py | 36 + .../augustus/scripts/uniqueness_gate.py | 32 +- CHANGELOG.md | 28 + README.md | 2 + docs/ecosystem.md | 3 + research/archive/findings.md | 44 + .../hourly/2026-09-20T16/meaning_bullets.json | 73 + .../2026-09-20T16/novel_high_this_run.json | 1540 +++++++++++++++++ .../hourly/2026-09-20T16/run_digest.json | 11 + research/changelog-hourly.md | 13 + research/notes.md | 355 ++++ research/refresh-log.md | 15 + research/sources.json | 144 ++ 28 files changed, 2457 insertions(+), 6 deletions(-) create mode 100644 research/archive/hourly/2026-09-20T16/meaning_bullets.json create mode 100644 research/archive/hourly/2026-09-20T16/novel_high_this_run.json create mode 100644 research/archive/hourly/2026-09-20T16/run_digest.json diff --git a/.agents/skills/augustus/SKILL.md b/.agents/skills/augustus/SKILL.md index c27a770b..6543f997 100644 --- a/.agents/skills/augustus/SKILL.md +++ b/.agents/skills/augustus/SKILL.md @@ -54,7 +54,7 @@ classical method you already trust, substitute it, classify the win "paraphrase brittleness", "allowlist then judge", "TOCTOU-of-Noul", "Jev inside the database / sqlite-jev", "Jev picks bitrate / join order / the model", "wait for Archer", "lint the request / missing - other", "training confronts Choice other / none-of-the-above", "soft AGENTS.md rules vs the linter", "screenshot Choice / omni System One", "extractive quotes / pointer not generator", "compaction summarize vs pointer", "encoder vs Jev compaction backend", "shadow-mode compaction rollout", "CI flaky-vs-real merge gate", "fail-open VOI wake/resume", "claim vs session evidence", "S1 indexer escalate-S2", "Harbor on/off routing", "fail-open vs fail-closed wake vs CI gate", "encoder vs Jev computer-use backend", "hybrid local decide + remote fill", "DONE vs verified success", "stdout prune vs session compaction", "OpenCode jev-pruner vs Claude jev-pruner", "zen-chat vs jev-zen Noul", "hard envelope then Noul prune", "Cua-S1 vs TypeSafe Jev", "plan vs execute dry-run", "specialist computer-use vs general agent", "local drop-in vs stub scorer", "route vs memory", "when does it hold / extractable from state", "decision model vs constrained LLM", "dual-process S1/S2", "combinatorial grid vs extractive", "uncalibrated local likelihoods", "decision-native RAG", "classify-first / read selectively", "living applied-mappings atlas / class patterns", "silence as safer / draft-gate heartbeat", "robotics text-state vs pixels", "verbatim ledger vs summary", "judgment as language primitive", "Stagehand extract pick-and-copy", "harness observe-score-act vs demo loop", "public judgment wall / six parallel questions", "meaning-search without embeddings", "attention ≠ correctness", "skills→oxlint / AST prove ∩ remainder", "session-sticky first-prompt routing", "measured RAG rerank vs generative rerank", "capability kernel / secrets never in the agent", "Jev is SENSOR not policy", "type-safe ≠ correct", "typed control plane around DSPy", "native vs verbalized confidence", "engine owns truth / Jev owns judgment", "human-confirmed kill gate", "train specialist vs few-shot hosted", "decide→policy→LLM leftover", "Noul 0.5 cannot-tell never rounded", "calibration ≠ sortable / ORDER BY", "pairwise inversion / Score ordinality / two-decimal ties", "wire-compat GLiFormer /v1/systemone", "class-backend economics", "loopback gateway hosted + local", "do not distill Jev as teacher", "active-learning triage", "evidence-packet explorer", "meaning-grep AND/OR/NOT", "closed-vote-only / no planner LLM", "Jev vs PCD Harbor", "PCD O(1) ≠ Noul", "host-owned handlers × System One", "OMP/pi fail-open gate", "permission vs probability / operator owns thresholds", "judgment ≠ permission / Jev never grants access", "eval integrity / instrument not score", "constrained optimizer + S1 features / never sole hot-path gate", "privilege ≠ verdict / effect contracts not tokens", "attention filter / VOI for human review / never blocks / never green unless sure", "measurement owns endorsement / evidence-gated question packs", "Jev supplies evidence / code owns authority", "ranking ≠ calibration / never hard-threshold raw p as frequency", "hot-click CU / indexed element table", "Jev judges relevance / code decides structure", "local rules first then remainder / never auto-train on own hides", "combinators / System One as control plane", "receipts not leaderboard / type-safe ≠ correct jaggedness", "VOI over skill library / skillranker abstention", "OOD calibration / AUC ≠ ECE", "Jev vs thinking-budget small models", "turnstile / replayable evidence≠authority", "MLX one-pass schema→JSON / Apple Silicon replica economics", "memory leases ended by new evidence", "never confidently wrong / TLA+ compose / escalate instead of hard-gate", "no seal no advance / coverage ledger / mint ≠ product brain", "skill-broker sibling / judgment ≠ permission", "sureness bands / max_prob is generous", "JevBench / calibration not in Main Score", "CI typed gate before expensive review", "Codex MCP host adapter", "judgment as attention redirect / jev-preflight", "compress-before-first-send / dizk jev-lens", "tools≠use / SessionStart over hoping", "observational memory / pi-om keep-kind", "open-Jev class / openvons / JevPick", "physical-world System One / HA-Jev / not for locks", "judgment outside the store / jevql", "landed-script trust / headless≠auto-approve", "digital-design combinators / extended five", "VOI cache admission / same-intent skip LLM", "BM25 vs Jev skill routing Harbor harness", "zeroshot vs BERT / contamination DiD", "typed escalate continue abort baton / inverted loop", "worth-your-attention VOI / ThinkyMiner Winnow", "Jev WHETHER Python HOW LLM WHAT", "conflict vs ignorance / named Choice escape", "Playwright executes Jev chooses", "OpenJev /v1/decide not drop-in", "SemIf wire-compat runoff; SemIf rename densify / MLX backend / 5.21× systems≠semantic / Softmax ≠ Noul (`notes.md` §117)", "decision-as-memory flywheel", "record/replay CI / jevassert", "failure-finding arena / jevarena ≠ jev-arena", "BBQ not a bias cert", "decider≠executor", "sentence-as-rule lint / jevlint", "sentence-as-rule lint / jev-lint is jevlint rename", "VOI hunk prune", "whole-repo intent VERIFIED/VIOLATION/UNKNOWN", "GLiNER2 spec ≠ replica", "open replica substrates / grande / laya-jolt / JEV-CPU", "ONNX local-jev not equivalent", "persist constraints across compaction / pi-heed", "calibration+cost first-class gates", "Harbor-shaped Jev vs SGR LLM-as-judge / jev-judge-bench ≠ jevarena ≠ jevbench", "hand no-text steps / jev-use / Vercel drops confidence", "Pi System-One control plane / pi-jev-control", "never free-generates / jev-gpt tree of Choices", "OpenRouter recipe atlas / samples not benches", "personal history feed / jevfeed / no social graph", "competing NAR claims / dual-channel ECE / openJev-verdict ≠ OpenJev", "empty compaction-proxy skip / IPECTER", "throughput ≠ latency / like-for-like ECE", "1-token logprob endpoint ≠ Noul / coverage ≠ correctness", "open replica engine / jevinf", "unofficial Elixir SDK ≠ OTP peer", "jevex n=16 files-to-read VOI", "commit pre-review attention≠verdict / middle band", "Hermes plugin is Agnes not TypeSafe", "pi-jev-compact ≠ pi-jev-compaction", "empty Codex-proxy skip / IPECTER runway", "decision-native inbox / mailordinal", "unofficial jev-cli not ready / ≠ jevql", "laya-multilingual / English checkpoint confident-wrong OOD", "schema-scorer peaked ranking ≠ calibration", "HF 401 / GitHub 404 Hub-only", "productized System One HTTP / classifier.dev", "escalate-under-threshold / smart tier / multi-label ignores", "silent FALLBACK / granite 0.546 vs advertised 0.800", "vs_jev tracked JSON / read eval/README", "choxos/jev-reviewer ≠ egma-ai / systematic-review pointer", "two-pass Choice+Noul evidence extraction", "not-found is an answer", "human check as productized judgment", "githubnext/localjev ≠ kunchenguid/local-jev", "wire-compat ≠ logit-equiv / prompted JSON ≠ structured read", "self-reported probs / entropy confidence", "GitHub Next local /v1/systemone", "LM Studio runner gap / structured-read primitives", "NandhaKishorM/laya packaging ≠ Hub-only / Router script-before-p", "post-T ECE ≠ raw ECE / Banking77 token-budget", "0.85 still soft / not TypeSafe drop-in", "external census ≠ scored bake-off", "GLiNER2+routers class-boundary", "incomplete openjev census vs watch", "Harbor honesty watch / silent fallback", "JevBench v1.2 geometric mean / cal ON rank / weight sensitivity", "option-order 72→21 / instruction models in the class table", "self-host latency ×2 assumption / est. costs", "Laya absent is a gap not a named exclusion", "Qwen3.8 27B ≠ Archer", "hourly already-folded watch / apply-the-five / skip thin noise", "hard-gate Noul as PR/quality gate is soundness theater", "S1 never stalls waiting / S2 one-use advisory", "Local controller ≠ githubnext/localjev", "purple telemetry = consumed not arrived", "seed = geometry not async replay", "20% starting gate still soft", "no pixels to either provider", "OCR+AX observe-score-act / typesafe-computer-use", "never send screenshot to frontier for the decision", "overlapping CU options = false low confidence", "split kind/item/site", "155× one-screenshot ≠ Harbor taskset", "decision ≠ answer-reader capture", "ASR observe-score-act / jev-voice-browser", "partial-speech VOI / free-text waits", "spoken confirm ≠ hard auth", "numbered overlay without another model", "wrap-as-execution / AgentGhost ALLOW ASK DENY", "rules first then Jev remainder / ASK throws / fail-closed", "reddpy/AgentGhost ≠ jwen5419807/agentghost ≠ vventirozos", "JP genre atlas / studio_yebisu / stars ephemeral ≠ eval", "Jev Clearly Explained / akshay_pachaar / LLM hammer", "schema-safe ≠ correct / 200× 400× TypeSafe ceiling", "questions-as-code / shadow first / not a TypeSafe how-to", "proposition ≠ embedding / contrast-set", "boolean composition of soft Nouls / AND OR NOT", "uehaj/jev-semgrep ≠ semgrep.dev", "meaning-grep dedicated fold / not a gate", "decision-validated UI / Jev never authors text", "decision-as-assert / jevtest ambiguous band", "typed decisions drive UI / jev2ui", "hybrid S1 closed verb menu / anima3", "pointer-not-generator search / JevFind", "jev-frontier-bench ≠ frontier-100", "product bakeoff ≠ architecture duel / GLiClass", "four engines same questions / majority floor", "authorship named escape / not evidence", "ha-switchboard HA remains execution", "n8n classify/route/score / Low Confidence", "fast-jev-compaction-pi ≠ pi-jev-compact ≠ pi-jev-compaction", "jevloop full-distribution optimizer / no LLM in the loop", "laya-vision SmolVLM / score untrained", "Cerebellum-2B /v1/decide ≠ TypeSafe / wire-compat vs agent-routing", "laya-grounded not drop-in / Platt not temperature", "GestaltLabs/Jeff-1 ≠ logan-markewich/jeff / acc vs ECE n=9730", "stanley-code empty findings ≠ approval / human promote", "findme ≠ JevFind / NL memory beam-search FS", "jevsubrouter price workers not conversation / counts ≠ dollars", "feelings .feels() default 0.5 is Noul-0.5-never-rounded / ≠ hunch ≠ Probably", "apa-agent-harness ≠ AntonioCoppe/jev-harness / unpublished npm", "grok-bot-jev skill cannot force a bot that ignores it / A/B proxies not tokens", "Essentiel-Jev never authority / human every action", "enzo-mcp independently falsifiable claims / ≠ jev-sift", "pigeonhole OTHER skip / decision-as-filing", "jev-reliability Nothing about accuracy", "clduab11/jev-test ≠ realZachi/jevtest / Nothing runs yet", "jev-rag-benchmark Jev wins is not an assumption", "dairui1/jev-lab ≠ BrendanH18/jev-lab", "jevmail gmail.readonly / mailjay archive/trash", "ZHUBoer/ego-jev reserved __none__", "runWorkflow completed ≠ success", "jsort scores are relative", "Noul not Choice for scale", "groundedness-judge-bench native vs schema-guided", "implicit_true included in yes", "jev_playground 0 promotions", "routing-backtest 0.0447%", "yuyang2230/jev-agent-skill jev-1.13-free", "jev-techstack-classifier stack_config.json", "s1_ruby collapse late", "undecided? abstain", "2389-research/judgement license null", "confidence ≠ winner p", "typesafeai-sdk-community not a new species", "tpellet/hunch exit 3", "never-execute list", "jev-file-search scores not calibrated accuracy", "jev-linkmap Jev never sees S2 prose", "muhammedilyasy/jev-mail metadata only", "tidy none-of-folders stay", "tab-bouncer pinned/audio/current never closed", "lkclean Show fail-open", "jev-yt-time-saver Show anyway", "ORIGIN pause-if-no-Jev", "validResponse sums-to-1", "jev-crawlers risk bands never raw boolean", "jevbrain AUTO_ACT is not a Noul", "judgekit YAML classify/score/route/verify", "typed-judge-kit verdict-in-code", "alsoleg89/decide packing VOI", "0.8 ≠ 80% accuracy", "Jev-Calibration Platt ECE 0.117→0.052", "jev-calibration-arena never acts", "ctmx/openrouter-jev-mcp Decision-as-Plugin", "FrancoisChastel/jev-code ≠ npm jev-code", "claudecode-jev-marketplace fail-open not hot path", "pedroknigge/mcp_jev packs not ask_jev", "cyrusasco/typesafe-mcp noul deadband 0.35–0.65", "codaaiteam/jev-skill jevtypesafeai.com ≠ TypeSafe", "hermes-switchyard ≠ hermes-jev-router ≠ hermes-plugin-jev", "nanoprune 2.8MB ECE 2.58%", "smartdio/jev-browser-agent ≠ ZHUBoer/ego-jev", "Dakai/omp-jev-web DONE ≠ proof", "hari007sh/jev ≠ dannote/jev", "0thernet/system-one-skills deterministic verify", "typed-gate band [0.40,0.60] is refusal", "pi-jev-gate fail-closed; choice is the verdict", "Foq ~25ms/2.2GB local", "rev prefill-only + HF jev-0.5b", "robfrase/jev planning memo", "typesafe_agent_gates 27/27 / 31/31", "EpicEric/safe-sh static remainder", "pastepilot Confirm before act", "Jev-Reranker live Jev not yet measured", "sessionwise opt-in relevance", "jev-search pointer sieve", "400ms Salesforce WebMCP", "typesafe-scheduler-diagnostics advisory", "droidjev screenshot-free", "Tewoto1 jevcu planner still writes", "ha-conversation-jev Jev→Grok", "dsh-jev can only gate", "jev-classification-benchmark specified not run", "jev-luna-pagerduty p≥0.50", "meldltd/meldecision laya-go ONNX", "laya-doom never pixels", "logixism/laya-api empty README", "akpsahan/laya ≠ Archer", "choxos/jevchess engine owns truth", "jev-drive sim not AV", "story-arc Jev never authors", "jev-hs-assistant HS6", "golergka/jev-plays-starcraft-2 UI-verified ≠ API Victory", "awesome-jev-use-cases catalog", "Nibir1/typesafe-go ≠ official", "fingerprint after redact", "recall vs decide", "publish fingerprints+answers", "CI replay as Harbor cousin", "Cache hit ≠ correctness", "hyperspaceai/jevcache ≠ kushals256/jevcache", "human labels only", "score never auto-accepts", "production capture flywheel", "sutro-sh/jev-align ≠ caiovicentino/jev-align", "guidance ≠ hook", "catalysts ≠ summaries", "compile-time System One", "unofficial ≠ TypeSafe", "format_version modernbert-jev/1", "Argos1111/jev_local ≠ us/jev-local ≠ kunchenguid/local-jev", "LFM default ≠ ModernBERT backend", "Nemotron ≠ TypeSafe Jev", "not a calibrated replacement", "djev-dev complements djev-spark", "images as Choice options", "Laya essay numbers *theirs*", "Router/OOD confidence", "hosted bootstrap ≠ silent TypeSafe", "difficulty + policy thresholds + JSONL trace", "jev-codex-pilot model + reasoning depth", "keep/shadow/hybrid/reject", "quarry evidence projection", "Frank-ZY-Dou/awesome-jev robotics/3D/control", "one-dollar-tahoe TypeSafe Jev defense eval", "jevguard calibrator/cache/escape", "jev-ci-selector CI shadow mode", "llama-jev llama.cpp replica", "petercr/jev-orchestrator ≠ FleeexCorp/jev-orchestrator", "seb4ez/jevguard ≠ AseemPrasad/JevGuard ≠ pablozr/JevGuard", "webNeat/llama-jev ≠ WiktorB2004/llama-index-jev", "OpenCode jev-pruner context sieve", "observe→score-candidates→prune", "jev-zen / jev-1.13-free", "zen-chat ≠ Noul", "fail-open original", "keepScore >0.1 floor", "host port of tamaratran/jev-pruner", "indiejoseph/opencode-jev-pruner ≠ nrdz-labs/fast-jev-opencode", "jev-webagent-bench empty stub", "Kiln-AI/jev_jsonschema noul_threshold 0.5", "NSStudent/JevSwiftSDK unofficial", "GLiNER2 native Apple path", "unofficial Swift/Core ML GLiNER 2.5-small", "entity spans + confidence", "not Choice/Score/Noul", "not TypeSafe", "label descriptions as schema", "on-device ANE economics", "honesty locks", "shershah1024/gliner-native-runtime ≠ Fastino", "≠ gliner25-compaction ≠ gliner2-ultrafast ≠ Eran-BA/Jev_from_GLiNER2 ≠ NSStudent/JevSwiftSDK ≠ jevmlx", "default threshold 0.1 still soft", "soft Noul ≠ hard safety", "Decision Graph Protocol frame→assess→commit", "app retains permissions/effects", "Jev-first assessor-neutral", "guarded commit / receipt/next frame", "assessment batching", "hard-gating DGP as safety theater", "numerous-com/dgp ≠ TypeSafe official", "jegrep calibrated path+range Nouls", "no embeddings/index/daemon", "~$0.01–0.03 typical", "agent --json", "can1357/jegrep ≠ Bentlybro/jevgrep ≠ uehaj/jev-semgrep", "Archer-arch fidelity", "kev family OOD 0.76–0.77 vs Jev 0.86", "block-causal isolation", "pointer/readout CE-trained", "/v1/systemone drop-in", "replica honesty", "cost-sensitive decision theory × System One probabilities → control flow", "thresholds derived from costs not hard-coded", "YES / NO / UNSURE from cost_false_yes / cost_false_no / cost_human", "auto-batching same-object questions", "Kungie/gut ≠ tpellet/hunch ≠ carldaws/hunch", "judgment vs generation", "deterministic execution after probabilistic judgment", "exactly one app-owned callback", "explicit uncertain branch", "Illusion47586/judge ≠ lexingtonhibiki/judgekit ≠ Ascurse/typed-judge-kit", "variable-N option scoring as the trainable object", "dynamic candidate bags not fixed label sets", "zwliJay/jev-forge ≠ NanoJev", "open replica economics / latency vs closed Jev", "NAR local drop-in", "wfzyx/von late-catch HIGH", "competing NAR claims / replica honesty", "typed judgments vs chat judges on guardrailing", "ishaannk/llm-vs-jev cross-note only", "deeper integrity fold is rh-guard", "nothing wins outright", "can be argued out of guarding"", "Jev IS the if-statement", "judgments/probabilities drive branches", "text model only writes prose", "interpreter owns variables/loops/budgets/replay", "otherwise maybe / confidence gate", "chaos samples after the gate", "southpolesteve/probably ≠ carldaws/hunch ≠ feelings ≠ Kungie/gut ≠ Illusion47586/judge ≠ tidymodels/probably", "133★ / forks 10 live", "build calibrated classifiers from human feedback", "retrieve by relevance not resemblance", "one calibrated yes/no per memory in one request", "pointer mode 17/18 19/20 *theirs*", "embedding resemblance misses the allergy", "samdotmak/jev-recall ≠ jev-search ≠ jev-sift ≠ carryforward ≠ chopratejas/invalidate", "memory leases ended by new evidence", "six Nouls then fixed rules in code", "0 of 157 false invalidations", "questions/plans/directives are not evidence", "unsure → review queue", "host keeps the store", "name↔body / comment truth / test-claims", "mizchi/jev-lint is mizchi/jevlint rename", "no shipped rule has severity error", "~1 in 5 findings wrong *theirs*", "mizchi/jev-lint ≠ huntedman/JevLint ≠ MichitoSugawara/jev-lint", "JSON Schema → typed JSON via Jev", "noul_threshold 0.5 decoder not a proof", "IncompatibleSchemaError lists every bad property", "on-device Laya CoreML ANE", "~5 ms P50 short decisions", "189/189 FP16 checkpoint parity", "10× not achieved", "mizorewww/laya-coreml ≠ gliner-native-runtime ≠ jevmlx ≠ NandhaKishorM/laya", "softmax over allowed tokens ≠ Noul", "question-first cache", "Micha0827/snapjudge ≠ githubnext/localjev ≠ jevmlx ≠ cendress/SnapJudge", "Jev-first Pi agent loop", "slow-LLM fallback", "explicit action menu / CandidateSource unimplemented", "62 tests wiring not quality", "direwolfiy/JevPi ≠ standardagents/jevpilot ≠ pi-jev-control", "resume-screening bias audit methodology", "name×resume factorial independent Nouls", "callback determined by resume quality", "mean-probability name gaps operationally negligible", "natemoo-re/bias-bench ≠ BBQ", "Plan/PRD panel → code-owned pass|review|block", "cheerleading out of scope", "austindixson/planalyzer ≠ single-goodness Noul", "cost-aware multi-model routing/escalation", "decide vs do", "successful-task cost", "cannacre8ive/switchboard-ai ≠ ha-switchboard ≠ hermes-switchyard", "frozen-protocol zero-shot bench", "TypeSafe Jev vs PrismNLI vs Laya", "contamination caveat", "elcronos/jev-vs-open-decision-models ≠ JevBench ≠ DMB", "context-window admission control", "VOI gate which tokens are worth the expensive model", "fail polarity per lens", "on small inputs lenses lose money", "cvsgireesh/jevusher ≠ jev-sift ≠ winnow", "typed decision control plane", "receipt ≠ authorization", "historical-v0 zero retained cases", "MokiMeow/jev-fabric ≠ jev-forge ≠ dgp", "live 15-dim typed rubric re-score per pause", "scoring economics exemplar", "OpenJev/Codiv ≠ TypeSafe hosted", "jose-troche/live-rubric ~$0.000004 desc / ~$0.000006 README", "adversarial pre-registered Jev eval", "28 predictions before data", "123,805 requests", "confidence does not track ignorance", "polite injection 65% / crude 0%", "willkelly/jev-evaluation ≠ jevals ≠ jev-baselines-eval", "provider-neutral Elixir/BEAM Noul/Choice/Score SDK", "class infrastructure", "nshkrdotcom/system_one_sdk ≠ typesafe_sdk ≠ dannote/jev", "question-linting of Jev questions themselves", "nine jaggedness rules, no API key, no labelled data", "static lint ≠ measured separation", "yodablocks/jevq ≠ tenbin ≠ JevLint ≠ commitjev", "open-weights Laya as class exemplar (binding)", "Nx/Bumblebee runtime", "host chooses backend", "ChristianAlexander/laya_ex ≠ system_one_sdk ≠ dannote/jev ≠ NandhaKishorM/laya", "on-chain/edge Laya deploy", "parity_verified stays false", "model output never grants Tx", "humandebri/IC-Laya ≠ laya_ex", "auditable weekend replica", "Jev outputs never used for training", "soft human-vote distributions", "unpaired 0.577 vs 0.727", "agilabs-ai/jev48 ≠ JevBench ≠ Mapika/decider", "adversarial dual-judge / framing attack surface", "comparative framing is the usable judgment", "prior injection crowds out evidence", "copyleftdev/ember ≠ ember.js", "Laya specialist fine-tune pipeline", "training still GPU-pending", "PIXELZX0/XERON ≠ convaiinnovations/laya", "Hub Laya replica drop", "daliborsb/laya ≠ convaiinnovations/laya ≠ NandhaKishorM/laya", "System One student distillation corpus", "gold is programmatic", "teacher is closed-API clone", "do not distill Jev as teacher of record", "MagaBitmex/jev-4b-distill-data ≠ missing student checkpoint", "non-LLM VIN System One", "planning depth not chat", "lewislululu/jevon ≠ douglance/jevon", "source-bound evidence checks", "local quote mismatch needs no API", "exit 0 ≠ claim truth", "WaynezProg/jev-kit ≠ jonathanavis96/jev-kit (Airlock) ≠ jev-use ≠ jev-mcp", "independent System One evidence catalog", "scores not one leaderboard", "no external record currently reproduced", "TokenTrim no-Jev matched hybrid 62.4%", "reachjalil/system-one-bench ≠ mallahyari/system-one-benchmark", "21 tasks · 134 items · 208 questions", "scenes from public GitHub contracts, not production logs", "SivletLabs/jev-eval ≠ willkelly/jev-evaluation ≠ 4esv/jev-eval ≠ xxkuboxx/jev-eval ≠ onlyoneaman/jev-eval ≠ dayhaysoos/jevals", "option isolation (sibling-blind)", "permutation-equivariant", "Hub OWNER not published", "nafisazizir/hev ≠ jaredpalmer/kev", "frozen local LLM logits, no trained decision head", "residual-head 9,222-param decreased 73/96→67/96", "confidence = 1−normalized entropy, not P(correct)", "yuki-oshio/mini-jev ≠ r-ms/mini-jev", "Jev classifier as autoregressive next-token predictor", "ChatJev-style soundness theater", "erik-dunteman/ChatJev ≠ dannote/jev ≠ jev-gpt", "calibrated decision head × AlphaProof value head", "implementation-layer isomorphism, semantic difference", "timeout = censoring", "do not launder Noul as proof", "parallel rank-prediction vs serial selection", "independent questions can conflict", "zzzzzec/jevsort ≠ keltokhy/jsort", "curated open System One ecosystem catalog", "rupeshpoojary9/awesome-open-system-one ≠ AnotiaWang/awesome-jev", "arXiv paper radar with Jev relevance scoring", "ranking ≠ calibration / 0.5 still soft", "fail-open failed evals not marked seen", "train calibrated ~27M from scratch", "typed Q→prob dist / one forward pass / no LLM decode", "hyusi2003/MiniSystemOne ≠ Colvin0315/MiniSystemOne", "description-only stub / size 5", "ESCI hard probe fails four of six", "jev_bool ECE 0.242 inversion 0.255", "do not re-fold §60 six-gates as new", "jobbyjev one-request-per-company from batch-size result", "find/design/evaluate TypeSafe Jev decision loops", "karanb192/jev-architect ≠ samtay32/jev-system-architect", "Jairik/jev-distiller size 1", "distill-Jev UI stub / do not distill Jev as teacher of record", "post-launch scored use-case map / Jev self-scores then human curation", "licensedsaucer9-web/jev-opportunities", "Jev-inize a use case into classifier/router", "gavinHuang/jevinize → simple-jev not TypeSafe", "featherless-ai/simple-jev", "compare saved decisions / same label can still change the branch", "VihaanAgarwal/jev-diff ≠ Saik0s/diffusiongemma-jev-macos", "not tested with a live Jev API key", "constrained logprob + temp/Platt ≠ Noul", "OpenJevPro pastes openjev-sglang JevBench as own", "zhangcy122/OpenJevPro ≠ IamBusy/OpenJev ≠ ekzhang/openjev-sglang", "PolyForm Noncommercial", "SmolLM-135M / sub-70ms / 0 output tokens", "demo P(True) 0.5052 / Choice conf 0.2872 / Score conf 0.0055", "README claims MIT / GitHub license null / no LICENSE file", "patelvishwa112/jev-system-one-rlcd ≠ arnabgho/rlcd-lite ≠ blackwood-rlcd", "source-backed Awesome Jev radar / 306+ commit-pinned", "logicrw/awesome-jev-projects ≠ AnotiaWang/awesome-jev ≠ yibie/awesome-jev ≠ cobanov/awesome-jev ≠ rupeshpoojary9/awesome-open-system-one", "auto GitHub sync / Issue-only submissions", "hashed n-gram encoder / rival-aware attention", "olanotolu/jevbetter vs jevlike starter", "synthetic hard menus top-1 0.916 vs 0.873 / ECE 0.0182 vs 0.0367 / 40 vs 4608 menus/sec", "shuffled-context control 0.335", "Turn any open LLM into System-One Jev", "uspraveen/Jevify ≠ Mintzs/jevify ≠ gulagala001/jevify", "Jevify-any-LLM architecture probe", "description-only stub / size 0", "Train encoder-only calibrated decision models from a task sentence", "Exu is a toolkit, not a method", "strictly proper scoring rule", "Pre-alpha", "Ruivalim/exu-base", "scratch-trained calibrated decision model", "typed Q → probability dists", "Colvin0315/MiniSystemOne ≠ hyusi2003/MiniSystemOne", "no published weights download URL", "90.5 seconds / 29.2% pipeline evidence", "p_i/p_j independent of other candidates", "Recipe for calibrated decision models — small model out", "init → synth → train → eval → serve", "91.1 % / ECE 0.022 *theirs*", "Jev zero-shot 75.1", "scienthoon/luce", "Put Jev's three headline claims on trial", "0.5B local GPU", "46x speedup / accuracy identical", "ECE 0.624 sentiment catastrophe", "bigger model worse calibration", "RichardoMrMu/jev-mini ≠ yuki-oshio/mini-jev ≠ r-ms/mini-jev", "System-1 decision engine for local LLMs", "structured choices only", "JSON parse of generated text ≠ Noul", "TypefAI JEV / Journal Entry Voucher", "tapsin/jev-local ≠ us/jev-local ≠ Argos1111/jev_local", "Jev 1.13 reward-model eval across 8 benchmark tracks", "40,940 examples / 0 API errors", "RewardBench v1 92.58%", "Precise IF 50.63%", "goya4140/jev-reward-model-evaluation", "Scaffolding in progress", "Jev vs LLM support-ticket routing", "static + live decision bench", "TypeSafe's own published benchmark", "illustrative simulations, not live API calls", "JevBench v1 — smart/cheap/fast/reliable", "I/C/S/K 25% geometric mean", "classifier.dev fast tier 84.8 is Jev behind its own API", "do not re-fold §78 v1.2 board as new", "Laya (421M) 70.1 now on board", "Zero-shot/few-shot LLM routing", "hard budget filter before Jev", "Jev never asked to perform budget arithmetic", "Jev judges the next state, XState enforces transitions", "simulation uses synthetic keyword fixtures", "catalog gravity", "v-modal/awesome-jev-tools", "★339 live REST", "curation is not endorsement", "crawler-maintained directory", "Daily GitHub + npm sweep, human-merged", "RadRebelSam/awesome-jev ≠ AnotiaWang ≠ yibie ≠ cobanov ≠ logicrw ≠ v-modal", "HF peft SPLADE/BGE reranker", "rdxtremity/jev-reranking ≠ carlaiau/jev-reranking", "query-side encoders, not a Jev replica", "ONNX System One Qwen3.5-4B scorer", "source:pngwn/system-one-qwen3.5-4b-scorer", "CC-BY-NC-4.0", "temperature 1.75", "transformers.js AutoModel cannot load this graph", "Consistency benchmark Space", "This Space contains no benchmark result yet", "12-case plumbing fixture", "Benchmark-driven Jev router and judge", "cheap alone is not success", "Jev does not write, sum prices, or claim accuracy %", "Sol 94.2 / Luna 83.9 / Jev path 89.7", "19.2% Sol / 62.3% cost save / 4.5pp miss of 2pp non-inferiority", "p50 latency worse than Sol due to routing overhead", "erendikmenn/jev-llm-router-benchmark ≠ jev-rag-benchmark ≠ ryantsai/jev-llm-router", "Express + node:sqlite", "mock and Jev decision engines", "previous_ticket_count >= 3 is code", "MIN_CONFIDENCE 0.6 still soft", "substring false positives", "aesaganda/jev-ticket-router ≠ SarathChandraBellam/jev-vs-llm-ticket-router", "Universal Figure & Diagram Router", "confidence ≥ 0.85 hard-gate is theater", "generative AI banned from scientific plots", "six visual branches", "hoangngochuong24947-gif/jev-figure-router", "human-labeled (state, question, label)", "166,054 rows / 22 configs", "soft_label for human uncertainty", "Praveenrajus/jev-bench ≠ fstandhartinger/jevbench", "ternary bonsai System One GGUF", "openjev's mechanism, Bonsai's weights", "Hub does not ship weights", "100/100 easy T/F is not Harbor", "label_mass ≠ correctness", "stock llama.cpp Q2_0 silently gibberish", "NicolaiMTLassen/open-bonzi-jev ≠ NicolaiLassen", "transformers.js DeBERTa ONNX", "source:com-kotobalabs/open-jev-deberta-v3-large", "temperature 1.05", "AutoModel from_pretrained works", "onnx-community/open-jev-deberta-v3-large-ONNX ≠ system-one-qwen3.5-4b-scorer-ONNX", "107★ densify", "GH 151M vs README 149.6M", "PR #1 now closed unmerged", "do not re-fold §71 claim-audit as a beat", "typed decisions, RLCD, confidence-gated routing", "structured ≠ correct", "mock not live API", "26 tests", "wjdjdakf17/jev-study ≠ baekenough/jev-study", "bonzi-27b-v2 / ternary-8b / 27b-v1 GGUF family densify", "WANLI-256 74.6% / 65.2% / 71.1% *theirs*", "Bonsai 1 27B Q1_0 runs on stock llama.cpp", "ternary still needs PrismML fork", "hf:heman10x/openJev-verdict-2.0 twin tokenizer-only", "OpenJev Vision image classification + uncertainty", "CLEVR-4 held-out joint 0%", "hfdataset:IamBusy/OpenJev-Vision-Research-v0.1 12,832", "294,912 derived targets not independent samples", "Laya multilingual ONNX WebGPU typed-decisions port", "63/63 selected answers / 5.1e-4 CPU / 1.2e-2 WebGPU", "UpHash-Network/mini-jev is yuki-oshio transfer", "jev-injection-bench 11,900 labelled prompts", "Jev best ranking / Haiku better ECE 0.021 vs 0.058", "0.5–0.9 band is where Jev's numbers do not mean what they say", "Prompt wording moves panic 28%", "manojlds/jev-dspy-bench ≠ dspachos/jev-dspy ≠ jmanhype/jev-dspy-lab", "Jev agreement is similarity, never ground truth", "no aggregate quality grade or merge gate", "AbstentionBench-on-Jev rank 1 of 20 vs 2025 field", "question-asymmetry", "forward-looking 0.465 never extreme", "openkev calibration layer not a runtime", "ECE vs coverage independent", "select_threshold returns inf", "escalation catches uncertainty not ignorance", "misakaikato/openkev ≠ jaredpalmer/kev", "pdf-race Docling→Jev vs Gemini", "parser owns the wall clock", "12/12 tie is a tie", "titles selected not generated", "flopcheck 16 calibrated tweet judgments", "mechanical tells in code", "ZeroX-01/jev-atlas ≠ Zaious/jev-capability-atlas ≠ gorock007/jev-atlas", "Laya calibration lab Gradio MCP", "T never changes argmax", "confidence ≠ top-label p", "easy probe set refused", "40–48 rows too small to ship T", "Gemma-4 26B-A4B jevify classification+calibration", "LoRA adapter twin not independent eval", "Gemma-4 E4B jevify", "E4B LoRA stub card", "kushalpatil/jevify-gemma4 ≠ Mintzs/jevify ≠ gulagala001/jevify ≠ uspraveen/Jevify", "GH kushalpatil07/jevify 404", "PAWS 0.580/ece 0.288 is the weak cell", "smaller E4B slightly better OOD ECE than 26B-A4B", "Hub jevify merged LoRA ships weights", "bonzi Bonsai-8B v1 GGUF densify", "Bonsai-1.7B v1", "Bonsai-4B v1", "WANLI-256 64.5% / 60.2% / 52.0% *theirs*", "rank #4 / #5 / #6 of 6", "JulesHuisman/jev-eval scaffolding / README SHA c356a584 (was empty e69de29b)", "JulesHuisman/jev-eval ≠ SivletLabs/jev-eval ≠ willkelly/jev-evaluation ≠ 4esv ≠ xxkuboxx ≠ onlyoneaman ≠ dayhaysoos/jevals", "7 bands 6/10 vs 40 bands 0/10", "source receipts + confidence slider re-policy without re-inference", "32/32 synthetic is smoke not production", "classify HF datasets across typed semantic dimensions", "roadus2 watch misspelling; lock roadius2/ultra_laya", "ultra_laya REVIEW defects", "default branch claude/laya-jev-review-gg5ppo", "XNLI EN 88.3% ECE 0.032 → RU 77.3% ECE 0.096", "Δ −11.0 pp [−14.2,−7.8]; ECE +0.063", "MASSIVE no detectable difference at n=600", "confidence is function of p_max (r=1.000)", "pointer-not-generator 400 human-authored responses", "proposed ≠ authorized", "FewRel 160: Jev 85.0% vs lexical 13.125%", "gated 100% (95/95) coverage 59.375%", "J++ composable semantic computation language", "judge-jev 0.5 still soft", "947 repos scored; A 273 / B 302 / C 372", "LLM rubric ≠ benches", "No benchmark winner is claimed", "phishing: naive 62.6% vs regex 91.8%; 5-atomic + LR 95.0% *theirs*", "AITuber tension ±15", "README npm global; repo is Rust", "git-confess code owns counting/blame/ratio", "httpx exhibit 11% (13/119) *theirs*", "90d trend +12.40% vs random +12.75% vs BH +41.71%", "5m win rate 25%", "Awesomejev 656 entries / 38,160 stars", "tracker likes 64 (+4) lastModified UNCHANGED", "Laya present; Blackwood ABSENT; Archer still promised_not_landed", "Blackwood tracker ABSENT; likes 2 gated manual", "r = c - p_a", "ECE 0.021; acc 0.807 vs warmup 0.746", "Independent primitive", "11.57s vs 54.10s · 4.67× · 120/128 *theirs*", "default path is pretrained Gemma probs not trained RLCD head", "GH Meanblock 404; lock leesk212/JEV-CPU", "softmax over letter slots ≠ Noul", "WANLI 0.741 vs openjev v2 0.77 *theirs*", "3-way NLI ≠ Noul", "priority 0.464 = majority floor", "banking77 contaminated", "raw margins not probabilities", "do not distill Jev as teacher of record (they distilled Haiku)", "“0.9 is not one number”", "ranking ≠ calibration", "banking77 0.8–0.9 stated 0.86 actual 0.73 over-confident *theirs*", "≠ Praveenrajus/jev-bench ≠ fstandhartinger/jevbench", "$0.0000153–$0.0000226 vs circulating $0.0004 (~20×)", "Score is 0..n-1 expectation not 0–1", "Noul has no confidence field", "TCP floor 198.8 ms", "type reliability is not a reason to choose Jev (json_schema 5/5)", "gateway tax not one number", "Function-only 5/8 vs hybrid 8/8", "4/8 without Jev", "8 designed cases not conversion lift", "200-row pilot Jev 86.5% 173/200 vs Gemini Flash-Lite 86.0% 172/200 vs Pro 87.0% 174/200 *theirs*", "not a ranking", "情緒測謊器", "8-example Jev vs GPT-5.6 Sol ~64× cost 5.4× latency *theirs*", "synthetic; no inference", "≠ JevBench v1.2 §78", "Judged 3317 / listed 2560", "Jev judges, code applies policy", "APA “microsecond policy / zero hallucination” overclaim", "Client-side quiz; pointer from held docs; scanned-PDF warn", "Jev judges / agent reasons / user decides", "selecting an option is not permission to implement", "pattern exact, judgement must clear floor", "no matching pattern → no model call", "not a correctness oracle", "Spec vs artifact remainder", "treating 0.85 as 85% / minProbability hard-gate as Harbor", "VERIFY acquires discriminating evidence, never same-pool confidence-only rescoring", "fast/full/max are ceilings not sizes", "Solar writes, Jev chooses NEXT ACTION", "do not reopen or amend PR #23 or #24 or #25 or #26 or #27", , "Calibration is not alpha", "NO CURRENT ALPHA CANDIDATE", "ΔR² approximately +0.00084", "Brier 0.2131387", "ECE 0.0421875", "Adding Jev probability to deterministic volatility improved Brier by only 1.4058e-05", "default 0.5 keeps zero non pinned", "keepResult median 0.14 to 0.17", "keepCall median 0.28 to 0.35", "usable range is about 0.10 to 0.25", "7.8% to 57.9%", "judges results it never sees", "task-finish eval not built yet", "$0.002 per compaction", "slavadubrov/sgr-judge-bench ≠ slavadubrov/jev-judge-bench", "Jev 108/120 $0.083 0.34 s", "Luna SGR 114/120", "paired Jev accuracy-difference intervals include zero", "not evidence of equivalence", "GLM SGR 26/120 93 format failures", "Terra-planned Jev hybrid 55/120", "rule-based by default, optionally Jev-backed", "empty README", "missing key cannot break the experience", "prefill plus exactly one decode", "softmax over A/B/C ≠ Noul", "BBQ 9,053/10,000 (90.53%)", "ECE 0.0890", "Mean confidence 0.9943", "overconfident", "score and noul not implemented", "DGUI 12 rows (was 6)", "INSTRUCT 119 rows likes 2", "encode the state once, decide everything in parallel", "0.740 accuracy against a 0.508 majority", "ECE 0.047", "fine-tune's advantage ends where its 384-token training data does", "jasonkneen/open-jev ≠ pngwn/open-jev", "same sha d41dc3cd", "Space does not call Jev", "recomputes routing from saved probabilities", "200-case Jev 97.0% / 100.0% / 95.0% / MAE 9.22", "synthetic repository benchmark", "Jev evaluations are advisory", "YehuiTang0316/jev-nlgrep ≠ Bentlybro/jevgrep ≠ can1357/jegrep ≠ uehaj/jev-semgrep", "default threshold 0.8 still soft", "40-line windows cannot prove whole function", "token-native sequential start/end Choice", "Gemini/Haiku stubs not configured yet", "handful of hand-written examples, not a benchmark", "Jev judged exactly what it was given", "laguagu/jev-skills ≠ laguagu/jev-evidence-lab ≠ Pleo2/awesome-jev-agent-skills", "contract_passed is not a claim of guaranteed factual truth", "Wilson lower bound 0.85 floor", "fixture mode no savings claim", "SemIf 2207★ (+21 vs §110 2186)", "jevlike 1043★ (+5 vs 1038)", "TypeAR 15★ (+1 vs 14)", "AnotiaWang 97★ (+1 vs 96)", "yibie/awesome-jev 506★ (+16 vs 490)", "Laya likes 822 (was 802)", "tracker likes 64 flat, lastModified UNCHANGED", "do not reopen or amend PR #23/#24/#25/#26/#27/#28", "Heman10x-NGU/Verdict-open-jev ≠ Heman10x-NGU/openJev-verdict-2.0", "TF-IDF + LogReg ECE 0.0207 vs Jev 0.1440", "Verdict-open-jev 48.07% vs Jev 90.80%", "abstention combined recall 10.00%", "p50 35.58 ms", "K=25 (maximum capacity) 72.00%", "0.85 coverage 84.60% selective risk 1.18%", "26.1× faster than standard Qwen JSON generation", "Jevify 90.0% / 167 ms CUDA graphs disabled", "Finding 1: Brier on stated confidence alone is a trap", "grpo_rlcr 0.78 / ECE 0.084", "reliability 0.007 but resolution 0.000", "27 900 schema-driven decisions", "13 600 / 13 600 questions", "candidate mass min 0.99999624", "22 configs · 166,054 rows · 4 calibration-gold", "sha a39eba3f", "Student B MAE 0.148 / Pearson 0.836 / 86.0%", "pngwn/open-jev-laya-bench README 404", "sha 9f69c742 likes 2", "HDFS 0.9933 (745/750) / retain 0.0084", "BGL ERROR/FATAL protection 1.0000", "2,479 / 2,500 HDFS uncertain", "cache hit 0.9648 (2412/2500)", "$0.153936 estimated", "E2 recomputes from saved probabilities", "Space sha eda59e0a", "MASSIVE English 0.783 / Khmer 0.033 / Hindi 0.133", "40–48 rows too small to ship T", "T never changes argmax", "siren2345/jev-single-decode-transformers ≠ siren2345/jev-single-decode", "Split Transformers experiment from llama.cpp runtime", "tanayvasishtha/jev-lab ≠ dairui1/jev-lab ≠ BrendanH18/jev-lab ≠ yibie/laya-jev-lab", "Four experiments stress-testing TypeSafe's Jev: calibration, bundle bias, label bias, and ensembling", "second pass must be $0.00 from cache", "The pages never call Jev", "Gemma 4 31B 77.0% / Jev 1.13.0 61.4% / Laya 322M 0.0%", "restriction state 95.0% against 84.4%", "None of the systems are particularly good at knowing when to stop and ask", "They skip the question and call a tool directly", "100% schema pass", "six-field joint 48.8% vs 72.8%", "ywchiu/jev_benchmark ≠ Running-Dolphins/jev-bench ≠ Praveenrajus/jev-bench", "ACT / REVIEW / FALLBACK", "A provider failure, timeout, malformed output, or missing answer is **not** a policy outcome", "confidence is descriptive provider output, not a substitute for probability", "Quality denominators include only valid scored answers", "an exact halfway tie chooses the lower level", "aiwithenoch/Jev-Skill ≠ simplosophy/jev-skill ≠ laguagu/jev-skills", "The local path does not claim to turn a smaller checkpoint into Jev", "Low support becomes decision: \"review\"", "MIT-0 SPDX NOASSERTION", "current-llm", "结构兼容,不是 Jev 模型能力", "altryne/jevify ≠ Mintzs/jevify ≠ gulagala001/jevify ≠ uspraveen/Jevify", "Find where Jev belongs. Design the questions. Measure the difference", "TypeAR-AI/TypeAR 301 → TypeLLM/TypeLLM", "TypeLLM/TypeLLM 16★", "SemIf 2241★ (+34 vs §111 2207)", "jevlike 1051★ (+8 vs 1043)", "AnotiaWang 98★ (+1 vs 97)", "yibie/awesome-jev 525★ (+19 vs 506)", "Laya likes 864 (was 822)", "tracker likes 67 (+3 vs 64)", "lastModified UNCHANGED `2026-09-20T04:29:16.000Z`", "do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#32", "hysteresis enter/exit / replay policy without inference", "calibration does not compose / hop-ECE permutation-invariant", "equal-width vs quantile ECE / ranking ≠ calibration", "Qwen2.5 ≠ Archer / Qwen 3.8 sparring ≠ Archer / Qwen/Qwen3.8-27B ≠ Archer", "Deferred Crispification / TCE / AMS", "g0runmezadam/what-is-jev IS tunahansahin897/what-is-jev", "pd.cut equal-width vs jeval quantile", "A hunch is a probability with a policy attached", "soundness theater / measurement theater / hourly 0843", , "Jev Capability Resolver / NiazMorshed2007/jcr", "one tool nested capability tree / returns context / does not execute", "skills vs capabilities / workflow+judgment vs operations", "format independent of Jev / proposed open standard", "JCR_BAND_RATIO 0.6 is application policy / soft scores ≠ hard gates", "routing ≠ permission / docs ≠ authority to run", "sol-vs-opus5-20 lookup+explain / n=1 / Not Harbor task-execution", "wall-time mixed / Sol slower with JCR in 19/20", "NiazMorshed2007/jcr ≠ skill-broker ≠ skillranker ≠ jev-sift ≠ jev-lens ≠ jevusher ≠ jev_select_capability", "do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33/#34", "notes.md §116", "copy the SemIf/MLX installer?", "quote 5.21× as beating Jev?", "treat 0.845 as a TypeSafe replica?", "collapse SemIf into kw2828/zhihz/semif-rs/semif-serve", "softmax over options as a Noul", "llm prompt to jev primitives", "conversion assistant not equivalent behavior", "heuristic conversion ≠ calibrated Noul", "alexwestco/llm-to-jev ≠ altryne/jevify", "user-provided 0940 / notes.md §118", "judge ≠ actuator", "candidate_mass", "softmax over A–H ≠ Noul", "hourly 0947 / notes.md §119", or "cascade sign-flip / calibration theater": read `references/faq.md`, + other", "training confronts Choice other / none-of-the-above", "soft AGENTS.md rules vs the linter", "screenshot Choice / omni System One", "extractive quotes / pointer not generator", "compaction summarize vs pointer", "encoder vs Jev compaction backend", "shadow-mode compaction rollout", "CI flaky-vs-real merge gate", "fail-open VOI wake/resume", "claim vs session evidence", "S1 indexer escalate-S2", "Harbor on/off routing", "fail-open vs fail-closed wake vs CI gate", "encoder vs Jev computer-use backend", "hybrid local decide + remote fill", "DONE vs verified success", "stdout prune vs session compaction", "OpenCode jev-pruner vs Claude jev-pruner", "zen-chat vs jev-zen Noul", "hard envelope then Noul prune", "Cua-S1 vs TypeSafe Jev", "plan vs execute dry-run", "specialist computer-use vs general agent", "local drop-in vs stub scorer", "route vs memory", "when does it hold / extractable from state", "decision model vs constrained LLM", "dual-process S1/S2", "combinatorial grid vs extractive", "uncalibrated local likelihoods", "decision-native RAG", "classify-first / read selectively", "living applied-mappings atlas / class patterns", "silence as safer / draft-gate heartbeat", "robotics text-state vs pixels", "verbatim ledger vs summary", "judgment as language primitive", "Stagehand extract pick-and-copy", "harness observe-score-act vs demo loop", "public judgment wall / six parallel questions", "meaning-search without embeddings", "attention ≠ correctness", "skills→oxlint / AST prove ∩ remainder", "session-sticky first-prompt routing", "measured RAG rerank vs generative rerank", "capability kernel / secrets never in the agent", "Jev is SENSOR not policy", "type-safe ≠ correct", "typed control plane around DSPy", "native vs verbalized confidence", "engine owns truth / Jev owns judgment", "human-confirmed kill gate", "train specialist vs few-shot hosted", "decide→policy→LLM leftover", "Noul 0.5 cannot-tell never rounded", "calibration ≠ sortable / ORDER BY", "pairwise inversion / Score ordinality / two-decimal ties", "wire-compat GLiFormer /v1/systemone", "class-backend economics", "loopback gateway hosted + local", "do not distill Jev as teacher", "active-learning triage", "evidence-packet explorer", "meaning-grep AND/OR/NOT", "closed-vote-only / no planner LLM", "Jev vs PCD Harbor", "PCD O(1) ≠ Noul", "host-owned handlers × System One", "OMP/pi fail-open gate", "permission vs probability / operator owns thresholds", "judgment ≠ permission / Jev never grants access", "eval integrity / instrument not score", "constrained optimizer + S1 features / never sole hot-path gate", "privilege ≠ verdict / effect contracts not tokens", "attention filter / VOI for human review / never blocks / never green unless sure", "measurement owns endorsement / evidence-gated question packs", "Jev supplies evidence / code owns authority", "ranking ≠ calibration / never hard-threshold raw p as frequency", "hot-click CU / indexed element table", "Jev judges relevance / code decides structure", "local rules first then remainder / never auto-train on own hides", "combinators / System One as control plane", "receipts not leaderboard / type-safe ≠ correct jaggedness", "VOI over skill library / skillranker abstention", "OOD calibration / AUC ≠ ECE", "Jev vs thinking-budget small models", "turnstile / replayable evidence≠authority", "MLX one-pass schema→JSON / Apple Silicon replica economics", "memory leases ended by new evidence", "never confidently wrong / TLA+ compose / escalate instead of hard-gate", "no seal no advance / coverage ledger / mint ≠ product brain", "skill-broker sibling / judgment ≠ permission", "sureness bands / max_prob is generous", "JevBench / calibration not in Main Score", "CI typed gate before expensive review", "Codex MCP host adapter", "judgment as attention redirect / jev-preflight", "compress-before-first-send / dizk jev-lens", "tools≠use / SessionStart over hoping", "observational memory / pi-om keep-kind", "open-Jev class / openvons / JevPick", "physical-world System One / HA-Jev / not for locks", "judgment outside the store / jevql", "landed-script trust / headless≠auto-approve", "digital-design combinators / extended five", "VOI cache admission / same-intent skip LLM", "BM25 vs Jev skill routing Harbor harness", "zeroshot vs BERT / contamination DiD", "typed escalate continue abort baton / inverted loop", "worth-your-attention VOI / ThinkyMiner Winnow", "Jev WHETHER Python HOW LLM WHAT", "conflict vs ignorance / named Choice escape", "Playwright executes Jev chooses", "OpenJev /v1/decide not drop-in", "SemIf wire-compat runoff; SemIf rename densify / MLX backend / 5.21× systems≠semantic / Softmax ≠ Noul (`notes.md` §117)", "decision-as-memory flywheel", "record/replay CI / jevassert", "failure-finding arena / jevarena ≠ jev-arena", "BBQ not a bias cert", "decider≠executor", "sentence-as-rule lint / jevlint", "sentence-as-rule lint / jev-lint is jevlint rename", "VOI hunk prune", "whole-repo intent VERIFIED/VIOLATION/UNKNOWN", "GLiNER2 spec ≠ replica", "open replica substrates / grande / laya-jolt / JEV-CPU", "ONNX local-jev not equivalent", "persist constraints across compaction / pi-heed", "calibration+cost first-class gates", "Harbor-shaped Jev vs SGR LLM-as-judge / jev-judge-bench ≠ jevarena ≠ jevbench", "hand no-text steps / jev-use / Vercel drops confidence", "Pi System-One control plane / pi-jev-control", "never free-generates / jev-gpt tree of Choices", "OpenRouter recipe atlas / samples not benches", "personal history feed / jevfeed / no social graph", "competing NAR claims / dual-channel ECE / openJev-verdict ≠ OpenJev", "empty compaction-proxy skip / IPECTER", "throughput ≠ latency / like-for-like ECE", "1-token logprob endpoint ≠ Noul / coverage ≠ correctness", "open replica engine / jevinf", "unofficial Elixir SDK ≠ OTP peer", "jevex n=16 files-to-read VOI", "commit pre-review attention≠verdict / middle band", "Hermes plugin is Agnes not TypeSafe", "pi-jev-compact ≠ pi-jev-compaction", "empty Codex-proxy skip / IPECTER runway", "decision-native inbox / mailordinal", "unofficial jev-cli not ready / ≠ jevql", "laya-multilingual / English checkpoint confident-wrong OOD", "schema-scorer peaked ranking ≠ calibration", "HF 401 / GitHub 404 Hub-only", "productized System One HTTP / classifier.dev", "escalate-under-threshold / smart tier / multi-label ignores", "silent FALLBACK / granite 0.546 vs advertised 0.800", "vs_jev tracked JSON / read eval/README", "choxos/jev-reviewer ≠ egma-ai / systematic-review pointer", "two-pass Choice+Noul evidence extraction", "not-found is an answer", "human check as productized judgment", "githubnext/localjev ≠ kunchenguid/local-jev", "wire-compat ≠ logit-equiv / prompted JSON ≠ structured read", "self-reported probs / entropy confidence", "GitHub Next local /v1/systemone", "LM Studio runner gap / structured-read primitives", "NandhaKishorM/laya packaging ≠ Hub-only / Router script-before-p", "post-T ECE ≠ raw ECE / Banking77 token-budget", "0.85 still soft / not TypeSafe drop-in", "external census ≠ scored bake-off", "GLiNER2+routers class-boundary", "incomplete openjev census vs watch", "Harbor honesty watch / silent fallback", "JevBench v1.2 geometric mean / cal ON rank / weight sensitivity", "option-order 72→21 / instruction models in the class table", "self-host latency ×2 assumption / est. costs", "Laya absent is a gap not a named exclusion", "Qwen3.8 27B ≠ Archer", "hourly already-folded watch / apply-the-five / skip thin noise", "hard-gate Noul as PR/quality gate is soundness theater", "S1 never stalls waiting / S2 one-use advisory", "Local controller ≠ githubnext/localjev", "purple telemetry = consumed not arrived", "seed = geometry not async replay", "20% starting gate still soft", "no pixels to either provider", "OCR+AX observe-score-act / typesafe-computer-use", "never send screenshot to frontier for the decision", "overlapping CU options = false low confidence", "split kind/item/site", "155× one-screenshot ≠ Harbor taskset", "decision ≠ answer-reader capture", "ASR observe-score-act / jev-voice-browser", "partial-speech VOI / free-text waits", "spoken confirm ≠ hard auth", "numbered overlay without another model", "wrap-as-execution / AgentGhost ALLOW ASK DENY", "rules first then Jev remainder / ASK throws / fail-closed", "reddpy/AgentGhost ≠ jwen5419807/agentghost ≠ vventirozos", "JP genre atlas / studio_yebisu / stars ephemeral ≠ eval", "Jev Clearly Explained / akshay_pachaar / LLM hammer", "schema-safe ≠ correct / 200× 400× TypeSafe ceiling", "questions-as-code / shadow first / not a TypeSafe how-to", "proposition ≠ embedding / contrast-set", "boolean composition of soft Nouls / AND OR NOT", "uehaj/jev-semgrep ≠ semgrep.dev", "meaning-grep dedicated fold / not a gate", "decision-validated UI / Jev never authors text", "decision-as-assert / jevtest ambiguous band", "typed decisions drive UI / jev2ui", "hybrid S1 closed verb menu / anima3", "pointer-not-generator search / JevFind", "jev-frontier-bench ≠ frontier-100", "product bakeoff ≠ architecture duel / GLiClass", "four engines same questions / majority floor", "authorship named escape / not evidence", "ha-switchboard HA remains execution", "n8n classify/route/score / Low Confidence", "fast-jev-compaction-pi ≠ pi-jev-compact ≠ pi-jev-compaction", "jevloop full-distribution optimizer / no LLM in the loop", "laya-vision SmolVLM / score untrained", "Cerebellum-2B /v1/decide ≠ TypeSafe / wire-compat vs agent-routing", "laya-grounded not drop-in / Platt not temperature", "GestaltLabs/Jeff-1 ≠ logan-markewich/jeff / acc vs ECE n=9730", "stanley-code empty findings ≠ approval / human promote", "findme ≠ JevFind / NL memory beam-search FS", "jevsubrouter price workers not conversation / counts ≠ dollars", "feelings .feels() default 0.5 is Noul-0.5-never-rounded / ≠ hunch ≠ Probably", "apa-agent-harness ≠ AntonioCoppe/jev-harness / unpublished npm", "grok-bot-jev skill cannot force a bot that ignores it / A/B proxies not tokens", "Essentiel-Jev never authority / human every action", "enzo-mcp independently falsifiable claims / ≠ jev-sift", "pigeonhole OTHER skip / decision-as-filing", "jev-reliability Nothing about accuracy", "clduab11/jev-test ≠ realZachi/jevtest / Nothing runs yet", "jev-rag-benchmark Jev wins is not an assumption", "dairui1/jev-lab ≠ BrendanH18/jev-lab", "jevmail gmail.readonly / mailjay archive/trash", "ZHUBoer/ego-jev reserved __none__", "runWorkflow completed ≠ success", "jsort scores are relative", "Noul not Choice for scale", "groundedness-judge-bench native vs schema-guided", "implicit_true included in yes", "jev_playground 0 promotions", "routing-backtest 0.0447%", "yuyang2230/jev-agent-skill jev-1.13-free", "jev-techstack-classifier stack_config.json", "s1_ruby collapse late", "undecided? abstain", "2389-research/judgement license null", "confidence ≠ winner p", "typesafeai-sdk-community not a new species", "tpellet/hunch exit 3", "never-execute list", "jev-file-search scores not calibrated accuracy", "jev-linkmap Jev never sees S2 prose", "muhammedilyasy/jev-mail metadata only", "tidy none-of-folders stay", "tab-bouncer pinned/audio/current never closed", "lkclean Show fail-open", "jev-yt-time-saver Show anyway", "ORIGIN pause-if-no-Jev", "validResponse sums-to-1", "jev-crawlers risk bands never raw boolean", "jevbrain AUTO_ACT is not a Noul", "judgekit YAML classify/score/route/verify", "typed-judge-kit verdict-in-code", "alsoleg89/decide packing VOI", "0.8 ≠ 80% accuracy", "Jev-Calibration Platt ECE 0.117→0.052", "jev-calibration-arena never acts", "ctmx/openrouter-jev-mcp Decision-as-Plugin", "FrancoisChastel/jev-code ≠ npm jev-code", "claudecode-jev-marketplace fail-open not hot path", "pedroknigge/mcp_jev packs not ask_jev", "cyrusasco/typesafe-mcp noul deadband 0.35–0.65", "codaaiteam/jev-skill jevtypesafeai.com ≠ TypeSafe", "hermes-switchyard ≠ hermes-jev-router ≠ hermes-plugin-jev", "nanoprune 2.8MB ECE 2.58%", "smartdio/jev-browser-agent ≠ ZHUBoer/ego-jev", "Dakai/omp-jev-web DONE ≠ proof", "hari007sh/jev ≠ dannote/jev", "0thernet/system-one-skills deterministic verify", "typed-gate band [0.40,0.60] is refusal", "pi-jev-gate fail-closed; choice is the verdict", "Foq ~25ms/2.2GB local", "rev prefill-only + HF jev-0.5b", "robfrase/jev planning memo", "typesafe_agent_gates 27/27 / 31/31", "EpicEric/safe-sh static remainder", "pastepilot Confirm before act", "Jev-Reranker live Jev not yet measured", "sessionwise opt-in relevance", "jev-search pointer sieve", "400ms Salesforce WebMCP", "typesafe-scheduler-diagnostics advisory", "droidjev screenshot-free", "Tewoto1 jevcu planner still writes", "ha-conversation-jev Jev→Grok", "dsh-jev can only gate", "jev-classification-benchmark specified not run", "jev-luna-pagerduty p≥0.50", "meldltd/meldecision laya-go ONNX", "laya-doom never pixels", "logixism/laya-api empty README", "akpsahan/laya ≠ Archer", "choxos/jevchess engine owns truth", "jev-drive sim not AV", "story-arc Jev never authors", "jev-hs-assistant HS6", "golergka/jev-plays-starcraft-2 UI-verified ≠ API Victory", "awesome-jev-use-cases catalog", "Nibir1/typesafe-go ≠ official", "fingerprint after redact", "recall vs decide", "publish fingerprints+answers", "CI replay as Harbor cousin", "Cache hit ≠ correctness", "hyperspaceai/jevcache ≠ kushals256/jevcache", "human labels only", "score never auto-accepts", "production capture flywheel", "sutro-sh/jev-align ≠ caiovicentino/jev-align", "guidance ≠ hook", "catalysts ≠ summaries", "compile-time System One", "unofficial ≠ TypeSafe", "format_version modernbert-jev/1", "Argos1111/jev_local ≠ us/jev-local ≠ kunchenguid/local-jev", "LFM default ≠ ModernBERT backend", "Nemotron ≠ TypeSafe Jev", "not a calibrated replacement", "djev-dev complements djev-spark", "images as Choice options", "Laya essay numbers *theirs*", "Router/OOD confidence", "hosted bootstrap ≠ silent TypeSafe", "difficulty + policy thresholds + JSONL trace", "jev-codex-pilot model + reasoning depth", "keep/shadow/hybrid/reject", "quarry evidence projection", "Frank-ZY-Dou/awesome-jev robotics/3D/control", "one-dollar-tahoe TypeSafe Jev defense eval", "jevguard calibrator/cache/escape", "jev-ci-selector CI shadow mode", "llama-jev llama.cpp replica", "petercr/jev-orchestrator ≠ FleeexCorp/jev-orchestrator", "seb4ez/jevguard ≠ AseemPrasad/JevGuard ≠ pablozr/JevGuard", "webNeat/llama-jev ≠ WiktorB2004/llama-index-jev", "OpenCode jev-pruner context sieve", "observe→score-candidates→prune", "jev-zen / jev-1.13-free", "zen-chat ≠ Noul", "fail-open original", "keepScore >0.1 floor", "host port of tamaratran/jev-pruner", "indiejoseph/opencode-jev-pruner ≠ nrdz-labs/fast-jev-opencode", "jev-webagent-bench empty stub", "Kiln-AI/jev_jsonschema noul_threshold 0.5", "NSStudent/JevSwiftSDK unofficial", "GLiNER2 native Apple path", "unofficial Swift/Core ML GLiNER 2.5-small", "entity spans + confidence", "not Choice/Score/Noul", "not TypeSafe", "label descriptions as schema", "on-device ANE economics", "honesty locks", "shershah1024/gliner-native-runtime ≠ Fastino", "≠ gliner25-compaction ≠ gliner2-ultrafast ≠ Eran-BA/Jev_from_GLiNER2 ≠ NSStudent/JevSwiftSDK ≠ jevmlx", "default threshold 0.1 still soft", "soft Noul ≠ hard safety", "Decision Graph Protocol frame→assess→commit", "app retains permissions/effects", "Jev-first assessor-neutral", "guarded commit / receipt/next frame", "assessment batching", "hard-gating DGP as safety theater", "numerous-com/dgp ≠ TypeSafe official", "jegrep calibrated path+range Nouls", "no embeddings/index/daemon", "~$0.01–0.03 typical", "agent --json", "can1357/jegrep ≠ Bentlybro/jevgrep ≠ uehaj/jev-semgrep", "Archer-arch fidelity", "kev family OOD 0.76–0.77 vs Jev 0.86", "block-causal isolation", "pointer/readout CE-trained", "/v1/systemone drop-in", "replica honesty", "cost-sensitive decision theory × System One probabilities → control flow", "thresholds derived from costs not hard-coded", "YES / NO / UNSURE from cost_false_yes / cost_false_no / cost_human", "auto-batching same-object questions", "Kungie/gut ≠ tpellet/hunch ≠ carldaws/hunch", "judgment vs generation", "deterministic execution after probabilistic judgment", "exactly one app-owned callback", "explicit uncertain branch", "Illusion47586/judge ≠ lexingtonhibiki/judgekit ≠ Ascurse/typed-judge-kit", "variable-N option scoring as the trainable object", "dynamic candidate bags not fixed label sets", "zwliJay/jev-forge ≠ NanoJev", "open replica economics / latency vs closed Jev", "NAR local drop-in", "wfzyx/von late-catch HIGH", "competing NAR claims / replica honesty", "typed judgments vs chat judges on guardrailing", "ishaannk/llm-vs-jev cross-note only", "deeper integrity fold is rh-guard", "nothing wins outright", "can be argued out of guarding"", "Jev IS the if-statement", "judgments/probabilities drive branches", "text model only writes prose", "interpreter owns variables/loops/budgets/replay", "otherwise maybe / confidence gate", "chaos samples after the gate", "southpolesteve/probably ≠ carldaws/hunch ≠ feelings ≠ Kungie/gut ≠ Illusion47586/judge ≠ tidymodels/probably", "133★ / forks 10 live", "build calibrated classifiers from human feedback", "retrieve by relevance not resemblance", "one calibrated yes/no per memory in one request", "pointer mode 17/18 19/20 *theirs*", "embedding resemblance misses the allergy", "samdotmak/jev-recall ≠ jev-search ≠ jev-sift ≠ carryforward ≠ chopratejas/invalidate", "memory leases ended by new evidence", "six Nouls then fixed rules in code", "0 of 157 false invalidations", "questions/plans/directives are not evidence", "unsure → review queue", "host keeps the store", "name↔body / comment truth / test-claims", "mizchi/jev-lint is mizchi/jevlint rename", "no shipped rule has severity error", "~1 in 5 findings wrong *theirs*", "mizchi/jev-lint ≠ huntedman/JevLint ≠ MichitoSugawara/jev-lint", "JSON Schema → typed JSON via Jev", "noul_threshold 0.5 decoder not a proof", "IncompatibleSchemaError lists every bad property", "on-device Laya CoreML ANE", "~5 ms P50 short decisions", "189/189 FP16 checkpoint parity", "10× not achieved", "mizorewww/laya-coreml ≠ gliner-native-runtime ≠ jevmlx ≠ NandhaKishorM/laya", "softmax over allowed tokens ≠ Noul", "question-first cache", "Micha0827/snapjudge ≠ githubnext/localjev ≠ jevmlx ≠ cendress/SnapJudge", "Jev-first Pi agent loop", "slow-LLM fallback", "explicit action menu / CandidateSource unimplemented", "62 tests wiring not quality", "direwolfiy/JevPi ≠ standardagents/jevpilot ≠ pi-jev-control", "resume-screening bias audit methodology", "name×resume factorial independent Nouls", "callback determined by resume quality", "mean-probability name gaps operationally negligible", "natemoo-re/bias-bench ≠ BBQ", "Plan/PRD panel → code-owned pass|review|block", "cheerleading out of scope", "austindixson/planalyzer ≠ single-goodness Noul", "cost-aware multi-model routing/escalation", "decide vs do", "successful-task cost", "cannacre8ive/switchboard-ai ≠ ha-switchboard ≠ hermes-switchyard", "frozen-protocol zero-shot bench", "TypeSafe Jev vs PrismNLI vs Laya", "contamination caveat", "elcronos/jev-vs-open-decision-models ≠ JevBench ≠ DMB", "context-window admission control", "VOI gate which tokens are worth the expensive model", "fail polarity per lens", "on small inputs lenses lose money", "cvsgireesh/jevusher ≠ jev-sift ≠ winnow", "typed decision control plane", "receipt ≠ authorization", "historical-v0 zero retained cases", "MokiMeow/jev-fabric ≠ jev-forge ≠ dgp", "live 15-dim typed rubric re-score per pause", "scoring economics exemplar", "OpenJev/Codiv ≠ TypeSafe hosted", "jose-troche/live-rubric ~$0.000004 desc / ~$0.000006 README", "adversarial pre-registered Jev eval", "28 predictions before data", "123,805 requests", "confidence does not track ignorance", "polite injection 65% / crude 0%", "willkelly/jev-evaluation ≠ jevals ≠ jev-baselines-eval", "provider-neutral Elixir/BEAM Noul/Choice/Score SDK", "class infrastructure", "nshkrdotcom/system_one_sdk ≠ typesafe_sdk ≠ dannote/jev", "question-linting of Jev questions themselves", "nine jaggedness rules, no API key, no labelled data", "static lint ≠ measured separation", "yodablocks/jevq ≠ tenbin ≠ JevLint ≠ commitjev", "open-weights Laya as class exemplar (binding)", "Nx/Bumblebee runtime", "host chooses backend", "ChristianAlexander/laya_ex ≠ system_one_sdk ≠ dannote/jev ≠ NandhaKishorM/laya", "on-chain/edge Laya deploy", "parity_verified stays false", "model output never grants Tx", "humandebri/IC-Laya ≠ laya_ex", "auditable weekend replica", "Jev outputs never used for training", "soft human-vote distributions", "unpaired 0.577 vs 0.727", "agilabs-ai/jev48 ≠ JevBench ≠ Mapika/decider", "adversarial dual-judge / framing attack surface", "comparative framing is the usable judgment", "prior injection crowds out evidence", "copyleftdev/ember ≠ ember.js", "Laya specialist fine-tune pipeline", "training still GPU-pending", "PIXELZX0/XERON ≠ convaiinnovations/laya", "Hub Laya replica drop", "daliborsb/laya ≠ convaiinnovations/laya ≠ NandhaKishorM/laya", "System One student distillation corpus", "gold is programmatic", "teacher is closed-API clone", "do not distill Jev as teacher of record", "MagaBitmex/jev-4b-distill-data ≠ missing student checkpoint", "non-LLM VIN System One", "planning depth not chat", "lewislululu/jevon ≠ douglance/jevon", "source-bound evidence checks", "local quote mismatch needs no API", "exit 0 ≠ claim truth", "WaynezProg/jev-kit ≠ jonathanavis96/jev-kit (Airlock) ≠ jev-use ≠ jev-mcp", "independent System One evidence catalog", "scores not one leaderboard", "no external record currently reproduced", "TokenTrim no-Jev matched hybrid 62.4%", "reachjalil/system-one-bench ≠ mallahyari/system-one-benchmark", "21 tasks · 134 items · 208 questions", "scenes from public GitHub contracts, not production logs", "SivletLabs/jev-eval ≠ willkelly/jev-evaluation ≠ 4esv/jev-eval ≠ xxkuboxx/jev-eval ≠ onlyoneaman/jev-eval ≠ dayhaysoos/jevals", "option isolation (sibling-blind)", "permutation-equivariant", "Hub OWNER not published", "nafisazizir/hev ≠ jaredpalmer/kev", "frozen local LLM logits, no trained decision head", "residual-head 9,222-param decreased 73/96→67/96", "confidence = 1−normalized entropy, not P(correct)", "yuki-oshio/mini-jev ≠ r-ms/mini-jev", "Jev classifier as autoregressive next-token predictor", "ChatJev-style soundness theater", "erik-dunteman/ChatJev ≠ dannote/jev ≠ jev-gpt", "calibrated decision head × AlphaProof value head", "implementation-layer isomorphism, semantic difference", "timeout = censoring", "do not launder Noul as proof", "parallel rank-prediction vs serial selection", "independent questions can conflict", "zzzzzec/jevsort ≠ keltokhy/jsort", "curated open System One ecosystem catalog", "rupeshpoojary9/awesome-open-system-one ≠ AnotiaWang/awesome-jev", "arXiv paper radar with Jev relevance scoring", "ranking ≠ calibration / 0.5 still soft", "fail-open failed evals not marked seen", "train calibrated ~27M from scratch", "typed Q→prob dist / one forward pass / no LLM decode", "hyusi2003/MiniSystemOne ≠ Colvin0315/MiniSystemOne", "description-only stub / size 5", "ESCI hard probe fails four of six", "jev_bool ECE 0.242 inversion 0.255", "do not re-fold §60 six-gates as new", "jobbyjev one-request-per-company from batch-size result", "find/design/evaluate TypeSafe Jev decision loops", "karanb192/jev-architect ≠ samtay32/jev-system-architect", "Jairik/jev-distiller size 1", "distill-Jev UI stub / do not distill Jev as teacher of record", "post-launch scored use-case map / Jev self-scores then human curation", "licensedsaucer9-web/jev-opportunities", "Jev-inize a use case into classifier/router", "gavinHuang/jevinize → simple-jev not TypeSafe", "featherless-ai/simple-jev", "compare saved decisions / same label can still change the branch", "VihaanAgarwal/jev-diff ≠ Saik0s/diffusiongemma-jev-macos", "not tested with a live Jev API key", "constrained logprob + temp/Platt ≠ Noul", "OpenJevPro pastes openjev-sglang JevBench as own", "zhangcy122/OpenJevPro ≠ IamBusy/OpenJev ≠ ekzhang/openjev-sglang", "PolyForm Noncommercial", "SmolLM-135M / sub-70ms / 0 output tokens", "demo P(True) 0.5052 / Choice conf 0.2872 / Score conf 0.0055", "README claims MIT / GitHub license null / no LICENSE file", "patelvishwa112/jev-system-one-rlcd ≠ arnabgho/rlcd-lite ≠ blackwood-rlcd", "source-backed Awesome Jev radar / 306+ commit-pinned", "logicrw/awesome-jev-projects ≠ AnotiaWang/awesome-jev ≠ yibie/awesome-jev ≠ cobanov/awesome-jev ≠ rupeshpoojary9/awesome-open-system-one", "auto GitHub sync / Issue-only submissions", "hashed n-gram encoder / rival-aware attention", "olanotolu/jevbetter vs jevlike starter", "synthetic hard menus top-1 0.916 vs 0.873 / ECE 0.0182 vs 0.0367 / 40 vs 4608 menus/sec", "shuffled-context control 0.335", "Turn any open LLM into System-One Jev", "uspraveen/Jevify ≠ Mintzs/jevify ≠ gulagala001/jevify", "Jevify-any-LLM architecture probe", "description-only stub / size 0", "Train encoder-only calibrated decision models from a task sentence", "Exu is a toolkit, not a method", "strictly proper scoring rule", "Pre-alpha", "Ruivalim/exu-base", "scratch-trained calibrated decision model", "typed Q → probability dists", "Colvin0315/MiniSystemOne ≠ hyusi2003/MiniSystemOne", "no published weights download URL", "90.5 seconds / 29.2% pipeline evidence", "p_i/p_j independent of other candidates", "Recipe for calibrated decision models — small model out", "init → synth → train → eval → serve", "91.1 % / ECE 0.022 *theirs*", "Jev zero-shot 75.1", "scienthoon/luce", "Put Jev's three headline claims on trial", "0.5B local GPU", "46x speedup / accuracy identical", "ECE 0.624 sentiment catastrophe", "bigger model worse calibration", "RichardoMrMu/jev-mini ≠ yuki-oshio/mini-jev ≠ r-ms/mini-jev", "System-1 decision engine for local LLMs", "structured choices only", "JSON parse of generated text ≠ Noul", "TypefAI JEV / Journal Entry Voucher", "tapsin/jev-local ≠ us/jev-local ≠ Argos1111/jev_local", "Jev 1.13 reward-model eval across 8 benchmark tracks", "40,940 examples / 0 API errors", "RewardBench v1 92.58%", "Precise IF 50.63%", "goya4140/jev-reward-model-evaluation", "Scaffolding in progress", "Jev vs LLM support-ticket routing", "static + live decision bench", "TypeSafe's own published benchmark", "illustrative simulations, not live API calls", "JevBench v1 — smart/cheap/fast/reliable", "I/C/S/K 25% geometric mean", "classifier.dev fast tier 84.8 is Jev behind its own API", "do not re-fold §78 v1.2 board as new", "Laya (421M) 70.1 now on board", "Zero-shot/few-shot LLM routing", "hard budget filter before Jev", "Jev never asked to perform budget arithmetic", "Jev judges the next state, XState enforces transitions", "simulation uses synthetic keyword fixtures", "catalog gravity", "v-modal/awesome-jev-tools", "★339 live REST", "curation is not endorsement", "crawler-maintained directory", "Daily GitHub + npm sweep, human-merged", "RadRebelSam/awesome-jev ≠ AnotiaWang ≠ yibie ≠ cobanov ≠ logicrw ≠ v-modal", "HF peft SPLADE/BGE reranker", "rdxtremity/jev-reranking ≠ carlaiau/jev-reranking", "query-side encoders, not a Jev replica", "ONNX System One Qwen3.5-4B scorer", "source:pngwn/system-one-qwen3.5-4b-scorer", "CC-BY-NC-4.0", "temperature 1.75", "transformers.js AutoModel cannot load this graph", "Consistency benchmark Space", "This Space contains no benchmark result yet", "12-case plumbing fixture", "Benchmark-driven Jev router and judge", "cheap alone is not success", "Jev does not write, sum prices, or claim accuracy %", "Sol 94.2 / Luna 83.9 / Jev path 89.7", "19.2% Sol / 62.3% cost save / 4.5pp miss of 2pp non-inferiority", "p50 latency worse than Sol due to routing overhead", "erendikmenn/jev-llm-router-benchmark ≠ jev-rag-benchmark ≠ ryantsai/jev-llm-router", "Express + node:sqlite", "mock and Jev decision engines", "previous_ticket_count >= 3 is code", "MIN_CONFIDENCE 0.6 still soft", "substring false positives", "aesaganda/jev-ticket-router ≠ SarathChandraBellam/jev-vs-llm-ticket-router", "Universal Figure & Diagram Router", "confidence ≥ 0.85 hard-gate is theater", "generative AI banned from scientific plots", "six visual branches", "hoangngochuong24947-gif/jev-figure-router", "human-labeled (state, question, label)", "166,054 rows / 22 configs", "soft_label for human uncertainty", "Praveenrajus/jev-bench ≠ fstandhartinger/jevbench", "ternary bonsai System One GGUF", "openjev's mechanism, Bonsai's weights", "Hub does not ship weights", "100/100 easy T/F is not Harbor", "label_mass ≠ correctness", "stock llama.cpp Q2_0 silently gibberish", "NicolaiMTLassen/open-bonzi-jev ≠ NicolaiLassen", "transformers.js DeBERTa ONNX", "source:com-kotobalabs/open-jev-deberta-v3-large", "temperature 1.05", "AutoModel from_pretrained works", "onnx-community/open-jev-deberta-v3-large-ONNX ≠ system-one-qwen3.5-4b-scorer-ONNX", "107★ densify", "GH 151M vs README 149.6M", "PR #1 now closed unmerged", "do not re-fold §71 claim-audit as a beat", "typed decisions, RLCD, confidence-gated routing", "structured ≠ correct", "mock not live API", "26 tests", "wjdjdakf17/jev-study ≠ baekenough/jev-study", "bonzi-27b-v2 / ternary-8b / 27b-v1 GGUF family densify", "WANLI-256 74.6% / 65.2% / 71.1% *theirs*", "Bonsai 1 27B Q1_0 runs on stock llama.cpp", "ternary still needs PrismML fork", "hf:heman10x/openJev-verdict-2.0 twin tokenizer-only", "OpenJev Vision image classification + uncertainty", "CLEVR-4 held-out joint 0%", "hfdataset:IamBusy/OpenJev-Vision-Research-v0.1 12,832", "294,912 derived targets not independent samples", "Laya multilingual ONNX WebGPU typed-decisions port", "63/63 selected answers / 5.1e-4 CPU / 1.2e-2 WebGPU", "UpHash-Network/mini-jev is yuki-oshio transfer", "jev-injection-bench 11,900 labelled prompts", "Jev best ranking / Haiku better ECE 0.021 vs 0.058", "0.5–0.9 band is where Jev's numbers do not mean what they say", "Prompt wording moves panic 28%", "manojlds/jev-dspy-bench ≠ dspachos/jev-dspy ≠ jmanhype/jev-dspy-lab", "Jev agreement is similarity, never ground truth", "no aggregate quality grade or merge gate", "AbstentionBench-on-Jev rank 1 of 20 vs 2025 field", "question-asymmetry", "forward-looking 0.465 never extreme", "openkev calibration layer not a runtime", "ECE vs coverage independent", "select_threshold returns inf", "escalation catches uncertainty not ignorance", "misakaikato/openkev ≠ jaredpalmer/kev", "pdf-race Docling→Jev vs Gemini", "parser owns the wall clock", "12/12 tie is a tie", "titles selected not generated", "flopcheck 16 calibrated tweet judgments", "mechanical tells in code", "ZeroX-01/jev-atlas ≠ Zaious/jev-capability-atlas ≠ gorock007/jev-atlas", "Laya calibration lab Gradio MCP", "T never changes argmax", "confidence ≠ top-label p", "easy probe set refused", "40–48 rows too small to ship T", "Gemma-4 26B-A4B jevify classification+calibration", "LoRA adapter twin not independent eval", "Gemma-4 E4B jevify", "E4B LoRA stub card", "kushalpatil/jevify-gemma4 ≠ Mintzs/jevify ≠ gulagala001/jevify ≠ uspraveen/Jevify", "GH kushalpatil07/jevify 404", "PAWS 0.580/ece 0.288 is the weak cell", "smaller E4B slightly better OOD ECE than 26B-A4B", "Hub jevify merged LoRA ships weights", "bonzi Bonsai-8B v1 GGUF densify", "Bonsai-1.7B v1", "Bonsai-4B v1", "WANLI-256 64.5% / 60.2% / 52.0% *theirs*", "rank #4 / #5 / #6 of 6", "JulesHuisman/jev-eval scaffolding / README SHA c356a584 (was empty e69de29b)", "JulesHuisman/jev-eval ≠ SivletLabs/jev-eval ≠ willkelly/jev-evaluation ≠ 4esv ≠ xxkuboxx ≠ onlyoneaman ≠ dayhaysoos/jevals", "7 bands 6/10 vs 40 bands 0/10", "source receipts + confidence slider re-policy without re-inference", "32/32 synthetic is smoke not production", "classify HF datasets across typed semantic dimensions", "roadus2 watch misspelling; lock roadius2/ultra_laya", "ultra_laya REVIEW defects", "default branch claude/laya-jev-review-gg5ppo", "XNLI EN 88.3% ECE 0.032 → RU 77.3% ECE 0.096", "Δ −11.0 pp [−14.2,−7.8]; ECE +0.063", "MASSIVE no detectable difference at n=600", "confidence is function of p_max (r=1.000)", "pointer-not-generator 400 human-authored responses", "proposed ≠ authorized", "FewRel 160: Jev 85.0% vs lexical 13.125%", "gated 100% (95/95) coverage 59.375%", "J++ composable semantic computation language", "judge-jev 0.5 still soft", "947 repos scored; A 273 / B 302 / C 372", "LLM rubric ≠ benches", "No benchmark winner is claimed", "phishing: naive 62.6% vs regex 91.8%; 5-atomic + LR 95.0% *theirs*", "AITuber tension ±15", "README npm global; repo is Rust", "git-confess code owns counting/blame/ratio", "httpx exhibit 11% (13/119) *theirs*", "90d trend +12.40% vs random +12.75% vs BH +41.71%", "5m win rate 25%", "Awesomejev 656 entries / 38,160 stars", "tracker likes 64 (+4) lastModified UNCHANGED", "Laya present; Blackwood ABSENT; Archer still promised_not_landed", "Blackwood tracker ABSENT; likes 2 gated manual", "r = c - p_a", "ECE 0.021; acc 0.807 vs warmup 0.746", "Independent primitive", "11.57s vs 54.10s · 4.67× · 120/128 *theirs*", "default path is pretrained Gemma probs not trained RLCD head", "GH Meanblock 404; lock leesk212/JEV-CPU", "softmax over letter slots ≠ Noul", "WANLI 0.741 vs openjev v2 0.77 *theirs*", "3-way NLI ≠ Noul", "priority 0.464 = majority floor", "banking77 contaminated", "raw margins not probabilities", "do not distill Jev as teacher of record (they distilled Haiku)", "“0.9 is not one number”", "ranking ≠ calibration", "banking77 0.8–0.9 stated 0.86 actual 0.73 over-confident *theirs*", "≠ Praveenrajus/jev-bench ≠ fstandhartinger/jevbench", "$0.0000153–$0.0000226 vs circulating $0.0004 (~20×)", "Score is 0..n-1 expectation not 0–1", "Noul has no confidence field", "TCP floor 198.8 ms", "type reliability is not a reason to choose Jev (json_schema 5/5)", "gateway tax not one number", "Function-only 5/8 vs hybrid 8/8", "4/8 without Jev", "8 designed cases not conversion lift", "200-row pilot Jev 86.5% 173/200 vs Gemini Flash-Lite 86.0% 172/200 vs Pro 87.0% 174/200 *theirs*", "not a ranking", "情緒測謊器", "8-example Jev vs GPT-5.6 Sol ~64× cost 5.4× latency *theirs*", "synthetic; no inference", "≠ JevBench v1.2 §78", "Judged 3317 / listed 2560", "Jev judges, code applies policy", "APA “microsecond policy / zero hallucination” overclaim", "Client-side quiz; pointer from held docs; scanned-PDF warn", "Jev judges / agent reasons / user decides", "selecting an option is not permission to implement", "pattern exact, judgement must clear floor", "no matching pattern → no model call", "not a correctness oracle", "Spec vs artifact remainder", "treating 0.85 as 85% / minProbability hard-gate as Harbor", "VERIFY acquires discriminating evidence, never same-pool confidence-only rescoring", "fast/full/max are ceilings not sizes", "Solar writes, Jev chooses NEXT ACTION", "do not reopen or amend PR #23 or #24 or #25 or #26 or #27", , "Calibration is not alpha", "NO CURRENT ALPHA CANDIDATE", "ΔR² approximately +0.00084", "Brier 0.2131387", "ECE 0.0421875", "Adding Jev probability to deterministic volatility improved Brier by only 1.4058e-05", "default 0.5 keeps zero non pinned", "keepResult median 0.14 to 0.17", "keepCall median 0.28 to 0.35", "usable range is about 0.10 to 0.25", "7.8% to 57.9%", "judges results it never sees", "task-finish eval not built yet", "$0.002 per compaction", "slavadubrov/sgr-judge-bench ≠ slavadubrov/jev-judge-bench", "Jev 108/120 $0.083 0.34 s", "Luna SGR 114/120", "paired Jev accuracy-difference intervals include zero", "not evidence of equivalence", "GLM SGR 26/120 93 format failures", "Terra-planned Jev hybrid 55/120", "rule-based by default, optionally Jev-backed", "empty README", "missing key cannot break the experience", "prefill plus exactly one decode", "softmax over A/B/C ≠ Noul", "BBQ 9,053/10,000 (90.53%)", "ECE 0.0890", "Mean confidence 0.9943", "overconfident", "score and noul not implemented", "DGUI 12 rows (was 6)", "INSTRUCT 119 rows likes 2", "encode the state once, decide everything in parallel", "0.740 accuracy against a 0.508 majority", "ECE 0.047", "fine-tune's advantage ends where its 384-token training data does", "jasonkneen/open-jev ≠ pngwn/open-jev", "same sha d41dc3cd", "Space does not call Jev", "recomputes routing from saved probabilities", "200-case Jev 97.0% / 100.0% / 95.0% / MAE 9.22", "synthetic repository benchmark", "Jev evaluations are advisory", "YehuiTang0316/jev-nlgrep ≠ Bentlybro/jevgrep ≠ can1357/jegrep ≠ uehaj/jev-semgrep", "default threshold 0.8 still soft", "40-line windows cannot prove whole function", "token-native sequential start/end Choice", "Gemini/Haiku stubs not configured yet", "handful of hand-written examples, not a benchmark", "Jev judged exactly what it was given", "laguagu/jev-skills ≠ laguagu/jev-evidence-lab ≠ Pleo2/awesome-jev-agent-skills", "contract_passed is not a claim of guaranteed factual truth", "Wilson lower bound 0.85 floor", "fixture mode no savings claim", "SemIf 2207★ (+21 vs §110 2186)", "jevlike 1043★ (+5 vs 1038)", "TypeAR 15★ (+1 vs 14)", "AnotiaWang 97★ (+1 vs 96)", "yibie/awesome-jev 506★ (+16 vs 490)", "Laya likes 822 (was 802)", "tracker likes 64 flat, lastModified UNCHANGED", "do not reopen or amend PR #23/#24/#25/#26/#27/#28", "Heman10x-NGU/Verdict-open-jev ≠ Heman10x-NGU/openJev-verdict-2.0", "TF-IDF + LogReg ECE 0.0207 vs Jev 0.1440", "Verdict-open-jev 48.07% vs Jev 90.80%", "abstention combined recall 10.00%", "p50 35.58 ms", "K=25 (maximum capacity) 72.00%", "0.85 coverage 84.60% selective risk 1.18%", "26.1× faster than standard Qwen JSON generation", "Jevify 90.0% / 167 ms CUDA graphs disabled", "Finding 1: Brier on stated confidence alone is a trap", "grpo_rlcr 0.78 / ECE 0.084", "reliability 0.007 but resolution 0.000", "27 900 schema-driven decisions", "13 600 / 13 600 questions", "candidate mass min 0.99999624", "22 configs · 166,054 rows · 4 calibration-gold", "sha a39eba3f", "Student B MAE 0.148 / Pearson 0.836 / 86.0%", "pngwn/open-jev-laya-bench README 404", "sha 9f69c742 likes 2", "HDFS 0.9933 (745/750) / retain 0.0084", "BGL ERROR/FATAL protection 1.0000", "2,479 / 2,500 HDFS uncertain", "cache hit 0.9648 (2412/2500)", "$0.153936 estimated", "E2 recomputes from saved probabilities", "Space sha eda59e0a", "MASSIVE English 0.783 / Khmer 0.033 / Hindi 0.133", "40–48 rows too small to ship T", "T never changes argmax", "siren2345/jev-single-decode-transformers ≠ siren2345/jev-single-decode", "Split Transformers experiment from llama.cpp runtime", "tanayvasishtha/jev-lab ≠ dairui1/jev-lab ≠ BrendanH18/jev-lab ≠ yibie/laya-jev-lab", "Four experiments stress-testing TypeSafe's Jev: calibration, bundle bias, label bias, and ensembling", "second pass must be $0.00 from cache", "The pages never call Jev", "Gemma 4 31B 77.0% / Jev 1.13.0 61.4% / Laya 322M 0.0%", "restriction state 95.0% against 84.4%", "None of the systems are particularly good at knowing when to stop and ask", "They skip the question and call a tool directly", "100% schema pass", "six-field joint 48.8% vs 72.8%", "ywchiu/jev_benchmark ≠ Running-Dolphins/jev-bench ≠ Praveenrajus/jev-bench", "ACT / REVIEW / FALLBACK", "A provider failure, timeout, malformed output, or missing answer is **not** a policy outcome", "confidence is descriptive provider output, not a substitute for probability", "Quality denominators include only valid scored answers", "an exact halfway tie chooses the lower level", "aiwithenoch/Jev-Skill ≠ simplosophy/jev-skill ≠ laguagu/jev-skills", "The local path does not claim to turn a smaller checkpoint into Jev", "Low support becomes decision: \"review\"", "MIT-0 SPDX NOASSERTION", "current-llm", "结构兼容,不是 Jev 模型能力", "altryne/jevify ≠ Mintzs/jevify ≠ gulagala001/jevify ≠ uspraveen/Jevify", "Find where Jev belongs. Design the questions. Measure the difference", "TypeAR-AI/TypeAR 301 → TypeLLM/TypeLLM", "TypeLLM/TypeLLM 16★", "SemIf 2241★ (+34 vs §111 2207)", "jevlike 1051★ (+8 vs 1043)", "AnotiaWang 98★ (+1 vs 97)", "yibie/awesome-jev 525★ (+19 vs 506)", "Laya likes 864 (was 822)", "tracker likes 67 (+3 vs 64)", "lastModified UNCHANGED `2026-09-20T04:29:16.000Z`", "do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#32", "hysteresis enter/exit / replay policy without inference", "calibration does not compose / hop-ECE permutation-invariant", "equal-width vs quantile ECE / ranking ≠ calibration", "Qwen2.5 ≠ Archer / Qwen 3.8 sparring ≠ Archer / Qwen/Qwen3.8-27B ≠ Archer", "Deferred Crispification / TCE / AMS", "g0runmezadam/what-is-jev IS tunahansahin897/what-is-jev", "pd.cut equal-width vs jeval quantile", "A hunch is a probability with a policy attached", "soundness theater / measurement theater / hourly 0843", , "Jev Capability Resolver / NiazMorshed2007/jcr", "one tool nested capability tree / returns context / does not execute", "skills vs capabilities / workflow+judgment vs operations", "format independent of Jev / proposed open standard", "JCR_BAND_RATIO 0.6 is application policy / soft scores ≠ hard gates", "routing ≠ permission / docs ≠ authority to run", "sol-vs-opus5-20 lookup+explain / n=1 / Not Harbor task-execution", "wall-time mixed / Sol slower with JCR in 19/20", "NiazMorshed2007/jcr ≠ skill-broker ≠ skillranker ≠ jev-sift ≠ jev-lens ≠ jevusher ≠ jev_select_capability", "do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33/#34", "notes.md §116", "copy the SemIf/MLX installer?", "quote 5.21× as beating Jev?", "treat 0.845 as a TypeSafe replica?", "collapse SemIf into kw2828/zhihz/semif-rs/semif-serve", "softmax over options as a Noul", "llm prompt to jev primitives", "conversion assistant not equivalent behavior", "heuristic conversion ≠ calibrated Noul", "alexwestco/llm-to-jev ≠ altryne/jevify", "user-provided 0940 / notes.md §118", "judge ≠ actuator", "candidate_mass", "softmax over A–H ≠ Noul", "hourly 0947 / notes.md §119", "ggmlc GGUF is not llama.cpp", "serving substrate ≠ calibrated replica", "Qwen3.5-9B ≠ Archer", "planner writes JEV selects", "hourly 1049 / notes.md §120", or "cascade sign-flip / calibration theater": read `references/faq.md`, then `references/mental-models.md`, then `references/mixed-architecture.md`, then `references/judgment-class.md` before any mapping. Proof, @@ -277,4 +277,15 @@ not a proof. Do not copy keys. Rebased onto `fb15455` reopen or amend PR #23–#40. Does not bump 0.5.0. Skip Archer. `invented_signal: false`. + +## Hourly 1049 HIGH (`notes.md` §120) + +ggmlc GGUF is not llama.cpp. serving substrate ≠ calibrated replica. +Qwen3.5-9B ≠ Archer. planner writes JEV selects. pick_by_id vs +pick_second. Soft scores ≠ hard gates. catalog ≠ endorsement. +Do not copy keys. Rebased onto `8f446c4` (merged #41) after +`fb15455` (merged #42 v0.5.0). Do not reopen or amend PR #23–#42. +Does not bump 0.5.0. Skip Archer. `invented_signal: false`. + Hourly 0947 uniqueness lock: Fast and cheap agent evals. jev as judge.; 18,041 skills from the 200 most-starred repos; Not a security scanner; 最简 Jev 调用演示器; confidence 不是正确率; q93304989-bit/jev-lab ≠ tanayvasishtha/jev-lab ≠ dairui1/jev-lab ≠ BrendanH18/jev-lab ≠ yibie/laya-jev-lab; 75% cheaper and 18% faster withdrawn; jev @0.15 100% recall 87% savings; 33Audits/jev-auto ≠ gargpratyush/jev-router; no Typesafe key, no PI_API_BASE, zero deps; tool-emitted Score/Noul ≠ calibrated Noul; semantic_compatibility: false; candidate_mass; Qwen3.5-2B ≠ Archer; Jev evaluates decisions; it cannot run a coding-agent session; Status: no model yet; S1LV3RJ1NX/openjev ≠ TheoLeeCJ/openjev; 28 accepted decisions; 3 targets; score 800; health 100; arcade game not a flight trainer; A successful live TypeSafe call has not been verified for v0.1.0; abhibansal60/tidy ≠ MANISH007700/tidy; No model, Jev included, predicted which channels its owner keeps; seed 1 selected on a held-out 400-item validation split; Brier 0.342 → 0.378; more accurate and more overconfident; Qwen3.5-4B ≠ Archer; static quants of kushalpatil/jevify-gemma4-26b-a4b; The labels were corrected, and one earlier result was retracted; zero of 23,869 eligible rows; Do not compare cost without checking task success; Exit 1 is not a proof; kisshan13/typesafe-ai-go ≠ Nibir1/typesafe-go ≠ official; 38 tests that cannot fail in a 356-model warehouse; if a parser can answer it, Jev is never asked; 359 of them; Games & Simulation 82; Education & Learning 1; Ratings are heuristics; syedabbasshaheer-art/jev-atlas ≠ ZeroX-01/jev-atlas ≠ Zaious/jev-capability-atlas ≠ gorock007/jev-atlas; anandi1989/awesome-jev-usecases ≠ whyashthakker/awesome-jev-use-cases ≠ walidboulanouar/awesome-jev-use-cases ≠ vamsikrishna2421/jev-usecases; Every headline result above is self-reported; Archer Hume 84.6% MMLU-Pro is a third-party probe not landed Archer; catalog ≠ endorsement; judge ≠ actuator; softmax over A–H ≠ Noul; SemIf 2270★; jevlike 1054★; TypeLLM/TypeLLM 16★; AnotiaWang 98★; yibie/awesome-jev 538★; Laya likes 889; tracker likes 68 lastModified UNCHANGED; Blackwood likes 2 gated manual; Archer still promised_not_landed; Hub archerhume/4rcherhume HTTP 401; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33/#35/#36/#37/#38/#40; do not push onto open #39; notes.md §119 +Hourly 1049 uniqueness lock: ggmlc GGUF is not llama.cpp; Loading them in llama.cpp will fail; one encoder pass; hf:mys/laya-GGUF sha 713ae6f6e39f likes 0 apache-2.0; hf:mys/laya-multilingual-GGUF sha 3b645ae54281; hf:mys/laya-typed-decisions-GGUF sha 1e9e8ba1f527; hf:tozp/laya-onnx sha 0862aeba1e65 Opset 14 FP32 and INT8; tozp/laya-onnx ≠ Mattepiu/laya-onnx ≠ gqgs/laya-onnx; docker-laya MIT HEAD 1b8239a51ddd README SHA 9cb7bdc3; laya.cpp RTX ggml CUDA HEAD 8590937c79a2 README SHA cdd429b9; serving substrate ≠ calibrated replica; Softmax over options ≠ calibrated Noul; Qwen3.5-4B ≠ Archer; Qwen3.5-9B ≠ Archer; jev-position-test n=6 HEAD 7a56ca1c2698 README SHA 23f194c9; jevmlx slots 5 of 6; hosted Jev 0 of 6; prior_correction made it worse; jevSweeper mean Spearman ρ −0.274; picked exact-optimal 1/25 (4%); 31 of 36 still logically decidable; 86% of the time we should not have been asking; game success ≠ calibrated Noul; LLM2Jev 64★ Apache-2.0 HEAD 924618721277 README SHA da35fe61; not affiliated with or endorsed by Jev or TypeSafe; No answer tokens are generated; OpenSourceJev llama.cpp Qwen3-1.7B HEAD 3c41fba3681d; JEV-MLX Qwen3.5-9B HEAD dec24cd929ea; decision-head-rlcd Qwen3.5-4B 4.9M LoRA; AUTO_ACT is not a Noul; closed-set fail-open stdlib-only; verified=False; soft scores ≠ hard gates; 22 to 40% cheaper *theirs*; first version 70% more expensive; 111-case benchmark *theirs*; CDC 5,000 Jev 72.2% AUC 0.7725 *theirs*; accuracy is a trap; 9.0% base rate always-no 91.0%; catalog ≠ endorsement; jev-skill 109★ 90 scenarios HEAD 4f6e899a24d4; awesome-jev-live 673 entries 4★; minecraft-agent 214★ 131 JEV decisions 35 Astra calls; nether-final-08 8 minutes 43.300 seconds; planner writes JEV selects; RoboJEV structured simulator state not images; ashare-trader 策略未通过自己的回测门槛; 36 组参数全部净期望为负; no positive expectation under real costs; typed_evals NOT an official TypeSafe AI product; jev-as-judge is a sensor; third-person-audit 40% & 60% watermarks still soft; The included experience uses a handwritten demo provider; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33/#34/#35/#36/#37/#38/#39/#40/#41/#42; notes.md §120 diff --git a/.agents/skills/augustus/references/agent-self-assessment.md b/.agents/skills/augustus/references/agent-self-assessment.md index f962aa7c..96579219 100644 --- a/.agents/skills/augustus/references/agent-self-assessment.md +++ b/.agents/skills/augustus/references/agent-self-assessment.md @@ -941,3 +941,7 @@ User-provided 0922 (`notes.md` §117 / items 330–336 / batch #100). SemIf was User-provided 0922 uniqueness lock: SemIf was formerly OpenJev; independent; not affiliated with Jev or TypeSafe; homepage openjev.com; default master; MIT; HEAD ca3ba65f1429; Tolerate float roundoff in MLX evidence verification; pushed 2026-09-19; live REST 2282★ / 140 forks; size 9177; README SHA 74ab7f7f; LICENSE SHA ca562883; interface pattern reproduction with open models; does not reproduce Jev undisclosed model/training; Direct option logits; 0 output tokens; shared-state parallel; MLX backend for Apple Silicon (`--backend mlx`); Qwen3.5-4B 3090 direct 1.023s vs AR JSON 5.332s (**5.21×**); argmax agree 18/21; systems comparison ≠ semantic equivalence; Parallel suffixes 20.03 dec/s on 777 decisions; Browser ladder Qwen3.5-4B authored BA 0.813, pert 0.766, TypeSafe subset agreement 0.845 vs Published Jev 0.883 (102 across 20 cases); Softmax over options ≠ calibrated Noul; typed output does not guarantee semantic correctness; wire/agreement ≠ replica of TypeSafe; SemIf ≠ kw2828/OpenJev playground ≠ zhihz/openjev ≠ apiplant/semif-rs port ≠ dddanielliu/semif-serve; rename is densify not a second census; JevBench 74.6 is §78 not this ladder; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33/#34/#35/#36/#38; do not push onto open #39/#40; notes.md §117 User-provided 0940 uniqueness lock: Turn decision-shaped LLM prompts into proposed Jev primitives; This is a conversion assistant, not an automatic guarantee of equivalent behavior; The compiler uses deterministic heuristics, not an LLM or evaluation model; It understands a deliberately small set of common prompt patterns; Generated instructions and criteria must be reviewed before production use; Score ranges such as 0 to 1 are translated into ordered Jev criteria; Prompts requiring open-ended prose are not a fit; suitability strong/partial/not_a_fit; compatibility full/partial/none; Writing new text stays with an LLM; Review the generated Score rubric; Jev scores ordered criteria, not an arbitrary 0-to-1 range; Everything runs locally in the browser; There is no framework, database, account, API, or server-side prompt processing; The key is read from the process environment and is never stored or printed; connect-src 'none'; alexwestco/llm-to-jev ≠ altryne/jevify ≠ ryana/jevify ≠ fidecastro/jevify ≠ Mintzs/jevify ≠ gulagala001/jevify ≠ uspraveen/Jevify; HEAD 234058ab372d; README SHA 43cd94fb; LICENSE SHA 5f334006; compiler SHA fdf235d0; 2★; MIT; JavaScript; size 29; Pages https://alexwestco.github.io/llm-to-jev/; invented_signal false; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33/#34/#35; notes.md §118 Hourly 0947 uniqueness lock: Fast and cheap agent evals. jev as judge.; 18,041 skills from the 200 most-starred repos; Not a security scanner; 最简 Jev 调用演示器; confidence 不是正确率; q93304989-bit/jev-lab ≠ tanayvasishtha/jev-lab ≠ dairui1/jev-lab ≠ BrendanH18/jev-lab ≠ yibie/laya-jev-lab; 75% cheaper and 18% faster withdrawn; jev @0.15 100% recall 87% savings; 33Audits/jev-auto ≠ gargpratyush/jev-router; no Typesafe key, no PI_API_BASE, zero deps; tool-emitted Score/Noul ≠ calibrated Noul; semantic_compatibility: false; candidate_mass; Qwen3.5-2B ≠ Archer; Jev evaluates decisions; it cannot run a coding-agent session; Status: no model yet; S1LV3RJ1NX/openjev ≠ TheoLeeCJ/openjev; 28 accepted decisions; 3 targets; score 800; health 100; arcade game not a flight trainer; A successful live TypeSafe call has not been verified for v0.1.0; abhibansal60/tidy ≠ MANISH007700/tidy; No model, Jev included, predicted which channels its owner keeps; seed 1 selected on a held-out 400-item validation split; Brier 0.342 → 0.378; more accurate and more overconfident; Qwen3.5-4B ≠ Archer; static quants of kushalpatil/jevify-gemma4-26b-a4b; The labels were corrected, and one earlier result was retracted; zero of 23,869 eligible rows; Do not compare cost without checking task success; Exit 1 is not a proof; kisshan13/typesafe-ai-go ≠ Nibir1/typesafe-go ≠ official; 38 tests that cannot fail in a 356-model warehouse; if a parser can answer it, Jev is never asked; 359 of them; Games & Simulation 82; Education & Learning 1; Ratings are heuristics; syedabbasshaheer-art/jev-atlas ≠ ZeroX-01/jev-atlas ≠ Zaious/jev-capability-atlas ≠ gorock007/jev-atlas; anandi1989/awesome-jev-usecases ≠ whyashthakker/awesome-jev-use-cases ≠ walidboulanouar/awesome-jev-use-cases ≠ vamsikrishna2421/jev-usecases; Every headline result above is self-reported; Archer Hume 84.6% MMLU-Pro is a third-party probe not landed Archer; catalog ≠ endorsement; judge ≠ actuator; softmax over A–H ≠ Noul; SemIf 2270★; jevlike 1054★; TypeLLM/TypeLLM 16★; AnotiaWang 98★; yibie/awesome-jev 538★; Laya likes 889; tracker likes 68 lastModified UNCHANGED; Blackwood likes 2 gated manual; Archer still promised_not_landed; Hub archerhume/4rcherhume HTTP 401; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33/#35/#36/#37/#38/#40; do not push onto open #39; notes.md §119 + + +**Hourly 1049 HIGH (`notes.md` §120).** ggmlc GGUF is not llama.cpp. serving substrate ≠ calibrated replica. Qwen3.5-9B ≠ Archer. planner writes JEV selects. pick_by_id vs pick_second. soft scores ≠ hard gates. catalog ≠ endorsement. Do not reopen or amend PR #23–#42. Does not bump 0.5.0. Skip Archer. `invented_signal: false`. +Hourly 1049 uniqueness lock: ggmlc GGUF is not llama.cpp; Loading them in llama.cpp will fail; one encoder pass; hf:mys/laya-GGUF sha 713ae6f6e39f likes 0 apache-2.0; hf:mys/laya-multilingual-GGUF sha 3b645ae54281; hf:mys/laya-typed-decisions-GGUF sha 1e9e8ba1f527; hf:tozp/laya-onnx sha 0862aeba1e65 Opset 14 FP32 and INT8; tozp/laya-onnx ≠ Mattepiu/laya-onnx ≠ gqgs/laya-onnx; docker-laya MIT HEAD 1b8239a51ddd README SHA 9cb7bdc3; laya.cpp RTX ggml CUDA HEAD 8590937c79a2 README SHA cdd429b9; serving substrate ≠ calibrated replica; Softmax over options ≠ calibrated Noul; Qwen3.5-4B ≠ Archer; Qwen3.5-9B ≠ Archer; jev-position-test n=6 HEAD 7a56ca1c2698 README SHA 23f194c9; jevmlx slots 5 of 6; hosted Jev 0 of 6; prior_correction made it worse; jevSweeper mean Spearman ρ −0.274; picked exact-optimal 1/25 (4%); 31 of 36 still logically decidable; 86% of the time we should not have been asking; game success ≠ calibrated Noul; LLM2Jev 64★ Apache-2.0 HEAD 924618721277 README SHA da35fe61; not affiliated with or endorsed by Jev or TypeSafe; No answer tokens are generated; OpenSourceJev llama.cpp Qwen3-1.7B HEAD 3c41fba3681d; JEV-MLX Qwen3.5-9B HEAD dec24cd929ea; decision-head-rlcd Qwen3.5-4B 4.9M LoRA; AUTO_ACT is not a Noul; closed-set fail-open stdlib-only; verified=False; soft scores ≠ hard gates; 22 to 40% cheaper *theirs*; first version 70% more expensive; 111-case benchmark *theirs*; CDC 5,000 Jev 72.2% AUC 0.7725 *theirs*; accuracy is a trap; 9.0% base rate always-no 91.0%; catalog ≠ endorsement; jev-skill 109★ 90 scenarios HEAD 4f6e899a24d4; awesome-jev-live 673 entries 4★; minecraft-agent 214★ 131 JEV decisions 35 Astra calls; nether-final-08 8 minutes 43.300 seconds; planner writes JEV selects; RoboJEV structured simulator state not images; ashare-trader 策略未通过自己的回测门槛; 36 组参数全部净期望为负; no positive expectation under real costs; typed_evals NOT an official TypeSafe AI product; jev-as-judge is a sensor; third-person-audit 40% & 60% watermarks still soft; The included experience uses a handwritten demo provider; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33/#34/#35/#36/#37/#38/#39/#40/#41/#42; notes.md §120 diff --git a/.agents/skills/augustus/references/applied-mappings.md b/.agents/skills/augustus/references/applied-mappings.md index 500f5649..52718558 100644 --- a/.agents/skills/augustus/references/applied-mappings.md +++ b/.agents/skills/augustus/references/applied-mappings.md @@ -2520,3 +2520,7 @@ User-provided 0922 (`notes.md` §117 / items 330–336 / batch #100). SemIf was User-provided 0922 uniqueness lock: SemIf was formerly OpenJev; independent; not affiliated with Jev or TypeSafe; homepage openjev.com; default master; MIT; HEAD ca3ba65f1429; Tolerate float roundoff in MLX evidence verification; pushed 2026-09-19; live REST 2282★ / 140 forks; size 9177; README SHA 74ab7f7f; LICENSE SHA ca562883; interface pattern reproduction with open models; does not reproduce Jev undisclosed model/training; Direct option logits; 0 output tokens; shared-state parallel; MLX backend for Apple Silicon (`--backend mlx`); Qwen3.5-4B 3090 direct 1.023s vs AR JSON 5.332s (**5.21×**); argmax agree 18/21; systems comparison ≠ semantic equivalence; Parallel suffixes 20.03 dec/s on 777 decisions; Browser ladder Qwen3.5-4B authored BA 0.813, pert 0.766, TypeSafe subset agreement 0.845 vs Published Jev 0.883 (102 across 20 cases); Softmax over options ≠ calibrated Noul; typed output does not guarantee semantic correctness; wire/agreement ≠ replica of TypeSafe; SemIf ≠ kw2828/OpenJev playground ≠ zhihz/openjev ≠ apiplant/semif-rs port ≠ dddanielliu/semif-serve; rename is densify not a second census; JevBench 74.6 is §78 not this ladder; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33/#34/#35/#36/#38; do not push onto open #39/#40; notes.md §117 User-provided 0940 uniqueness lock: Turn decision-shaped LLM prompts into proposed Jev primitives; This is a conversion assistant, not an automatic guarantee of equivalent behavior; The compiler uses deterministic heuristics, not an LLM or evaluation model; It understands a deliberately small set of common prompt patterns; Generated instructions and criteria must be reviewed before production use; Score ranges such as 0 to 1 are translated into ordered Jev criteria; Prompts requiring open-ended prose are not a fit; suitability strong/partial/not_a_fit; compatibility full/partial/none; Writing new text stays with an LLM; Review the generated Score rubric; Jev scores ordered criteria, not an arbitrary 0-to-1 range; Everything runs locally in the browser; There is no framework, database, account, API, or server-side prompt processing; The key is read from the process environment and is never stored or printed; connect-src 'none'; alexwestco/llm-to-jev ≠ altryne/jevify ≠ ryana/jevify ≠ fidecastro/jevify ≠ Mintzs/jevify ≠ gulagala001/jevify ≠ uspraveen/Jevify; HEAD 234058ab372d; README SHA 43cd94fb; LICENSE SHA 5f334006; compiler SHA fdf235d0; 2★; MIT; JavaScript; size 29; Pages https://alexwestco.github.io/llm-to-jev/; invented_signal false; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33/#34/#35; notes.md §118 Hourly 0947 uniqueness lock: Fast and cheap agent evals. jev as judge.; 18,041 skills from the 200 most-starred repos; Not a security scanner; 最简 Jev 调用演示器; confidence 不是正确率; q93304989-bit/jev-lab ≠ tanayvasishtha/jev-lab ≠ dairui1/jev-lab ≠ BrendanH18/jev-lab ≠ yibie/laya-jev-lab; 75% cheaper and 18% faster withdrawn; jev @0.15 100% recall 87% savings; 33Audits/jev-auto ≠ gargpratyush/jev-router; no Typesafe key, no PI_API_BASE, zero deps; tool-emitted Score/Noul ≠ calibrated Noul; semantic_compatibility: false; candidate_mass; Qwen3.5-2B ≠ Archer; Jev evaluates decisions; it cannot run a coding-agent session; Status: no model yet; S1LV3RJ1NX/openjev ≠ TheoLeeCJ/openjev; 28 accepted decisions; 3 targets; score 800; health 100; arcade game not a flight trainer; A successful live TypeSafe call has not been verified for v0.1.0; abhibansal60/tidy ≠ MANISH007700/tidy; No model, Jev included, predicted which channels its owner keeps; seed 1 selected on a held-out 400-item validation split; Brier 0.342 → 0.378; more accurate and more overconfident; Qwen3.5-4B ≠ Archer; static quants of kushalpatil/jevify-gemma4-26b-a4b; The labels were corrected, and one earlier result was retracted; zero of 23,869 eligible rows; Do not compare cost without checking task success; Exit 1 is not a proof; kisshan13/typesafe-ai-go ≠ Nibir1/typesafe-go ≠ official; 38 tests that cannot fail in a 356-model warehouse; if a parser can answer it, Jev is never asked; 359 of them; Games & Simulation 82; Education & Learning 1; Ratings are heuristics; syedabbasshaheer-art/jev-atlas ≠ ZeroX-01/jev-atlas ≠ Zaious/jev-capability-atlas ≠ gorock007/jev-atlas; anandi1989/awesome-jev-usecases ≠ whyashthakker/awesome-jev-use-cases ≠ walidboulanouar/awesome-jev-use-cases ≠ vamsikrishna2421/jev-usecases; Every headline result above is self-reported; Archer Hume 84.6% MMLU-Pro is a third-party probe not landed Archer; catalog ≠ endorsement; judge ≠ actuator; softmax over A–H ≠ Noul; SemIf 2270★; jevlike 1054★; TypeLLM/TypeLLM 16★; AnotiaWang 98★; yibie/awesome-jev 538★; Laya likes 889; tracker likes 68 lastModified UNCHANGED; Blackwood likes 2 gated manual; Archer still promised_not_landed; Hub archerhume/4rcherhume HTTP 401; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33/#35/#36/#37/#38/#40; do not push onto open #39; notes.md §119 + + +**Hourly 1049 HIGH (`notes.md` §120).** ggmlc GGUF is not llama.cpp. serving substrate ≠ calibrated replica. Qwen3.5-9B ≠ Archer. planner writes JEV selects. pick_by_id vs pick_second. soft scores ≠ hard gates. catalog ≠ endorsement. Do not reopen or amend PR #23–#42. Does not bump 0.5.0. Skip Archer. `invented_signal: false`. +Hourly 1049 uniqueness lock: ggmlc GGUF is not llama.cpp; Loading them in llama.cpp will fail; one encoder pass; hf:mys/laya-GGUF sha 713ae6f6e39f likes 0 apache-2.0; hf:mys/laya-multilingual-GGUF sha 3b645ae54281; hf:mys/laya-typed-decisions-GGUF sha 1e9e8ba1f527; hf:tozp/laya-onnx sha 0862aeba1e65 Opset 14 FP32 and INT8; tozp/laya-onnx ≠ Mattepiu/laya-onnx ≠ gqgs/laya-onnx; docker-laya MIT HEAD 1b8239a51ddd README SHA 9cb7bdc3; laya.cpp RTX ggml CUDA HEAD 8590937c79a2 README SHA cdd429b9; serving substrate ≠ calibrated replica; Softmax over options ≠ calibrated Noul; Qwen3.5-4B ≠ Archer; Qwen3.5-9B ≠ Archer; jev-position-test n=6 HEAD 7a56ca1c2698 README SHA 23f194c9; jevmlx slots 5 of 6; hosted Jev 0 of 6; prior_correction made it worse; jevSweeper mean Spearman ρ −0.274; picked exact-optimal 1/25 (4%); 31 of 36 still logically decidable; 86% of the time we should not have been asking; game success ≠ calibrated Noul; LLM2Jev 64★ Apache-2.0 HEAD 924618721277 README SHA da35fe61; not affiliated with or endorsed by Jev or TypeSafe; No answer tokens are generated; OpenSourceJev llama.cpp Qwen3-1.7B HEAD 3c41fba3681d; JEV-MLX Qwen3.5-9B HEAD dec24cd929ea; decision-head-rlcd Qwen3.5-4B 4.9M LoRA; AUTO_ACT is not a Noul; closed-set fail-open stdlib-only; verified=False; soft scores ≠ hard gates; 22 to 40% cheaper *theirs*; first version 70% more expensive; 111-case benchmark *theirs*; CDC 5,000 Jev 72.2% AUC 0.7725 *theirs*; accuracy is a trap; 9.0% base rate always-no 91.0%; catalog ≠ endorsement; jev-skill 109★ 90 scenarios HEAD 4f6e899a24d4; awesome-jev-live 673 entries 4★; minecraft-agent 214★ 131 JEV decisions 35 Astra calls; nether-final-08 8 minutes 43.300 seconds; planner writes JEV selects; RoboJEV structured simulator state not images; ashare-trader 策略未通过自己的回测门槛; 36 组参数全部净期望为负; no positive expectation under real costs; typed_evals NOT an official TypeSafe AI product; jev-as-judge is a sensor; third-person-audit 40% & 60% watermarks still soft; The included experience uses a handwritten demo provider; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33/#34/#35/#36/#37/#38/#39/#40/#41/#42; notes.md §120 diff --git a/.agents/skills/augustus/references/composition-algebra.md b/.agents/skills/augustus/references/composition-algebra.md index 3eaf54f9..8bf2a3fe 100644 --- a/.agents/skills/augustus/references/composition-algebra.md +++ b/.agents/skills/augustus/references/composition-algebra.md @@ -2250,6 +2250,78 @@ Hourly 0743 uniqueness lock: Heman10x-NGU/Verdict-open-jev ≠ Heman10x-NGU/open syedabbasshaheer-art/jev-atlas ≠ ZeroX-01/jev-atlas ≠ Zaious/jev-capability-atlas ≠ gorock007/jev-atlas. Full cards: `faq.md`, `mixed-architecture.md`. + +353. **ggmlc GGUF serving substrate PRIMARY** (mys/laya-GGUF family): + positions 1 (Operand) × 10 (Discretizer). ggmlc GGUF is not llama.cpp. + Loading them in llama.cpp will fail. one encoder pass. + serving substrate ≠ calibrated replica. + Full cards: `judgment-class.md`, `faq.md`. +354. **ONNX cousins** (tozp/laya-onnx): + namesake lock. Opset 14 FP32 and INT8. + tozp/laya-onnx ≠ Mattepiu/laya-onnx ≠ gqgs/laya-onnx. + Softmax over options ≠ calibrated Noul. + Full cards: `faq.md`, `judgment-class.md`. +355. **docker-laya serving** (chneau/docker-laya): + position 7 (Policy of the tool itself). Dockerized FastAPI typed decisions. + serving substrate ≠ calibrated replica. Do not copy `docker pull`. + Full cards: `mixed-architecture.md`. +356. **laya.cpp RTX ggml CUDA** (lkarlslund/laya.cpp): + position 7. Native C++ inference. Systems throughput ≠ semantic equivalence. + Full cards: `judgment-class.md`, `validation.md`. +357. **option-order measurement PRIMARY** (imaddde867/jev-position-test): + position 8 (Metric). n=6. jevmlx slots 5 of 6. hosted Jev 0 of 6. + prior_correction made it worse. pick_by_id vs pick_second. + Full cards: `validation.md`, `faq.md`. +358. **exact-p Minesweeper** (lvk901/jevSweeper): + position 8 (Metric). mean Spearman ρ −0.274. picked exact-optimal 1/25 (4%). + 31 of 36 still logically decidable. 86% of the time we should not have been asking. + game success ≠ calibrated Noul. Full cards: `validation.md`, `formal-methods.md`. +359. **LLM2Jev adapter** (Yinsongxu/LLM2Jev): + position 1 (Operand). 64★ Apache-2.0. not affiliated with or endorsed by Jev or TypeSafe. + No answer tokens are generated. Full cards: `judgment-class.md`, `faq.md`. +360. **OpenSourceJev llama.cpp** (sabeel111/OpenSourceJev): + namesake lock. llama.cpp Qwen3-1.7B. Candidate-only softmax ≠ Noul. + Full cards: `judgment-class.md`. +361. **JEV-MLX Qwen3.5-9B** (CoderInPajamas/JEV-MLX): + position 1. Qwen3.5-9B ≠ Archer. scores not calibrated probabilities of correctness. + Full cards: `judgment-class.md`, `validation.md`. +362. **decision-head-rlcd densify** (Astro-Han/decision-head-rlcd): + densify §119. Qwen3.5-4B 4.9M LoRA. Brier 0.342 → 0.378. + Qwen3.5-4B ≠ Archer. Full cards: `validation.md`. +363. **fail-open harness + handwritten demo** (litshing/jevcore + ashleyotooligan/jevbrain): + position 3 (Guard). AUTO_ACT is not a Noul. closed-set fail-open stdlib-only. + verified=False. The included experience uses a handwritten demo provider. + Full cards: `mixed-architecture.md`, `faq.md`. +364. **nitro tool-gate economics** (daniel-farina/nitro): + position 6 (Router). 22 to 40% cheaper *theirs*. first version 70% more expensive. + 0.30 keep-set is application policy. Full cards: `validation.md`, `mixed-architecture.md`. +365. **soft watermarks / BLOCK bands** (healthcare + guardian + TPA): + position 3 (Guard). urgency 0.92 still soft. BLOCK / QUARANTINE still soft. + third-person-audit 40% & 60% watermarks still soft. + Full cards: `mixed-architecture.md`, `agent-self-assessment.md`. +366. **catalogs** (wuyoscar/jev-skill + wh000wh000/awesome-jev-live): + catalog ≠ endorsement. 109★ 90 scenarios. 673 entries 4★. + Full cards: `faq.md`, `applied-mappings.md`. +367. **planner writes JEV selects** (rmalde/minecraft-agent + lykycy123/RoboJEV): + positions 4 (Selector) × 7 (Controller). 214★. 131 JEV decisions. 35 Astra calls. + nether-final-08 8 minutes 43.300 seconds. structured simulator state not images. + Full cards: `mixed-architecture.md`, `applied-mappings.md`. +368. **negative-EV honesty + jev-as-judge sensor** (xuboboo/ashare-trader + TrustifAI/typed_evals + heart-risk + fabric): + 策略未通过自己的回测门槛. 36 组参数全部净期望为负. no positive expectation under real costs. + typed_evals NOT an official TypeSafe AI product. jev-as-judge is a sensor. + CDC 5,000 Jev 72.2% AUC 0.7725 *theirs*. accuracy is a trap. 9.0% base rate always-no 91.0%. + 111-case benchmark *theirs*. Full cards: `validation.md`, `faq.md`. + +Hourly 1049 items 353–368 (`notes.md` §120). Do **not** +re-fold §119 items 337–352 / §118 items 322–329 / §117 items 330–336 +/ §116 items 309–316 / §115 items 303–308 / §114 items 289–302. +Skip Archer rewrite. +Qwen3.5-4B ≠ Archer. Qwen3.5-9B ≠ Archer. +serving substrate ≠ calibrated replica; planner writes JEV selects; +catalog ≠ endorsement; pick_by_id vs pick_second. +do not reopen or amend PR #23–#42. +Soft Noul ≠ hard safety. + Hourly 0947 items 337–352 (`notes.md` §119). Do **not** re-fold §118 items 322–329 / §117 items 330–336 / §116 items 309–316 / §115 items 303–308 / §114 items 289–302 / §113 items 273–288. @@ -2345,3 +2417,4 @@ User-provided 0920 jcr uniqueness lock: NiazMorshed2007/jcr MIT; site https://jc User-provided 0922 uniqueness lock: SemIf was formerly OpenJev; independent; not affiliated with Jev or TypeSafe; homepage openjev.com; default master; MIT; HEAD ca3ba65f1429; Tolerate float roundoff in MLX evidence verification; pushed 2026-09-19; live REST 2282★ / 140 forks; size 9177; README SHA 74ab7f7f; LICENSE SHA ca562883; interface pattern reproduction with open models; does not reproduce Jev undisclosed model/training; Direct option logits; 0 output tokens; shared-state parallel; MLX backend for Apple Silicon (`--backend mlx`); Qwen3.5-4B 3090 direct 1.023s vs AR JSON 5.332s (**5.21×**); argmax agree 18/21; systems comparison ≠ semantic equivalence; Parallel suffixes 20.03 dec/s on 777 decisions; Browser ladder Qwen3.5-4B authored BA 0.813, pert 0.766, TypeSafe subset agreement 0.845 vs Published Jev 0.883 (102 across 20 cases); Softmax over options ≠ calibrated Noul; typed output does not guarantee semantic correctness; wire/agreement ≠ replica of TypeSafe; SemIf ≠ kw2828/OpenJev playground ≠ zhihz/openjev ≠ apiplant/semif-rs port ≠ dddanielliu/semif-serve; rename is densify not a second census; JevBench 74.6 is §78 not this ladder; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33/#34/#35/#36/#38; do not push onto open #39/#40; notes.md §117 User-provided 0940 uniqueness lock: Turn decision-shaped LLM prompts into proposed Jev primitives; This is a conversion assistant, not an automatic guarantee of equivalent behavior; The compiler uses deterministic heuristics, not an LLM or evaluation model; It understands a deliberately small set of common prompt patterns; Generated instructions and criteria must be reviewed before production use; Score ranges such as 0 to 1 are translated into ordered Jev criteria; Prompts requiring open-ended prose are not a fit; suitability strong/partial/not_a_fit; compatibility full/partial/none; Writing new text stays with an LLM; Review the generated Score rubric; Jev scores ordered criteria, not an arbitrary 0-to-1 range; Everything runs locally in the browser; There is no framework, database, account, API, or server-side prompt processing; The key is read from the process environment and is never stored or printed; connect-src 'none'; alexwestco/llm-to-jev ≠ altryne/jevify ≠ ryana/jevify ≠ fidecastro/jevify ≠ Mintzs/jevify ≠ gulagala001/jevify ≠ uspraveen/Jevify; HEAD 234058ab372d; README SHA 43cd94fb; LICENSE SHA 5f334006; compiler SHA fdf235d0; 2★; MIT; JavaScript; size 29; Pages https://alexwestco.github.io/llm-to-jev/; invented_signal false; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33/#34/#35; notes.md §118 Hourly 0947 uniqueness lock: Fast and cheap agent evals. jev as judge.; 18,041 skills from the 200 most-starred repos; Not a security scanner; 最简 Jev 调用演示器; confidence 不是正确率; q93304989-bit/jev-lab ≠ tanayvasishtha/jev-lab ≠ dairui1/jev-lab ≠ BrendanH18/jev-lab ≠ yibie/laya-jev-lab; 75% cheaper and 18% faster withdrawn; jev @0.15 100% recall 87% savings; 33Audits/jev-auto ≠ gargpratyush/jev-router; no Typesafe key, no PI_API_BASE, zero deps; tool-emitted Score/Noul ≠ calibrated Noul; semantic_compatibility: false; candidate_mass; Qwen3.5-2B ≠ Archer; Jev evaluates decisions; it cannot run a coding-agent session; Status: no model yet; S1LV3RJ1NX/openjev ≠ TheoLeeCJ/openjev; 28 accepted decisions; 3 targets; score 800; health 100; arcade game not a flight trainer; A successful live TypeSafe call has not been verified for v0.1.0; abhibansal60/tidy ≠ MANISH007700/tidy; No model, Jev included, predicted which channels its owner keeps; seed 1 selected on a held-out 400-item validation split; Brier 0.342 → 0.378; more accurate and more overconfident; Qwen3.5-4B ≠ Archer; static quants of kushalpatil/jevify-gemma4-26b-a4b; The labels were corrected, and one earlier result was retracted; zero of 23,869 eligible rows; Do not compare cost without checking task success; Exit 1 is not a proof; kisshan13/typesafe-ai-go ≠ Nibir1/typesafe-go ≠ official; 38 tests that cannot fail in a 356-model warehouse; if a parser can answer it, Jev is never asked; 359 of them; Games & Simulation 82; Education & Learning 1; Ratings are heuristics; syedabbasshaheer-art/jev-atlas ≠ ZeroX-01/jev-atlas ≠ Zaious/jev-capability-atlas ≠ gorock007/jev-atlas; anandi1989/awesome-jev-usecases ≠ whyashthakker/awesome-jev-use-cases ≠ walidboulanouar/awesome-jev-use-cases ≠ vamsikrishna2421/jev-usecases; Every headline result above is self-reported; Archer Hume 84.6% MMLU-Pro is a third-party probe not landed Archer; catalog ≠ endorsement; judge ≠ actuator; softmax over A–H ≠ Noul; SemIf 2270★; jevlike 1054★; TypeLLM/TypeLLM 16★; AnotiaWang 98★; yibie/awesome-jev 538★; Laya likes 889; tracker likes 68 lastModified UNCHANGED; Blackwood likes 2 gated manual; Archer still promised_not_landed; Hub archerhume/4rcherhume HTTP 401; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33/#35/#36/#37/#38/#40; do not push onto open #39; notes.md §119 +Hourly 1049 uniqueness lock: ggmlc GGUF is not llama.cpp; Loading them in llama.cpp will fail; one encoder pass; hf:mys/laya-GGUF sha 713ae6f6e39f likes 0 apache-2.0; hf:mys/laya-multilingual-GGUF sha 3b645ae54281; hf:mys/laya-typed-decisions-GGUF sha 1e9e8ba1f527; hf:tozp/laya-onnx sha 0862aeba1e65 Opset 14 FP32 and INT8; tozp/laya-onnx ≠ Mattepiu/laya-onnx ≠ gqgs/laya-onnx; docker-laya MIT HEAD 1b8239a51ddd README SHA 9cb7bdc3; laya.cpp RTX ggml CUDA HEAD 8590937c79a2 README SHA cdd429b9; serving substrate ≠ calibrated replica; Softmax over options ≠ calibrated Noul; Qwen3.5-4B ≠ Archer; Qwen3.5-9B ≠ Archer; jev-position-test n=6 HEAD 7a56ca1c2698 README SHA 23f194c9; jevmlx slots 5 of 6; hosted Jev 0 of 6; prior_correction made it worse; jevSweeper mean Spearman ρ −0.274; picked exact-optimal 1/25 (4%); 31 of 36 still logically decidable; 86% of the time we should not have been asking; game success ≠ calibrated Noul; LLM2Jev 64★ Apache-2.0 HEAD 924618721277 README SHA da35fe61; not affiliated with or endorsed by Jev or TypeSafe; No answer tokens are generated; OpenSourceJev llama.cpp Qwen3-1.7B HEAD 3c41fba3681d; JEV-MLX Qwen3.5-9B HEAD dec24cd929ea; decision-head-rlcd Qwen3.5-4B 4.9M LoRA; AUTO_ACT is not a Noul; closed-set fail-open stdlib-only; verified=False; soft scores ≠ hard gates; 22 to 40% cheaper *theirs*; first version 70% more expensive; 111-case benchmark *theirs*; CDC 5,000 Jev 72.2% AUC 0.7725 *theirs*; accuracy is a trap; 9.0% base rate always-no 91.0%; catalog ≠ endorsement; jev-skill 109★ 90 scenarios HEAD 4f6e899a24d4; awesome-jev-live 673 entries 4★; minecraft-agent 214★ 131 JEV decisions 35 Astra calls; nether-final-08 8 minutes 43.300 seconds; planner writes JEV selects; RoboJEV structured simulator state not images; ashare-trader 策略未通过自己的回测门槛; 36 组参数全部净期望为负; no positive expectation under real costs; typed_evals NOT an official TypeSafe AI product; jev-as-judge is a sensor; third-person-audit 40% & 60% watermarks still soft; The included experience uses a handwritten demo provider; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33/#34/#35/#36/#37/#38/#39/#40/#41/#42; notes.md §120 diff --git a/.agents/skills/augustus/references/faq.md b/.agents/skills/augustus/references/faq.md index 3508c5dd..8ae37ce4 100644 --- a/.agents/skills/augustus/references/faq.md +++ b/.agents/skills/augustus/references/faq.md @@ -3660,4 +3660,27 @@ No. SemIf ≠ kw2828/OpenJev playground ≠ zhihz/openjev ≠ apiplant/semif-rs User-provided 0920 jcr uniqueness lock: NiazMorshed2007/jcr MIT; site https://jcr.niazmorshed.dev; topics ai-agents,jev,mcp; **4★**; HEAD `138b3832`; README SHA `2a49dbc1`; LICENSE SHA `46231303`; size **14850**; Jev Capability Resolver; one tool to find documented deterministic commands in a nested capability tree; returns context; **does not execute**; skills = workflow+judgment; capabilities = individual operations; format independent of Jev; proposed open standard exploration; classify (Jev) → optional OpenAI decompose compound → beam search geometric mean of routing probs; keep up to 3 paths ≥60% of best (JCR_BAND_RATIO 0.6); ambiguity / no-match / depth-limit explicit; soft scores ≠ hard gates; 0.6 band is application policy; routing ≠ permission; docs ≠ authority to run; sol-vs-opus5-20 *theirs*: 20 scenarios × 4 variants = 80 runs; lookup+explain only, no execution; Claude Opus 5: agent input 108,585→15,819 (−85%), cost $0.3700→$0.1222 (−67%), wall 105.5s→77.7s; Codex GPT-5.6-Sol: 61,952→47,669 (−23%), $0.1377→$0.1151 (−16%), wall 25.3s→62.4s (Sol slower with JCR in 19/20); One Sol outlier 372.6s / 193 Jev calls; n=1 per cell; Not Harbor task-execution; Claude/Codex harnesses; compare mode; 50 scenarios bundled; 11 groups, 960 nodes, 11,360 items; 16 routing rounds per step; NiazMorshed2007/jcr ≠ skill-broker ≠ skillranker ≠ jev-sift ≠ jev-lens ≠ jevusher ≠ jev_select_capability; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33/#34; notes.md §116 User-provided 0922 uniqueness lock: SemIf was formerly OpenJev; independent; not affiliated with Jev or TypeSafe; homepage openjev.com; default master; MIT; HEAD ca3ba65f1429; Tolerate float roundoff in MLX evidence verification; pushed 2026-09-19; live REST 2282★ / 140 forks; size 9177; README SHA 74ab7f7f; LICENSE SHA ca562883; interface pattern reproduction with open models; does not reproduce Jev undisclosed model/training; Direct option logits; 0 output tokens; shared-state parallel; MLX backend for Apple Silicon (`--backend mlx`); Qwen3.5-4B 3090 direct 1.023s vs AR JSON 5.332s (**5.21×**); argmax agree 18/21; systems comparison ≠ semantic equivalence; Parallel suffixes 20.03 dec/s on 777 decisions; Browser ladder Qwen3.5-4B authored BA 0.813, pert 0.766, TypeSafe subset agreement 0.845 vs Published Jev 0.883 (102 across 20 cases); Softmax over options ≠ calibrated Noul; typed output does not guarantee semantic correctness; wire/agreement ≠ replica of TypeSafe; SemIf ≠ kw2828/OpenJev playground ≠ zhihz/openjev ≠ apiplant/semif-rs port ≠ dddanielliu/semif-serve; rename is densify not a second census; JevBench 74.6 is §78 not this ladder; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33/#34/#35/#36/#38; do not push onto open #39/#40; notes.md §117 User-provided 0940 uniqueness lock: Turn decision-shaped LLM prompts into proposed Jev primitives; This is a conversion assistant, not an automatic guarantee of equivalent behavior; The compiler uses deterministic heuristics, not an LLM or evaluation model; It understands a deliberately small set of common prompt patterns; Generated instructions and criteria must be reviewed before production use; Score ranges such as 0 to 1 are translated into ordered Jev criteria; Prompts requiring open-ended prose are not a fit; suitability strong/partial/not_a_fit; compatibility full/partial/none; Writing new text stays with an LLM; Review the generated Score rubric; Jev scores ordered criteria, not an arbitrary 0-to-1 range; Everything runs locally in the browser; There is no framework, database, account, API, or server-side prompt processing; The key is read from the process environment and is never stored or printed; connect-src 'none'; alexwestco/llm-to-jev ≠ altryne/jevify ≠ ryana/jevify ≠ fidecastro/jevify ≠ Mintzs/jevify ≠ gulagala001/jevify ≠ uspraveen/Jevify; HEAD 234058ab372d; README SHA 43cd94fb; LICENSE SHA 5f334006; compiler SHA fdf235d0; 2★; MIT; JavaScript; size 29; Pages https://alexwestco.github.io/llm-to-jev/; invented_signal false; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33/#34/#35; notes.md §118 + + +## Treat ggmlc GGUF as llama.cpp? Treat ONNX / Docker / C++ serving as a calibrated replica? + +No and no. ggmlc GGUF is not llama.cpp. Loading them in llama.cpp will fail. one encoder pass. serving substrate ≠ calibrated replica. Softmax over options ≠ calibrated Noul. tozp/laya-onnx ≠ Mattepiu/laya-onnx ≠ gqgs/laya-onnx. `notes.md` §120. + +## Treat option-order flips as a Noul? Treat Minesweeper ρ as "the model cannot play"? Treat pick_second as pick_by_id? + +No, no, and no. jev-position-test n=6. jevmlx slots 5 of 6. hosted Jev 0 of 6. prior_correction made it worse. jevSweeper mean Spearman ρ −0.274. picked exact-optimal 1/25 (4%). 31 of 36 still logically decidable. 86% of the time we should not have been asking. game success ≠ calibrated Noul. pick_by_id vs pick_second. `notes.md` §120. + +## Treat LLM2Jev / OpenSourceJev / JEV-MLX / RLCD LoRA as TypeSafe Jev? Treat Qwen3.5-9B as Archer? + +No and no. not affiliated with or endorsed by Jev or TypeSafe. No answer tokens are generated. OpenSourceJev llama.cpp Qwen3-1.7B. JEV-MLX Qwen3.5-9B. decision-head-rlcd Qwen3.5-4B 4.9M LoRA. Qwen3.5-4B ≠ Archer. Qwen3.5-9B ≠ Archer. `notes.md` §120. + +## Treat AUTO_ACT as a Noul? Treat 0.92 urgency / BLOCK / 40% watermarks as hard gates? Treat 22 to 40% cheaper as Harbor? + +No across the board. AUTO_ACT is not a Noul. closed-set fail-open stdlib-only. verified=False. soft scores ≠ hard gates. 22 to 40% cheaper *theirs*. first version 70% more expensive. third-person-audit 40% & 60% watermarks still soft. The included experience uses a handwritten demo provider. `notes.md` §120. + +## Treat jev-skill / awesome-jev-live as endorsement? Treat planner writes as JEV writes? Treat ashare-trader as a working edge? Treat typed_evals as official TypeSafe? + +No. catalog ≠ endorsement. jev-skill 109★ 90 scenarios. awesome-jev-live 673 entries 4★. planner writes JEV selects. minecraft-agent 214★ 131 JEV decisions 35 Astra calls. nether-final-08 8 minutes 43.300 seconds. RoboJEV structured simulator state not images. 策略未通过自己的回测门槛. 36 组参数全部净期望为负. no positive expectation under real costs. typed_evals NOT an official TypeSafe AI product. jev-as-judge is a sensor. CDC 5,000 Jev 72.2% AUC 0.7725 *theirs*. accuracy is a trap. 9.0% base rate always-no 91.0%. 111-case benchmark *theirs*. Do not reopen or amend PR #23–#42. `invented_signal: false`. `notes.md` §120. + Hourly 0947 uniqueness lock: Fast and cheap agent evals. jev as judge.; 18,041 skills from the 200 most-starred repos; Not a security scanner; 最简 Jev 调用演示器; confidence 不是正确率; q93304989-bit/jev-lab ≠ tanayvasishtha/jev-lab ≠ dairui1/jev-lab ≠ BrendanH18/jev-lab ≠ yibie/laya-jev-lab; 75% cheaper and 18% faster withdrawn; jev @0.15 100% recall 87% savings; 33Audits/jev-auto ≠ gargpratyush/jev-router; no Typesafe key, no PI_API_BASE, zero deps; tool-emitted Score/Noul ≠ calibrated Noul; semantic_compatibility: false; candidate_mass; Qwen3.5-2B ≠ Archer; Jev evaluates decisions; it cannot run a coding-agent session; Status: no model yet; S1LV3RJ1NX/openjev ≠ TheoLeeCJ/openjev; 28 accepted decisions; 3 targets; score 800; health 100; arcade game not a flight trainer; A successful live TypeSafe call has not been verified for v0.1.0; abhibansal60/tidy ≠ MANISH007700/tidy; No model, Jev included, predicted which channels its owner keeps; seed 1 selected on a held-out 400-item validation split; Brier 0.342 → 0.378; more accurate and more overconfident; Qwen3.5-4B ≠ Archer; static quants of kushalpatil/jevify-gemma4-26b-a4b; The labels were corrected, and one earlier result was retracted; zero of 23,869 eligible rows; Do not compare cost without checking task success; Exit 1 is not a proof; kisshan13/typesafe-ai-go ≠ Nibir1/typesafe-go ≠ official; 38 tests that cannot fail in a 356-model warehouse; if a parser can answer it, Jev is never asked; 359 of them; Games & Simulation 82; Education & Learning 1; Ratings are heuristics; syedabbasshaheer-art/jev-atlas ≠ ZeroX-01/jev-atlas ≠ Zaious/jev-capability-atlas ≠ gorock007/jev-atlas; anandi1989/awesome-jev-usecases ≠ whyashthakker/awesome-jev-use-cases ≠ walidboulanouar/awesome-jev-use-cases ≠ vamsikrishna2421/jev-usecases; Every headline result above is self-reported; Archer Hume 84.6% MMLU-Pro is a third-party probe not landed Archer; catalog ≠ endorsement; judge ≠ actuator; softmax over A–H ≠ Noul; SemIf 2270★; jevlike 1054★; TypeLLM/TypeLLM 16★; AnotiaWang 98★; yibie/awesome-jev 538★; Laya likes 889; tracker likes 68 lastModified UNCHANGED; Blackwood likes 2 gated manual; Archer still promised_not_landed; Hub archerhume/4rcherhume HTTP 401; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33/#35/#36/#37/#38/#40; do not push onto open #39; notes.md §119 +Hourly 1049 uniqueness lock: ggmlc GGUF is not llama.cpp; Loading them in llama.cpp will fail; one encoder pass; hf:mys/laya-GGUF sha 713ae6f6e39f likes 0 apache-2.0; hf:mys/laya-multilingual-GGUF sha 3b645ae54281; hf:mys/laya-typed-decisions-GGUF sha 1e9e8ba1f527; hf:tozp/laya-onnx sha 0862aeba1e65 Opset 14 FP32 and INT8; tozp/laya-onnx ≠ Mattepiu/laya-onnx ≠ gqgs/laya-onnx; docker-laya MIT HEAD 1b8239a51ddd README SHA 9cb7bdc3; laya.cpp RTX ggml CUDA HEAD 8590937c79a2 README SHA cdd429b9; serving substrate ≠ calibrated replica; Softmax over options ≠ calibrated Noul; Qwen3.5-4B ≠ Archer; Qwen3.5-9B ≠ Archer; jev-position-test n=6 HEAD 7a56ca1c2698 README SHA 23f194c9; jevmlx slots 5 of 6; hosted Jev 0 of 6; prior_correction made it worse; jevSweeper mean Spearman ρ −0.274; picked exact-optimal 1/25 (4%); 31 of 36 still logically decidable; 86% of the time we should not have been asking; game success ≠ calibrated Noul; LLM2Jev 64★ Apache-2.0 HEAD 924618721277 README SHA da35fe61; not affiliated with or endorsed by Jev or TypeSafe; No answer tokens are generated; OpenSourceJev llama.cpp Qwen3-1.7B HEAD 3c41fba3681d; JEV-MLX Qwen3.5-9B HEAD dec24cd929ea; decision-head-rlcd Qwen3.5-4B 4.9M LoRA; AUTO_ACT is not a Noul; closed-set fail-open stdlib-only; verified=False; soft scores ≠ hard gates; 22 to 40% cheaper *theirs*; first version 70% more expensive; 111-case benchmark *theirs*; CDC 5,000 Jev 72.2% AUC 0.7725 *theirs*; accuracy is a trap; 9.0% base rate always-no 91.0%; catalog ≠ endorsement; jev-skill 109★ 90 scenarios HEAD 4f6e899a24d4; awesome-jev-live 673 entries 4★; minecraft-agent 214★ 131 JEV decisions 35 Astra calls; nether-final-08 8 minutes 43.300 seconds; planner writes JEV selects; RoboJEV structured simulator state not images; ashare-trader 策略未通过自己的回测门槛; 36 组参数全部净期望为负; no positive expectation under real costs; typed_evals NOT an official TypeSafe AI product; jev-as-judge is a sensor; third-person-audit 40% & 60% watermarks still soft; The included experience uses a handwritten demo provider; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33/#34/#35/#36/#37/#38/#39/#40/#41/#42; notes.md §120 diff --git a/.agents/skills/augustus/references/formal-methods.md b/.agents/skills/augustus/references/formal-methods.md index cf09ce9d..30870cec 100644 --- a/.agents/skills/augustus/references/formal-methods.md +++ b/.agents/skills/augustus/references/formal-methods.md @@ -1370,3 +1370,7 @@ User-provided 0922 (`notes.md` §117 / items 330–336 / batch #100). SemIf was User-provided 0922 uniqueness lock: SemIf was formerly OpenJev; independent; not affiliated with Jev or TypeSafe; homepage openjev.com; default master; MIT; HEAD ca3ba65f1429; Tolerate float roundoff in MLX evidence verification; pushed 2026-09-19; live REST 2282★ / 140 forks; size 9177; README SHA 74ab7f7f; LICENSE SHA ca562883; interface pattern reproduction with open models; does not reproduce Jev undisclosed model/training; Direct option logits; 0 output tokens; shared-state parallel; MLX backend for Apple Silicon (`--backend mlx`); Qwen3.5-4B 3090 direct 1.023s vs AR JSON 5.332s (**5.21×**); argmax agree 18/21; systems comparison ≠ semantic equivalence; Parallel suffixes 20.03 dec/s on 777 decisions; Browser ladder Qwen3.5-4B authored BA 0.813, pert 0.766, TypeSafe subset agreement 0.845 vs Published Jev 0.883 (102 across 20 cases); Softmax over options ≠ calibrated Noul; typed output does not guarantee semantic correctness; wire/agreement ≠ replica of TypeSafe; SemIf ≠ kw2828/OpenJev playground ≠ zhihz/openjev ≠ apiplant/semif-rs port ≠ dddanielliu/semif-serve; rename is densify not a second census; JevBench 74.6 is §78 not this ladder; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33/#34/#35/#36/#38; do not push onto open #39/#40; notes.md §117 User-provided 0940 uniqueness lock: Turn decision-shaped LLM prompts into proposed Jev primitives; This is a conversion assistant, not an automatic guarantee of equivalent behavior; The compiler uses deterministic heuristics, not an LLM or evaluation model; It understands a deliberately small set of common prompt patterns; Generated instructions and criteria must be reviewed before production use; Score ranges such as 0 to 1 are translated into ordered Jev criteria; Prompts requiring open-ended prose are not a fit; suitability strong/partial/not_a_fit; compatibility full/partial/none; Writing new text stays with an LLM; Review the generated Score rubric; Jev scores ordered criteria, not an arbitrary 0-to-1 range; Everything runs locally in the browser; There is no framework, database, account, API, or server-side prompt processing; The key is read from the process environment and is never stored or printed; connect-src 'none'; alexwestco/llm-to-jev ≠ altryne/jevify ≠ ryana/jevify ≠ fidecastro/jevify ≠ Mintzs/jevify ≠ gulagala001/jevify ≠ uspraveen/Jevify; HEAD 234058ab372d; README SHA 43cd94fb; LICENSE SHA 5f334006; compiler SHA fdf235d0; 2★; MIT; JavaScript; size 29; Pages https://alexwestco.github.io/llm-to-jev/; invented_signal false; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33/#34/#35; notes.md §118 Hourly 0947 uniqueness lock: Fast and cheap agent evals. jev as judge.; 18,041 skills from the 200 most-starred repos; Not a security scanner; 最简 Jev 调用演示器; confidence 不是正确率; q93304989-bit/jev-lab ≠ tanayvasishtha/jev-lab ≠ dairui1/jev-lab ≠ BrendanH18/jev-lab ≠ yibie/laya-jev-lab; 75% cheaper and 18% faster withdrawn; jev @0.15 100% recall 87% savings; 33Audits/jev-auto ≠ gargpratyush/jev-router; no Typesafe key, no PI_API_BASE, zero deps; tool-emitted Score/Noul ≠ calibrated Noul; semantic_compatibility: false; candidate_mass; Qwen3.5-2B ≠ Archer; Jev evaluates decisions; it cannot run a coding-agent session; Status: no model yet; S1LV3RJ1NX/openjev ≠ TheoLeeCJ/openjev; 28 accepted decisions; 3 targets; score 800; health 100; arcade game not a flight trainer; A successful live TypeSafe call has not been verified for v0.1.0; abhibansal60/tidy ≠ MANISH007700/tidy; No model, Jev included, predicted which channels its owner keeps; seed 1 selected on a held-out 400-item validation split; Brier 0.342 → 0.378; more accurate and more overconfident; Qwen3.5-4B ≠ Archer; static quants of kushalpatil/jevify-gemma4-26b-a4b; The labels were corrected, and one earlier result was retracted; zero of 23,869 eligible rows; Do not compare cost without checking task success; Exit 1 is not a proof; kisshan13/typesafe-ai-go ≠ Nibir1/typesafe-go ≠ official; 38 tests that cannot fail in a 356-model warehouse; if a parser can answer it, Jev is never asked; 359 of them; Games & Simulation 82; Education & Learning 1; Ratings are heuristics; syedabbasshaheer-art/jev-atlas ≠ ZeroX-01/jev-atlas ≠ Zaious/jev-capability-atlas ≠ gorock007/jev-atlas; anandi1989/awesome-jev-usecases ≠ whyashthakker/awesome-jev-use-cases ≠ walidboulanouar/awesome-jev-use-cases ≠ vamsikrishna2421/jev-usecases; Every headline result above is self-reported; Archer Hume 84.6% MMLU-Pro is a third-party probe not landed Archer; catalog ≠ endorsement; judge ≠ actuator; softmax over A–H ≠ Noul; SemIf 2270★; jevlike 1054★; TypeLLM/TypeLLM 16★; AnotiaWang 98★; yibie/awesome-jev 538★; Laya likes 889; tracker likes 68 lastModified UNCHANGED; Blackwood likes 2 gated manual; Archer still promised_not_landed; Hub archerhume/4rcherhume HTTP 401; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33/#35/#36/#37/#38/#40; do not push onto open #39; notes.md §119 + + +**Hourly 1049 HIGH (`notes.md` §120).** ggmlc GGUF is not llama.cpp. serving substrate ≠ calibrated replica. Qwen3.5-9B ≠ Archer. planner writes JEV selects. pick_by_id vs pick_second. soft scores ≠ hard gates. catalog ≠ endorsement. Do not reopen or amend PR #23–#42. Does not bump 0.5.0. Skip Archer. `invented_signal: false`. +Hourly 1049 uniqueness lock: ggmlc GGUF is not llama.cpp; Loading them in llama.cpp will fail; one encoder pass; hf:mys/laya-GGUF sha 713ae6f6e39f likes 0 apache-2.0; hf:mys/laya-multilingual-GGUF sha 3b645ae54281; hf:mys/laya-typed-decisions-GGUF sha 1e9e8ba1f527; hf:tozp/laya-onnx sha 0862aeba1e65 Opset 14 FP32 and INT8; tozp/laya-onnx ≠ Mattepiu/laya-onnx ≠ gqgs/laya-onnx; docker-laya MIT HEAD 1b8239a51ddd README SHA 9cb7bdc3; laya.cpp RTX ggml CUDA HEAD 8590937c79a2 README SHA cdd429b9; serving substrate ≠ calibrated replica; Softmax over options ≠ calibrated Noul; Qwen3.5-4B ≠ Archer; Qwen3.5-9B ≠ Archer; jev-position-test n=6 HEAD 7a56ca1c2698 README SHA 23f194c9; jevmlx slots 5 of 6; hosted Jev 0 of 6; prior_correction made it worse; jevSweeper mean Spearman ρ −0.274; picked exact-optimal 1/25 (4%); 31 of 36 still logically decidable; 86% of the time we should not have been asking; game success ≠ calibrated Noul; LLM2Jev 64★ Apache-2.0 HEAD 924618721277 README SHA da35fe61; not affiliated with or endorsed by Jev or TypeSafe; No answer tokens are generated; OpenSourceJev llama.cpp Qwen3-1.7B HEAD 3c41fba3681d; JEV-MLX Qwen3.5-9B HEAD dec24cd929ea; decision-head-rlcd Qwen3.5-4B 4.9M LoRA; AUTO_ACT is not a Noul; closed-set fail-open stdlib-only; verified=False; soft scores ≠ hard gates; 22 to 40% cheaper *theirs*; first version 70% more expensive; 111-case benchmark *theirs*; CDC 5,000 Jev 72.2% AUC 0.7725 *theirs*; accuracy is a trap; 9.0% base rate always-no 91.0%; catalog ≠ endorsement; jev-skill 109★ 90 scenarios HEAD 4f6e899a24d4; awesome-jev-live 673 entries 4★; minecraft-agent 214★ 131 JEV decisions 35 Astra calls; nether-final-08 8 minutes 43.300 seconds; planner writes JEV selects; RoboJEV structured simulator state not images; ashare-trader 策略未通过自己的回测门槛; 36 组参数全部净期望为负; no positive expectation under real costs; typed_evals NOT an official TypeSafe AI product; jev-as-judge is a sensor; third-person-audit 40% & 60% watermarks still soft; The included experience uses a handwritten demo provider; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33/#34/#35/#36/#37/#38/#39/#40/#41/#42; notes.md §120 diff --git a/.agents/skills/augustus/references/formal-semi-formal.md b/.agents/skills/augustus/references/formal-semi-formal.md index 79f6f791..108461d7 100644 --- a/.agents/skills/augustus/references/formal-semi-formal.md +++ b/.agents/skills/augustus/references/formal-semi-formal.md @@ -83,3 +83,7 @@ User-provided 0922 (`notes.md` §117 / items 330–336 / batch #100). SemIf was User-provided 0922 uniqueness lock: SemIf was formerly OpenJev; independent; not affiliated with Jev or TypeSafe; homepage openjev.com; default master; MIT; HEAD ca3ba65f1429; Tolerate float roundoff in MLX evidence verification; pushed 2026-09-19; live REST 2282★ / 140 forks; size 9177; README SHA 74ab7f7f; LICENSE SHA ca562883; interface pattern reproduction with open models; does not reproduce Jev undisclosed model/training; Direct option logits; 0 output tokens; shared-state parallel; MLX backend for Apple Silicon (`--backend mlx`); Qwen3.5-4B 3090 direct 1.023s vs AR JSON 5.332s (**5.21×**); argmax agree 18/21; systems comparison ≠ semantic equivalence; Parallel suffixes 20.03 dec/s on 777 decisions; Browser ladder Qwen3.5-4B authored BA 0.813, pert 0.766, TypeSafe subset agreement 0.845 vs Published Jev 0.883 (102 across 20 cases); Softmax over options ≠ calibrated Noul; typed output does not guarantee semantic correctness; wire/agreement ≠ replica of TypeSafe; SemIf ≠ kw2828/OpenJev playground ≠ zhihz/openjev ≠ apiplant/semif-rs port ≠ dddanielliu/semif-serve; rename is densify not a second census; JevBench 74.6 is §78 not this ladder; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33/#34/#35/#36/#38; do not push onto open #39/#40; notes.md §117 User-provided 0940 uniqueness lock: Turn decision-shaped LLM prompts into proposed Jev primitives; This is a conversion assistant, not an automatic guarantee of equivalent behavior; The compiler uses deterministic heuristics, not an LLM or evaluation model; It understands a deliberately small set of common prompt patterns; Generated instructions and criteria must be reviewed before production use; Score ranges such as 0 to 1 are translated into ordered Jev criteria; Prompts requiring open-ended prose are not a fit; suitability strong/partial/not_a_fit; compatibility full/partial/none; Writing new text stays with an LLM; Review the generated Score rubric; Jev scores ordered criteria, not an arbitrary 0-to-1 range; Everything runs locally in the browser; There is no framework, database, account, API, or server-side prompt processing; The key is read from the process environment and is never stored or printed; connect-src 'none'; alexwestco/llm-to-jev ≠ altryne/jevify ≠ ryana/jevify ≠ fidecastro/jevify ≠ Mintzs/jevify ≠ gulagala001/jevify ≠ uspraveen/Jevify; HEAD 234058ab372d; README SHA 43cd94fb; LICENSE SHA 5f334006; compiler SHA fdf235d0; 2★; MIT; JavaScript; size 29; Pages https://alexwestco.github.io/llm-to-jev/; invented_signal false; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33/#34/#35; notes.md §118 Hourly 0947 uniqueness lock: Fast and cheap agent evals. jev as judge.; 18,041 skills from the 200 most-starred repos; Not a security scanner; 最简 Jev 调用演示器; confidence 不是正确率; q93304989-bit/jev-lab ≠ tanayvasishtha/jev-lab ≠ dairui1/jev-lab ≠ BrendanH18/jev-lab ≠ yibie/laya-jev-lab; 75% cheaper and 18% faster withdrawn; jev @0.15 100% recall 87% savings; 33Audits/jev-auto ≠ gargpratyush/jev-router; no Typesafe key, no PI_API_BASE, zero deps; tool-emitted Score/Noul ≠ calibrated Noul; semantic_compatibility: false; candidate_mass; Qwen3.5-2B ≠ Archer; Jev evaluates decisions; it cannot run a coding-agent session; Status: no model yet; S1LV3RJ1NX/openjev ≠ TheoLeeCJ/openjev; 28 accepted decisions; 3 targets; score 800; health 100; arcade game not a flight trainer; A successful live TypeSafe call has not been verified for v0.1.0; abhibansal60/tidy ≠ MANISH007700/tidy; No model, Jev included, predicted which channels its owner keeps; seed 1 selected on a held-out 400-item validation split; Brier 0.342 → 0.378; more accurate and more overconfident; Qwen3.5-4B ≠ Archer; static quants of kushalpatil/jevify-gemma4-26b-a4b; The labels were corrected, and one earlier result was retracted; zero of 23,869 eligible rows; Do not compare cost without checking task success; Exit 1 is not a proof; kisshan13/typesafe-ai-go ≠ Nibir1/typesafe-go ≠ official; 38 tests that cannot fail in a 356-model warehouse; if a parser can answer it, Jev is never asked; 359 of them; Games & Simulation 82; Education & Learning 1; Ratings are heuristics; syedabbasshaheer-art/jev-atlas ≠ ZeroX-01/jev-atlas ≠ Zaious/jev-capability-atlas ≠ gorock007/jev-atlas; anandi1989/awesome-jev-usecases ≠ whyashthakker/awesome-jev-use-cases ≠ walidboulanouar/awesome-jev-use-cases ≠ vamsikrishna2421/jev-usecases; Every headline result above is self-reported; Archer Hume 84.6% MMLU-Pro is a third-party probe not landed Archer; catalog ≠ endorsement; judge ≠ actuator; softmax over A–H ≠ Noul; SemIf 2270★; jevlike 1054★; TypeLLM/TypeLLM 16★; AnotiaWang 98★; yibie/awesome-jev 538★; Laya likes 889; tracker likes 68 lastModified UNCHANGED; Blackwood likes 2 gated manual; Archer still promised_not_landed; Hub archerhume/4rcherhume HTTP 401; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33/#35/#36/#37/#38/#40; do not push onto open #39; notes.md §119 + + +**Hourly 1049 HIGH (`notes.md` §120).** ggmlc GGUF is not llama.cpp. serving substrate ≠ calibrated replica. Qwen3.5-9B ≠ Archer. planner writes JEV selects. pick_by_id vs pick_second. soft scores ≠ hard gates. catalog ≠ endorsement. Do not reopen or amend PR #23–#42. Does not bump 0.5.0. Skip Archer. `invented_signal: false`. +Hourly 1049 uniqueness lock: ggmlc GGUF is not llama.cpp; Loading them in llama.cpp will fail; one encoder pass; hf:mys/laya-GGUF sha 713ae6f6e39f likes 0 apache-2.0; hf:mys/laya-multilingual-GGUF sha 3b645ae54281; hf:mys/laya-typed-decisions-GGUF sha 1e9e8ba1f527; hf:tozp/laya-onnx sha 0862aeba1e65 Opset 14 FP32 and INT8; tozp/laya-onnx ≠ Mattepiu/laya-onnx ≠ gqgs/laya-onnx; docker-laya MIT HEAD 1b8239a51ddd README SHA 9cb7bdc3; laya.cpp RTX ggml CUDA HEAD 8590937c79a2 README SHA cdd429b9; serving substrate ≠ calibrated replica; Softmax over options ≠ calibrated Noul; Qwen3.5-4B ≠ Archer; Qwen3.5-9B ≠ Archer; jev-position-test n=6 HEAD 7a56ca1c2698 README SHA 23f194c9; jevmlx slots 5 of 6; hosted Jev 0 of 6; prior_correction made it worse; jevSweeper mean Spearman ρ −0.274; picked exact-optimal 1/25 (4%); 31 of 36 still logically decidable; 86% of the time we should not have been asking; game success ≠ calibrated Noul; LLM2Jev 64★ Apache-2.0 HEAD 924618721277 README SHA da35fe61; not affiliated with or endorsed by Jev or TypeSafe; No answer tokens are generated; OpenSourceJev llama.cpp Qwen3-1.7B HEAD 3c41fba3681d; JEV-MLX Qwen3.5-9B HEAD dec24cd929ea; decision-head-rlcd Qwen3.5-4B 4.9M LoRA; AUTO_ACT is not a Noul; closed-set fail-open stdlib-only; verified=False; soft scores ≠ hard gates; 22 to 40% cheaper *theirs*; first version 70% more expensive; 111-case benchmark *theirs*; CDC 5,000 Jev 72.2% AUC 0.7725 *theirs*; accuracy is a trap; 9.0% base rate always-no 91.0%; catalog ≠ endorsement; jev-skill 109★ 90 scenarios HEAD 4f6e899a24d4; awesome-jev-live 673 entries 4★; minecraft-agent 214★ 131 JEV decisions 35 Astra calls; nether-final-08 8 minutes 43.300 seconds; planner writes JEV selects; RoboJEV structured simulator state not images; ashare-trader 策略未通过自己的回测门槛; 36 组参数全部净期望为负; no positive expectation under real costs; typed_evals NOT an official TypeSafe AI product; jev-as-judge is a sensor; third-person-audit 40% & 60% watermarks still soft; The included experience uses a handwritten demo provider; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33/#34/#35/#36/#37/#38/#39/#40/#41/#42; notes.md §120 diff --git a/.agents/skills/augustus/references/judgment-class.md b/.agents/skills/augustus/references/judgment-class.md index e5466e06..8873b056 100644 --- a/.agents/skills/augustus/references/judgment-class.md +++ b/.agents/skills/augustus/references/judgment-class.md @@ -1466,3 +1466,7 @@ User-provided 0922 (`notes.md` §117 / items 330–336 / batch #100). SemIf was User-provided 0922 uniqueness lock: SemIf was formerly OpenJev; independent; not affiliated with Jev or TypeSafe; homepage openjev.com; default master; MIT; HEAD ca3ba65f1429; Tolerate float roundoff in MLX evidence verification; pushed 2026-09-19; live REST 2282★ / 140 forks; size 9177; README SHA 74ab7f7f; LICENSE SHA ca562883; interface pattern reproduction with open models; does not reproduce Jev undisclosed model/training; Direct option logits; 0 output tokens; shared-state parallel; MLX backend for Apple Silicon (`--backend mlx`); Qwen3.5-4B 3090 direct 1.023s vs AR JSON 5.332s (**5.21×**); argmax agree 18/21; systems comparison ≠ semantic equivalence; Parallel suffixes 20.03 dec/s on 777 decisions; Browser ladder Qwen3.5-4B authored BA 0.813, pert 0.766, TypeSafe subset agreement 0.845 vs Published Jev 0.883 (102 across 20 cases); Softmax over options ≠ calibrated Noul; typed output does not guarantee semantic correctness; wire/agreement ≠ replica of TypeSafe; SemIf ≠ kw2828/OpenJev playground ≠ zhihz/openjev ≠ apiplant/semif-rs port ≠ dddanielliu/semif-serve; rename is densify not a second census; JevBench 74.6 is §78 not this ladder; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33/#34/#35/#36/#38; do not push onto open #39/#40; notes.md §117 User-provided 0940 uniqueness lock: Turn decision-shaped LLM prompts into proposed Jev primitives; This is a conversion assistant, not an automatic guarantee of equivalent behavior; The compiler uses deterministic heuristics, not an LLM or evaluation model; It understands a deliberately small set of common prompt patterns; Generated instructions and criteria must be reviewed before production use; Score ranges such as 0 to 1 are translated into ordered Jev criteria; Prompts requiring open-ended prose are not a fit; suitability strong/partial/not_a_fit; compatibility full/partial/none; Writing new text stays with an LLM; Review the generated Score rubric; Jev scores ordered criteria, not an arbitrary 0-to-1 range; Everything runs locally in the browser; There is no framework, database, account, API, or server-side prompt processing; The key is read from the process environment and is never stored or printed; connect-src 'none'; alexwestco/llm-to-jev ≠ altryne/jevify ≠ ryana/jevify ≠ fidecastro/jevify ≠ Mintzs/jevify ≠ gulagala001/jevify ≠ uspraveen/Jevify; HEAD 234058ab372d; README SHA 43cd94fb; LICENSE SHA 5f334006; compiler SHA fdf235d0; 2★; MIT; JavaScript; size 29; Pages https://alexwestco.github.io/llm-to-jev/; invented_signal false; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33/#34/#35; notes.md §118 Hourly 0947 uniqueness lock: Fast and cheap agent evals. jev as judge.; 18,041 skills from the 200 most-starred repos; Not a security scanner; 最简 Jev 调用演示器; confidence 不是正确率; q93304989-bit/jev-lab ≠ tanayvasishtha/jev-lab ≠ dairui1/jev-lab ≠ BrendanH18/jev-lab ≠ yibie/laya-jev-lab; 75% cheaper and 18% faster withdrawn; jev @0.15 100% recall 87% savings; 33Audits/jev-auto ≠ gargpratyush/jev-router; no Typesafe key, no PI_API_BASE, zero deps; tool-emitted Score/Noul ≠ calibrated Noul; semantic_compatibility: false; candidate_mass; Qwen3.5-2B ≠ Archer; Jev evaluates decisions; it cannot run a coding-agent session; Status: no model yet; S1LV3RJ1NX/openjev ≠ TheoLeeCJ/openjev; 28 accepted decisions; 3 targets; score 800; health 100; arcade game not a flight trainer; A successful live TypeSafe call has not been verified for v0.1.0; abhibansal60/tidy ≠ MANISH007700/tidy; No model, Jev included, predicted which channels its owner keeps; seed 1 selected on a held-out 400-item validation split; Brier 0.342 → 0.378; more accurate and more overconfident; Qwen3.5-4B ≠ Archer; static quants of kushalpatil/jevify-gemma4-26b-a4b; The labels were corrected, and one earlier result was retracted; zero of 23,869 eligible rows; Do not compare cost without checking task success; Exit 1 is not a proof; kisshan13/typesafe-ai-go ≠ Nibir1/typesafe-go ≠ official; 38 tests that cannot fail in a 356-model warehouse; if a parser can answer it, Jev is never asked; 359 of them; Games & Simulation 82; Education & Learning 1; Ratings are heuristics; syedabbasshaheer-art/jev-atlas ≠ ZeroX-01/jev-atlas ≠ Zaious/jev-capability-atlas ≠ gorock007/jev-atlas; anandi1989/awesome-jev-usecases ≠ whyashthakker/awesome-jev-use-cases ≠ walidboulanouar/awesome-jev-use-cases ≠ vamsikrishna2421/jev-usecases; Every headline result above is self-reported; Archer Hume 84.6% MMLU-Pro is a third-party probe not landed Archer; catalog ≠ endorsement; judge ≠ actuator; softmax over A–H ≠ Noul; SemIf 2270★; jevlike 1054★; TypeLLM/TypeLLM 16★; AnotiaWang 98★; yibie/awesome-jev 538★; Laya likes 889; tracker likes 68 lastModified UNCHANGED; Blackwood likes 2 gated manual; Archer still promised_not_landed; Hub archerhume/4rcherhume HTTP 401; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33/#35/#36/#37/#38/#40; do not push onto open #39; notes.md §119 + + +**Hourly 1049 HIGH (`notes.md` §120).** ggmlc GGUF is not llama.cpp. serving substrate ≠ calibrated replica. Qwen3.5-9B ≠ Archer. planner writes JEV selects. pick_by_id vs pick_second. soft scores ≠ hard gates. catalog ≠ endorsement. Do not reopen or amend PR #23–#42. Does not bump 0.5.0. Skip Archer. `invented_signal: false`. +Hourly 1049 uniqueness lock: ggmlc GGUF is not llama.cpp; Loading them in llama.cpp will fail; one encoder pass; hf:mys/laya-GGUF sha 713ae6f6e39f likes 0 apache-2.0; hf:mys/laya-multilingual-GGUF sha 3b645ae54281; hf:mys/laya-typed-decisions-GGUF sha 1e9e8ba1f527; hf:tozp/laya-onnx sha 0862aeba1e65 Opset 14 FP32 and INT8; tozp/laya-onnx ≠ Mattepiu/laya-onnx ≠ gqgs/laya-onnx; docker-laya MIT HEAD 1b8239a51ddd README SHA 9cb7bdc3; laya.cpp RTX ggml CUDA HEAD 8590937c79a2 README SHA cdd429b9; serving substrate ≠ calibrated replica; Softmax over options ≠ calibrated Noul; Qwen3.5-4B ≠ Archer; Qwen3.5-9B ≠ Archer; jev-position-test n=6 HEAD 7a56ca1c2698 README SHA 23f194c9; jevmlx slots 5 of 6; hosted Jev 0 of 6; prior_correction made it worse; jevSweeper mean Spearman ρ −0.274; picked exact-optimal 1/25 (4%); 31 of 36 still logically decidable; 86% of the time we should not have been asking; game success ≠ calibrated Noul; LLM2Jev 64★ Apache-2.0 HEAD 924618721277 README SHA da35fe61; not affiliated with or endorsed by Jev or TypeSafe; No answer tokens are generated; OpenSourceJev llama.cpp Qwen3-1.7B HEAD 3c41fba3681d; JEV-MLX Qwen3.5-9B HEAD dec24cd929ea; decision-head-rlcd Qwen3.5-4B 4.9M LoRA; AUTO_ACT is not a Noul; closed-set fail-open stdlib-only; verified=False; soft scores ≠ hard gates; 22 to 40% cheaper *theirs*; first version 70% more expensive; 111-case benchmark *theirs*; CDC 5,000 Jev 72.2% AUC 0.7725 *theirs*; accuracy is a trap; 9.0% base rate always-no 91.0%; catalog ≠ endorsement; jev-skill 109★ 90 scenarios HEAD 4f6e899a24d4; awesome-jev-live 673 entries 4★; minecraft-agent 214★ 131 JEV decisions 35 Astra calls; nether-final-08 8 minutes 43.300 seconds; planner writes JEV selects; RoboJEV structured simulator state not images; ashare-trader 策略未通过自己的回测门槛; 36 组参数全部净期望为负; no positive expectation under real costs; typed_evals NOT an official TypeSafe AI product; jev-as-judge is a sensor; third-person-audit 40% & 60% watermarks still soft; The included experience uses a handwritten demo provider; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33/#34/#35/#36/#37/#38/#39/#40/#41/#42; notes.md §120 diff --git a/.agents/skills/augustus/references/mappings.md b/.agents/skills/augustus/references/mappings.md index 44c3cc69..2e7feeed 100644 --- a/.agents/skills/augustus/references/mappings.md +++ b/.agents/skills/augustus/references/mappings.md @@ -2475,3 +2475,7 @@ User-provided 0922 (`notes.md` §117 / items 330–336 / batch #100). SemIf was User-provided 0922 uniqueness lock: SemIf was formerly OpenJev; independent; not affiliated with Jev or TypeSafe; homepage openjev.com; default master; MIT; HEAD ca3ba65f1429; Tolerate float roundoff in MLX evidence verification; pushed 2026-09-19; live REST 2282★ / 140 forks; size 9177; README SHA 74ab7f7f; LICENSE SHA ca562883; interface pattern reproduction with open models; does not reproduce Jev undisclosed model/training; Direct option logits; 0 output tokens; shared-state parallel; MLX backend for Apple Silicon (`--backend mlx`); Qwen3.5-4B 3090 direct 1.023s vs AR JSON 5.332s (**5.21×**); argmax agree 18/21; systems comparison ≠ semantic equivalence; Parallel suffixes 20.03 dec/s on 777 decisions; Browser ladder Qwen3.5-4B authored BA 0.813, pert 0.766, TypeSafe subset agreement 0.845 vs Published Jev 0.883 (102 across 20 cases); Softmax over options ≠ calibrated Noul; typed output does not guarantee semantic correctness; wire/agreement ≠ replica of TypeSafe; SemIf ≠ kw2828/OpenJev playground ≠ zhihz/openjev ≠ apiplant/semif-rs port ≠ dddanielliu/semif-serve; rename is densify not a second census; JevBench 74.6 is §78 not this ladder; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33/#34/#35/#36/#38; do not push onto open #39/#40; notes.md §117 User-provided 0940 uniqueness lock: Turn decision-shaped LLM prompts into proposed Jev primitives; This is a conversion assistant, not an automatic guarantee of equivalent behavior; The compiler uses deterministic heuristics, not an LLM or evaluation model; It understands a deliberately small set of common prompt patterns; Generated instructions and criteria must be reviewed before production use; Score ranges such as 0 to 1 are translated into ordered Jev criteria; Prompts requiring open-ended prose are not a fit; suitability strong/partial/not_a_fit; compatibility full/partial/none; Writing new text stays with an LLM; Review the generated Score rubric; Jev scores ordered criteria, not an arbitrary 0-to-1 range; Everything runs locally in the browser; There is no framework, database, account, API, or server-side prompt processing; The key is read from the process environment and is never stored or printed; connect-src 'none'; alexwestco/llm-to-jev ≠ altryne/jevify ≠ ryana/jevify ≠ fidecastro/jevify ≠ Mintzs/jevify ≠ gulagala001/jevify ≠ uspraveen/Jevify; HEAD 234058ab372d; README SHA 43cd94fb; LICENSE SHA 5f334006; compiler SHA fdf235d0; 2★; MIT; JavaScript; size 29; Pages https://alexwestco.github.io/llm-to-jev/; invented_signal false; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33/#34/#35; notes.md §118 Hourly 0947 uniqueness lock: Fast and cheap agent evals. jev as judge.; 18,041 skills from the 200 most-starred repos; Not a security scanner; 最简 Jev 调用演示器; confidence 不是正确率; q93304989-bit/jev-lab ≠ tanayvasishtha/jev-lab ≠ dairui1/jev-lab ≠ BrendanH18/jev-lab ≠ yibie/laya-jev-lab; 75% cheaper and 18% faster withdrawn; jev @0.15 100% recall 87% savings; 33Audits/jev-auto ≠ gargpratyush/jev-router; no Typesafe key, no PI_API_BASE, zero deps; tool-emitted Score/Noul ≠ calibrated Noul; semantic_compatibility: false; candidate_mass; Qwen3.5-2B ≠ Archer; Jev evaluates decisions; it cannot run a coding-agent session; Status: no model yet; S1LV3RJ1NX/openjev ≠ TheoLeeCJ/openjev; 28 accepted decisions; 3 targets; score 800; health 100; arcade game not a flight trainer; A successful live TypeSafe call has not been verified for v0.1.0; abhibansal60/tidy ≠ MANISH007700/tidy; No model, Jev included, predicted which channels its owner keeps; seed 1 selected on a held-out 400-item validation split; Brier 0.342 → 0.378; more accurate and more overconfident; Qwen3.5-4B ≠ Archer; static quants of kushalpatil/jevify-gemma4-26b-a4b; The labels were corrected, and one earlier result was retracted; zero of 23,869 eligible rows; Do not compare cost without checking task success; Exit 1 is not a proof; kisshan13/typesafe-ai-go ≠ Nibir1/typesafe-go ≠ official; 38 tests that cannot fail in a 356-model warehouse; if a parser can answer it, Jev is never asked; 359 of them; Games & Simulation 82; Education & Learning 1; Ratings are heuristics; syedabbasshaheer-art/jev-atlas ≠ ZeroX-01/jev-atlas ≠ Zaious/jev-capability-atlas ≠ gorock007/jev-atlas; anandi1989/awesome-jev-usecases ≠ whyashthakker/awesome-jev-use-cases ≠ walidboulanouar/awesome-jev-use-cases ≠ vamsikrishna2421/jev-usecases; Every headline result above is self-reported; Archer Hume 84.6% MMLU-Pro is a third-party probe not landed Archer; catalog ≠ endorsement; judge ≠ actuator; softmax over A–H ≠ Noul; SemIf 2270★; jevlike 1054★; TypeLLM/TypeLLM 16★; AnotiaWang 98★; yibie/awesome-jev 538★; Laya likes 889; tracker likes 68 lastModified UNCHANGED; Blackwood likes 2 gated manual; Archer still promised_not_landed; Hub archerhume/4rcherhume HTTP 401; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33/#35/#36/#37/#38/#40; do not push onto open #39; notes.md §119 + + +**Hourly 1049 HIGH (`notes.md` §120).** ggmlc GGUF is not llama.cpp. serving substrate ≠ calibrated replica. Qwen3.5-9B ≠ Archer. planner writes JEV selects. pick_by_id vs pick_second. soft scores ≠ hard gates. catalog ≠ endorsement. Do not reopen or amend PR #23–#42. Does not bump 0.5.0. Skip Archer. `invented_signal: false`. +Hourly 1049 uniqueness lock: ggmlc GGUF is not llama.cpp; Loading them in llama.cpp will fail; one encoder pass; hf:mys/laya-GGUF sha 713ae6f6e39f likes 0 apache-2.0; hf:mys/laya-multilingual-GGUF sha 3b645ae54281; hf:mys/laya-typed-decisions-GGUF sha 1e9e8ba1f527; hf:tozp/laya-onnx sha 0862aeba1e65 Opset 14 FP32 and INT8; tozp/laya-onnx ≠ Mattepiu/laya-onnx ≠ gqgs/laya-onnx; docker-laya MIT HEAD 1b8239a51ddd README SHA 9cb7bdc3; laya.cpp RTX ggml CUDA HEAD 8590937c79a2 README SHA cdd429b9; serving substrate ≠ calibrated replica; Softmax over options ≠ calibrated Noul; Qwen3.5-4B ≠ Archer; Qwen3.5-9B ≠ Archer; jev-position-test n=6 HEAD 7a56ca1c2698 README SHA 23f194c9; jevmlx slots 5 of 6; hosted Jev 0 of 6; prior_correction made it worse; jevSweeper mean Spearman ρ −0.274; picked exact-optimal 1/25 (4%); 31 of 36 still logically decidable; 86% of the time we should not have been asking; game success ≠ calibrated Noul; LLM2Jev 64★ Apache-2.0 HEAD 924618721277 README SHA da35fe61; not affiliated with or endorsed by Jev or TypeSafe; No answer tokens are generated; OpenSourceJev llama.cpp Qwen3-1.7B HEAD 3c41fba3681d; JEV-MLX Qwen3.5-9B HEAD dec24cd929ea; decision-head-rlcd Qwen3.5-4B 4.9M LoRA; AUTO_ACT is not a Noul; closed-set fail-open stdlib-only; verified=False; soft scores ≠ hard gates; 22 to 40% cheaper *theirs*; first version 70% more expensive; 111-case benchmark *theirs*; CDC 5,000 Jev 72.2% AUC 0.7725 *theirs*; accuracy is a trap; 9.0% base rate always-no 91.0%; catalog ≠ endorsement; jev-skill 109★ 90 scenarios HEAD 4f6e899a24d4; awesome-jev-live 673 entries 4★; minecraft-agent 214★ 131 JEV decisions 35 Astra calls; nether-final-08 8 minutes 43.300 seconds; planner writes JEV selects; RoboJEV structured simulator state not images; ashare-trader 策略未通过自己的回测门槛; 36 组参数全部净期望为负; no positive expectation under real costs; typed_evals NOT an official TypeSafe AI product; jev-as-judge is a sensor; third-person-audit 40% & 60% watermarks still soft; The included experience uses a handwritten demo provider; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33/#34/#35/#36/#37/#38/#39/#40/#41/#42; notes.md §120 diff --git a/.agents/skills/augustus/references/mental-models.md b/.agents/skills/augustus/references/mental-models.md index 5defe20f..a7f7d03c 100644 --- a/.agents/skills/augustus/references/mental-models.md +++ b/.agents/skills/augustus/references/mental-models.md @@ -2973,4 +2973,18 @@ Hourly 0843 uniqueness lock: A hunch is a probability with a policy attached; { User-provided 0920 jcr uniqueness lock: NiazMorshed2007/jcr MIT; site https://jcr.niazmorshed.dev; topics ai-agents,jev,mcp; **4★**; HEAD `138b3832`; README SHA `2a49dbc1`; LICENSE SHA `46231303`; size **14850**; Jev Capability Resolver; one tool to find documented deterministic commands in a nested capability tree; returns context; **does not execute**; skills = workflow+judgment; capabilities = individual operations; format independent of Jev; proposed open standard exploration; classify (Jev) → optional OpenAI decompose compound → beam search geometric mean of routing probs; keep up to 3 paths ≥60% of best (JCR_BAND_RATIO 0.6); ambiguity / no-match / depth-limit explicit; soft scores ≠ hard gates; 0.6 band is application policy; routing ≠ permission; docs ≠ authority to run; sol-vs-opus5-20 *theirs*: 20 scenarios × 4 variants = 80 runs; lookup+explain only, no execution; Claude Opus 5: agent input 108,585→15,819 (−85%), cost $0.3700→$0.1222 (−67%), wall 105.5s→77.7s; Codex GPT-5.6-Sol: 61,952→47,669 (−23%), $0.1377→$0.1151 (−16%), wall 25.3s→62.4s (Sol slower with JCR in 19/20); One Sol outlier 372.6s / 193 Jev calls; n=1 per cell; Not Harbor task-execution; Claude/Codex harnesses; compare mode; 50 scenarios bundled; 11 groups, 960 nodes, 11,360 items; 16 routing rounds per step; NiazMorshed2007/jcr ≠ skill-broker ≠ skillranker ≠ jev-sift ≠ jev-lens ≠ jevusher ≠ jev_select_capability; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33/#34; notes.md §116 User-provided 0922 uniqueness lock: SemIf was formerly OpenJev; independent; not affiliated with Jev or TypeSafe; homepage openjev.com; default master; MIT; HEAD ca3ba65f1429; Tolerate float roundoff in MLX evidence verification; pushed 2026-09-19; live REST 2282★ / 140 forks; size 9177; README SHA 74ab7f7f; LICENSE SHA ca562883; interface pattern reproduction with open models; does not reproduce Jev undisclosed model/training; Direct option logits; 0 output tokens; shared-state parallel; MLX backend for Apple Silicon (`--backend mlx`); Qwen3.5-4B 3090 direct 1.023s vs AR JSON 5.332s (**5.21×**); argmax agree 18/21; systems comparison ≠ semantic equivalence; Parallel suffixes 20.03 dec/s on 777 decisions; Browser ladder Qwen3.5-4B authored BA 0.813, pert 0.766, TypeSafe subset agreement 0.845 vs Published Jev 0.883 (102 across 20 cases); Softmax over options ≠ calibrated Noul; typed output does not guarantee semantic correctness; wire/agreement ≠ replica of TypeSafe; SemIf ≠ kw2828/OpenJev playground ≠ zhihz/openjev ≠ apiplant/semif-rs port ≠ dddanielliu/semif-serve; rename is densify not a second census; JevBench 74.6 is §78 not this ladder; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33/#34/#35/#36/#38; do not push onto open #39/#40; notes.md §117 User-provided 0940 uniqueness lock: Turn decision-shaped LLM prompts into proposed Jev primitives; This is a conversion assistant, not an automatic guarantee of equivalent behavior; The compiler uses deterministic heuristics, not an LLM or evaluation model; It understands a deliberately small set of common prompt patterns; Generated instructions and criteria must be reviewed before production use; Score ranges such as 0 to 1 are translated into ordered Jev criteria; Prompts requiring open-ended prose are not a fit; suitability strong/partial/not_a_fit; compatibility full/partial/none; Writing new text stays with an LLM; Review the generated Score rubric; Jev scores ordered criteria, not an arbitrary 0-to-1 range; Everything runs locally in the browser; There is no framework, database, account, API, or server-side prompt processing; The key is read from the process environment and is never stored or printed; connect-src 'none'; alexwestco/llm-to-jev ≠ altryne/jevify ≠ ryana/jevify ≠ fidecastro/jevify ≠ Mintzs/jevify ≠ gulagala001/jevify ≠ uspraveen/Jevify; HEAD 234058ab372d; README SHA 43cd94fb; LICENSE SHA 5f334006; compiler SHA fdf235d0; 2★; MIT; JavaScript; size 29; Pages https://alexwestco.github.io/llm-to-jev/; invented_signal false; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33/#34/#35; notes.md §118 + + +## Apply 1049 (`notes.md` §120) + +Serving substrate ≠ calibrated replica. ggmlc GGUF is not llama.cpp. +Option-order and exact-p games are measurements: pick_by_id vs pick_second. +Adapters / RLCD heads are class members, not TypeSafe drop-ins. +Qwen3.5-9B ≠ Archer. Qwen3.5-4B ≠ Archer. +Soft scores ≠ hard gates. planner writes JEV selects. +catalog ≠ endorsement. Life analogue: shuffle the menu; if dinner +changes, you ranked seats, not dishes. A new bottle does not recertify +the medicine. Negative EV stays negative. Do not copy keys. + Hourly 0947 uniqueness lock: Fast and cheap agent evals. jev as judge.; 18,041 skills from the 200 most-starred repos; Not a security scanner; 最简 Jev 调用演示器; confidence 不是正确率; q93304989-bit/jev-lab ≠ tanayvasishtha/jev-lab ≠ dairui1/jev-lab ≠ BrendanH18/jev-lab ≠ yibie/laya-jev-lab; 75% cheaper and 18% faster withdrawn; jev @0.15 100% recall 87% savings; 33Audits/jev-auto ≠ gargpratyush/jev-router; no Typesafe key, no PI_API_BASE, zero deps; tool-emitted Score/Noul ≠ calibrated Noul; semantic_compatibility: false; candidate_mass; Qwen3.5-2B ≠ Archer; Jev evaluates decisions; it cannot run a coding-agent session; Status: no model yet; S1LV3RJ1NX/openjev ≠ TheoLeeCJ/openjev; 28 accepted decisions; 3 targets; score 800; health 100; arcade game not a flight trainer; A successful live TypeSafe call has not been verified for v0.1.0; abhibansal60/tidy ≠ MANISH007700/tidy; No model, Jev included, predicted which channels its owner keeps; seed 1 selected on a held-out 400-item validation split; Brier 0.342 → 0.378; more accurate and more overconfident; Qwen3.5-4B ≠ Archer; static quants of kushalpatil/jevify-gemma4-26b-a4b; The labels were corrected, and one earlier result was retracted; zero of 23,869 eligible rows; Do not compare cost without checking task success; Exit 1 is not a proof; kisshan13/typesafe-ai-go ≠ Nibir1/typesafe-go ≠ official; 38 tests that cannot fail in a 356-model warehouse; if a parser can answer it, Jev is never asked; 359 of them; Games & Simulation 82; Education & Learning 1; Ratings are heuristics; syedabbasshaheer-art/jev-atlas ≠ ZeroX-01/jev-atlas ≠ Zaious/jev-capability-atlas ≠ gorock007/jev-atlas; anandi1989/awesome-jev-usecases ≠ whyashthakker/awesome-jev-use-cases ≠ walidboulanouar/awesome-jev-use-cases ≠ vamsikrishna2421/jev-usecases; Every headline result above is self-reported; Archer Hume 84.6% MMLU-Pro is a third-party probe not landed Archer; catalog ≠ endorsement; judge ≠ actuator; softmax over A–H ≠ Noul; SemIf 2270★; jevlike 1054★; TypeLLM/TypeLLM 16★; AnotiaWang 98★; yibie/awesome-jev 538★; Laya likes 889; tracker likes 68 lastModified UNCHANGED; Blackwood likes 2 gated manual; Archer still promised_not_landed; Hub archerhume/4rcherhume HTTP 401; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33/#35/#36/#37/#38/#40; do not push onto open #39; notes.md §119 +Hourly 1049 uniqueness lock: ggmlc GGUF is not llama.cpp; Loading them in llama.cpp will fail; one encoder pass; hf:mys/laya-GGUF sha 713ae6f6e39f likes 0 apache-2.0; hf:mys/laya-multilingual-GGUF sha 3b645ae54281; hf:mys/laya-typed-decisions-GGUF sha 1e9e8ba1f527; hf:tozp/laya-onnx sha 0862aeba1e65 Opset 14 FP32 and INT8; tozp/laya-onnx ≠ Mattepiu/laya-onnx ≠ gqgs/laya-onnx; docker-laya MIT HEAD 1b8239a51ddd README SHA 9cb7bdc3; laya.cpp RTX ggml CUDA HEAD 8590937c79a2 README SHA cdd429b9; serving substrate ≠ calibrated replica; Softmax over options ≠ calibrated Noul; Qwen3.5-4B ≠ Archer; Qwen3.5-9B ≠ Archer; jev-position-test n=6 HEAD 7a56ca1c2698 README SHA 23f194c9; jevmlx slots 5 of 6; hosted Jev 0 of 6; prior_correction made it worse; jevSweeper mean Spearman ρ −0.274; picked exact-optimal 1/25 (4%); 31 of 36 still logically decidable; 86% of the time we should not have been asking; game success ≠ calibrated Noul; LLM2Jev 64★ Apache-2.0 HEAD 924618721277 README SHA da35fe61; not affiliated with or endorsed by Jev or TypeSafe; No answer tokens are generated; OpenSourceJev llama.cpp Qwen3-1.7B HEAD 3c41fba3681d; JEV-MLX Qwen3.5-9B HEAD dec24cd929ea; decision-head-rlcd Qwen3.5-4B 4.9M LoRA; AUTO_ACT is not a Noul; closed-set fail-open stdlib-only; verified=False; soft scores ≠ hard gates; 22 to 40% cheaper *theirs*; first version 70% more expensive; 111-case benchmark *theirs*; CDC 5,000 Jev 72.2% AUC 0.7725 *theirs*; accuracy is a trap; 9.0% base rate always-no 91.0%; catalog ≠ endorsement; jev-skill 109★ 90 scenarios HEAD 4f6e899a24d4; awesome-jev-live 673 entries 4★; minecraft-agent 214★ 131 JEV decisions 35 Astra calls; nether-final-08 8 minutes 43.300 seconds; planner writes JEV selects; RoboJEV structured simulator state not images; ashare-trader 策略未通过自己的回测门槛; 36 组参数全部净期望为负; no positive expectation under real costs; typed_evals NOT an official TypeSafe AI product; jev-as-judge is a sensor; third-person-audit 40% & 60% watermarks still soft; The included experience uses a handwritten demo provider; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33/#34/#35/#36/#37/#38/#39/#40/#41/#42; notes.md §120 diff --git a/.agents/skills/augustus/references/methods-catalog.md b/.agents/skills/augustus/references/methods-catalog.md index dc54b804..60e0a039 100644 --- a/.agents/skills/augustus/references/methods-catalog.md +++ b/.agents/skills/augustus/references/methods-catalog.md @@ -277,3 +277,7 @@ User-provided 0922 (`notes.md` §117 / items 330–336 / batch #100). SemIf was User-provided 0922 uniqueness lock: SemIf was formerly OpenJev; independent; not affiliated with Jev or TypeSafe; homepage openjev.com; default master; MIT; HEAD ca3ba65f1429; Tolerate float roundoff in MLX evidence verification; pushed 2026-09-19; live REST 2282★ / 140 forks; size 9177; README SHA 74ab7f7f; LICENSE SHA ca562883; interface pattern reproduction with open models; does not reproduce Jev undisclosed model/training; Direct option logits; 0 output tokens; shared-state parallel; MLX backend for Apple Silicon (`--backend mlx`); Qwen3.5-4B 3090 direct 1.023s vs AR JSON 5.332s (**5.21×**); argmax agree 18/21; systems comparison ≠ semantic equivalence; Parallel suffixes 20.03 dec/s on 777 decisions; Browser ladder Qwen3.5-4B authored BA 0.813, pert 0.766, TypeSafe subset agreement 0.845 vs Published Jev 0.883 (102 across 20 cases); Softmax over options ≠ calibrated Noul; typed output does not guarantee semantic correctness; wire/agreement ≠ replica of TypeSafe; SemIf ≠ kw2828/OpenJev playground ≠ zhihz/openjev ≠ apiplant/semif-rs port ≠ dddanielliu/semif-serve; rename is densify not a second census; JevBench 74.6 is §78 not this ladder; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33/#34/#35/#36/#38; do not push onto open #39/#40; notes.md §117 User-provided 0940 uniqueness lock: Turn decision-shaped LLM prompts into proposed Jev primitives; This is a conversion assistant, not an automatic guarantee of equivalent behavior; The compiler uses deterministic heuristics, not an LLM or evaluation model; It understands a deliberately small set of common prompt patterns; Generated instructions and criteria must be reviewed before production use; Score ranges such as 0 to 1 are translated into ordered Jev criteria; Prompts requiring open-ended prose are not a fit; suitability strong/partial/not_a_fit; compatibility full/partial/none; Writing new text stays with an LLM; Review the generated Score rubric; Jev scores ordered criteria, not an arbitrary 0-to-1 range; Everything runs locally in the browser; There is no framework, database, account, API, or server-side prompt processing; The key is read from the process environment and is never stored or printed; connect-src 'none'; alexwestco/llm-to-jev ≠ altryne/jevify ≠ ryana/jevify ≠ fidecastro/jevify ≠ Mintzs/jevify ≠ gulagala001/jevify ≠ uspraveen/Jevify; HEAD 234058ab372d; README SHA 43cd94fb; LICENSE SHA 5f334006; compiler SHA fdf235d0; 2★; MIT; JavaScript; size 29; Pages https://alexwestco.github.io/llm-to-jev/; invented_signal false; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33/#34/#35; notes.md §118 Hourly 0947 uniqueness lock: Fast and cheap agent evals. jev as judge.; 18,041 skills from the 200 most-starred repos; Not a security scanner; 最简 Jev 调用演示器; confidence 不是正确率; q93304989-bit/jev-lab ≠ tanayvasishtha/jev-lab ≠ dairui1/jev-lab ≠ BrendanH18/jev-lab ≠ yibie/laya-jev-lab; 75% cheaper and 18% faster withdrawn; jev @0.15 100% recall 87% savings; 33Audits/jev-auto ≠ gargpratyush/jev-router; no Typesafe key, no PI_API_BASE, zero deps; tool-emitted Score/Noul ≠ calibrated Noul; semantic_compatibility: false; candidate_mass; Qwen3.5-2B ≠ Archer; Jev evaluates decisions; it cannot run a coding-agent session; Status: no model yet; S1LV3RJ1NX/openjev ≠ TheoLeeCJ/openjev; 28 accepted decisions; 3 targets; score 800; health 100; arcade game not a flight trainer; A successful live TypeSafe call has not been verified for v0.1.0; abhibansal60/tidy ≠ MANISH007700/tidy; No model, Jev included, predicted which channels its owner keeps; seed 1 selected on a held-out 400-item validation split; Brier 0.342 → 0.378; more accurate and more overconfident; Qwen3.5-4B ≠ Archer; static quants of kushalpatil/jevify-gemma4-26b-a4b; The labels were corrected, and one earlier result was retracted; zero of 23,869 eligible rows; Do not compare cost without checking task success; Exit 1 is not a proof; kisshan13/typesafe-ai-go ≠ Nibir1/typesafe-go ≠ official; 38 tests that cannot fail in a 356-model warehouse; if a parser can answer it, Jev is never asked; 359 of them; Games & Simulation 82; Education & Learning 1; Ratings are heuristics; syedabbasshaheer-art/jev-atlas ≠ ZeroX-01/jev-atlas ≠ Zaious/jev-capability-atlas ≠ gorock007/jev-atlas; anandi1989/awesome-jev-usecases ≠ whyashthakker/awesome-jev-use-cases ≠ walidboulanouar/awesome-jev-use-cases ≠ vamsikrishna2421/jev-usecases; Every headline result above is self-reported; Archer Hume 84.6% MMLU-Pro is a third-party probe not landed Archer; catalog ≠ endorsement; judge ≠ actuator; softmax over A–H ≠ Noul; SemIf 2270★; jevlike 1054★; TypeLLM/TypeLLM 16★; AnotiaWang 98★; yibie/awesome-jev 538★; Laya likes 889; tracker likes 68 lastModified UNCHANGED; Blackwood likes 2 gated manual; Archer still promised_not_landed; Hub archerhume/4rcherhume HTTP 401; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33/#35/#36/#37/#38/#40; do not push onto open #39; notes.md §119 + + +**Hourly 1049 HIGH (`notes.md` §120).** ggmlc GGUF is not llama.cpp. serving substrate ≠ calibrated replica. Qwen3.5-9B ≠ Archer. planner writes JEV selects. pick_by_id vs pick_second. soft scores ≠ hard gates. catalog ≠ endorsement. Do not reopen or amend PR #23–#42. Does not bump 0.5.0. Skip Archer. `invented_signal: false`. +Hourly 1049 uniqueness lock: ggmlc GGUF is not llama.cpp; Loading them in llama.cpp will fail; one encoder pass; hf:mys/laya-GGUF sha 713ae6f6e39f likes 0 apache-2.0; hf:mys/laya-multilingual-GGUF sha 3b645ae54281; hf:mys/laya-typed-decisions-GGUF sha 1e9e8ba1f527; hf:tozp/laya-onnx sha 0862aeba1e65 Opset 14 FP32 and INT8; tozp/laya-onnx ≠ Mattepiu/laya-onnx ≠ gqgs/laya-onnx; docker-laya MIT HEAD 1b8239a51ddd README SHA 9cb7bdc3; laya.cpp RTX ggml CUDA HEAD 8590937c79a2 README SHA cdd429b9; serving substrate ≠ calibrated replica; Softmax over options ≠ calibrated Noul; Qwen3.5-4B ≠ Archer; Qwen3.5-9B ≠ Archer; jev-position-test n=6 HEAD 7a56ca1c2698 README SHA 23f194c9; jevmlx slots 5 of 6; hosted Jev 0 of 6; prior_correction made it worse; jevSweeper mean Spearman ρ −0.274; picked exact-optimal 1/25 (4%); 31 of 36 still logically decidable; 86% of the time we should not have been asking; game success ≠ calibrated Noul; LLM2Jev 64★ Apache-2.0 HEAD 924618721277 README SHA da35fe61; not affiliated with or endorsed by Jev or TypeSafe; No answer tokens are generated; OpenSourceJev llama.cpp Qwen3-1.7B HEAD 3c41fba3681d; JEV-MLX Qwen3.5-9B HEAD dec24cd929ea; decision-head-rlcd Qwen3.5-4B 4.9M LoRA; AUTO_ACT is not a Noul; closed-set fail-open stdlib-only; verified=False; soft scores ≠ hard gates; 22 to 40% cheaper *theirs*; first version 70% more expensive; 111-case benchmark *theirs*; CDC 5,000 Jev 72.2% AUC 0.7725 *theirs*; accuracy is a trap; 9.0% base rate always-no 91.0%; catalog ≠ endorsement; jev-skill 109★ 90 scenarios HEAD 4f6e899a24d4; awesome-jev-live 673 entries 4★; minecraft-agent 214★ 131 JEV decisions 35 Astra calls; nether-final-08 8 minutes 43.300 seconds; planner writes JEV selects; RoboJEV structured simulator state not images; ashare-trader 策略未通过自己的回测门槛; 36 组参数全部净期望为负; no positive expectation under real costs; typed_evals NOT an official TypeSafe AI product; jev-as-judge is a sensor; third-person-audit 40% & 60% watermarks still soft; The included experience uses a handwritten demo provider; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33/#34/#35/#36/#37/#38/#39/#40/#41/#42; notes.md §120 diff --git a/.agents/skills/augustus/references/mixed-architecture.md b/.agents/skills/augustus/references/mixed-architecture.md index 78353227..d0fbb623 100644 --- a/.agents/skills/augustus/references/mixed-architecture.md +++ b/.agents/skills/augustus/references/mixed-architecture.md @@ -1444,3 +1444,7 @@ User-provided 0922 (`notes.md` §117 / items 330–336 / batch #100). SemIf was User-provided 0922 uniqueness lock: SemIf was formerly OpenJev; independent; not affiliated with Jev or TypeSafe; homepage openjev.com; default master; MIT; HEAD ca3ba65f1429; Tolerate float roundoff in MLX evidence verification; pushed 2026-09-19; live REST 2282★ / 140 forks; size 9177; README SHA 74ab7f7f; LICENSE SHA ca562883; interface pattern reproduction with open models; does not reproduce Jev undisclosed model/training; Direct option logits; 0 output tokens; shared-state parallel; MLX backend for Apple Silicon (`--backend mlx`); Qwen3.5-4B 3090 direct 1.023s vs AR JSON 5.332s (**5.21×**); argmax agree 18/21; systems comparison ≠ semantic equivalence; Parallel suffixes 20.03 dec/s on 777 decisions; Browser ladder Qwen3.5-4B authored BA 0.813, pert 0.766, TypeSafe subset agreement 0.845 vs Published Jev 0.883 (102 across 20 cases); Softmax over options ≠ calibrated Noul; typed output does not guarantee semantic correctness; wire/agreement ≠ replica of TypeSafe; SemIf ≠ kw2828/OpenJev playground ≠ zhihz/openjev ≠ apiplant/semif-rs port ≠ dddanielliu/semif-serve; rename is densify not a second census; JevBench 74.6 is §78 not this ladder; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33/#34/#35/#36/#38; do not push onto open #39/#40; notes.md §117 User-provided 0940 uniqueness lock: Turn decision-shaped LLM prompts into proposed Jev primitives; This is a conversion assistant, not an automatic guarantee of equivalent behavior; The compiler uses deterministic heuristics, not an LLM or evaluation model; It understands a deliberately small set of common prompt patterns; Generated instructions and criteria must be reviewed before production use; Score ranges such as 0 to 1 are translated into ordered Jev criteria; Prompts requiring open-ended prose are not a fit; suitability strong/partial/not_a_fit; compatibility full/partial/none; Writing new text stays with an LLM; Review the generated Score rubric; Jev scores ordered criteria, not an arbitrary 0-to-1 range; Everything runs locally in the browser; There is no framework, database, account, API, or server-side prompt processing; The key is read from the process environment and is never stored or printed; connect-src 'none'; alexwestco/llm-to-jev ≠ altryne/jevify ≠ ryana/jevify ≠ fidecastro/jevify ≠ Mintzs/jevify ≠ gulagala001/jevify ≠ uspraveen/Jevify; HEAD 234058ab372d; README SHA 43cd94fb; LICENSE SHA 5f334006; compiler SHA fdf235d0; 2★; MIT; JavaScript; size 29; Pages https://alexwestco.github.io/llm-to-jev/; invented_signal false; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33/#34/#35; notes.md §118 Hourly 0947 uniqueness lock: Fast and cheap agent evals. jev as judge.; 18,041 skills from the 200 most-starred repos; Not a security scanner; 最简 Jev 调用演示器; confidence 不是正确率; q93304989-bit/jev-lab ≠ tanayvasishtha/jev-lab ≠ dairui1/jev-lab ≠ BrendanH18/jev-lab ≠ yibie/laya-jev-lab; 75% cheaper and 18% faster withdrawn; jev @0.15 100% recall 87% savings; 33Audits/jev-auto ≠ gargpratyush/jev-router; no Typesafe key, no PI_API_BASE, zero deps; tool-emitted Score/Noul ≠ calibrated Noul; semantic_compatibility: false; candidate_mass; Qwen3.5-2B ≠ Archer; Jev evaluates decisions; it cannot run a coding-agent session; Status: no model yet; S1LV3RJ1NX/openjev ≠ TheoLeeCJ/openjev; 28 accepted decisions; 3 targets; score 800; health 100; arcade game not a flight trainer; A successful live TypeSafe call has not been verified for v0.1.0; abhibansal60/tidy ≠ MANISH007700/tidy; No model, Jev included, predicted which channels its owner keeps; seed 1 selected on a held-out 400-item validation split; Brier 0.342 → 0.378; more accurate and more overconfident; Qwen3.5-4B ≠ Archer; static quants of kushalpatil/jevify-gemma4-26b-a4b; The labels were corrected, and one earlier result was retracted; zero of 23,869 eligible rows; Do not compare cost without checking task success; Exit 1 is not a proof; kisshan13/typesafe-ai-go ≠ Nibir1/typesafe-go ≠ official; 38 tests that cannot fail in a 356-model warehouse; if a parser can answer it, Jev is never asked; 359 of them; Games & Simulation 82; Education & Learning 1; Ratings are heuristics; syedabbasshaheer-art/jev-atlas ≠ ZeroX-01/jev-atlas ≠ Zaious/jev-capability-atlas ≠ gorock007/jev-atlas; anandi1989/awesome-jev-usecases ≠ whyashthakker/awesome-jev-use-cases ≠ walidboulanouar/awesome-jev-use-cases ≠ vamsikrishna2421/jev-usecases; Every headline result above is self-reported; Archer Hume 84.6% MMLU-Pro is a third-party probe not landed Archer; catalog ≠ endorsement; judge ≠ actuator; softmax over A–H ≠ Noul; SemIf 2270★; jevlike 1054★; TypeLLM/TypeLLM 16★; AnotiaWang 98★; yibie/awesome-jev 538★; Laya likes 889; tracker likes 68 lastModified UNCHANGED; Blackwood likes 2 gated manual; Archer still promised_not_landed; Hub archerhume/4rcherhume HTTP 401; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33/#35/#36/#37/#38/#40; do not push onto open #39; notes.md §119 + + +**Hourly 1049 HIGH (`notes.md` §120).** ggmlc GGUF is not llama.cpp. serving substrate ≠ calibrated replica. Qwen3.5-9B ≠ Archer. planner writes JEV selects. pick_by_id vs pick_second. soft scores ≠ hard gates. catalog ≠ endorsement. Do not reopen or amend PR #23–#42. Does not bump 0.5.0. Skip Archer. `invented_signal: false`. +Hourly 1049 uniqueness lock: ggmlc GGUF is not llama.cpp; Loading them in llama.cpp will fail; one encoder pass; hf:mys/laya-GGUF sha 713ae6f6e39f likes 0 apache-2.0; hf:mys/laya-multilingual-GGUF sha 3b645ae54281; hf:mys/laya-typed-decisions-GGUF sha 1e9e8ba1f527; hf:tozp/laya-onnx sha 0862aeba1e65 Opset 14 FP32 and INT8; tozp/laya-onnx ≠ Mattepiu/laya-onnx ≠ gqgs/laya-onnx; docker-laya MIT HEAD 1b8239a51ddd README SHA 9cb7bdc3; laya.cpp RTX ggml CUDA HEAD 8590937c79a2 README SHA cdd429b9; serving substrate ≠ calibrated replica; Softmax over options ≠ calibrated Noul; Qwen3.5-4B ≠ Archer; Qwen3.5-9B ≠ Archer; jev-position-test n=6 HEAD 7a56ca1c2698 README SHA 23f194c9; jevmlx slots 5 of 6; hosted Jev 0 of 6; prior_correction made it worse; jevSweeper mean Spearman ρ −0.274; picked exact-optimal 1/25 (4%); 31 of 36 still logically decidable; 86% of the time we should not have been asking; game success ≠ calibrated Noul; LLM2Jev 64★ Apache-2.0 HEAD 924618721277 README SHA da35fe61; not affiliated with or endorsed by Jev or TypeSafe; No answer tokens are generated; OpenSourceJev llama.cpp Qwen3-1.7B HEAD 3c41fba3681d; JEV-MLX Qwen3.5-9B HEAD dec24cd929ea; decision-head-rlcd Qwen3.5-4B 4.9M LoRA; AUTO_ACT is not a Noul; closed-set fail-open stdlib-only; verified=False; soft scores ≠ hard gates; 22 to 40% cheaper *theirs*; first version 70% more expensive; 111-case benchmark *theirs*; CDC 5,000 Jev 72.2% AUC 0.7725 *theirs*; accuracy is a trap; 9.0% base rate always-no 91.0%; catalog ≠ endorsement; jev-skill 109★ 90 scenarios HEAD 4f6e899a24d4; awesome-jev-live 673 entries 4★; minecraft-agent 214★ 131 JEV decisions 35 Astra calls; nether-final-08 8 minutes 43.300 seconds; planner writes JEV selects; RoboJEV structured simulator state not images; ashare-trader 策略未通过自己的回测门槛; 36 组参数全部净期望为负; no positive expectation under real costs; typed_evals NOT an official TypeSafe AI product; jev-as-judge is a sensor; third-person-audit 40% & 60% watermarks still soft; The included experience uses a handwritten demo provider; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33/#34/#35/#36/#37/#38/#39/#40/#41/#42; notes.md §120 diff --git a/.agents/skills/augustus/references/question-design.md b/.agents/skills/augustus/references/question-design.md index f7639be9..8a87b195 100644 --- a/.agents/skills/augustus/references/question-design.md +++ b/.agents/skills/augustus/references/question-design.md @@ -424,3 +424,7 @@ User-provided 0922 (`notes.md` §117 / items 330–336 / batch #100). SemIf was User-provided 0922 uniqueness lock: SemIf was formerly OpenJev; independent; not affiliated with Jev or TypeSafe; homepage openjev.com; default master; MIT; HEAD ca3ba65f1429; Tolerate float roundoff in MLX evidence verification; pushed 2026-09-19; live REST 2282★ / 140 forks; size 9177; README SHA 74ab7f7f; LICENSE SHA ca562883; interface pattern reproduction with open models; does not reproduce Jev undisclosed model/training; Direct option logits; 0 output tokens; shared-state parallel; MLX backend for Apple Silicon (`--backend mlx`); Qwen3.5-4B 3090 direct 1.023s vs AR JSON 5.332s (**5.21×**); argmax agree 18/21; systems comparison ≠ semantic equivalence; Parallel suffixes 20.03 dec/s on 777 decisions; Browser ladder Qwen3.5-4B authored BA 0.813, pert 0.766, TypeSafe subset agreement 0.845 vs Published Jev 0.883 (102 across 20 cases); Softmax over options ≠ calibrated Noul; typed output does not guarantee semantic correctness; wire/agreement ≠ replica of TypeSafe; SemIf ≠ kw2828/OpenJev playground ≠ zhihz/openjev ≠ apiplant/semif-rs port ≠ dddanielliu/semif-serve; rename is densify not a second census; JevBench 74.6 is §78 not this ladder; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33/#34/#35/#36/#38; do not push onto open #39/#40; notes.md §117 User-provided 0940 uniqueness lock: Turn decision-shaped LLM prompts into proposed Jev primitives; This is a conversion assistant, not an automatic guarantee of equivalent behavior; The compiler uses deterministic heuristics, not an LLM or evaluation model; It understands a deliberately small set of common prompt patterns; Generated instructions and criteria must be reviewed before production use; Score ranges such as 0 to 1 are translated into ordered Jev criteria; Prompts requiring open-ended prose are not a fit; suitability strong/partial/not_a_fit; compatibility full/partial/none; Writing new text stays with an LLM; Review the generated Score rubric; Jev scores ordered criteria, not an arbitrary 0-to-1 range; Everything runs locally in the browser; There is no framework, database, account, API, or server-side prompt processing; The key is read from the process environment and is never stored or printed; connect-src 'none'; alexwestco/llm-to-jev ≠ altryne/jevify ≠ ryana/jevify ≠ fidecastro/jevify ≠ Mintzs/jevify ≠ gulagala001/jevify ≠ uspraveen/Jevify; HEAD 234058ab372d; README SHA 43cd94fb; LICENSE SHA 5f334006; compiler SHA fdf235d0; 2★; MIT; JavaScript; size 29; Pages https://alexwestco.github.io/llm-to-jev/; invented_signal false; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33/#34/#35; notes.md §118 Hourly 0947 uniqueness lock: Fast and cheap agent evals. jev as judge.; 18,041 skills from the 200 most-starred repos; Not a security scanner; 最简 Jev 调用演示器; confidence 不是正确率; q93304989-bit/jev-lab ≠ tanayvasishtha/jev-lab ≠ dairui1/jev-lab ≠ BrendanH18/jev-lab ≠ yibie/laya-jev-lab; 75% cheaper and 18% faster withdrawn; jev @0.15 100% recall 87% savings; 33Audits/jev-auto ≠ gargpratyush/jev-router; no Typesafe key, no PI_API_BASE, zero deps; tool-emitted Score/Noul ≠ calibrated Noul; semantic_compatibility: false; candidate_mass; Qwen3.5-2B ≠ Archer; Jev evaluates decisions; it cannot run a coding-agent session; Status: no model yet; S1LV3RJ1NX/openjev ≠ TheoLeeCJ/openjev; 28 accepted decisions; 3 targets; score 800; health 100; arcade game not a flight trainer; A successful live TypeSafe call has not been verified for v0.1.0; abhibansal60/tidy ≠ MANISH007700/tidy; No model, Jev included, predicted which channels its owner keeps; seed 1 selected on a held-out 400-item validation split; Brier 0.342 → 0.378; more accurate and more overconfident; Qwen3.5-4B ≠ Archer; static quants of kushalpatil/jevify-gemma4-26b-a4b; The labels were corrected, and one earlier result was retracted; zero of 23,869 eligible rows; Do not compare cost without checking task success; Exit 1 is not a proof; kisshan13/typesafe-ai-go ≠ Nibir1/typesafe-go ≠ official; 38 tests that cannot fail in a 356-model warehouse; if a parser can answer it, Jev is never asked; 359 of them; Games & Simulation 82; Education & Learning 1; Ratings are heuristics; syedabbasshaheer-art/jev-atlas ≠ ZeroX-01/jev-atlas ≠ Zaious/jev-capability-atlas ≠ gorock007/jev-atlas; anandi1989/awesome-jev-usecases ≠ whyashthakker/awesome-jev-use-cases ≠ walidboulanouar/awesome-jev-use-cases ≠ vamsikrishna2421/jev-usecases; Every headline result above is self-reported; Archer Hume 84.6% MMLU-Pro is a third-party probe not landed Archer; catalog ≠ endorsement; judge ≠ actuator; softmax over A–H ≠ Noul; SemIf 2270★; jevlike 1054★; TypeLLM/TypeLLM 16★; AnotiaWang 98★; yibie/awesome-jev 538★; Laya likes 889; tracker likes 68 lastModified UNCHANGED; Blackwood likes 2 gated manual; Archer still promised_not_landed; Hub archerhume/4rcherhume HTTP 401; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33/#35/#36/#37/#38/#40; do not push onto open #39; notes.md §119 + + +**Hourly 1049 HIGH (`notes.md` §120).** ggmlc GGUF is not llama.cpp. serving substrate ≠ calibrated replica. Qwen3.5-9B ≠ Archer. planner writes JEV selects. pick_by_id vs pick_second. soft scores ≠ hard gates. catalog ≠ endorsement. Do not reopen or amend PR #23–#42. Does not bump 0.5.0. Skip Archer. `invented_signal: false`. +Hourly 1049 uniqueness lock: ggmlc GGUF is not llama.cpp; Loading them in llama.cpp will fail; one encoder pass; hf:mys/laya-GGUF sha 713ae6f6e39f likes 0 apache-2.0; hf:mys/laya-multilingual-GGUF sha 3b645ae54281; hf:mys/laya-typed-decisions-GGUF sha 1e9e8ba1f527; hf:tozp/laya-onnx sha 0862aeba1e65 Opset 14 FP32 and INT8; tozp/laya-onnx ≠ Mattepiu/laya-onnx ≠ gqgs/laya-onnx; docker-laya MIT HEAD 1b8239a51ddd README SHA 9cb7bdc3; laya.cpp RTX ggml CUDA HEAD 8590937c79a2 README SHA cdd429b9; serving substrate ≠ calibrated replica; Softmax over options ≠ calibrated Noul; Qwen3.5-4B ≠ Archer; Qwen3.5-9B ≠ Archer; jev-position-test n=6 HEAD 7a56ca1c2698 README SHA 23f194c9; jevmlx slots 5 of 6; hosted Jev 0 of 6; prior_correction made it worse; jevSweeper mean Spearman ρ −0.274; picked exact-optimal 1/25 (4%); 31 of 36 still logically decidable; 86% of the time we should not have been asking; game success ≠ calibrated Noul; LLM2Jev 64★ Apache-2.0 HEAD 924618721277 README SHA da35fe61; not affiliated with or endorsed by Jev or TypeSafe; No answer tokens are generated; OpenSourceJev llama.cpp Qwen3-1.7B HEAD 3c41fba3681d; JEV-MLX Qwen3.5-9B HEAD dec24cd929ea; decision-head-rlcd Qwen3.5-4B 4.9M LoRA; AUTO_ACT is not a Noul; closed-set fail-open stdlib-only; verified=False; soft scores ≠ hard gates; 22 to 40% cheaper *theirs*; first version 70% more expensive; 111-case benchmark *theirs*; CDC 5,000 Jev 72.2% AUC 0.7725 *theirs*; accuracy is a trap; 9.0% base rate always-no 91.0%; catalog ≠ endorsement; jev-skill 109★ 90 scenarios HEAD 4f6e899a24d4; awesome-jev-live 673 entries 4★; minecraft-agent 214★ 131 JEV decisions 35 Astra calls; nether-final-08 8 minutes 43.300 seconds; planner writes JEV selects; RoboJEV structured simulator state not images; ashare-trader 策略未通过自己的回测门槛; 36 组参数全部净期望为负; no positive expectation under real costs; typed_evals NOT an official TypeSafe AI product; jev-as-judge is a sensor; third-person-audit 40% & 60% watermarks still soft; The included experience uses a handwritten demo provider; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33/#34/#35/#36/#37/#38/#39/#40/#41/#42; notes.md §120 diff --git a/.agents/skills/augustus/references/toolbox-mapping.md b/.agents/skills/augustus/references/toolbox-mapping.md index b0224f02..128c0d42 100644 --- a/.agents/skills/augustus/references/toolbox-mapping.md +++ b/.agents/skills/augustus/references/toolbox-mapping.md @@ -353,3 +353,7 @@ User-provided 0922 (`notes.md` §117 / items 330–336 / batch #100). SemIf was User-provided 0922 uniqueness lock: SemIf was formerly OpenJev; independent; not affiliated with Jev or TypeSafe; homepage openjev.com; default master; MIT; HEAD ca3ba65f1429; Tolerate float roundoff in MLX evidence verification; pushed 2026-09-19; live REST 2282★ / 140 forks; size 9177; README SHA 74ab7f7f; LICENSE SHA ca562883; interface pattern reproduction with open models; does not reproduce Jev undisclosed model/training; Direct option logits; 0 output tokens; shared-state parallel; MLX backend for Apple Silicon (`--backend mlx`); Qwen3.5-4B 3090 direct 1.023s vs AR JSON 5.332s (**5.21×**); argmax agree 18/21; systems comparison ≠ semantic equivalence; Parallel suffixes 20.03 dec/s on 777 decisions; Browser ladder Qwen3.5-4B authored BA 0.813, pert 0.766, TypeSafe subset agreement 0.845 vs Published Jev 0.883 (102 across 20 cases); Softmax over options ≠ calibrated Noul; typed output does not guarantee semantic correctness; wire/agreement ≠ replica of TypeSafe; SemIf ≠ kw2828/OpenJev playground ≠ zhihz/openjev ≠ apiplant/semif-rs port ≠ dddanielliu/semif-serve; rename is densify not a second census; JevBench 74.6 is §78 not this ladder; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33/#34/#35/#36/#38; do not push onto open #39/#40; notes.md §117 User-provided 0940 uniqueness lock: Turn decision-shaped LLM prompts into proposed Jev primitives; This is a conversion assistant, not an automatic guarantee of equivalent behavior; The compiler uses deterministic heuristics, not an LLM or evaluation model; It understands a deliberately small set of common prompt patterns; Generated instructions and criteria must be reviewed before production use; Score ranges such as 0 to 1 are translated into ordered Jev criteria; Prompts requiring open-ended prose are not a fit; suitability strong/partial/not_a_fit; compatibility full/partial/none; Writing new text stays with an LLM; Review the generated Score rubric; Jev scores ordered criteria, not an arbitrary 0-to-1 range; Everything runs locally in the browser; There is no framework, database, account, API, or server-side prompt processing; The key is read from the process environment and is never stored or printed; connect-src 'none'; alexwestco/llm-to-jev ≠ altryne/jevify ≠ ryana/jevify ≠ fidecastro/jevify ≠ Mintzs/jevify ≠ gulagala001/jevify ≠ uspraveen/Jevify; HEAD 234058ab372d; README SHA 43cd94fb; LICENSE SHA 5f334006; compiler SHA fdf235d0; 2★; MIT; JavaScript; size 29; Pages https://alexwestco.github.io/llm-to-jev/; invented_signal false; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33/#34/#35; notes.md §118 Hourly 0947 uniqueness lock: Fast and cheap agent evals. jev as judge.; 18,041 skills from the 200 most-starred repos; Not a security scanner; 最简 Jev 调用演示器; confidence 不是正确率; q93304989-bit/jev-lab ≠ tanayvasishtha/jev-lab ≠ dairui1/jev-lab ≠ BrendanH18/jev-lab ≠ yibie/laya-jev-lab; 75% cheaper and 18% faster withdrawn; jev @0.15 100% recall 87% savings; 33Audits/jev-auto ≠ gargpratyush/jev-router; no Typesafe key, no PI_API_BASE, zero deps; tool-emitted Score/Noul ≠ calibrated Noul; semantic_compatibility: false; candidate_mass; Qwen3.5-2B ≠ Archer; Jev evaluates decisions; it cannot run a coding-agent session; Status: no model yet; S1LV3RJ1NX/openjev ≠ TheoLeeCJ/openjev; 28 accepted decisions; 3 targets; score 800; health 100; arcade game not a flight trainer; A successful live TypeSafe call has not been verified for v0.1.0; abhibansal60/tidy ≠ MANISH007700/tidy; No model, Jev included, predicted which channels its owner keeps; seed 1 selected on a held-out 400-item validation split; Brier 0.342 → 0.378; more accurate and more overconfident; Qwen3.5-4B ≠ Archer; static quants of kushalpatil/jevify-gemma4-26b-a4b; The labels were corrected, and one earlier result was retracted; zero of 23,869 eligible rows; Do not compare cost without checking task success; Exit 1 is not a proof; kisshan13/typesafe-ai-go ≠ Nibir1/typesafe-go ≠ official; 38 tests that cannot fail in a 356-model warehouse; if a parser can answer it, Jev is never asked; 359 of them; Games & Simulation 82; Education & Learning 1; Ratings are heuristics; syedabbasshaheer-art/jev-atlas ≠ ZeroX-01/jev-atlas ≠ Zaious/jev-capability-atlas ≠ gorock007/jev-atlas; anandi1989/awesome-jev-usecases ≠ whyashthakker/awesome-jev-use-cases ≠ walidboulanouar/awesome-jev-use-cases ≠ vamsikrishna2421/jev-usecases; Every headline result above is self-reported; Archer Hume 84.6% MMLU-Pro is a third-party probe not landed Archer; catalog ≠ endorsement; judge ≠ actuator; softmax over A–H ≠ Noul; SemIf 2270★; jevlike 1054★; TypeLLM/TypeLLM 16★; AnotiaWang 98★; yibie/awesome-jev 538★; Laya likes 889; tracker likes 68 lastModified UNCHANGED; Blackwood likes 2 gated manual; Archer still promised_not_landed; Hub archerhume/4rcherhume HTTP 401; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33/#35/#36/#37/#38/#40; do not push onto open #39; notes.md §119 + + +**Hourly 1049 HIGH (`notes.md` §120).** ggmlc GGUF is not llama.cpp. serving substrate ≠ calibrated replica. Qwen3.5-9B ≠ Archer. planner writes JEV selects. pick_by_id vs pick_second. soft scores ≠ hard gates. catalog ≠ endorsement. Do not reopen or amend PR #23–#42. Does not bump 0.5.0. Skip Archer. `invented_signal: false`. +Hourly 1049 uniqueness lock: ggmlc GGUF is not llama.cpp; Loading them in llama.cpp will fail; one encoder pass; hf:mys/laya-GGUF sha 713ae6f6e39f likes 0 apache-2.0; hf:mys/laya-multilingual-GGUF sha 3b645ae54281; hf:mys/laya-typed-decisions-GGUF sha 1e9e8ba1f527; hf:tozp/laya-onnx sha 0862aeba1e65 Opset 14 FP32 and INT8; tozp/laya-onnx ≠ Mattepiu/laya-onnx ≠ gqgs/laya-onnx; docker-laya MIT HEAD 1b8239a51ddd README SHA 9cb7bdc3; laya.cpp RTX ggml CUDA HEAD 8590937c79a2 README SHA cdd429b9; serving substrate ≠ calibrated replica; Softmax over options ≠ calibrated Noul; Qwen3.5-4B ≠ Archer; Qwen3.5-9B ≠ Archer; jev-position-test n=6 HEAD 7a56ca1c2698 README SHA 23f194c9; jevmlx slots 5 of 6; hosted Jev 0 of 6; prior_correction made it worse; jevSweeper mean Spearman ρ −0.274; picked exact-optimal 1/25 (4%); 31 of 36 still logically decidable; 86% of the time we should not have been asking; game success ≠ calibrated Noul; LLM2Jev 64★ Apache-2.0 HEAD 924618721277 README SHA da35fe61; not affiliated with or endorsed by Jev or TypeSafe; No answer tokens are generated; OpenSourceJev llama.cpp Qwen3-1.7B HEAD 3c41fba3681d; JEV-MLX Qwen3.5-9B HEAD dec24cd929ea; decision-head-rlcd Qwen3.5-4B 4.9M LoRA; AUTO_ACT is not a Noul; closed-set fail-open stdlib-only; verified=False; soft scores ≠ hard gates; 22 to 40% cheaper *theirs*; first version 70% more expensive; 111-case benchmark *theirs*; CDC 5,000 Jev 72.2% AUC 0.7725 *theirs*; accuracy is a trap; 9.0% base rate always-no 91.0%; catalog ≠ endorsement; jev-skill 109★ 90 scenarios HEAD 4f6e899a24d4; awesome-jev-live 673 entries 4★; minecraft-agent 214★ 131 JEV decisions 35 Astra calls; nether-final-08 8 minutes 43.300 seconds; planner writes JEV selects; RoboJEV structured simulator state not images; ashare-trader 策略未通过自己的回测门槛; 36 组参数全部净期望为负; no positive expectation under real costs; typed_evals NOT an official TypeSafe AI product; jev-as-judge is a sensor; third-person-audit 40% & 60% watermarks still soft; The included experience uses a handwritten demo provider; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33/#34/#35/#36/#37/#38/#39/#40/#41/#42; notes.md §120 diff --git a/.agents/skills/augustus/references/validation.md b/.agents/skills/augustus/references/validation.md index a9aa40e5..47d0a0d1 100644 --- a/.agents/skills/augustus/references/validation.md +++ b/.agents/skills/augustus/references/validation.md @@ -1157,3 +1157,7 @@ User-provided 0922 (`notes.md` §117 / items 330–336 / batch #100). SemIf was User-provided 0922 uniqueness lock: SemIf was formerly OpenJev; independent; not affiliated with Jev or TypeSafe; homepage openjev.com; default master; MIT; HEAD ca3ba65f1429; Tolerate float roundoff in MLX evidence verification; pushed 2026-09-19; live REST 2282★ / 140 forks; size 9177; README SHA 74ab7f7f; LICENSE SHA ca562883; interface pattern reproduction with open models; does not reproduce Jev undisclosed model/training; Direct option logits; 0 output tokens; shared-state parallel; MLX backend for Apple Silicon (`--backend mlx`); Qwen3.5-4B 3090 direct 1.023s vs AR JSON 5.332s (**5.21×**); argmax agree 18/21; systems comparison ≠ semantic equivalence; Parallel suffixes 20.03 dec/s on 777 decisions; Browser ladder Qwen3.5-4B authored BA 0.813, pert 0.766, TypeSafe subset agreement 0.845 vs Published Jev 0.883 (102 across 20 cases); Softmax over options ≠ calibrated Noul; typed output does not guarantee semantic correctness; wire/agreement ≠ replica of TypeSafe; SemIf ≠ kw2828/OpenJev playground ≠ zhihz/openjev ≠ apiplant/semif-rs port ≠ dddanielliu/semif-serve; rename is densify not a second census; JevBench 74.6 is §78 not this ladder; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33/#34/#35/#36/#38; do not push onto open #39/#40; notes.md §117 User-provided 0940 uniqueness lock: Turn decision-shaped LLM prompts into proposed Jev primitives; This is a conversion assistant, not an automatic guarantee of equivalent behavior; The compiler uses deterministic heuristics, not an LLM or evaluation model; It understands a deliberately small set of common prompt patterns; Generated instructions and criteria must be reviewed before production use; Score ranges such as 0 to 1 are translated into ordered Jev criteria; Prompts requiring open-ended prose are not a fit; suitability strong/partial/not_a_fit; compatibility full/partial/none; Writing new text stays with an LLM; Review the generated Score rubric; Jev scores ordered criteria, not an arbitrary 0-to-1 range; Everything runs locally in the browser; There is no framework, database, account, API, or server-side prompt processing; The key is read from the process environment and is never stored or printed; connect-src 'none'; alexwestco/llm-to-jev ≠ altryne/jevify ≠ ryana/jevify ≠ fidecastro/jevify ≠ Mintzs/jevify ≠ gulagala001/jevify ≠ uspraveen/Jevify; HEAD 234058ab372d; README SHA 43cd94fb; LICENSE SHA 5f334006; compiler SHA fdf235d0; 2★; MIT; JavaScript; size 29; Pages https://alexwestco.github.io/llm-to-jev/; invented_signal false; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33/#34/#35; notes.md §118 Hourly 0947 uniqueness lock: Fast and cheap agent evals. jev as judge.; 18,041 skills from the 200 most-starred repos; Not a security scanner; 最简 Jev 调用演示器; confidence 不是正确率; q93304989-bit/jev-lab ≠ tanayvasishtha/jev-lab ≠ dairui1/jev-lab ≠ BrendanH18/jev-lab ≠ yibie/laya-jev-lab; 75% cheaper and 18% faster withdrawn; jev @0.15 100% recall 87% savings; 33Audits/jev-auto ≠ gargpratyush/jev-router; no Typesafe key, no PI_API_BASE, zero deps; tool-emitted Score/Noul ≠ calibrated Noul; semantic_compatibility: false; candidate_mass; Qwen3.5-2B ≠ Archer; Jev evaluates decisions; it cannot run a coding-agent session; Status: no model yet; S1LV3RJ1NX/openjev ≠ TheoLeeCJ/openjev; 28 accepted decisions; 3 targets; score 800; health 100; arcade game not a flight trainer; A successful live TypeSafe call has not been verified for v0.1.0; abhibansal60/tidy ≠ MANISH007700/tidy; No model, Jev included, predicted which channels its owner keeps; seed 1 selected on a held-out 400-item validation split; Brier 0.342 → 0.378; more accurate and more overconfident; Qwen3.5-4B ≠ Archer; static quants of kushalpatil/jevify-gemma4-26b-a4b; The labels were corrected, and one earlier result was retracted; zero of 23,869 eligible rows; Do not compare cost without checking task success; Exit 1 is not a proof; kisshan13/typesafe-ai-go ≠ Nibir1/typesafe-go ≠ official; 38 tests that cannot fail in a 356-model warehouse; if a parser can answer it, Jev is never asked; 359 of them; Games & Simulation 82; Education & Learning 1; Ratings are heuristics; syedabbasshaheer-art/jev-atlas ≠ ZeroX-01/jev-atlas ≠ Zaious/jev-capability-atlas ≠ gorock007/jev-atlas; anandi1989/awesome-jev-usecases ≠ whyashthakker/awesome-jev-use-cases ≠ walidboulanouar/awesome-jev-use-cases ≠ vamsikrishna2421/jev-usecases; Every headline result above is self-reported; Archer Hume 84.6% MMLU-Pro is a third-party probe not landed Archer; catalog ≠ endorsement; judge ≠ actuator; softmax over A–H ≠ Noul; SemIf 2270★; jevlike 1054★; TypeLLM/TypeLLM 16★; AnotiaWang 98★; yibie/awesome-jev 538★; Laya likes 889; tracker likes 68 lastModified UNCHANGED; Blackwood likes 2 gated manual; Archer still promised_not_landed; Hub archerhume/4rcherhume HTTP 401; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33/#35/#36/#37/#38/#40; do not push onto open #39; notes.md §119 + + +**Hourly 1049 HIGH (`notes.md` §120).** ggmlc GGUF is not llama.cpp. serving substrate ≠ calibrated replica. Qwen3.5-9B ≠ Archer. planner writes JEV selects. pick_by_id vs pick_second. soft scores ≠ hard gates. catalog ≠ endorsement. Do not reopen or amend PR #23–#42. Does not bump 0.5.0. Skip Archer. `invented_signal: false`. +Hourly 1049 uniqueness lock: ggmlc GGUF is not llama.cpp; Loading them in llama.cpp will fail; one encoder pass; hf:mys/laya-GGUF sha 713ae6f6e39f likes 0 apache-2.0; hf:mys/laya-multilingual-GGUF sha 3b645ae54281; hf:mys/laya-typed-decisions-GGUF sha 1e9e8ba1f527; hf:tozp/laya-onnx sha 0862aeba1e65 Opset 14 FP32 and INT8; tozp/laya-onnx ≠ Mattepiu/laya-onnx ≠ gqgs/laya-onnx; docker-laya MIT HEAD 1b8239a51ddd README SHA 9cb7bdc3; laya.cpp RTX ggml CUDA HEAD 8590937c79a2 README SHA cdd429b9; serving substrate ≠ calibrated replica; Softmax over options ≠ calibrated Noul; Qwen3.5-4B ≠ Archer; Qwen3.5-9B ≠ Archer; jev-position-test n=6 HEAD 7a56ca1c2698 README SHA 23f194c9; jevmlx slots 5 of 6; hosted Jev 0 of 6; prior_correction made it worse; jevSweeper mean Spearman ρ −0.274; picked exact-optimal 1/25 (4%); 31 of 36 still logically decidable; 86% of the time we should not have been asking; game success ≠ calibrated Noul; LLM2Jev 64★ Apache-2.0 HEAD 924618721277 README SHA da35fe61; not affiliated with or endorsed by Jev or TypeSafe; No answer tokens are generated; OpenSourceJev llama.cpp Qwen3-1.7B HEAD 3c41fba3681d; JEV-MLX Qwen3.5-9B HEAD dec24cd929ea; decision-head-rlcd Qwen3.5-4B 4.9M LoRA; AUTO_ACT is not a Noul; closed-set fail-open stdlib-only; verified=False; soft scores ≠ hard gates; 22 to 40% cheaper *theirs*; first version 70% more expensive; 111-case benchmark *theirs*; CDC 5,000 Jev 72.2% AUC 0.7725 *theirs*; accuracy is a trap; 9.0% base rate always-no 91.0%; catalog ≠ endorsement; jev-skill 109★ 90 scenarios HEAD 4f6e899a24d4; awesome-jev-live 673 entries 4★; minecraft-agent 214★ 131 JEV decisions 35 Astra calls; nether-final-08 8 minutes 43.300 seconds; planner writes JEV selects; RoboJEV structured simulator state not images; ashare-trader 策略未通过自己的回测门槛; 36 组参数全部净期望为负; no positive expectation under real costs; typed_evals NOT an official TypeSafe AI product; jev-as-judge is a sensor; third-person-audit 40% & 60% watermarks still soft; The included experience uses a handwritten demo provider; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33/#34/#35/#36/#37/#38/#39/#40/#41/#42; notes.md §120 diff --git a/.agents/skills/augustus/scripts/evaluate_decisions.py b/.agents/skills/augustus/scripts/evaluate_decisions.py index ea7890ec..945a3585 100755 --- a/.agents/skills/augustus/scripts/evaluate_decisions.py +++ b/.agents/skills/augustus/scripts/evaluate_decisions.py @@ -15,6 +15,8 @@ * hop-ECE permutation invariance (shuffle the stream; ECE does not move) * candidate_mass vs renormalized bag: softmax over allowed tokens is a peaked ranking, not a Noul (mass outside the bag can be hidden) + * pick_by_id vs pick_second: the same scores, shuffled option order; + slot-two is ranking theater, not a Noul Select thresholds on one split, evaluate on another: run twice with different files. Missing labels or costs produce a stated limitation, not defaults. @@ -175,6 +177,26 @@ def _softmax(xs): return [e / total for e in exps] + +def pick_by_id(scores_by_id): + """Max score wins regardless of presentation order.""" + if not scores_by_id: + raise ValueError("scores_by_id empty") + return max(scores_by_id.items(), key=lambda kv: kv[1])[0] + + +def pick_second(ordered_ids, scores_by_id=None): + """Whatever sits in slot two. Option order is the answer. + + pick_by_id vs pick_second: the same scores can yield four + dinners if the menu is shuffled. That is ranking theater, + not a Noul. + """ + if len(ordered_ids) < 2: + raise ValueError("need at least two options") + return ordered_ids[1] + + def candidate_mass_renorm(full_logits, allowed_indices): """Softmax over allowed tokens is a peaked ranking, not a Noul. @@ -264,6 +286,20 @@ def self_test(): mass_in, _ = candidate_mass_renorm(inside, [0, 1, 2]) assert mass_in > 0.99, mass_in + + # pick_by_id vs pick_second: same scores, four orders. + scores = {"A": 0.2, "B": 0.7, "C": 0.1} + orders = [ + ("A", "B", "C"), + ("B", "A", "C"), + ("C", "B", "A"), + ("A", "C", "B"), + ] + by_id = [pick_by_id({k: scores[k] for k in order}) for order in orders] + second = [pick_second(order) for order in orders] + assert by_id == ["B", "B", "B", "B"], by_id + assert second == ["B", "A", "B", "C"], second + print("self-test ok") diff --git a/.agents/skills/augustus/scripts/uniqueness_gate.py b/.agents/skills/augustus/scripts/uniqueness_gate.py index 0ef961ab..abc5c808 100644 --- a/.agents/skills/augustus/scripts/uniqueness_gate.py +++ b/.agents/skills/augustus/scripts/uniqueness_gate.py @@ -1,14 +1,14 @@ #!/usr/bin/env python3 """Uniqueness gate for merged 0843 (§114), merged 0915 NanoJev (§115), merged 0920 jcr (§116), merged 0922 SemIf (§117), merged 0940 -llm-to-jev (§118), and hourly 0947 HIGH (§119). +llm-to-jev (§118), hourly 0947 HIGH (§119), and hourly 1049 HIGH (§120). Each lock must appear as one consecutive substring in every listed overlay. Fragments scattered across files do not count. -Also: YAML-parse SKILL.md frontmatter; notes.md owns §114–§119; -composition items 289–316, 322–329, 330–336, and 337–352 exist; -findings batches #97–#102 exist. Items 317–321 stay unused. +Also: YAML-parse SKILL.md frontmatter; notes.md owns §114–§120; +composition items 289–316, 322–329, 330–336, 337–352, and 353–368 exist; +findings batches #97–#103 exist. Items 317–321 stay unused. CHANGELOG.md must not hold uniqueness dump walls (dumps live in changelog-hourly.md). README.md must not hold the 0743 dump wall. Pages greps stay in docs/index.md and docs/_layouts/default.html. @@ -119,6 +119,11 @@ 'Hourly 0947 uniqueness lock: Fast and cheap agent evals. jev as judge.; 18,041 skills from the 200 most-starred repos; Not a security scanner; 最简 Jev 调用演示器; confidence 不是正确率; q93304989-bit/jev-lab ≠ tanayvasishtha/jev-lab ≠ dairui1/jev-lab ≠ BrendanH18/jev-lab ≠ yibie/laya-jev-lab; 75% cheaper and 18% faster withdrawn; jev @0.15 100% recall 87% savings; 33Audits/jev-auto ≠ gargpratyush/jev-router; no Typesafe key, no PI_API_BASE, zero deps; tool-emitted Score/Noul ≠ calibrated Noul; semantic_compatibility: false; candidate_mass; Qwen3.5-2B ≠ Archer; Jev evaluates decisions; it cannot run a coding-agent session; Status: no model yet; S1LV3RJ1NX/openjev ≠ TheoLeeCJ/openjev; 28 accepted decisions; 3 targets; score 800; health 100; arcade game not a flight trainer; A successful live TypeSafe call has not been verified for v0.1.0; abhibansal60/tidy ≠ MANISH007700/tidy; No model, Jev included, predicted which channels its owner keeps; seed 1 selected on a held-out 400-item validation split; Brier 0.342 → 0.378; more accurate and more overconfident; Qwen3.5-4B ≠ Archer; static quants of kushalpatil/jevify-gemma4-26b-a4b; The labels were corrected, and one earlier result was retracted; zero of 23,869 eligible rows; Do not compare cost without checking task success; Exit 1 is not a proof; kisshan13/typesafe-ai-go ≠ Nibir1/typesafe-go ≠ official; 38 tests that cannot fail in a 356-model warehouse; if a parser can answer it, Jev is never asked; 359 of them; Games & Simulation 82; Education & Learning 1; Ratings are heuristics; syedabbasshaheer-art/jev-atlas ≠ ZeroX-01/jev-atlas ≠ Zaious/jev-capability-atlas ≠ gorock007/jev-atlas; anandi1989/awesome-jev-usecases ≠ whyashthakker/awesome-jev-use-cases ≠ walidboulanouar/awesome-jev-use-cases ≠ vamsikrishna2421/jev-usecases; Every headline result above is self-reported; Archer Hume 84.6% MMLU-Pro is a third-party probe not landed Archer; catalog ≠ endorsement; judge ≠ actuator; softmax over A–H ≠ Noul; SemIf 2270★; jevlike 1054★; TypeLLM/TypeLLM 16★; AnotiaWang 98★; yibie/awesome-jev 538★; Laya likes 889; tracker likes 68 lastModified UNCHANGED; Blackwood likes 2 gated manual; Archer still promised_not_landed; Hub archerhume/4rcherhume HTTP 401; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33/#35/#36/#37/#38/#40; do not push onto open #39; notes.md §119' ) + +UNIQ_1049 = ( + 'Hourly 1049 uniqueness lock: ggmlc GGUF is not llama.cpp; Loading them in llama.cpp will fail; one encoder pass; hf:mys/laya-GGUF sha 713ae6f6e39f likes 0 apache-2.0; hf:mys/laya-multilingual-GGUF sha 3b645ae54281; hf:mys/laya-typed-decisions-GGUF sha 1e9e8ba1f527; hf:tozp/laya-onnx sha 0862aeba1e65 Opset 14 FP32 and INT8; tozp/laya-onnx ≠ Mattepiu/laya-onnx ≠ gqgs/laya-onnx; docker-laya MIT HEAD 1b8239a51ddd README SHA 9cb7bdc3; laya.cpp RTX ggml CUDA HEAD 8590937c79a2 README SHA cdd429b9; serving substrate ≠ calibrated replica; Softmax over options ≠ calibrated Noul; Qwen3.5-4B ≠ Archer; Qwen3.5-9B ≠ Archer; jev-position-test n=6 HEAD 7a56ca1c2698 README SHA 23f194c9; jevmlx slots 5 of 6; hosted Jev 0 of 6; prior_correction made it worse; jevSweeper mean Spearman ρ −0.274; picked exact-optimal 1/25 (4%); 31 of 36 still logically decidable; 86% of the time we should not have been asking; game success ≠ calibrated Noul; LLM2Jev 64★ Apache-2.0 HEAD 924618721277 README SHA da35fe61; not affiliated with or endorsed by Jev or TypeSafe; No answer tokens are generated; OpenSourceJev llama.cpp Qwen3-1.7B HEAD 3c41fba3681d; JEV-MLX Qwen3.5-9B HEAD dec24cd929ea; decision-head-rlcd Qwen3.5-4B 4.9M LoRA; AUTO_ACT is not a Noul; closed-set fail-open stdlib-only; verified=False; soft scores ≠ hard gates; 22 to 40% cheaper *theirs*; first version 70% more expensive; 111-case benchmark *theirs*; CDC 5,000 Jev 72.2% AUC 0.7725 *theirs*; accuracy is a trap; 9.0% base rate always-no 91.0%; catalog ≠ endorsement; jev-skill 109★ 90 scenarios HEAD 4f6e899a24d4; awesome-jev-live 673 entries 4★; minecraft-agent 214★ 131 JEV decisions 35 Astra calls; nether-final-08 8 minutes 43.300 seconds; planner writes JEV selects; RoboJEV structured simulator state not images; ashare-trader 策略未通过自己的回测门槛; 36 组参数全部净期望为负; no positive expectation under real costs; typed_evals NOT an official TypeSafe AI product; jev-as-judge is a sensor; third-person-audit 40% & 60% watermarks still soft; The included experience uses a handwritten demo provider; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33/#34/#35/#36/#37/#38/#39/#40/#41/#42; notes.md §120' +) + OVERLAYS = [ "research/notes.md", "research/changelog-hourly.md", @@ -176,6 +181,8 @@ def main() -> int: failed.append(f"0940 lock missing as one substring: {rel}") if UNIQ_0947 not in body: failed.append(f"0947 lock missing as one substring: {rel}") + if UNIQ_1049 not in body: + failed.append(f"1049 lock missing as one substring: {rel}") notes = (ROOT / "research/notes.md").read_text(encoding="utf-8") if "## 114. Hourly 0843 HIGH" not in notes: failed.append("notes.md missing §114 heading") @@ -189,10 +196,12 @@ def main() -> int: failed.append("notes.md missing §118 heading") if "## 119. Hourly 0947 HIGH" not in notes: failed.append("notes.md missing §119 heading") + if "## 120. Hourly 1049 HIGH" not in notes: + failed.append("notes.md missing §120 heading") algebra = (ROOT / ".agents/skills/augustus/references/composition-algebra.md").read_text( encoding="utf-8" ) - for n in list(range(289, 317)) + list(range(322, 330)) + list(range(330, 337)) + list(range(337, 353)): + for n in list(range(289, 317)) + list(range(322, 330)) + list(range(330, 337)) + list(range(337, 353)) + list(range(353, 369)): needle = f"{n}. **" if needle not in algebra: failed.append(f"composition-algebra missing item {n}") @@ -208,6 +217,7 @@ def main() -> int: "## Batch #100", "## Batch #101", "## Batch #102", + "## Batch #103", ): if batch not in findings: failed.append(f"findings.md missing {batch}") @@ -251,6 +261,11 @@ def main() -> int: "candidate_mass", "Qwen3.5-2B ≠ Archer", "Qwen3.5-4B ≠ Archer", + "ggmlc GGUF is not llama.cpp", + "serving substrate ≠ calibrated replica", + "Qwen3.5-9B ≠ Archer", + "planner writes JEV selects", + "pick_by_id vs pick_second", ): if frag not in haystack: failed.append(f"SKILL.md missing fragment {frag!r}") @@ -272,6 +287,11 @@ def main() -> int: "candidate_mass", "softmax over A–H ≠ Noul", "hourly 0947 / notes.md §119", + "ggmlc GGUF is not llama.cpp", + "serving substrate ≠ calibrated replica", + "Qwen3.5-9B ≠ Archer", + "planner writes JEV selects", + "hourly 1049 / notes.md §120", ): if frag not in proto_line: failed.append(f"SKILL.md protocol missing {frag!r}") @@ -283,6 +303,7 @@ def main() -> int: ("0922", UNIQ_0922), ("0940", UNIQ_0940), ("0947", UNIQ_0947), + ("1049", UNIQ_1049), ): if lock in changelog: failed.append( @@ -314,6 +335,7 @@ def main() -> int: f"0843 chars={len(UNIQ_0843)} 0915 chars={len(UNIQ_0915)} " f"jcr chars={len(UNIQ_JCR)} lock0922 chars={len(UNIQ_0922)} " f"0940 chars={len(UNIQ_0940)} 0947 chars={len(UNIQ_0947)} " + f"1049 chars={len(UNIQ_1049)} " f"overlays={len(OVERLAYS)}" ) return 0 diff --git a/CHANGELOG.md b/CHANGELOG.md index d0006d3e..0decb9ae 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -16,6 +16,34 @@ folds: `research/notes.md`. ## [Unreleased] +Hourly 1049 HIGH (`research/notes.md` §120 / composition items +353–368 / findings batch #103). Does **not** bump the 0.5.0 pin. +Uniqueness dumps live in +[`research/changelog-hourly.md`](research/changelog-hourly.md). +Do not reopen or amend PR #23–#42. Do not amend released 0.5.0 +(#42). Merged #41 owns §119. + +### Added + +- **Hourly 1049 HIGH (`notes.md` §120).** ggmlc GGUF serving / + option-order / exact-p Minesweeper / adapters / catalogs. + ggmlc GGUF is not llama.cpp. Serving substrate is not a + calibrated replica. Qwen3.5-9B ≠ Archer. planner writes + JEV selects. pick_by_id vs pick_second. Soft scores ≠ hard + gates. catalog ≠ endorsement. Evaluator: option-order + trap. uniqueness_gate.py now checks 0843 + 0915 + jcr + + 0922 + 0940 + 0947 + 1049. Composition items 353–368 / + batch #103. **HARD RULE:** do not reopen or amend PR + #23–#42. Does **not** bump 0.5.0. + +- **Recipe (class, not Jev-only).** Without Augustus: load a + GGUF or ONNX graph and treat the new bottle as a replica, + or pick whatever sits second in a shuffled menu. With + Augustus: serving substrate ≠ calibrated replica; measure + option-order with pick_by_id vs pick_second; keep exact-p + oracles in code; adapters stay class members. Same split + for any Choice/Score/Noul-style head, not only hosted Jev. + Hourly 0947 HIGH (`research/notes.md` §119 / composition items 337–352 / findings batch #102). Does **not** bump the 0.5.0 pin. Uniqueness dumps live in diff --git a/README.md b/README.md index b077c828..c321b8fc 100644 --- a/README.md +++ b/README.md @@ -142,4 +142,6 @@ User-provided 0920 jcr uniqueness lock: NiazMorshed2007/jcr MIT; site https://jc User-provided 0922 uniqueness lock: SemIf was formerly OpenJev; independent; not affiliated with Jev or TypeSafe; homepage openjev.com; default master; MIT; HEAD ca3ba65f1429; Tolerate float roundoff in MLX evidence verification; pushed 2026-09-19; live REST 2282★ / 140 forks; size 9177; README SHA 74ab7f7f; LICENSE SHA ca562883; interface pattern reproduction with open models; does not reproduce Jev undisclosed model/training; Direct option logits; 0 output tokens; shared-state parallel; MLX backend for Apple Silicon (`--backend mlx`); Qwen3.5-4B 3090 direct 1.023s vs AR JSON 5.332s (**5.21×**); argmax agree 18/21; systems comparison ≠ semantic equivalence; Parallel suffixes 20.03 dec/s on 777 decisions; Browser ladder Qwen3.5-4B authored BA 0.813, pert 0.766, TypeSafe subset agreement 0.845 vs Published Jev 0.883 (102 across 20 cases); Softmax over options ≠ calibrated Noul; typed output does not guarantee semantic correctness; wire/agreement ≠ replica of TypeSafe; SemIf ≠ kw2828/OpenJev playground ≠ zhihz/openjev ≠ apiplant/semif-rs port ≠ dddanielliu/semif-serve; rename is densify not a second census; JevBench 74.6 is §78 not this ladder; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33/#34/#35/#36/#38; do not push onto open #39/#40; notes.md §117 User-provided 0940 uniqueness lock: Turn decision-shaped LLM prompts into proposed Jev primitives; This is a conversion assistant, not an automatic guarantee of equivalent behavior; The compiler uses deterministic heuristics, not an LLM or evaluation model; It understands a deliberately small set of common prompt patterns; Generated instructions and criteria must be reviewed before production use; Score ranges such as 0 to 1 are translated into ordered Jev criteria; Prompts requiring open-ended prose are not a fit; suitability strong/partial/not_a_fit; compatibility full/partial/none; Writing new text stays with an LLM; Review the generated Score rubric; Jev scores ordered criteria, not an arbitrary 0-to-1 range; Everything runs locally in the browser; There is no framework, database, account, API, or server-side prompt processing; The key is read from the process environment and is never stored or printed; connect-src 'none'; alexwestco/llm-to-jev ≠ altryne/jevify ≠ ryana/jevify ≠ fidecastro/jevify ≠ Mintzs/jevify ≠ gulagala001/jevify ≠ uspraveen/Jevify; HEAD 234058ab372d; README SHA 43cd94fb; LICENSE SHA 5f334006; compiler SHA fdf235d0; 2★; MIT; JavaScript; size 29; Pages https://alexwestco.github.io/llm-to-jev/; invented_signal false; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33/#34/#35; notes.md §118 Hourly 0947 uniqueness lock: Fast and cheap agent evals. jev as judge.; 18,041 skills from the 200 most-starred repos; Not a security scanner; 最简 Jev 调用演示器; confidence 不是正确率; q93304989-bit/jev-lab ≠ tanayvasishtha/jev-lab ≠ dairui1/jev-lab ≠ BrendanH18/jev-lab ≠ yibie/laya-jev-lab; 75% cheaper and 18% faster withdrawn; jev @0.15 100% recall 87% savings; 33Audits/jev-auto ≠ gargpratyush/jev-router; no Typesafe key, no PI_API_BASE, zero deps; tool-emitted Score/Noul ≠ calibrated Noul; semantic_compatibility: false; candidate_mass; Qwen3.5-2B ≠ Archer; Jev evaluates decisions; it cannot run a coding-agent session; Status: no model yet; S1LV3RJ1NX/openjev ≠ TheoLeeCJ/openjev; 28 accepted decisions; 3 targets; score 800; health 100; arcade game not a flight trainer; A successful live TypeSafe call has not been verified for v0.1.0; abhibansal60/tidy ≠ MANISH007700/tidy; No model, Jev included, predicted which channels its owner keeps; seed 1 selected on a held-out 400-item validation split; Brier 0.342 → 0.378; more accurate and more overconfident; Qwen3.5-4B ≠ Archer; static quants of kushalpatil/jevify-gemma4-26b-a4b; The labels were corrected, and one earlier result was retracted; zero of 23,869 eligible rows; Do not compare cost without checking task success; Exit 1 is not a proof; kisshan13/typesafe-ai-go ≠ Nibir1/typesafe-go ≠ official; 38 tests that cannot fail in a 356-model warehouse; if a parser can answer it, Jev is never asked; 359 of them; Games & Simulation 82; Education & Learning 1; Ratings are heuristics; syedabbasshaheer-art/jev-atlas ≠ ZeroX-01/jev-atlas ≠ Zaious/jev-capability-atlas ≠ gorock007/jev-atlas; anandi1989/awesome-jev-usecases ≠ whyashthakker/awesome-jev-use-cases ≠ walidboulanouar/awesome-jev-use-cases ≠ vamsikrishna2421/jev-usecases; Every headline result above is self-reported; Archer Hume 84.6% MMLU-Pro is a third-party probe not landed Archer; catalog ≠ endorsement; judge ≠ actuator; softmax over A–H ≠ Noul; SemIf 2270★; jevlike 1054★; TypeLLM/TypeLLM 16★; AnotiaWang 98★; yibie/awesome-jev 538★; Laya likes 889; tracker likes 68 lastModified UNCHANGED; Blackwood likes 2 gated manual; Archer still promised_not_landed; Hub archerhume/4rcherhume HTTP 401; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33/#35/#36/#37/#38/#40; do not push onto open #39; notes.md §119 +- Hourly 1049 HIGH (`research/notes.md` §120 / items 353–368 / batch #103). ggmlc GGUF serving / option-order / exact-p Minesweeper / adapters / catalogs. uniqueness_gate 0843+0915+jcr+0922+0940+0947+1049. Does not bump 0.5.0. Merged #41 owns §119. Merged #42 is the 0.5.0 release. - Hourly 0947 HIGH (`research/notes.md` §119 / items 337–352 / batch #102). jev-as-judge / OneForward candidate_mass / catalogs / replay / RLCD heads. uniqueness_gate 0843+0915+jcr+0922+0940+0947. Does not bump 0.5.0. Merged #42 is the 0.5.0 release. Merged #40 owns §118. Merged #39 is §114 hygiene. Merged #37 owns §117. Merged #38 owns §116. +Hourly 1049 uniqueness lock: ggmlc GGUF is not llama.cpp; Loading them in llama.cpp will fail; one encoder pass; hf:mys/laya-GGUF sha 713ae6f6e39f likes 0 apache-2.0; hf:mys/laya-multilingual-GGUF sha 3b645ae54281; hf:mys/laya-typed-decisions-GGUF sha 1e9e8ba1f527; hf:tozp/laya-onnx sha 0862aeba1e65 Opset 14 FP32 and INT8; tozp/laya-onnx ≠ Mattepiu/laya-onnx ≠ gqgs/laya-onnx; docker-laya MIT HEAD 1b8239a51ddd README SHA 9cb7bdc3; laya.cpp RTX ggml CUDA HEAD 8590937c79a2 README SHA cdd429b9; serving substrate ≠ calibrated replica; Softmax over options ≠ calibrated Noul; Qwen3.5-4B ≠ Archer; Qwen3.5-9B ≠ Archer; jev-position-test n=6 HEAD 7a56ca1c2698 README SHA 23f194c9; jevmlx slots 5 of 6; hosted Jev 0 of 6; prior_correction made it worse; jevSweeper mean Spearman ρ −0.274; picked exact-optimal 1/25 (4%); 31 of 36 still logically decidable; 86% of the time we should not have been asking; game success ≠ calibrated Noul; LLM2Jev 64★ Apache-2.0 HEAD 924618721277 README SHA da35fe61; not affiliated with or endorsed by Jev or TypeSafe; No answer tokens are generated; OpenSourceJev llama.cpp Qwen3-1.7B HEAD 3c41fba3681d; JEV-MLX Qwen3.5-9B HEAD dec24cd929ea; decision-head-rlcd Qwen3.5-4B 4.9M LoRA; AUTO_ACT is not a Noul; closed-set fail-open stdlib-only; verified=False; soft scores ≠ hard gates; 22 to 40% cheaper *theirs*; first version 70% more expensive; 111-case benchmark *theirs*; CDC 5,000 Jev 72.2% AUC 0.7725 *theirs*; accuracy is a trap; 9.0% base rate always-no 91.0%; catalog ≠ endorsement; jev-skill 109★ 90 scenarios HEAD 4f6e899a24d4; awesome-jev-live 673 entries 4★; minecraft-agent 214★ 131 JEV decisions 35 Astra calls; nether-final-08 8 minutes 43.300 seconds; planner writes JEV selects; RoboJEV structured simulator state not images; ashare-trader 策略未通过自己的回测门槛; 36 组参数全部净期望为负; no positive expectation under real costs; typed_evals NOT an official TypeSafe AI product; jev-as-judge is a sensor; third-person-audit 40% & 60% watermarks still soft; The included experience uses a handwritten demo provider; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33/#34/#35/#36/#37/#38/#39/#40/#41/#42; notes.md §120 diff --git a/docs/ecosystem.md b/docs/ecosystem.md index 85962adf..bb7d10ed 100644 --- a/docs/ecosystem.md +++ b/docs/ecosystem.md @@ -18,6 +18,8 @@ weekdays. Jev is the densest public corpus, not the class monopoly. ## Categories +- **Hourly 1049 serving / measurement (`notes.md` §120).** ggmlc GGUF is not llama.cpp. serving substrate ≠ calibrated replica. pick_by_id vs pick_second. Qwen3.5-9B ≠ Archer. planner writes JEV selects. catalog ≠ endorsement. + ### Decision engines & search - **paulobueno164/jev-mcts** — MCTS with typed fidelity (grounded depth 24, speculative capped at 2), "only probes concede," calibrated thresholds, seeded debias. 24/24 vs greedy 1/24. - **lhemerly/mcts-agent** — batched Noul prune + Choice PUCT priors + Score/10 leaf value (no random rollouts). @@ -1138,3 +1140,4 @@ User-provided 0920 jcr uniqueness lock: NiazMorshed2007/jcr MIT; site https://jc User-provided 0922 uniqueness lock: SemIf was formerly OpenJev; independent; not affiliated with Jev or TypeSafe; homepage openjev.com; default master; MIT; HEAD ca3ba65f1429; Tolerate float roundoff in MLX evidence verification; pushed 2026-09-19; live REST 2282★ / 140 forks; size 9177; README SHA 74ab7f7f; LICENSE SHA ca562883; interface pattern reproduction with open models; does not reproduce Jev undisclosed model/training; Direct option logits; 0 output tokens; shared-state parallel; MLX backend for Apple Silicon (`--backend mlx`); Qwen3.5-4B 3090 direct 1.023s vs AR JSON 5.332s (**5.21×**); argmax agree 18/21; systems comparison ≠ semantic equivalence; Parallel suffixes 20.03 dec/s on 777 decisions; Browser ladder Qwen3.5-4B authored BA 0.813, pert 0.766, TypeSafe subset agreement 0.845 vs Published Jev 0.883 (102 across 20 cases); Softmax over options ≠ calibrated Noul; typed output does not guarantee semantic correctness; wire/agreement ≠ replica of TypeSafe; SemIf ≠ kw2828/OpenJev playground ≠ zhihz/openjev ≠ apiplant/semif-rs port ≠ dddanielliu/semif-serve; rename is densify not a second census; JevBench 74.6 is §78 not this ladder; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33/#34/#35/#36/#38; do not push onto open #39/#40; notes.md §117 User-provided 0940 uniqueness lock: Turn decision-shaped LLM prompts into proposed Jev primitives; This is a conversion assistant, not an automatic guarantee of equivalent behavior; The compiler uses deterministic heuristics, not an LLM or evaluation model; It understands a deliberately small set of common prompt patterns; Generated instructions and criteria must be reviewed before production use; Score ranges such as 0 to 1 are translated into ordered Jev criteria; Prompts requiring open-ended prose are not a fit; suitability strong/partial/not_a_fit; compatibility full/partial/none; Writing new text stays with an LLM; Review the generated Score rubric; Jev scores ordered criteria, not an arbitrary 0-to-1 range; Everything runs locally in the browser; There is no framework, database, account, API, or server-side prompt processing; The key is read from the process environment and is never stored or printed; connect-src 'none'; alexwestco/llm-to-jev ≠ altryne/jevify ≠ ryana/jevify ≠ fidecastro/jevify ≠ Mintzs/jevify ≠ gulagala001/jevify ≠ uspraveen/Jevify; HEAD 234058ab372d; README SHA 43cd94fb; LICENSE SHA 5f334006; compiler SHA fdf235d0; 2★; MIT; JavaScript; size 29; Pages https://alexwestco.github.io/llm-to-jev/; invented_signal false; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33/#34/#35; notes.md §118 Hourly 0947 uniqueness lock: Fast and cheap agent evals. jev as judge.; 18,041 skills from the 200 most-starred repos; Not a security scanner; 最简 Jev 调用演示器; confidence 不是正确率; q93304989-bit/jev-lab ≠ tanayvasishtha/jev-lab ≠ dairui1/jev-lab ≠ BrendanH18/jev-lab ≠ yibie/laya-jev-lab; 75% cheaper and 18% faster withdrawn; jev @0.15 100% recall 87% savings; 33Audits/jev-auto ≠ gargpratyush/jev-router; no Typesafe key, no PI_API_BASE, zero deps; tool-emitted Score/Noul ≠ calibrated Noul; semantic_compatibility: false; candidate_mass; Qwen3.5-2B ≠ Archer; Jev evaluates decisions; it cannot run a coding-agent session; Status: no model yet; S1LV3RJ1NX/openjev ≠ TheoLeeCJ/openjev; 28 accepted decisions; 3 targets; score 800; health 100; arcade game not a flight trainer; A successful live TypeSafe call has not been verified for v0.1.0; abhibansal60/tidy ≠ MANISH007700/tidy; No model, Jev included, predicted which channels its owner keeps; seed 1 selected on a held-out 400-item validation split; Brier 0.342 → 0.378; more accurate and more overconfident; Qwen3.5-4B ≠ Archer; static quants of kushalpatil/jevify-gemma4-26b-a4b; The labels were corrected, and one earlier result was retracted; zero of 23,869 eligible rows; Do not compare cost without checking task success; Exit 1 is not a proof; kisshan13/typesafe-ai-go ≠ Nibir1/typesafe-go ≠ official; 38 tests that cannot fail in a 356-model warehouse; if a parser can answer it, Jev is never asked; 359 of them; Games & Simulation 82; Education & Learning 1; Ratings are heuristics; syedabbasshaheer-art/jev-atlas ≠ ZeroX-01/jev-atlas ≠ Zaious/jev-capability-atlas ≠ gorock007/jev-atlas; anandi1989/awesome-jev-usecases ≠ whyashthakker/awesome-jev-use-cases ≠ walidboulanouar/awesome-jev-use-cases ≠ vamsikrishna2421/jev-usecases; Every headline result above is self-reported; Archer Hume 84.6% MMLU-Pro is a third-party probe not landed Archer; catalog ≠ endorsement; judge ≠ actuator; softmax over A–H ≠ Noul; SemIf 2270★; jevlike 1054★; TypeLLM/TypeLLM 16★; AnotiaWang 98★; yibie/awesome-jev 538★; Laya likes 889; tracker likes 68 lastModified UNCHANGED; Blackwood likes 2 gated manual; Archer still promised_not_landed; Hub archerhume/4rcherhume HTTP 401; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33/#35/#36/#37/#38/#40; do not push onto open #39; notes.md §119 +Hourly 1049 uniqueness lock: ggmlc GGUF is not llama.cpp; Loading them in llama.cpp will fail; one encoder pass; hf:mys/laya-GGUF sha 713ae6f6e39f likes 0 apache-2.0; hf:mys/laya-multilingual-GGUF sha 3b645ae54281; hf:mys/laya-typed-decisions-GGUF sha 1e9e8ba1f527; hf:tozp/laya-onnx sha 0862aeba1e65 Opset 14 FP32 and INT8; tozp/laya-onnx ≠ Mattepiu/laya-onnx ≠ gqgs/laya-onnx; docker-laya MIT HEAD 1b8239a51ddd README SHA 9cb7bdc3; laya.cpp RTX ggml CUDA HEAD 8590937c79a2 README SHA cdd429b9; serving substrate ≠ calibrated replica; Softmax over options ≠ calibrated Noul; Qwen3.5-4B ≠ Archer; Qwen3.5-9B ≠ Archer; jev-position-test n=6 HEAD 7a56ca1c2698 README SHA 23f194c9; jevmlx slots 5 of 6; hosted Jev 0 of 6; prior_correction made it worse; jevSweeper mean Spearman ρ −0.274; picked exact-optimal 1/25 (4%); 31 of 36 still logically decidable; 86% of the time we should not have been asking; game success ≠ calibrated Noul; LLM2Jev 64★ Apache-2.0 HEAD 924618721277 README SHA da35fe61; not affiliated with or endorsed by Jev or TypeSafe; No answer tokens are generated; OpenSourceJev llama.cpp Qwen3-1.7B HEAD 3c41fba3681d; JEV-MLX Qwen3.5-9B HEAD dec24cd929ea; decision-head-rlcd Qwen3.5-4B 4.9M LoRA; AUTO_ACT is not a Noul; closed-set fail-open stdlib-only; verified=False; soft scores ≠ hard gates; 22 to 40% cheaper *theirs*; first version 70% more expensive; 111-case benchmark *theirs*; CDC 5,000 Jev 72.2% AUC 0.7725 *theirs*; accuracy is a trap; 9.0% base rate always-no 91.0%; catalog ≠ endorsement; jev-skill 109★ 90 scenarios HEAD 4f6e899a24d4; awesome-jev-live 673 entries 4★; minecraft-agent 214★ 131 JEV decisions 35 Astra calls; nether-final-08 8 minutes 43.300 seconds; planner writes JEV selects; RoboJEV structured simulator state not images; ashare-trader 策略未通过自己的回测门槛; 36 组参数全部净期望为负; no positive expectation under real costs; typed_evals NOT an official TypeSafe AI product; jev-as-judge is a sensor; third-person-audit 40% & 60% watermarks still soft; The included experience uses a handwritten demo provider; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33/#34/#35/#36/#37/#38/#39/#40/#41/#42; notes.md §120 diff --git a/research/archive/findings.md b/research/archive/findings.md index 482ff5b8..c1c249ae 100644 --- a/research/archive/findings.md +++ b/research/archive/findings.md @@ -1,6 +1,50 @@ # Deep-read findings (evidence for research/notes.md) +## Batch #103 (2026-09-20 ~10:49 Boise / ~16:49 UTC) — hourly 1049 HIGH + +Hourly 1049 uniqueness lock: ggmlc GGUF is not llama.cpp; Loading them in llama.cpp will fail; one encoder pass; hf:mys/laya-GGUF sha 713ae6f6e39f likes 0 apache-2.0; hf:mys/laya-multilingual-GGUF sha 3b645ae54281; hf:mys/laya-typed-decisions-GGUF sha 1e9e8ba1f527; hf:tozp/laya-onnx sha 0862aeba1e65 Opset 14 FP32 and INT8; tozp/laya-onnx ≠ Mattepiu/laya-onnx ≠ gqgs/laya-onnx; docker-laya MIT HEAD 1b8239a51ddd README SHA 9cb7bdc3; laya.cpp RTX ggml CUDA HEAD 8590937c79a2 README SHA cdd429b9; serving substrate ≠ calibrated replica; Softmax over options ≠ calibrated Noul; Qwen3.5-4B ≠ Archer; Qwen3.5-9B ≠ Archer; jev-position-test n=6 HEAD 7a56ca1c2698 README SHA 23f194c9; jevmlx slots 5 of 6; hosted Jev 0 of 6; prior_correction made it worse; jevSweeper mean Spearman ρ −0.274; picked exact-optimal 1/25 (4%); 31 of 36 still logically decidable; 86% of the time we should not have been asking; game success ≠ calibrated Noul; LLM2Jev 64★ Apache-2.0 HEAD 924618721277 README SHA da35fe61; not affiliated with or endorsed by Jev or TypeSafe; No answer tokens are generated; OpenSourceJev llama.cpp Qwen3-1.7B HEAD 3c41fba3681d; JEV-MLX Qwen3.5-9B HEAD dec24cd929ea; decision-head-rlcd Qwen3.5-4B 4.9M LoRA; AUTO_ACT is not a Noul; closed-set fail-open stdlib-only; verified=False; soft scores ≠ hard gates; 22 to 40% cheaper *theirs*; first version 70% more expensive; 111-case benchmark *theirs*; CDC 5,000 Jev 72.2% AUC 0.7725 *theirs*; accuracy is a trap; 9.0% base rate always-no 91.0%; catalog ≠ endorsement; jev-skill 109★ 90 scenarios HEAD 4f6e899a24d4; awesome-jev-live 673 entries 4★; minecraft-agent 214★ 131 JEV decisions 35 Astra calls; nether-final-08 8 minutes 43.300 seconds; planner writes JEV selects; RoboJEV structured simulator state not images; ashare-trader 策略未通过自己的回测门槛; 36 组参数全部净期望为负; no positive expectation under real costs; typed_evals NOT an official TypeSafe AI product; jev-as-judge is a sensor; third-person-audit 40% & 60% watermarks still soft; The included experience uses a handwritten demo provider; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33/#34/#35/#36/#37/#38/#39/#40/#41/#42; notes.md §120 + +Note: `research/notes.md` §120. Docs + evaluator, rebased +onto latest `main` (`8f446c4` / merged #41 0947) after +`fb15455` / merged #42 v0.5.0 after `38e4e92` / merged #39 +§114 hygiene. Merged #41 owns §119 / 337–352 / #102. +Merged #40 owns §118 / 322–329 / #101. Merged #42 is the +0.5.0 release. This fold stays §120 / items 353–368 / +batch #103. +**HARD RULE:** do not reopen or amend PR #23–#42. +Quote READMEs. Soft Noul ≠ hard safety. Augustus owns +placement. `invented_signal: false`. + +- **Serving substrate PRIMARY.** mys/laya-GGUF family. docker-laya. laya.cpp. tozp ONNX. + ggmlc GGUF is not llama.cpp. Loading them in llama.cpp will fail. one encoder pass. + serving substrate ≠ calibrated replica. tozp/laya-onnx ≠ Mattepiu/laya-onnx ≠ gqgs/laya-onnx. +- **Option-order / exact-p.** imaddde867/jev-position-test. lvk901/jevSweeper. + n=6. jevmlx slots 5 of 6. hosted Jev 0 of 6. prior_correction made it worse. + mean Spearman ρ −0.274. picked exact-optimal 1/25 (4%). 31 of 36 still logically decidable. + 86% of the time we should not have been asking. game success ≠ calibrated Noul. + Evaluator: pick_by_id vs pick_second. +- **Adapters / RLCD.** LLM2Jev 64★. OpenSourceJev llama.cpp Qwen3-1.7B. JEV-MLX Qwen3.5-9B. + decision-head-rlcd Qwen3.5-4B 4.9M LoRA densify. not affiliated with or endorsed by Jev or TypeSafe. + No answer tokens are generated. Qwen3.5-9B ≠ Archer. Qwen3.5-4B ≠ Archer. +- **Soft ≠ hard.** AUTO_ACT is not a Noul. closed-set fail-open stdlib-only. verified=False. + 22 to 40% cheaper *theirs*. first version 70% more expensive. + third-person-audit 40% & 60% watermarks still soft. + The included experience uses a handwritten demo provider. +- **Catalogs / life / honesty.** jev-skill 109★ 90 scenarios. awesome-jev-live 673 entries 4★. + minecraft-agent 214★ 131 JEV decisions 35 Astra calls. nether-final-08 8 minutes 43.300 seconds. + planner writes JEV selects. RoboJEV structured simulator state not images. + ashare-trader 策略未通过自己的回测门槛; 36 组参数全部净期望为负. + typed_evals NOT an official TypeSafe AI product. jev-as-judge is a sensor. + CDC 5,000 Jev 72.2% AUC 0.7725 *theirs*. accuracy is a trap. 9.0% base rate always-no 91.0%. + 111-case benchmark *theirs*. catalog ≠ endorsement. + +Pulse: Archer still NOT landed. Hub archerhume/4rcherhume HTTP **401**. +minecraft-agent **214★**; jev-skill **109★**; LLM2Jev **64★**; RoboJEV **8★**. +`invented_signal: false`. + + + ## Batch #102 (2026-09-20 ~09:47 Boise / ~15:47 UTC) — hourly 0947 HIGH diff --git a/research/archive/hourly/2026-09-20T16/meaning_bullets.json b/research/archive/hourly/2026-09-20T16/meaning_bullets.json new file mode 100644 index 00000000..4364a86b --- /dev/null +++ b/research/archive/hourly/2026-09-20T16/meaning_bullets.json @@ -0,0 +1,73 @@ +{ + "boise_label": "1049", + "novel_high_count": 98, + "bullets": [ + "Laya ports surge: GGUF/ONNX (mys/*, tozp), docker-laya, laya.cpp, laya-pong, laya-tetris-finetune", + "Measurement/judges: jev-field-tests, eval-diff, typed_evals, jev-heart-risk-bench, jev-position-test, third-person-audit", + "CU/browser/MCP: jev-browser, playwright-jev, jev-mcp-spring, dsh-jev tools, paseo-supervision", + "Gate cousins: hermes-jev tool gate, pi-jev-router, jev-regime-gate, Prompt-Injection-Guardian, healthcare router", + "Open replicas/adapters: LLM2Jev, OpenSourceJev, JEV-MLX, decision-head-rlcd, Minecraft/RoboJEV controllers" + ], + "top_starred_novel": [ + { + "id": "rmalde/minecraft-agent", + "stars": 206, + "snip": "Astra planner and JEV controller for Minecraft, with native recording, tested routes, and run verifi" + }, + { + "id": "wuyoscar/jev-skill", + "stars": 108, + "snip": "An awesome collection of Jev use cases, workflows, and agent skills." + }, + { + "id": "Yinsongxu/LLM2Jev", + "stars": 64, + "snip": "Adapt local language models into Jev-compatible structured decision engines with Choice, Score, and " + }, + { + "id": "FerryCorleone/crush-monitor", + "stars": 28, + "snip": "Crush \u597d\u611f\u76d1\u63a7\u5668\uff1a\u7528 Jev \u5206\u6790\u5fae\u4fe1\u804a\u5929\u7684\u60c5\u7eea\u3001\u610f\u56fe\u548c\u56de\u590d\u8868\u73b0\u3002\u672c\u673a\u90e8\u7f72\uff0c\u4f7f\u7528\u81ea\u5df1\u7684 API Key\u3002" + }, + { + "id": "sabeel111/OpenSourceJev", + "stars": 10, + "snip": "Turning an LLM model into a Jev like System." + }, + { + "id": "shitianfang/jev-use", + "stars": 10, + "snip": "Claude Code / Codex / pi plugin that hands agent steps needing no text output to Jev (TypeSafe's jud" + }, + { + "id": "keeltrace/hermes-jev", + "stars": 9, + "snip": "Typed System One decisions, ranking, verification, and an opt-in Hermes tool gate using TypeSafe Jev" + }, + { + "id": "lykycy123/RoboJEV", + "stars": 6, + "snip": "Two-stage JEV control of a Franka Panda in MuJoCo: pick/place, surface pushing, and fixed-pedestal s" + }, + { + "id": "hoangnb24/paseo-supervision", + "stars": 4, + "snip": "Paseo plugin for supervising Lead\u2013Peer communication protocol drift with Jev" + }, + { + "id": "wh000wh000/awesome-jev-live", + "stars": 4, + "snip": "Evidence-graded index of the Jev / TypeSafe System One ecosystem. Rebuilt every 2 hours in 20 langua" + }, + { + "id": "ChosenXu/newsletter-link-harvester", + "stars": 3, + "snip": "Agent Skill: harvest links from newsletter emails into Raindrop.io with the author editorial context" + }, + { + "id": "hraness/algal", + "stars": 3, + "snip": "ALGAL is a new take on the agent graph: a language and runtime for agentic program evolution." + } + ] +} \ No newline at end of file diff --git a/research/archive/hourly/2026-09-20T16/novel_high_this_run.json b/research/archive/hourly/2026-09-20T16/novel_high_this_run.json new file mode 100644 index 00000000..43706f69 --- /dev/null +++ b/research/archive/hourly/2026-09-20T16/novel_high_this_run.json @@ -0,0 +1,1540 @@ +[ + { + "id": "ashleyotooligan/jevbrain", + "source": "github", + "why_high": "s1_family,decision_scoring", + "url": "https://github.com/ashleyotooligan/jevbrain", + "description_snip": "Add the decision engine, optional Jev AI integration, offline replay workbench, original pixel-brain logo, 18 reproducible demo experiments, visualizations, tes", + "stars": 1, + "created": "2026-09-20T16:29:19Z", + "recent": true, + "tier": "HIGH", + "description": "Add the decision engine, optional Jev AI integration, offline replay workbench, original pixel-brain logo, 18 reproducible demo experiments, visualizations, tests and detailed documentation.", + "fold_targets": [ + "Augustus" + ] + }, + { + "id": "vladzima/jev-x", + "source": "github", + "why_high": "s1_family", + "url": "https://github.com/vladzima/jev-x", + "description_snip": "Cut the noise on your X timeline: Jev (TypeSafe System One) scores posts on firsthand experience, promo, bait, depth, and relevance; your sliders decide what ge", + "stars": 1, + "created": "2026-09-20T16:30:45Z", + "recent": true, + "tier": "HIGH", + "description": "Cut the noise on your X timeline: Jev (TypeSafe System One) scores posts on firsthand experience, promo, bait, depth, and relevance; your sliders decide what gets dimmed or collapsed.", + "fold_targets": [ + "Augustus", + "Jev-omni" + ] + }, + { + "id": "AliZareh-CoE/JevRev", + "source": "github", + "why_high": "s1_family", + "url": "https://github.com/AliZareh-CoE/JevRev", + "description_snip": "Narrow a literature review with Jev: typed, calibrated relevance judgments over abstracts and paragraphs.", + "stars": 0, + "created": "2026-09-20T16:17:09Z", + "recent": true, + "tier": "HIGH", + "description": "Narrow a literature review with Jev: typed, calibrated relevance judgments over abstracts and paragraphs.", + "fold_targets": [ + "Augustus", + "Jev-omni" + ] + }, + { + "id": "Ashfaqbs/jev-mcp-spring", + "source": "github", + "why_high": "s1_family,kit_runtime", + "url": "https://github.com/Ashfaqbs/jev-mcp-spring", + "description_snip": "Java/Spring Boot MCP server for TypeSafe Jev", + "stars": 0, + "created": "2026-09-20T16:24:45Z", + "recent": true, + "tier": "HIGH", + "description": "Java/Spring Boot MCP server for TypeSafe Jev", + "fold_targets": [ + "Jev-omni" + ] + }, + { + "id": "HorusJiang/dsh-jev-tools", + "source": "github", + "why_high": "s1_family,guardrail,kit_runtime", + "url": "https://github.com/HorusJiang/dsh-jev-tools", + "description_snip": "Jev judgment, not generation: prune long tool output, screen fetched pages for injected instructions, and gate completion claims inside DeepSeek Harness.", + "stars": 0, + "created": "2026-09-20T16:32:55Z", + "recent": true, + "tier": "HIGH", + "description": "Jev judgment, not generation: prune long tool output, screen fetched pages for injected instructions, and gate completion claims inside DeepSeek Harness.", + "fold_targets": [ + "rh-guard", + "Jev-omni" + ] + }, + { + "id": "Hugeliz/jev.mixrnation.ch", + "source": "github", + "why_high": "s1_family", + "url": "https://github.com/Hugeliz/jev.mixrnation.ch", + "description_snip": "A test with Jev - rate a website", + "stars": 0, + "created": "2026-09-20T16:23:54Z", + "recent": true, + "tier": "HIGH", + "description": "A test with Jev - rate a website", + "fold_targets": [ + "Augustus", + "Jev-omni" + ] + }, + { + "id": "Md-Zaid-Ahmed/GMAIL-JEV-DEV", + "source": "github", + "why_high": "s1_family", + "url": "https://github.com/Md-Zaid-Ahmed/GMAIL-JEV-DEV", + "description_snip": "Email classification using jev - Gmail (dev)", + "stars": 0, + "created": "2026-09-20T16:43:28Z", + "recent": true, + "tier": "HIGH", + "description": "Email classification using jev - Gmail (dev)", + "fold_targets": [ + "Augustus", + "Jev-omni" + ] + }, + { + "id": "MrSAO666/jev-file-explorer", + "source": "github", + "why_high": "s1_family", + "url": "https://github.com/MrSAO666/jev-file-explorer", + "description_snip": "Windows \u6587\u4ef6\u68c0\u7d22\u4e0e\u6e05\u7406\u684c\u9762\u8f6f\u4ef6 \u00b7 Jev \u8bed\u4e49\u68c0\u7d22 \u00b7 SHA-256 \u91cd\u590d\u68c0\u6d4b \u00b7 PySide6 + Anaconda", + "stars": 0, + "created": "2026-09-20T16:21:38Z", + "recent": true, + "tier": "HIGH", + "description": "Windows \u6587\u4ef6\u68c0\u7d22\u4e0e\u6e05\u7406\u684c\u9762\u8f6f\u4ef6 \u00b7 Jev \u8bed\u4e49\u68c0\u7d22 \u00b7 SHA-256 \u91cd\u590d\u68c0\u6d4b \u00b7 PySide6 + Anaconda", + "fold_targets": [ + "Augustus", + "Jev-omni" + ] + }, + { + "id": "Steven04hub/qq_clinet", + "source": "github", + "why_high": "s1_family", + "url": "https://github.com/Steven04hub/qq_clinet", + "description_snip": "\u57fa\u4e8e NapCat \u7684 QQ AI \u7fa4\u804a\u673a\u5668\u4eba\uff0c\u652f\u6301 Web \u63a7\u5236\u53f0\u3001\u7fa4\u52a9\u624b\u4e0e\u804a\u5929\u6a21\u5f0f\u3001\u8bb0\u5fc6\u68c0\u7d22\u548c Jev \u51b3\u7b56\u3002", + "stars": 0, + "created": "2026-09-20T16:22:24Z", + "recent": true, + "tier": "HIGH", + "description": "\u57fa\u4e8e NapCat \u7684 QQ AI \u7fa4\u804a\u673a\u5668\u4eba\uff0c\u652f\u6301 Web \u63a7\u5236\u53f0\u3001\u7fa4\u52a9\u624b\u4e0e\u804a\u5929\u6a21\u5f0f\u3001\u8bb0\u5fc6\u68c0\u7d22\u548c Jev \u51b3\u7b56\u3002", + "fold_targets": [ + "Augustus", + "Jev-omni" + ] + }, + { + "id": "VarSamLewis/eval-diff", + "source": "github", + "why_high": "s1_family,measurement_bench", + "url": "https://github.com/VarSamLewis/eval-diff", + "description_snip": "CI tool to send PR diff to a system one model for change evaluation", + "stars": 0, + "created": "2026-09-20T16:27:16Z", + "recent": true, + "tier": "HIGH", + "description": "CI tool to send PR diff to a system one model for change evaluation", + "fold_targets": [ + "Augustus" + ] + }, + { + "id": "YuyaForest/JEV-Prompt-Injection-Guardian", + "source": "github", + "why_high": "s1_family", + "url": "https://github.com/YuyaForest/JEV-Prompt-Injection-Guardian", + "description_snip": "JEV Prompt Injection Guardian is a prompt injection quarantine and risk-scoring system for LLMs powered by JEV (jev-1.13.0), TypeSafe AI's innovative \"System On", + "stars": 0, + "created": "2026-09-20T16:11:09Z", + "recent": true, + "tier": "HIGH", + "description": "JEV Prompt Injection Guardian is a prompt injection quarantine and risk-scoring system for LLMs powered by JEV (jev-1.13.0), TypeSafe AI's innovative \"System One\" model, as its primary engine. By eliminating ambiguous, long-winded natural language explanations, it instantly computes rigorously calibrated probability risk-scores.", + "fold_targets": [ + "rh-guard", + "Augustus", + "Jev-omni" + ] + }, + { + "id": "bhaskarpraveen/jev-healthcare-support-router", + "source": "github", + "why_high": "s1_family,guardrail,decision_scoring", + "url": "https://github.com/bhaskarpraveen/jev-healthcare-support-router", + "description_snip": "TypeScript demo using Jev as a decision layer for healthcare customer-support routing, urgency detection, and human escalation.", + "stars": 0, + "created": "2026-09-20T16:44:50Z", + "recent": true, + "tier": "HIGH", + "description": "TypeScript demo using Jev as a decision layer for healthcare customer-support routing, urgency detection, and human escalation.", + "fold_targets": [ + "rh-guard", + "Augustus" + ] + }, + { + "id": "chneau/docker-laya", + "source": "github", + "why_high": "s1_family,guardrail,decision_scoring", + "url": "https://github.com/chneau/docker-laya", + "description_snip": "Dockerized FastAPI service for Laya typed-decision predictions: multi-checkpoint routing, API-key/Basic auth, presets and bulk inference.", + "stars": 0, + "created": "2026-09-20T16:29:10Z", + "recent": true, + "tier": "HIGH", + "description": "Dockerized FastAPI service for Laya typed-decision predictions: multi-checkpoint routing, API-key/Basic auth, presets and bulk inference.", + "fold_targets": [ + "Augustus", + "Jev-omni" + ] + }, + { + "id": "daniel-farina/nitro", + "source": "github", + "why_high": "s1_family,guardrail", + "url": "https://github.com/daniel-farina/nitro", + "description_snip": "Grok Build with TypeSafe Jev routing tool selection once per turn: 22 to 40% cheaper on the same tasks", + "stars": 0, + "created": "2026-09-20T16:40:44Z", + "recent": true, + "tier": "HIGH", + "description": "Grok Build with TypeSafe Jev routing tool selection once per turn: 22 to 40% cheaper on the same tasks", + "fold_targets": [ + "Augustus", + "Jev-omni" + ] + }, + { + "id": "dillera/prMonster", + "source": "github", + "why_high": "s1_family,kit_runtime", + "url": "https://github.com/dillera/prMonster", + "description_snip": "Jev-based PR triage harness for FujiNet firmware: gates, typed model questions, human-confirmed actions", + "stars": 0, + "created": "2026-09-20T16:31:56Z", + "recent": true, + "tier": "HIGH", + "description": "Jev-based PR triage harness for FujiNet firmware: gates, typed model questions, human-confirmed actions", + "fold_targets": [ + "rh-guard", + "Jev-omni" + ] + }, + { + "id": "dingw530/playwright-jev", + "source": "github", + "why_high": "s1_family,kit_runtime", + "url": "https://github.com/dingw530/playwright-jev", + "description_snip": "\u57fa\u4e8e Jev + playwright-cli \u7684\u81ea\u7136\u8bed\u8a00 Web E2E \u6d4b\u8bd5\u5de5\u5177\uff1aJev \u8d1f\u8d23\u51b3\u7b56\uff0cPlaywright \u8d1f\u8d23\u6267\u884c\uff0c\u4ee3\u7801\u8d1f\u8d23\u65ad\u8a00\u4e0e\u5b89\u5168\u8fb9\u754c\u3002Goal-driven web E2E testing with Jev + playwright-cli: bounded AI decisions, rea", + "stars": 0, + "created": "2026-09-20T15:54:04Z", + "recent": true, + "tier": "HIGH", + "description": "\u57fa\u4e8e Jev + playwright-cli \u7684\u81ea\u7136\u8bed\u8a00 Web E2E \u6d4b\u8bd5\u5de5\u5177\uff1aJev \u8d1f\u8d23\u51b3\u7b56\uff0cPlaywright \u8d1f\u8d23\u6267\u884c\uff0c\u4ee3\u7801\u8d1f\u8d23\u65ad\u8a00\u4e0e\u5b89\u5168\u8fb9\u754c\u3002Goal-driven web E2E testing with Jev + playwright-cli: bounded AI decisions, real browser execution, and deterministic assertions.", + "fold_targets": [ + "Jev-omni" + ] + }, + { + "id": "fly2abhishek/jev-field-tests", + "source": "github", + "why_high": "s1_family,guardrail,measurement_bench", + "url": "https://github.com/fly2abhishek/jev-field-tests", + "description_snip": "Twelve field tests for TypeSafe's Jev model: calibration, guardrails, r\u00e9sum\u00e9 screening, interview rubrics and its failure modes. Bring your own API key.", + "stars": 0, + "created": "2026-09-20T16:22:16Z", + "recent": true, + "tier": "HIGH", + "description": "Twelve field tests for TypeSafe's Jev model: calibration, guardrails, r\u00e9sum\u00e9 screening, interview rubrics and its failure modes. Bring your own API key.", + "fold_targets": [ + "Augustus" + ] + }, + { + "id": "fredrsat/stil-lint", + "source": "github", + "why_high": "s1_family,kit_runtime", + "url": "https://github.com/fredrsat/stil-lint", + "description_snip": "Style and quality linter for Norwegian and English text - MCP server for agents, with TypeSafe's Jev as the judgment layer. Reports findings, never authorship v", + "stars": 0, + "created": "2026-09-20T16:05:22Z", + "recent": true, + "tier": "HIGH", + "description": "Style and quality linter for Norwegian and English text - MCP server for agents, with TypeSafe's Jev as the judgment layer. Reports findings, never authorship verdicts.", + "fold_targets": [ + "Jev-omni" + ] + }, + { + "id": "gyu-don/jev-othello", + "source": "github", + "why_high": "s1_family", + "url": "https://github.com/gyu-don/jev-othello", + "description_snip": "TypeSafe AI \u306e\u8a55\u4fa1\u30e2\u30c7\u30eb Jev \u3068\u30ed\u30fc\u30ab\u30eb\u3067\u5bfe\u6226\u3059\u308b\u30aa\u30bb\u30ed", + "stars": 0, + "created": "2026-09-20T16:37:09Z", + "recent": true, + "tier": "HIGH", + "description": "TypeSafe AI \u306e\u8a55\u4fa1\u30e2\u30c7\u30eb Jev \u3068\u30ed\u30fc\u30ab\u30eb\u3067\u5bfe\u6226\u3059\u308b\u30aa\u30bb\u30ed", + "fold_targets": [ + "Augustus", + "Jev-omni" + ] + }, + { + "id": "havietkok-sys/BizzJev", + "source": "github", + "why_high": "s1_family", + "url": "https://github.com/havietkok-sys/BizzJev", + "description_snip": "Experiments with TypeSafe/Jev semantic gates and a Semantic Operations Lab demo.", + "stars": 0, + "created": "2026-09-20T15:49:29Z", + "recent": true, + "tier": "HIGH", + "description": "Experiments with TypeSafe/Jev semantic gates and a Semantic Operations Lab demo.", + "fold_targets": [ + "Augustus", + "Jev-omni" + ] + }, + { + "id": "hf:mys/laya-GGUF", + "source": "hf_model", + "why_high": "s1_family,open_weights_or_port,decision_scoring", + "url": "https://huggingface.co/mys/laya-GGUF", + "description_snip": "gguf ggmlc laya jev modernbert decision system-1 en base_model:convaiinnovations/laya base_model:quantized:convaiinnovations/laya license:apache-2.0 region:us", + "stars": 0, + "created": "2026-09-20T15:54:02.000Z", + "recent": true, + "tier": "HIGH", + "description": "gguf ggmlc laya jev modernbert decision system-1 en base_model:convaiinnovations/laya base_model:quantized:convaiinnovations/laya license:apache-2.0 region:us", + "fold_targets": [ + "Augustus", + "Jev-omni" + ] + }, + { + "id": "hf:mys/laya-multilingual-GGUF", + "source": "hf_model", + "why_high": "s1_family,open_weights_or_port,decision_scoring", + "url": "https://huggingface.co/mys/laya-multilingual-GGUF", + "description_snip": "gguf ggmlc laya jev mmbert multilingual decision system-1 base_model:convaiinnovations/laya-multilingual base_model:quantized:convaiinnovations/laya-multilingua", + "stars": 0, + "created": "2026-09-20T16:06:41.000Z", + "recent": true, + "tier": "HIGH", + "description": "gguf ggmlc laya jev mmbert multilingual decision system-1 base_model:convaiinnovations/laya-multilingual base_model:quantized:convaiinnovations/laya-multilingual license:apache-2.0 region:us", + "fold_targets": [ + "Augustus", + "Jev-omni" + ] + }, + { + "id": "hf:mys/laya-typed-decisions-GGUF", + "source": "hf_model", + "why_high": "s1_family,open_weights_or_port,decision_scoring", + "url": "https://huggingface.co/mys/laya-typed-decisions-GGUF", + "description_snip": "gguf ggmlc laya jev modernbert decision system-1 en base_model:convaiinnovations/laya-typed-decisions base_model:quantized:convaiinnovations/laya-typed-decision", + "stars": 0, + "created": "2026-09-20T16:16:06.000Z", + "recent": true, + "tier": "HIGH", + "description": "gguf ggmlc laya jev modernbert decision system-1 en base_model:convaiinnovations/laya-typed-decisions base_model:quantized:convaiinnovations/laya-typed-decisions license:apache-2.0 region:us", + "fold_targets": [ + "Augustus", + "Jev-omni" + ] + }, + { + "id": "hf:tozp/laya-onnx", + "source": "hf_model", + "why_high": "s1_family,open_weights_or_port,decision_scoring", + "url": "https://huggingface.co/tozp/laya-onnx", + "description_snip": "onnx modernbert onnxruntime laya system-one decision-making text-classification en base_model:convaiinnovations/laya base_model:quantized:convaiinnovations/laya", + "stars": 0, + "created": "2026-09-20T16:23:55.000Z", + "recent": true, + "tier": "HIGH", + "description": "onnx modernbert onnxruntime laya system-one decision-making text-classification en base_model:convaiinnovations/laya base_model:quantized:convaiinnovations/laya region:us", + "fold_targets": [ + "Augustus", + "Jev-omni" + ] + }, + { + "id": "hoaphm/jev-decision-maker", + "source": "github", + "why_high": "s1_family,kit_runtime,decision_scoring", + "url": "https://github.com/hoaphm/jev-decision-maker", + "description_snip": "omp plugin: JEV (TypeSafe System One) picks the next coding step from agent-supplied candidates", + "stars": 0, + "created": "2026-09-20T16:04:55Z", + "recent": true, + "tier": "HIGH", + "description": "omp plugin: JEV (TypeSafe System One) picks the next coding step from agent-supplied candidates", + "fold_targets": [ + "Jev-omni" + ] + }, + { + "id": "javierdv7/calibrated-decisions-iot-demo", + "source": "github", + "why_high": "s1_family,guardrail,open_weights_or_port", + "url": "https://github.com/javierdv7/calibrated-decisions-iot-demo", + "description_snip": "Interactive smart home demo comparing local Laya (MLX) with Jev (TypeSafe API) using shared routing and safety rules.", + "stars": 0, + "created": "2026-09-20T16:12:16Z", + "recent": true, + "tier": "HIGH", + "description": "Interactive smart home demo comparing local Laya (MLX) with Jev (TypeSafe API) using shared routing and safety rules.", + "fold_targets": [ + "Jev-omni", + "Augustus" + ] + }, + { + "id": "kuromoka/garmin-fuel-guide", + "source": "github", + "why_high": "s1_family", + "url": "https://github.com/kuromoka/garmin-fuel-guide", + "description_snip": "Experimental Garmin Connect IQ data field for Forerunner 265 with a Jev relay for fueling-plan review", + "stars": 0, + "created": "2026-09-20T16:39:04Z", + "recent": true, + "tier": "HIGH", + "description": "Experimental Garmin Connect IQ data field for Forerunner 265 with a Jev relay for fueling-plan review", + "fold_targets": [ + "Augustus", + "Jev-omni" + ] + }, + { + "id": "litshing/jevcore", + "source": "github", + "why_high": "s1_family,kit_runtime,measurement_bench", + "url": "https://github.com/litshing/jevcore", + "description_snip": "JEV core \u2014 the judgement primitive and harness for TypeSafe System One (Jev). Closed-set, fail-open, stdlib-only.", + "stars": 0, + "created": "2026-09-20T16:33:34Z", + "recent": true, + "tier": "HIGH", + "description": "JEV core \u2014 the judgement primitive and harness for TypeSafe System One (Jev). Closed-set, fail-open, stdlib-only.", + "fold_targets": [ + "Augustus", + "Jev-omni" + ] + }, + { + "id": "lvk901/jevSweeper", + "source": "github", + "why_high": "s1_family", + "url": "https://github.com/lvk901/jevSweeper", + "description_snip": "Jev (TypeSafe System One) plays Minesweeper, scored against an exact probability oracle. Its judgments came out worse than random \u2014 and 86% of the model calls t", + "stars": 0, + "created": "2026-09-20T16:32:37Z", + "recent": true, + "tier": "HIGH", + "description": "Jev (TypeSafe System One) plays Minesweeper, scored against an exact probability oracle. Its judgments came out worse than random \u2014 and 86% of the model calls turned out to be unnecessary.", + "fold_targets": [ + "Augustus", + "Jev-omni" + ] + }, + { + "id": "netiqus/sema-jev-framework", + "source": "github", + "why_high": "s1_family,decision_scoring", + "url": "https://github.com/netiqus/sema-jev-framework", + "description_snip": "Semantic decision gates for Python: explicit outcomes, versioned policies and per-request accounting with Jev.", + "stars": 0, + "created": "2026-09-20T16:18:49Z", + "recent": true, + "tier": "HIGH", + "description": "Semantic decision gates for Python: explicit outcomes, versioned policies and per-request accounting with Jev.", + "fold_targets": [ + "Augustus", + "Jev-omni" + ] + }, + { + "id": "rubinagentagi-tech/jev-heart-risk-bench", + "source": "github", + "why_high": "s1_family,measurement_bench", + "url": "https://github.com/rubinagentagi-tech/jev-heart-risk-bench", + "description_snip": "Benchmarking Jev (TypeSafe System One) on 5,000 real CDC survey respondents, with an interactive demo where every profile has a real model answer", + "stars": 0, + "created": "2026-09-20T15:54:21Z", + "recent": true, + "tier": "HIGH", + "description": "Benchmarking Jev (TypeSafe System One) on 5,000 real CDC survey respondents, with an interactive demo where every profile has a real model answer", + "fold_targets": [ + "Augustus" + ] + }, + { + "id": "sriannamalai/Jev.UI", + "source": "github", + "why_high": "s1_family", + "url": "https://github.com/sriannamalai/Jev.UI", + "description_snip": "Easy to use User Interface for System One's Jev Model interaction.", + "stars": 0, + "created": "2026-09-20T16:11:46Z", + "recent": true, + "tier": "HIGH", + "description": "Easy to use User Interface for System One's Jev Model interaction.", + "fold_targets": [ + "Augustus", + "Jev-omni" + ] + }, + { + "id": "timsamart/jev-graph-walk", + "source": "github", + "why_high": "s1_family", + "url": "https://github.com/timsamart/jev-graph-walk", + "description_snip": "Context-carrying graph retrieval with Jev: branching walks, reproducible ablations, and a visual replay.", + "stars": 0, + "created": "2026-09-20T15:56:25Z", + "recent": true, + "tier": "HIGH", + "description": "Context-carrying graph retrieval with Jev: branching walks, reproducible ablations, and a visual replay.", + "fold_targets": [ + "Augustus" + ] + }, + { + "id": "tr1v3r/dsh-jev", + "source": "github", + "why_high": "s1_family,guardrail,kit_runtime,decision_scoring", + "url": "https://github.com/tr1v3r/dsh-jev", + "description_snip": "jev \u00d7 DeepSeek Harness: System One decision client, MCP server, per-turn router and effort plugins", + "stars": 0, + "created": "2026-09-20T15:59:47Z", + "recent": true, + "tier": "HIGH", + "description": "jev \u00d7 DeepSeek Harness: System One decision client, MCP server, per-turn router and effort plugins", + "fold_targets": [ + "rh-guard", + "Jev-omni" + ] + }, + { + "id": "rmalde/minecraft-agent", + "source": "github", + "why_high": "s1_family", + "url": "https://github.com/rmalde/minecraft-agent", + "description_snip": "Astra planner and JEV controller for Minecraft, with native recording, tested routes, and run verification.", + "stars": 206, + "created": "2026-09-20T05:05:12Z", + "recent": true, + "tier": "HIGH", + "description": "Astra planner and JEV controller for Minecraft, with native recording, tested routes, and run verification.", + "fold_targets": [ + "Augustus", + "Jev-omni" + ] + }, + { + "id": "wuyoscar/jev-skill", + "source": "github", + "why_high": "s1_family,kit_runtime,measurement_bench", + "url": "https://github.com/wuyoscar/jev-skill", + "description_snip": "An awesome collection of Jev use cases, workflows, and agent skills.", + "stars": 108, + "created": "2026-09-20T02:35:26Z", + "recent": true, + "tier": "HIGH", + "description": "An awesome collection of Jev use cases, workflows, and agent skills.", + "fold_targets": [ + "Augustus", + "Jev-omni" + ] + }, + { + "id": "Yinsongxu/LLM2Jev", + "source": "github", + "why_high": "s1_family,decision_scoring", + "url": "https://github.com/Yinsongxu/LLM2Jev", + "description_snip": "Adapt local language models into Jev-compatible structured decision engines with Choice, Score, and Noul outputs powered by prefill-only binary inference.", + "stars": 64, + "created": "2026-09-19T10:59:33Z", + "recent": true, + "tier": "HIGH", + "description": "Adapt local language models into Jev-compatible structured decision engines with Choice, Score, and Noul outputs powered by prefill-only binary inference.", + "fold_targets": [ + "Augustus", + "Jev-omni" + ] + }, + { + "id": "FerryCorleone/crush-monitor", + "source": "github", + "why_high": "s1_family", + "url": "https://github.com/FerryCorleone/crush-monitor", + "description_snip": "Crush \u597d\u611f\u76d1\u63a7\u5668\uff1a\u7528 Jev \u5206\u6790\u5fae\u4fe1\u804a\u5929\u7684\u60c5\u7eea\u3001\u610f\u56fe\u548c\u56de\u590d\u8868\u73b0\u3002\u672c\u673a\u90e8\u7f72\uff0c\u4f7f\u7528\u81ea\u5df1\u7684 API Key\u3002", + "stars": 28, + "created": "2026-09-20T12:46:10Z", + "recent": true, + "tier": "HIGH", + "description": "Crush \u597d\u611f\u76d1\u63a7\u5668\uff1a\u7528 Jev \u5206\u6790\u5fae\u4fe1\u804a\u5929\u7684\u60c5\u7eea\u3001\u610f\u56fe\u548c\u56de\u590d\u8868\u73b0\u3002\u672c\u673a\u90e8\u7f72\uff0c\u4f7f\u7528\u81ea\u5df1\u7684 API Key\u3002", + "fold_targets": [ + "Augustus", + "Jev-omni" + ] + }, + { + "id": "sabeel111/OpenSourceJev", + "source": "github", + "why_high": "s1_family", + "url": "https://github.com/sabeel111/OpenSourceJev", + "description_snip": "Turning an LLM model into a Jev like System.", + "stars": 10, + "created": "2026-09-19T18:44:34Z", + "recent": true, + "tier": "HIGH", + "description": "Turning an LLM model into a Jev like System.", + "fold_targets": [ + "Augustus", + "Jev-omni" + ] + }, + { + "id": "shitianfang/jev-use", + "source": "github", + "why_high": "s1_family,guardrail,kit_runtime", + "url": "https://github.com/shitianfang/jev-use", + "description_snip": "Claude Code / Codex / pi plugin that hands agent steps needing no text output to Jev (TypeSafe's judgment model) \u2014 measured p50 ~230 ms and ~$0.02 per 1,000 jud", + "stars": 10, + "created": "2026-09-19T10:13:56Z", + "recent": true, + "tier": "HIGH", + "description": "Claude Code / Codex / pi plugin that hands agent steps needing no text output to Jev (TypeSafe's judgment model) \u2014 measured p50 ~230 ms and ~$0.02 per 1,000 judgments, with typed escalation back to the LLM", + "fold_targets": [ + "Jev-omni" + ] + }, + { + "id": "keeltrace/hermes-jev", + "source": "github", + "why_high": "s1_family,guardrail", + "url": "https://github.com/keeltrace/hermes-jev", + "description_snip": "Typed System One decisions, ranking, verification, and an opt-in Hermes tool gate using TypeSafe Jev.", + "stars": 9, + "created": "2026-09-18T02:10:22Z", + "recent": true, + "tier": "HIGH", + "description": "Typed System One decisions, ranking, verification, and an opt-in Hermes tool gate using TypeSafe Jev.", + "fold_targets": [ + "rh-guard", + "Jev-omni" + ] + }, + { + "id": "lykycy123/RoboJEV", + "source": "github", + "why_high": "s1_family", + "url": "https://github.com/lykycy123/RoboJEV", + "description_snip": "Two-stage JEV control of a Franka Panda in MuJoCo: pick/place, surface pushing, and fixed-pedestal stacking.", + "stars": 6, + "created": "2026-09-20T09:36:44Z", + "recent": true, + "tier": "HIGH", + "description": "Two-stage JEV control of a Franka Panda in MuJoCo: pick/place, surface pushing, and fixed-pedestal stacking.", + "fold_targets": [ + "Augustus", + "Jev-omni" + ] + }, + { + "id": "hoangnb24/paseo-supervision", + "source": "github", + "why_high": "s1_family,kit_runtime", + "url": "https://github.com/hoangnb24/paseo-supervision", + "description_snip": "Paseo plugin for supervising Lead\u2013Peer communication protocol drift with Jev", + "stars": 4, + "created": "2026-09-20T01:49:27Z", + "recent": true, + "tier": "HIGH", + "description": "Paseo plugin for supervising Lead\u2013Peer communication protocol drift with Jev", + "fold_targets": [ + "Jev-omni" + ] + }, + { + "id": "wh000wh000/awesome-jev-live", + "source": "github", + "why_high": "s1_family", + "url": "https://github.com/wh000wh000/awesome-jev-live", + "description_snip": "Evidence-graded index of the Jev / TypeSafe System One ecosystem. Rebuilt every 2 hours in 20 languages.", + "stars": 4, + "created": "2026-09-18T13:37:03Z", + "recent": true, + "tier": "HIGH", + "description": "Evidence-graded index of the Jev / TypeSafe System One ecosystem. Rebuilt every 2 hours in 20 languages.", + "fold_targets": [ + "Augustus", + "Jev-omni" + ] + }, + { + "id": "ChosenXu/newsletter-link-harvester", + "source": "github", + "why_high": "s1_family,kit_runtime", + "url": "https://github.com/ChosenXu/newsletter-link-harvester", + "description_snip": "Agent Skill: harvest links from newsletter emails into Raindrop.io with the author editorial context attached - Gmail read-only, three-layer dedup, zero-context", + "stars": 3, + "created": "2026-09-18T05:40:35Z", + "recent": true, + "tier": "HIGH", + "description": "Agent Skill: harvest links from newsletter emails into Raindrop.io with the author editorial context attached - Gmail read-only, three-layer dedup, zero-context library compare", + "fold_targets": [ + "Jev-omni" + ] + }, + { + "id": "hraness/algal", + "source": "github", + "why_high": "s1_family,kit_runtime", + "url": "https://github.com/hraness/algal", + "description_snip": "ALGAL is a new take on the agent graph: a language and runtime for agentic program evolution.", + "stars": 3, + "created": "2026-09-18T19:19:06Z", + "recent": true, + "tier": "HIGH", + "description": "ALGAL is a new take on the agent graph: a language and runtime for agentic program evolution.", + "fold_targets": [ + "Jev-omni" + ] + }, + { + "id": "imaddde867/jev-position-test", + "source": "github", + "why_high": "s1_family,measurement_bench", + "url": "https://github.com/imaddde867/jev-position-test", + "description_snip": "Reorder the enum options: a Jev clone changes its answer, Jev doesn't. n=6, raw data included.", + "stars": 3, + "created": "2026-09-20T15:04:23Z", + "recent": true, + "tier": "HIGH", + "description": "Reorder the enum options: a Jev clone changes its answer, Jev doesn't. n=6, raw data included.", + "fold_targets": [ + "Augustus", + "Jev-omni" + ] + }, + { + "id": "mizchi/jev-lexer", + "source": "github", + "why_high": "s1_family", + "url": "https://github.com/mizchi/jev-lexer", + "description_snip": "Language-agnostic syntax highlighter: split like gpu-lexer, classify every part with Jev, render Shiki-compatible tokens, HTML and ANSI", + "stars": 3, + "created": "2026-09-20T10:25:42Z", + "recent": true, + "tier": "HIGH", + "description": "Language-agnostic syntax highlighter: split like gpu-lexer, classify every part with Jev, render Shiki-compatible tokens, HTML and ANSI", + "fold_targets": [ + "Augustus", + "Jev-omni" + ] + }, + { + "id": "455-dIAO/windows-save-token-jev-setup", + "source": "github", + "why_high": "s1_family,guardrail,kit_runtime", + "url": "https://github.com/455-dIAO/windows-save-token-jev-setup", + "description_snip": "Windows Codex Skill\uff1a\u901a\u8fc7 npx \u6216 Git \u5b89\u88c5\uff0c\u5b89\u5168\u914d\u7f6e save-token-jev \u7684 PreCompact/SessionStart Hooks\uff0c\u5e76\u63d0\u4f9b\u4fe1\u4efb\u3001\u539f\u751f\u538b\u7f29\u4e0e\u65e7\u5185\u5bb9\u9694\u79bb\u9a8c\u8bc1\u3002", + "stars": 2, + "created": "2026-09-20T07:47:44Z", + "recent": true, + "tier": "HIGH", + "description": "Windows Codex Skill\uff1a\u901a\u8fc7 npx \u6216 Git \u5b89\u88c5\uff0c\u5b89\u5168\u914d\u7f6e save-token-jev \u7684 PreCompact/SessionStart Hooks\uff0c\u5e76\u63d0\u4f9b\u4fe1\u4efb\u3001\u539f\u751f\u538b\u7f29\u4e0e\u65e7\u5185\u5bb9\u9694\u79bb\u9a8c\u8bc1\u3002", + "fold_targets": [ + "Jev-omni" + ] + }, + { + "id": "Bald0Wang/jev-docs-zh", + "source": "github", + "why_high": "s1_family", + "url": "https://github.com/Bald0Wang/jev-docs-zh", + "description_snip": "Jev \u6a21\u578b\uff08TypeSafe AI\uff09\u5b98\u65b9\u4f7f\u7528\u6587\u6863\u7684\u4e2d\u6587\u7ffb\u8bd1 | Unofficial Chinese translation of the official Jev (TypeSafe AI) docs \u2014 https://docs.typesafe.ai", + "stars": 2, + "created": "2026-09-20T06:39:31Z", + "recent": true, + "tier": "HIGH", + "description": "Jev \u6a21\u578b\uff08TypeSafe AI\uff09\u5b98\u65b9\u4f7f\u7528\u6587\u6863\u7684\u4e2d\u6587\u7ffb\u8bd1 | Unofficial Chinese translation of the official Jev (TypeSafe AI) docs \u2014 https://docs.typesafe.ai", + "fold_targets": [ + "Augustus", + "Jev-omni" + ] + }, + { + "id": "ChristianAlexander/effect-jev-cwe", + "source": "github", + "why_high": "s1_family", + "url": "https://github.com/ChristianAlexander/effect-jev-cwe", + "description_snip": "A demonstration of the Jev System 1 model in Effect, matching vulnerabilities to their underlying CWEs", + "stars": 2, + "created": "2026-09-20T04:03:06Z", + "recent": true, + "tier": "HIGH", + "description": "A demonstration of the Jev System 1 model in Effect, matching vulnerabilities to their underlying CWEs", + "fold_targets": [ + "Augustus", + "Jev-omni" + ] + }, + { + "id": "CompleteTech-LLC-AI-Research/jev-311-heatmap", + "source": "github", + "why_high": "s1_family", + "url": "https://github.com/CompleteTech-LLC-AI-Research/jev-311-heatmap", + "description_snip": "NYC 311 complaint heatmaps with TypeSafe JEV: reproducible pipeline, live research results, and interactive geographic visualizations.", + "stars": 2, + "created": "2026-09-20T00:20:04Z", + "recent": true, + "tier": "HIGH", + "description": "NYC 311 complaint heatmaps with TypeSafe JEV: reproducible pipeline, live research results, and interactive geographic visualizations.", + "fold_targets": [ + "Augustus", + "Jev-omni" + ] + }, + { + "id": "aovestdipaperino/laya-pong", + "source": "github", + "why_high": "s1_family,kit_runtime,decision_scoring", + "url": "https://github.com/aovestdipaperino/laya-pong", + "description_snip": "Browser pong whose paddle is decided by a Laya typed decision, one question per frame at 18.7 ms", + "stars": 2, + "created": "2026-09-20T13:48:34Z", + "recent": true, + "tier": "HIGH", + "description": "Browser pong whose paddle is decided by a Laya typed decision, one question per frame at 18.7 ms", + "fold_targets": [ + "Jev-omni" + ] + }, + { + "id": "hoangngochuong24947-gif/jev-patent-disclosure", + "source": "github", + "why_high": "s1_family,kit_runtime,decision_scoring", + "url": "https://github.com/hoangngochuong24947-gif/jev-patent-disclosure", + "description_snip": "Patent disclosure and application drafting skill powered by TypeSafe Jev / Jeb System-1", + "stars": 2, + "created": "2026-09-20T06:55:05Z", + "recent": true, + "tier": "HIGH", + "description": "Patent disclosure and application drafting skill powered by TypeSafe Jev / Jeb System-1", + "fold_targets": [ + "Jev-omni" + ] + }, + { + "id": "minorun365/jev-cloud-quiz", + "source": "github", + "why_high": "s1_family,kit_runtime", + "url": "https://github.com/minorun365/jev-cloud-quiz", + "description_snip": "\u4e09\u5927\u30af\u30e9\u30a6\u30c9\u306e\u6a5f\u80fd\u540d\u3092\u3001TypeSafe AI \u306e System One \u30e2\u30c7\u30eb Jev \u304c\u78ba\u7387\u3064\u304d\u3067\u5224\u5b9a\u3059\u308b\u30c7\u30e2", + "stars": 2, + "created": "2026-09-20T03:54:01Z", + "recent": true, + "tier": "HIGH", + "description": "\u4e09\u5927\u30af\u30e9\u30a6\u30c9\u306e\u6a5f\u80fd\u540d\u3092\u3001TypeSafe AI \u306e System One \u30e2\u30c7\u30eb Jev \u304c\u78ba\u7387\u3064\u304d\u3067\u5224\u5b9a\u3059\u308b\u30c7\u30e2", + "fold_targets": [ + "Jev-omni" + ] + }, + { + "id": "nabendu82/jev-reflex", + "source": "github", + "why_high": "s1_family", + "url": "https://github.com/nabendu82/jev-reflex", + "description_snip": "Multimodal Mac controller using Hand gestures and voice", + "stars": 2, + "created": "2026-09-20T07:59:15Z", + "recent": true, + "tier": "HIGH", + "description": "Multimodal Mac controller using Hand gestures and voice", + "fold_targets": [ + "Augustus", + "Jev-omni" + ] + }, + { + "id": "win4r/pi-jev-router", + "source": "github", + "why_high": "s1_family,guardrail", + "url": "https://github.com/win4r/pi-jev-router", + "description_snip": "Task-boundary model routing for Pi Coding Agent, powered by TypeSafe Jev. Conservative policies, exact caching, and observable failover.", + "stars": 2, + "created": "2026-09-20T13:51:07Z", + "recent": true, + "tier": "HIGH", + "description": "Task-boundary model routing for Pi Coding Agent, powered by TypeSafe Jev. Conservative policies, exact caching, and observable failover.", + "fold_targets": [ + "rh-guard" + ] + }, + { + "id": "cheeaun/jevmoji", + "source": "github", + "why_high": "s1_family", + "url": "https://github.com/cheeaun/jevmoji", + "description_snip": "Type anything. Get related emojis scored 0\u20133 with Jev.", + "stars": 1, + "created": "2026-09-20T10:49:23Z", + "recent": true, + "tier": "HIGH", + "description": "Type anything. Get related emojis scored 0\u20133 with Jev.", + "fold_targets": [ + "Augustus", + "Jev-omni" + ] + }, + { + "id": "glamboyosa/docket", + "source": "github", + "why_high": "s1_family", + "url": "https://github.com/glamboyosa/docket", + "description_snip": "A Go TUI that uses Jev to classify documents, assess sensitivity and urgency, and determine whether action is required.", + "stars": 1, + "created": "2026-09-19T15:52:34Z", + "recent": true, + "tier": "HIGH", + "description": "A Go TUI that uses Jev to classify documents, assess sensitivity and urgency, and determine whether action is required.", + "fold_targets": [ + "Augustus", + "Jev-omni" + ] + }, + { + "id": "hfdataset:roskosmos19/SystemOne", + "source": "hf_dataset", + "why_high": "s1_family,open_weights_or_port", + "url": "https://huggingface.co/datasets/roskosmos19/SystemOne", + "description_snip": "region:us", + "stars": 1, + "created": "2026-09-19T07:20:09.000Z", + "recent": true, + "tier": "HIGH", + "description": "region:us", + "fold_targets": [ + "Jev-omni" + ] + }, + { + "id": "Ashadeepa/typesafe-showcase", + "source": "github", + "why_high": "s1_family,decision_scoring", + "url": "https://github.com/Ashadeepa/typesafe-showcase", + "description_snip": "Next.js UI showing off TypeSafe's System One model (Jev) \u2014 parallel Noul judgments and a Choice-based citation checker, deployable to Vercel", + "stars": 0, + "created": "2026-09-19T05:04:11Z", + "recent": true, + "tier": "HIGH", + "description": "Next.js UI showing off TypeSafe's System One model (Jev) \u2014 parallel Noul judgments and a Choice-based citation checker, deployable to Vercel", + "fold_targets": [ + "Augustus", + "Jev-omni" + ] + }, + { + "id": "Astro-Han/decision-head-rlcd", + "source": "github", + "why_high": "s1_family,open_weights_or_port,decision_scoring", + "url": "https://github.com/Astro-Han/decision-head-rlcd", + "description_snip": "Where does a decision model's generalisation come from? RLCD on Qwen3.5-4B, held-out sets grouped by training-data coverage, JevBench and three external suites.", + "stars": 0, + "created": "2026-09-20T15:18:34Z", + "recent": true, + "tier": "HIGH", + "description": "Where does a decision model's generalisation come from? RLCD on Qwen3.5-4B, held-out sets grouped by training-data coverage, JevBench and three external suites.", + "fold_targets": [ + "Augustus", + "Jev-omni" + ] + }, + { + "id": "Bigthap/canvas-quiz-ai-solver", + "source": "github", + "why_high": "s1_family,decision_scoring", + "url": "https://github.com/Bigthap/canvas-quiz-ai-solver", + "description_snip": "High-performance Canvas LMS quiz assistant & scraper designed for TypeSafe AI System One decision primitives", + "stars": 0, + "created": "2026-09-19T10:28:06Z", + "recent": true, + "tier": "HIGH", + "description": "High-performance Canvas LMS quiz assistant & scraper designed for TypeSafe AI System One decision primitives", + "fold_targets": [ + "Augustus", + "Jev-omni" + ] + }, + { + "id": "CoderInPajamas/JEV-MLX", + "source": "github", + "why_high": "s1_family,guardrail,open_weights_or_port", + "url": "https://github.com/CoderInPajamas/JEV-MLX", + "description_snip": "JEV-inspired local decisions for Apple Silicon, powered by MLX.", + "stars": 0, + "created": "2026-09-20T13:09:46Z", + "recent": true, + "tier": "HIGH", + "description": "JEV-inspired local decisions for Apple Silicon, powered by MLX.", + "fold_targets": [ + "Jev-omni", + "Augustus" + ] + }, + { + "id": "JJRPF/antigravity-auto-mode", + "source": "github", + "why_high": "s1_family", + "url": "https://github.com/JJRPF/antigravity-auto-mode", + "description_snip": "Claude Code Auto Mode emulation for Google Antigravity (AGY) powered by TypeSafe AI System One", + "stars": 0, + "created": "2026-09-20T14:00:20Z", + "recent": true, + "tier": "HIGH", + "description": "Claude Code Auto Mode emulation for Google Antigravity (AGY) powered by TypeSafe AI System One", + "fold_targets": [ + "Augustus", + "Jev-omni" + ] + }, + { + "id": "LingyeNBird/codesafe", + "source": "github", + "why_high": "s1_family", + "url": "https://github.com/LingyeNBird/codesafe", + "description_snip": "Fast per-file safety/bug triage via TypeSafe System One API", + "stars": 0, + "created": "2026-09-20T14:49:46Z", + "recent": true, + "tier": "HIGH", + "description": "Fast per-file safety/bug triage via TypeSafe System One API", + "fold_targets": [ + "Augustus", + "Jev-omni" + ] + }, + { + "id": "Mrmimee/hermes-plugin-jev", + "source": "github", + "why_high": "s1_family,kit_runtime,decision_scoring", + "url": "https://github.com/Mrmimee/hermes-plugin-jev", + "description_snip": "Jev (TypeSafe AI) System One decision engine plugin for Hermes Agent, backed by Agnes AI Flash.", + "stars": 0, + "created": "2026-09-19T13:42:27Z", + "recent": true, + "tier": "HIGH", + "description": "Jev (TypeSafe AI) System One decision engine plugin for Hermes Agent, backed by Agnes AI Flash.", + "fold_targets": [ + "Jev-omni" + ] + }, + { + "id": "StephenChan-1/Jev-usecases", + "source": "github", + "why_high": "s1_family,measurement_bench", + "url": "https://github.com/StephenChan-1/Jev-usecases", + "description_snip": "open source hub for Jev use cases", + "stars": 0, + "created": "2026-09-20T12:38:22Z", + "recent": true, + "tier": "HIGH", + "description": "open source hub for Jev use cases", + "fold_targets": [ + "Augustus" + ] + }, + { + "id": "Sunwood-ai-labs/jev-colab-lab", + "source": "github", + "why_high": "s1_family,decision_scoring", + "url": "https://github.com/Sunwood-ai-labs/jev-colab-lab", + "description_snip": "Reproducible Google Colab GPU experiments for decision-model inference", + "stars": 0, + "created": "2026-09-20T14:33:33Z", + "recent": true, + "tier": "HIGH", + "description": "Reproducible Google Colab GPU experiments for decision-model inference", + "fold_targets": [ + "Augustus", + "Jev-omni" + ] + }, + { + "id": "TrustifAI/typed_evals", + "source": "github", + "why_high": "s1_family,measurement_bench", + "url": "https://github.com/TrustifAI/typed_evals", + "description_snip": "Fast, typed, calibrated evaluations for LLM and agent outputs, powered by Jev \u2014 with simple, framework-agnostic Python APIs", + "stars": 0, + "created": "2026-09-20T07:58:36Z", + "recent": true, + "tier": "HIGH", + "description": "Fast, typed, calibrated evaluations for LLM and agent outputs, powered by Jev \u2014 with simple, framework-agnostic Python APIs", + "fold_targets": [ + "Augustus" + ] + }, + { + "id": "WiredMind2/jev", + "source": "github", + "why_high": "s1_family,measurement_bench,open_weights_or_port,decision_scoring", + "url": "https://github.com/WiredMind2/jev", + "description_snip": "Independent research notes toward an open Jev-like decision model: public facts, API contract, training and eval plan.", + "stars": 0, + "created": "2026-09-20T00:59:36Z", + "recent": true, + "tier": "HIGH", + "description": "Independent research notes toward an open Jev-like decision model: public facts, API contract, training and eval plan.", + "fold_targets": [ + "Augustus", + "Jev-omni" + ] + }, + { + "id": "az9713/jev-projects", + "source": "github", + "why_high": "s1_family,kit_runtime", + "url": "https://github.com/az9713/jev-projects", + "description_snip": "Small demos of Jev (TypeSafe) through the Vercel AI Gateway: wiki race, town of agents, bullet chess, and more", + "stars": 0, + "created": "2026-09-19T16:50:01Z", + "recent": true, + "tier": "HIGH", + "description": "Small demos of Jev (TypeSafe) through the Vercel AI Gateway: wiki race, town of agents, bullet chess, and more", + "fold_targets": [ + "Jev-omni" + ] + }, + { + "id": "breymander/grocery-finder", + "source": "github", + "why_high": "s1_family", + "url": "https://github.com/breymander/grocery-finder", + "description_snip": "Ingredient-matching PoC: Jev ranks catalog items across stores", + "stars": 0, + "created": "2026-09-19T20:22:47Z", + "recent": true, + "tier": "HIGH", + "description": "Ingredient-matching PoC: Jev ranks catalog items across stores", + "fold_targets": [ + "Augustus", + "Jev-omni" + ] + }, + { + "id": "chalkychalk42/jev", + "source": "github", + "why_high": "s1_family", + "url": "https://github.com/chalkychalk42/jev", + "description_snip": "A guide-directed leveling agent for a private TBC 2.4.3 server, built so it gets cheaper to run the longer it runs", + "stars": 0, + "created": "2026-09-20T12:47:49Z", + "recent": true, + "tier": "HIGH", + "description": "A guide-directed leveling agent for a private TBC 2.4.3 server, built so it gets cheaper to run the longer it runs", + "fold_targets": [ + "Augustus", + "Jev-omni" + ] + }, + { + "id": "ddlaws0n/jevportfolio", + "source": "github", + "why_high": "s1_family,guardrail", + "url": "https://github.com/ddlaws0n/jevportfolio", + "description_snip": "1,000 synthetic SaaS accounts, 6,000 constrained judgments from TypeSafe's Jev, and ordinary TypeScript deciding who needs a human today. TanStack Start on Bun.", + "stars": 0, + "created": "2026-09-20T12:50:21Z", + "recent": true, + "tier": "HIGH", + "description": "1,000 synthetic SaaS accounts, 6,000 constrained judgments from TypeSafe's Jev, and ordinary TypeScript deciding who needs a human today. TanStack Start on Bun.", + "fold_targets": [ + "Augustus", + "Jev-omni" + ] + }, + { + "id": "eclecticv/jev-adcp-decision-economics", + "source": "github", + "why_high": "s1_family,decision_scoring", + "url": "https://github.com/eclecticv/jev-adcp-decision-economics", + "description_snip": "Jev System One vs LLM cost/latency estimates for AdCP buyer and seller agent decisions (verified 2026-09-20)", + "stars": 0, + "created": "2026-09-20T15:29:46Z", + "recent": true, + "tier": "HIGH", + "description": "Jev System One vs LLM cost/latency estimates for AdCP buyer and seller agent decisions (verified 2026-09-20)", + "fold_targets": [ + "Augustus", + "Jev-omni" + ] + }, + { + "id": "ghubnab99/jev-enterprise-decision-fabric", + "source": "github", + "why_high": "s1_family,guardrail,measurement_bench,decision_scoring", + "url": "https://github.com/ghubnab99/jev-enterprise-decision-fabric", + "description_snip": "Architecture for running many semantic decisions through one validated path, with a labelled 111-case benchmark comparing TypeSafe Jev against a Claude baseline", + "stars": 0, + "created": "2026-09-19T11:04:08Z", + "recent": true, + "tier": "HIGH", + "description": "Architecture for running many semantic decisions through one validated path, with a labelled 111-case benchmark comparing TypeSafe Jev against a Claude baseline, and a dashboard for inspecting any single decision. Experimental, not production.", + "fold_targets": [ + "Augustus" + ] + }, + { + "id": "hama-jp/laya-tetris-finetuning", + "source": "github", + "why_high": "s1_family", + "url": "https://github.com/hama-jp/laya-tetris-finetuning", + "description_snip": "Fine-tuning 421M Laya for real-time Tetris decisions: code, reproduction guide and experiment results", + "stars": 0, + "created": "2026-09-20T14:10:43Z", + "recent": true, + "tier": "HIGH", + "description": "Fine-tuning 421M Laya for real-time Tetris decisions: code, reproduction guide and experiment results", + "fold_targets": [ + "Augustus", + "Jev-omni" + ] + }, + { + "id": "hfspace:AlexWortega/openjev", + "source": "hf_space", + "why_high": "s1_family,open_weights_or_port,kit_runtime", + "url": "https://huggingface.co/spaces/AlexWortega/openjev", + "description_snip": "gradio region:us", + "stars": 0, + "created": "2026-09-19T20:41:59.000Z", + "recent": true, + "tier": "HIGH", + "description": "gradio region:us", + "fold_targets": [ + "Jev-omni" + ] + }, + { + "id": "hraness/sys1", + "source": "github", + "why_high": "s1_family,guardrail,kit_runtime", + "url": "https://github.com/hraness/sys1", + "description_snip": "Typed decisions for agents: a local model router, Jev-compatible gateway, and Node/Bun client.", + "stars": 0, + "created": "2026-09-19T02:32:00Z", + "recent": true, + "tier": "HIGH", + "description": "Typed decisions for agents: a local model router, Jev-compatible gateway, and Node/Bun client.", + "fold_targets": [ + "rh-guard", + "Jev-omni" + ] + }, + { + "id": "hraness/system-one-skills", + "source": "github", + "why_high": "s1_family,kit_runtime", + "url": "https://github.com/hraness/system-one-skills", + "description_snip": "System One skills for Devin, Claude Code and Codex. Cut noisy validation-log tokens with one deterministic skill, private full logs, and measured evidence.", + "stars": 0, + "created": "2026-09-19T06:04:33Z", + "recent": true, + "tier": "HIGH", + "description": "System One skills for Devin, Claude Code and Codex. Cut noisy validation-log tokens with one deterministic skill, private full logs, and measured evidence.", + "fold_targets": [ + "Jev-omni" + ] + }, + { + "id": "ignitewala/system-one", + "source": "github", + "why_high": "s1_family,measurement_bench", + "url": "https://github.com/ignitewala/system-one", + "description_snip": "System-one model evaluation and example", + "stars": 0, + "created": "2026-09-20T14:52:21Z", + "recent": true, + "tier": "HIGH", + "description": "System-one model evaluation and example", + "fold_targets": [ + "Augustus" + ] + }, + { + "id": "lkarlslund/laya.cpp", + "source": "github", + "why_high": "s1_family", + "url": "https://github.com/lkarlslund/laya.cpp", + "description_snip": "RTX-optimized C++ inference for Laya typed decisions", + "stars": 0, + "created": "2026-09-20T09:18:04Z", + "recent": true, + "tier": "HIGH", + "description": "RTX-optimized C++ inference for Laya typed decisions", + "fold_targets": [ + "Augustus", + "Jev-omni" + ] + }, + { + "id": "matchstick-trading/jev-regime-gate", + "source": "github", + "why_high": "s1_family,guardrail", + "url": "https://github.com/matchstick-trading/jev-regime-gate", + "description_snip": "Jev-powered regime gate for trading strategy backtests. Research experiment, not investment advice.", + "stars": 0, + "created": "2026-09-20T00:40:51Z", + "recent": true, + "tier": "HIGH", + "description": "Jev-powered regime gate for trading strategy backtests. Research experiment, not investment advice.", + "fold_targets": [ + "rh-guard" + ] + }, + { + "id": "minhlucvan/dsh-plugin-system-one", + "source": "github", + "why_high": "s1_family,kit_runtime,measurement_bench", + "url": "https://github.com/minhlucvan/dsh-plugin-system-one", + "description_snip": "TypeSafe Jev (System One) as DeepSeek Harness agent tools, with the token accounting and benchmark to prove what they cost", + "stars": 0, + "created": "2026-09-20T10:17:06Z", + "recent": true, + "tier": "HIGH", + "description": "TypeSafe Jev (System One) as DeepSeek Harness agent tools, with the token accounting and benchmark to prove what they cost", + "fold_targets": [ + "Augustus", + "Jev-omni" + ] + }, + { + "id": "muratmirgun/owncode", + "source": "github", + "why_high": "s1_family,kit_runtime", + "url": "https://github.com/muratmirgun/owncode", + "description_snip": "An experimental terminal coding agent with configurable models, Witch orchestration, and multiple context compaction methods.", + "stars": 0, + "created": "2026-09-18T22:26:06Z", + "recent": true, + "tier": "HIGH", + "description": "An experimental terminal coding agent with configurable models, Witch orchestration, and multiple context compaction methods.", + "fold_targets": [ + "Jev-omni" + ] + }, + { + "id": "n0nuser/battlesnake-jev", + "source": "github", + "why_high": "s1_family", + "url": "https://github.com/n0nuser/battlesnake-jev", + "description_snip": "A Battlesnake in Go where deterministic code owns tactics and a TypeSafe Jev classifier gets the judgment calls \u2014 with controls measuring whether that actually ", + "stars": 0, + "created": "2026-09-19T22:58:49Z", + "recent": true, + "tier": "HIGH", + "description": "A Battlesnake in Go where deterministic code owns tactics and a TypeSafe Jev classifier gets the judgment calls \u2014 with controls measuring whether that actually helped.", + "fold_targets": [ + "Augustus", + "Jev-omni" + ] + }, + { + "id": "noelserdna/cartas-ciberseguridad-jev", + "source": "github", + "why_high": "s1_family", + "url": "https://github.com/noelserdna/cartas-ciberseguridad-jev", + "description_snip": "Puntuar curr\u00edculums de ciberseguridad con JEV, el modelo System One de TypeSafe, sobre Cloudflare Workers. Clasificaci\u00f3n contra NICE, ECSF y SFIA, con cada n\u00fame", + "stars": 0, + "created": "2026-09-20T14:06:32Z", + "recent": true, + "tier": "HIGH", + "description": "Puntuar curr\u00edculums de ciberseguridad con JEV, el modelo System One de TypeSafe, sobre Cloudflare Workers. Clasificaci\u00f3n contra NICE, ECSF y SFIA, con cada n\u00famero trazable hasta la pregunta que lo produjo.", + "fold_targets": [ + "Augustus", + "Jev-omni" + ] + }, + { + "id": "nyattoh/model-effort-router", + "source": "github", + "why_high": "s1_family,guardrail,decision_scoring", + "url": "https://github.com/nyattoh/model-effort-router", + "description_snip": "Provider-neutral task decomposition and model-effort routing with optional Jev decision signals", + "stars": 0, + "created": "2026-09-20T14:25:12Z", + "recent": true, + "tier": "HIGH", + "description": "Provider-neutral task decomposition and model-effort routing with optional Jev decision signals", + "fold_targets": [ + "rh-guard", + "Augustus" + ] + }, + { + "id": "osuki-dev/opencode-osuki-agent", + "source": "github", + "why_high": "s1_family,guardrail", + "url": "https://github.com/osuki-dev/opencode-osuki-agent", + "description_snip": "Effect-native OpenCode coordinator with Jev routing and persistent goals", + "stars": 0, + "created": "2026-09-20T05:21:02Z", + "recent": true, + "tier": "HIGH", + "description": "Effect-native OpenCode coordinator with Jev routing and persistent goals", + "fold_targets": [ + "Augustus", + "Jev-omni" + ] + }, + { + "id": "royosherove/graphlin", + "source": "github", + "why_high": "s1_family", + "url": "https://github.com/royosherove/graphlin", + "description_snip": "Live architecture and activity diagrams for coding agents using JEV.", + "stars": 0, + "created": "2026-09-20T12:24:22Z", + "recent": true, + "tier": "HIGH", + "description": "Live architecture and activity diagrams for coding agents using JEV.", + "fold_targets": [ + "Augustus", + "Jev-omni" + ] + }, + { + "id": "saifullahshafin/system1-third-person-audit", + "source": "github", + "why_high": "s1_family,measurement_bench,decision_scoring", + "url": "https://github.com/saifullahshafin/system1-third-person-audit", + "description_snip": "System One Deterministic Decision Layer & Third-Person Audit (TPA) for Autonomous AI Agents. Arrests human-agent echo chambers and false completion at mid-fligh", + "stars": 0, + "created": "2026-09-20T14:54:32Z", + "recent": true, + "tier": "HIGH", + "description": "System One Deterministic Decision Layer & Third-Person Audit (TPA) for Autonomous AI Agents. Arrests human-agent echo chambers and false completion at mid-flight (40% & 60% watermarks).", + "fold_targets": [ + "Augustus" + ] + }, + { + "id": "suidouble/let-jev-speak", + "source": "github", + "why_high": "s1_family", + "url": "https://github.com/suidouble/let-jev-speak", + "description_snip": "Experiment to trick Typesafe\u2019s Jev, aka \u201cthe language model that won\u2019t talk\u201d into actually talking.", + "stars": 0, + "created": "2026-09-20T02:51:44Z", + "recent": true, + "tier": "HIGH", + "description": "Experiment to trick Typesafe\u2019s Jev, aka \u201cthe language model that won\u2019t talk\u201d into actually talking.", + "fold_targets": [ + "Augustus", + "Jev-omni" + ] + }, + { + "id": "swarooppatilx/oxox", + "source": "github", + "why_high": "s1_family", + "url": "https://github.com/swarooppatilx/oxox", + "description_snip": "Probably the least useful thing you can build with Jev", + "stars": 0, + "created": "2026-09-19T21:52:53Z", + "recent": true, + "tier": "HIGH", + "description": "Probably the least useful thing you can build with Jev", + "fold_targets": [ + "Augustus", + "Jev-omni" + ] + }, + { + "id": "wquguru/dasheng", + "source": "github", + "why_high": "s1_family", + "url": "https://github.com/wquguru/dasheng", + "description_snip": "\u5927\u58f0\u8bfb \u2014 R2T2 \u6d41\u5f0f ASR \u542c\uff0cJev \u9010\u8bcd\u5224\uff0c\u82f1\u6587\u6717\u8bfb\u8bc4\u5206", + "stars": 0, + "created": "2026-09-20T15:33:25Z", + "recent": true, + "tier": "HIGH", + "description": "\u5927\u58f0\u8bfb \u2014 R2T2 \u6d41\u5f0f ASR \u542c\uff0cJev \u9010\u8bcd\u5224\uff0c\u82f1\u6587\u6717\u8bfb\u8bc4\u5206", + "fold_targets": [ + "Augustus", + "Jev-omni" + ] + }, + { + "id": "xpressabhi/jev-browser", + "source": "github", + "why_high": "s1_family,kit_runtime", + "url": "https://github.com/xpressabhi/jev-browser", + "description_snip": "Jev decides. The harness acts.", + "stars": 0, + "created": "2026-09-18T07:37:02Z", + "recent": true, + "tier": "HIGH", + "description": "Jev decides. The harness acts.", + "fold_targets": [ + "Jev-omni" + ] + }, + { + "id": "xuboboo/ashare-trader", + "source": "github", + "why_high": "s1_family", + "url": "https://github.com/xuboboo/ashare-trader", + "description_snip": "\u9996\u4e2a\u57fa\u4e8eJev\u6a21\u578b\u7684A \u80a1 T+1 \u5f00\u6e90\u91cf\u5316\u51b3\u7b56\u53f0\uff1a\u5c3e\u76d8\u9009\u80a1 + \u4eba\u5de5\u6267\u884c\u56de\u586b + \u5f71\u5b50\u8bb0\u8d26 + \u65e5\u7ebf\u56de\u6d4b\u3002\u5f53\u524d\u7b56\u7565\u5728\u771f\u5b9e\u6210\u672c\u4e0b\u6ca1\u6709\u6b63\u671f\u671b\uff0c\u5df2\u505c\u5728\u56de\u6d4b\u5c42\uff08README \u6709\u5168\u90e8\u6570\u636e\uff09\u3002", + "stars": 0, + "created": "2026-09-20T15:06:31Z", + "recent": true, + "tier": "HIGH", + "description": "\u9996\u4e2a\u57fa\u4e8eJev\u6a21\u578b\u7684A \u80a1 T+1 \u5f00\u6e90\u91cf\u5316\u51b3\u7b56\u53f0\uff1a\u5c3e\u76d8\u9009\u80a1 + \u4eba\u5de5\u6267\u884c\u56de\u586b + \u5f71\u5b50\u8bb0\u8d26 + \u65e5\u7ebf\u56de\u6d4b\u3002\u5f53\u524d\u7b56\u7565\u5728\u771f\u5b9e\u6210\u672c\u4e0b\u6ca1\u6709\u6b63\u671f\u671b\uff0c\u5df2\u505c\u5728\u56de\u6d4b\u5c42\uff08README \u6709\u5168\u90e8\u6570\u636e\uff09\u3002", + "fold_targets": [ + "Augustus", + "Jev-omni" + ] + }, + { + "id": "zebedelu/sudoku-vs-jev", + "source": "github", + "why_high": "s1_family,decision_scoring", + "url": "https://github.com/zebedelu/sudoku-vs-jev", + "description_snip": "A terminal Sudoku game where TypeSafe's Jev model plays the game, built to probe its decision-making move by move", + "stars": 0, + "created": "2026-09-19T19:13:18Z", + "recent": true, + "tier": "HIGH", + "description": "A terminal Sudoku game where TypeSafe's Jev model plays the game, built to probe its decision-making move by move", + "fold_targets": [ + "Augustus", + "Jev-omni" + ] + } +] \ No newline at end of file diff --git a/research/archive/hourly/2026-09-20T16/run_digest.json b/research/archive/hourly/2026-09-20T16/run_digest.json new file mode 100644 index 00000000..9aed909e --- /dev/null +++ b/research/archive/hourly/2026-09-20T16/run_digest.json @@ -0,0 +1,11 @@ +{ + "boise_label": "1049", + "novel_high_count": 98, + "archer_status": "promised_not_landed", + "notify_recommended": true, + "folds": { + "Augustus": "bc-bc005425-b0a6-5047-8168-5cbc20030187", + "Jev-omni": "bc-7d34ba61-5b34-5bd6-90a2-7edbc8a62e94", + "rh-guard": "bc-23b2d4e0-708e-59fc-adbe-7f9beee209ba" + } +} \ No newline at end of file diff --git a/research/changelog-hourly.md b/research/changelog-hourly.md index 3eb42602..494f168d 100644 --- a/research/changelog-hourly.md +++ b/research/changelog-hourly.md @@ -13,6 +13,19 @@ This is the uniqueness-lock archive after hourly folds (#2–#40 / notes --- +## Hourly 1049 HIGH (notes.md §120 / items 353–368 / batch #103) + +- Fresh PR off latest `main` after merged #41 (0947 / §119) and + merged #42 (v0.5.0). **HARD RULE:** do not reopen or amend + PR #23–#42. Skip Archer. Quote *theirs*. No wrappers. + `invented_signal: false`. +- ggmlc GGUF is not llama.cpp. serving substrate ≠ calibrated replica. + Qwen3.5-9B ≠ Archer. planner writes JEV selects. + pick_by_id vs pick_second. +- Hourly 1049 uniqueness lock: ggmlc GGUF is not llama.cpp; Loading them in llama.cpp will fail; one encoder pass; hf:mys/laya-GGUF sha 713ae6f6e39f likes 0 apache-2.0; hf:mys/laya-multilingual-GGUF sha 3b645ae54281; hf:mys/laya-typed-decisions-GGUF sha 1e9e8ba1f527; hf:tozp/laya-onnx sha 0862aeba1e65 Opset 14 FP32 and INT8; tozp/laya-onnx ≠ Mattepiu/laya-onnx ≠ gqgs/laya-onnx; docker-laya MIT HEAD 1b8239a51ddd README SHA 9cb7bdc3; laya.cpp RTX ggml CUDA HEAD 8590937c79a2 README SHA cdd429b9; serving substrate ≠ calibrated replica; Softmax over options ≠ calibrated Noul; Qwen3.5-4B ≠ Archer; Qwen3.5-9B ≠ Archer; jev-position-test n=6 HEAD 7a56ca1c2698 README SHA 23f194c9; jevmlx slots 5 of 6; hosted Jev 0 of 6; prior_correction made it worse; jevSweeper mean Spearman ρ −0.274; picked exact-optimal 1/25 (4%); 31 of 36 still logically decidable; 86% of the time we should not have been asking; game success ≠ calibrated Noul; LLM2Jev 64★ Apache-2.0 HEAD 924618721277 README SHA da35fe61; not affiliated with or endorsed by Jev or TypeSafe; No answer tokens are generated; OpenSourceJev llama.cpp Qwen3-1.7B HEAD 3c41fba3681d; JEV-MLX Qwen3.5-9B HEAD dec24cd929ea; decision-head-rlcd Qwen3.5-4B 4.9M LoRA; AUTO_ACT is not a Noul; closed-set fail-open stdlib-only; verified=False; soft scores ≠ hard gates; 22 to 40% cheaper *theirs*; first version 70% more expensive; 111-case benchmark *theirs*; CDC 5,000 Jev 72.2% AUC 0.7725 *theirs*; accuracy is a trap; 9.0% base rate always-no 91.0%; catalog ≠ endorsement; jev-skill 109★ 90 scenarios HEAD 4f6e899a24d4; awesome-jev-live 673 entries 4★; minecraft-agent 214★ 131 JEV decisions 35 Astra calls; nether-final-08 8 minutes 43.300 seconds; planner writes JEV selects; RoboJEV structured simulator state not images; ashare-trader 策略未通过自己的回测门槛; 36 组参数全部净期望为负; no positive expectation under real costs; typed_evals NOT an official TypeSafe AI product; jev-as-judge is a sensor; third-person-audit 40% & 60% watermarks still soft; The included experience uses a handwritten demo provider; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33/#34/#35/#36/#37/#38/#39/#40/#41/#42; notes.md §120 + +--- + ## User-provided 0922 HIGH (SemIf densify, notes.md §117) - Docs-only reconstructed onto latest main after merged #34/#35/#36/#38. diff --git a/research/notes.md b/research/notes.md index 2c40ad86..4ec2d399 100644 --- a/research/notes.md +++ b/research/notes.md @@ -29632,3 +29632,358 @@ Parent merge only after **CLEAN** adversarial review key. No wrappers. Hourly 0947 uniqueness lock: Fast and cheap agent evals. jev as judge.; 18,041 skills from the 200 most-starred repos; Not a security scanner; 最简 Jev 调用演示器; confidence 不是正确率; q93304989-bit/jev-lab ≠ tanayvasishtha/jev-lab ≠ dairui1/jev-lab ≠ BrendanH18/jev-lab ≠ yibie/laya-jev-lab; 75% cheaper and 18% faster withdrawn; jev @0.15 100% recall 87% savings; 33Audits/jev-auto ≠ gargpratyush/jev-router; no Typesafe key, no PI_API_BASE, zero deps; tool-emitted Score/Noul ≠ calibrated Noul; semantic_compatibility: false; candidate_mass; Qwen3.5-2B ≠ Archer; Jev evaluates decisions; it cannot run a coding-agent session; Status: no model yet; S1LV3RJ1NX/openjev ≠ TheoLeeCJ/openjev; 28 accepted decisions; 3 targets; score 800; health 100; arcade game not a flight trainer; A successful live TypeSafe call has not been verified for v0.1.0; abhibansal60/tidy ≠ MANISH007700/tidy; No model, Jev included, predicted which channels its owner keeps; seed 1 selected on a held-out 400-item validation split; Brier 0.342 → 0.378; more accurate and more overconfident; Qwen3.5-4B ≠ Archer; static quants of kushalpatil/jevify-gemma4-26b-a4b; The labels were corrected, and one earlier result was retracted; zero of 23,869 eligible rows; Do not compare cost without checking task success; Exit 1 is not a proof; kisshan13/typesafe-ai-go ≠ Nibir1/typesafe-go ≠ official; 38 tests that cannot fail in a 356-model warehouse; if a parser can answer it, Jev is never asked; 359 of them; Games & Simulation 82; Education & Learning 1; Ratings are heuristics; syedabbasshaheer-art/jev-atlas ≠ ZeroX-01/jev-atlas ≠ Zaious/jev-capability-atlas ≠ gorock007/jev-atlas; anandi1989/awesome-jev-usecases ≠ whyashthakker/awesome-jev-use-cases ≠ walidboulanouar/awesome-jev-use-cases ≠ vamsikrishna2421/jev-usecases; Every headline result above is self-reported; Archer Hume 84.6% MMLU-Pro is a third-party probe not landed Archer; catalog ≠ endorsement; judge ≠ actuator; softmax over A–H ≠ Noul; SemIf 2270★; jevlike 1054★; TypeLLM/TypeLLM 16★; AnotiaWang 98★; yibie/awesome-jev 538★; Laya likes 889; tracker likes 68 lastModified UNCHANGED; Blackwood likes 2 gated manual; Archer still promised_not_landed; Hub archerhume/4rcherhume HTTP 401; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33/#35/#36/#37/#38/#40; do not push onto open #39; notes.md §119 + + +## 120. Hourly 1049 HIGH (2026-09-20 ~10:49 Boise / 2026-09-20T16:49Z) + +Measurement / serving / class-member fold on a **fresh PR off +latest `main`** (`cursor/fold-hourly-1049-high-1067`), rebased +onto `8f446c4` (merged #41 hourly 0947, `notes.md` §119 / +items 337–352 / batch #102) after `fb15455` (merged #42 v0.5.0) +after `38e4e92` (merged #39 §114 hygiene) after `9824aa2` +(merged #40 llm-to-jev, `notes.md` §118). **HARD RULE:** do not +reopen or amend PR #23–#42. Do **not** push onto the 0947 +rebase track (now merged #41). This fold's IDs: +`notes.md` §120 / composition 353–368 / findings batch #103. + +Never reopen merged #7–**#42**. Do **not** re-fold §119 0947 / +§118 llm-to-jev / §117 SemIf / §116 jcr / §115 NanoJev / §114 +0843 *as a second census* (this hour densifies serving +substrates of Laya, option-order vs exact-p measurement, and +adapter heads already adjacent to §119 RLCD). Densify +`ashleyotooligan/jevbrain` as handwritten demo provider (AUTO_ACT +is not a Noul already catalogued). Skip Archer rewrite. +Quote READMEs. Mark *theirs*. No wrappers, keys, `npm` / +`pip` / `uv` / `docker pull` install recipes. +`invented_signal: false`. Hunches labeled. + +Lane is Augustus: **mathematical / logical / algorithmic +mental models** for Jev-class categorization/scoring across +AI / SWE / **business / knowledge work / life**, not +SWE-only. PRIMARY this hour is **serving substrate ≠ +calibrated replica**: ggmlc GGUF is not llama.cpp; Loading +them in llama.cpp will fail; one encoder pass. Option-order +and exact-p Minesweeper are measurements. Adapters / RLCD +heads are class members, not TypeSafe drop-ins. Soft scores ≠ +hard gates. Catalogs are indexes. Planner writes, JEV +selects. Qwen3.5-4B ≠ Archer. Qwen3.5-9B ≠ Archer. +Archer still **promised_not_landed**. + +Unique consecutive fragments (this hour) must appear as +**one substring** in overlays (see uniqueness gate): +Hourly 1049 uniqueness lock: ggmlc GGUF is not llama.cpp; Loading them in llama.cpp will fail; one encoder pass; hf:mys/laya-GGUF sha 713ae6f6e39f likes 0 apache-2.0; hf:mys/laya-multilingual-GGUF sha 3b645ae54281; hf:mys/laya-typed-decisions-GGUF sha 1e9e8ba1f527; hf:tozp/laya-onnx sha 0862aeba1e65 Opset 14 FP32 and INT8; tozp/laya-onnx ≠ Mattepiu/laya-onnx ≠ gqgs/laya-onnx; docker-laya MIT HEAD 1b8239a51ddd README SHA 9cb7bdc3; laya.cpp RTX ggml CUDA HEAD 8590937c79a2 README SHA cdd429b9; serving substrate ≠ calibrated replica; Softmax over options ≠ calibrated Noul; Qwen3.5-4B ≠ Archer; Qwen3.5-9B ≠ Archer; jev-position-test n=6 HEAD 7a56ca1c2698 README SHA 23f194c9; jevmlx slots 5 of 6; hosted Jev 0 of 6; prior_correction made it worse; jevSweeper mean Spearman ρ −0.274; picked exact-optimal 1/25 (4%); 31 of 36 still logically decidable; 86% of the time we should not have been asking; game success ≠ calibrated Noul; LLM2Jev 64★ Apache-2.0 HEAD 924618721277 README SHA da35fe61; not affiliated with or endorsed by Jev or TypeSafe; No answer tokens are generated; OpenSourceJev llama.cpp Qwen3-1.7B HEAD 3c41fba3681d; JEV-MLX Qwen3.5-9B HEAD dec24cd929ea; decision-head-rlcd Qwen3.5-4B 4.9M LoRA; AUTO_ACT is not a Noul; closed-set fail-open stdlib-only; verified=False; soft scores ≠ hard gates; 22 to 40% cheaper *theirs*; first version 70% more expensive; 111-case benchmark *theirs*; CDC 5,000 Jev 72.2% AUC 0.7725 *theirs*; accuracy is a trap; 9.0% base rate always-no 91.0%; catalog ≠ endorsement; jev-skill 109★ 90 scenarios HEAD 4f6e899a24d4; awesome-jev-live 673 entries 4★; minecraft-agent 214★ 131 JEV decisions 35 Astra calls; nether-final-08 8 minutes 43.300 seconds; planner writes JEV selects; RoboJEV structured simulator state not images; ashare-trader 策略未通过自己的回测门槛; 36 组参数全部净期望为负; no positive expectation under real costs; typed_evals NOT an official TypeSafe AI product; jev-as-judge is a sensor; third-person-audit 40% & 60% watermarks still soft; The included experience uses a handwritten demo provider; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33/#34/#35/#36/#37/#38/#39/#40/#41/#42; notes.md §120 + +### How-to-apply (five placements / measurement lenses) + +These are *class* lenses, not vendor tutorials. Same +discipline as §119 (judge ≠ actuator) and §114 (calibration +does not compose). Formal methods **compose** with scoring: +a Noul is a SENSOR; serving is a substrate; policy / parser / +replay / cost table are exact work. + +1. **Serving substrate ≠ calibrated replica** + (*theirs*, ggmlc GGUF PRIMARY + docker-laya + laya.cpp + + ONNX cousins). ggmlc GGUF is not llama.cpp. Loading them + in llama.cpp will fail. one encoder pass. Opset 14 FP32 + and INT8. tozp/laya-onnx ≠ Mattepiu/laya-onnx ≠ + gqgs/laya-onnx. Softmax over options ≠ calibrated Noul. + Wire-compat / ggml / ONNX / FastAPI is packaging, not + Harbor. Life analogue: a new bottle does not recertify + the medicine. +2. **Option-order and exact-p games are measurements** + (*theirs*, jev-position-test + jevSweeper). n=6. jevmlx + slots 5 of 6. hosted Jev 0 of 6. prior_correction made + it worse. mean Spearman ρ −0.274. picked exact-optimal + 1/25 (4%). 31 of 36 still logically decidable. 86% of + the time we should not have been asking. game success ≠ + calibrated Noul. Evaluator locks pick_by_id vs + pick_second: the same scores, four orders, four + second-slot winners. Life analogue: shuffle the menu; + if dinner changes, you were ranking seats, not dishes. +3. **Adapters / RLCD heads are class members, not drop-ins** + (*theirs*, LLM2Jev + OpenSourceJev + JEV-MLX + + decision-head-rlcd). not affiliated with or endorsed by + Jev or TypeSafe. No answer tokens are generated. + llama.cpp Qwen3-1.7B. Qwen3.5-9B ≠ Archer. Qwen3.5-4B + 4.9M LoRA. Candidate scores rank choices and are not + calibrated probabilities of correctness. Densify §119 + RLCD Brier climb; do not mint a sibling census. +4. **Soft scores ≠ hard gates** + (AUTO_ACT / fail-open / healthcare / nitro / TPA / + guardian). AUTO_ACT is not a Noul. closed-set fail-open + stdlib-only. verified=False. 22 to 40% cheaper *theirs*. + first version 70% more expensive. healthcare urgency + 0.92 is application policy. BLOCK / QUARANTINE bands + are application policy. third-person-audit 40% & 60% + watermarks still soft. The included experience uses a + handwritten demo provider. +5. **Catalogs and life placements are indexes, not + endorsements.** jev-skill 109★ 90 scenarios. + awesome-jev-live 673 entries 4★. catalog ≠ endorsement. + planner writes JEV selects (minecraft-agent 214★; 131 + JEV decisions; 35 Astra calls; nether-final-08 8 minutes + 43.300 seconds). RoboJEV structured simulator state not + images. ashare-trader 策略未通过自己的回测门槛; 36 组参数全部净期望为负; + no positive expectation under real costs. typed_evals + NOT an official TypeSafe AI product. jev-as-judge is a + sensor. CDC 5,000 Jev 72.2% AUC 0.7725 *theirs*; + accuracy is a trap; 9.0% base rate always-no 91.0%. + 111-case benchmark *theirs*. + +### HIGH + +1. **[hf:mys/laya-GGUF](https://huggingface.co/mys/laya-GGUF)** + — NEW HIGH PRIMARY (ggmlc GGUF; apache-2.0; likes **0**; + sha `713ae6f6e39f`). Quote *theirs*: These files are not + llama.cpp / llama-cli GGUFs. Loading them in llama.cpp + will fail. Typed questions scored in one encoder pass. + ggmlc GGUF is not llama.cpp. serving substrate ≠ + calibrated replica. Do **not** copy `huggingface-cli` + / `laya serve`. +2. **[hf:mys/laya-multilingual-GGUF](https://huggingface.co/mys/laya-multilingual-GGUF)** + — NEW HIGH (ggmlc GGUF; apache-2.0; likes **0**; sha + `3b645ae54281`). Same compiler family. English + checkpoint does not degrade gracefully off English. + Quant ≠ calibration. +3. **[hf:mys/laya-typed-decisions-GGUF](https://huggingface.co/mys/laya-typed-decisions-GGUF)** + — NEW HIGH (ggmlc GGUF; apache-2.0; likes **0**; sha + `1e9e8ba1f527`). Same one-encoder-pass substrate. Not a + new training run. +4. **[hf:tozp/laya-onnx](https://huggingface.co/tozp/laya-onnx)** + — NEW HIGH (ONNX; likes **0**; sha `0862aeba1e65`). + Quote *theirs*: Opset Version 14 (TorchScript exporter). + Precision FP32 and INT8. tozp/laya-onnx ≠ + Mattepiu/laya-onnx (HF likes **9**) ≠ gqgs/laya-onnx + (GH ★1; Hub HTTP 401). ONNX ≠ Noul. Do **not** copy + `pip install laya`. +5. **[chneau/docker-laya](https://github.com/chneau/docker-laya)** + — NEW HIGH (Python MIT; **0★**; size **0** WITH + CONTENTS; HEAD `1b8239a51ddd`; README SHA `9cb7bdc3`). + Dockerized FastAPI typed-decision service. Serving + substrate ≠ calibrated replica. Do **not** copy + `docker pull` / `API_KEYS`. +6. **[lkarlslund/laya.cpp](https://github.com/lkarlslund/laya.cpp)** + — NEW HIGH (C++ MIT; **0★**; HEAD `8590937c79a2`; + README SHA `cdd429b9`). RTX ggml CUDA native inference. + Throughput *theirs* is a systems comparison, not + semantic equivalence. Do **not** copy CMake/CUDA + recipes. +7. **[imaddde867/jev-position-test](https://github.com/imaddde867/jev-position-test)** + — NEW HIGH PRIMARY measurement (Python MIT; **3★**; + HEAD `7a56ca1c2698`; README SHA `23f194c9`). Quote + *theirs*: n=6. jevmlx slots 5 of 6. hosted Jev 0 of 6. + prior_correction made it worse (the fix made it worse). + A method demonstration, not a benchmark. Evaluator: + pick_by_id vs pick_second. +8. **[lvk901/jevSweeper](https://github.com/lvk901/jevSweeper)** + — NEW HIGH measurement (Python MIT; **0★**; HEAD + `c31685473536`; README SHA `a6a0a416`). Quote *theirs*: + mean Spearman ρ −0.274. picked exact-optimal 1/25 (4%). + 31 of 36 still logically decidable. 86% of the time we + should not have been asking. game success ≠ calibrated + Noul. Exact work first; remainder judged. +9. **[Yinsongxu/LLM2Jev](https://github.com/Yinsongxu/LLM2Jev)** + — NEW HIGH (Python Apache-2.0; **64★**; HEAD + `924618721277`; README SHA `da35fe61`). Quote *theirs*: + not affiliated with or endorsed by Jev or TypeSafe. No + answer tokens are generated. Adapter ≠ TypeSafe drop-in. + Do **not** copy `uv sync`. +10. **[sabeel111/OpenSourceJev](https://github.com/sabeel111/OpenSourceJev)** + — NEW HIGH (Python MIT; **10★**; HEAD `3c41fba3681d`). + llama.cpp + Qwen3-1.7B. Candidate-only softmax ≠ Noul. + BoolQ T≈9.47 ECE 0.20→0.09 *theirs*. Qwen3-1.7B ≠ + Archer. Do **not** copy `pip` / DLL setup. +11. **[CoderInPajamas/JEV-MLX](https://github.com/CoderInPajamas/JEV-MLX)** + — NEW HIGH (Python MIT; **0★**; HEAD `dec24cd929ea`). + Quote *theirs*: Qwen3.5-9B. Candidate scores rank + choices and are not calibrated probabilities of + correctness. Qwen3.5-9B ≠ Archer. No affiliation or + endorsement. Do **not** copy `pip install`. +12. **[Astro-Han/decision-head-rlcd](https://github.com/Astro-Han/decision-head-rlcd)** + — DENSIFY §119 HF card (Python; license **null**; + **0★**; HEAD `a720133f5302`; README SHA `1913c765`). + Quote *theirs*: Qwen3.5-4B, 4.9M-parameter LoRA. + Brier 0.342 → 0.378; more accurate and more + overconfident. Qwen3.5-4B ≠ Archer. Do **not** paste + JevBench 0.779 as class Harbor. +13. **[litshing/jevcore](https://github.com/litshing/jevcore)** + — NEW HIGH (Python; **0★**; HEAD `25fdfea90d87`; + README SHA `970ad68b`). Quote *theirs*: closed-set, + fail-open, stdlib-only. marks every answer + verified=False. soft scores ≠ hard gates. Do **not** + copy `ln -sf` / keys. +14. **[daniel-farina/nitro](https://github.com/daniel-farina/nitro)** + — NEW HIGH (Rust; **0★**; HEAD `99b8749d3474`; README + SHA `ad5db53e`). Quote *theirs*: 22 to 40% cheaper. + first version 70% more expensive. 0.30 keep-set is + application policy. Fail-open on timeout. Do **not** + copy `./install.sh`. +15. **[rubinagentagi-tech/jev-heart-risk-bench](https://github.com/rubinagentagi-tech/jev-heart-risk-bench)** + — NEW HIGH (Python; **0★**; HEAD `bf75f8a1a7fc`; + README SHA `1ccfb64e`). Quote *theirs*: CDC 5,000. + Jev 72.2% AUC 0.7725. accuracy is a trap. 9.0% base + rate always-no 91.0%. Not a medical device. Do **not** + paste AUC as Harbor. +16. **[wuyoscar/jev-skill](https://github.com/wuyoscar/jev-skill)** + — NEW HIGH catalog (Python MIT; **109★**; HEAD + `4f6e899a24d4`; README SHA `e5573025`). Quote *theirs*: + 90 scenarios. catalog ≠ endorsement. wuyoscar/jev-skill + ≠ aiwithenoch/Jev-Skill ≠ simplosophy/jev-skill ≠ + laguagu/jev-skills. Do **not** paste listed numbers as + ours. +17. **[wh000wh000/awesome-jev-live](https://github.com/wh000wh000/awesome-jev-live)** + — NEW HIGH catalog (Python; **4★**; HEAD `9cf39e12`; + README SHA `d16a4649`). Quote *theirs*: 673 entries. + catalog ≠ endorsement. Ratings are heuristics. +18. **[rmalde/minecraft-agent](https://github.com/rmalde/minecraft-agent)** + — NEW HIGH life/control (JavaScript; **214★**; HEAD + `78b40ed59514`; README SHA `46bae5ab`). Quote *theirs*: + 131 JEV decisions and 35 Astra calls. nether-final-08 + 8 minutes 43.300 seconds. planner writes JEV selects. + Structured-state control, not screenshots. Do **not** + copy `npm` / Secret Manager recipes. +19. **[lykycy123/RoboJEV](https://github.com/lykycy123/RoboJEV)** + — NEW HIGH robotics (Python Apache-2.0; **8★**; HEAD + `aa82be096030`; README SHA `aafd62c0`). Quote *theirs*: + structured simulator state, not images. Model answers + cannot declare success. planner writes JEV selects. +20. **[xuboboo/ashare-trader](https://github.com/xuboboo/ashare-trader)** + — NEW HIGH honesty (TypeScript; **0★**; HEAD + `048fd921b3d8`; README SHA `9691a9c1`). Quote *theirs*: + 策略未通过自己的回测门槛. 36 组参数全部净期望为负. no + positive expectation under real costs. Negative EV is + a finding, not a product claim. +21. **[TrustifAI/typed_evals](https://github.com/TrustifAI/typed_evals)** + — NEW HIGH jev-as-judge (Python MIT; **0★**; HEAD + `0d54b0ab6121`; README SHA `d0927206`). Quote *theirs*: + This is NOT an official TypeSafe AI product. + jev-as-judge is a sensor. threshold 0.9 still soft. + Do **not** copy `pip install`. +22. **[saifullahshafin/system1-third-person-audit](https://github.com/saifullahshafin/system1-third-person-audit)** + — NEW HIGH (Python MIT; **0★**; HEAD `8dc897fd2fcb`; + README SHA `5fbb1e78`). Quote *theirs*: 40% Mid-Flight + Gate. 60% Convergence Gate. watermarks still soft. + HALT is application policy. Do **not** copy `pip`. +23. **[ashleyotooligan/jevbrain](https://github.com/ashleyotooligan/jevbrain)** + — DENSIFY (MIT; **1★**; HEAD `672478592ba4`; README + SHA `56d74078`). Quote *theirs*: The included + experience uses a handwritten demo provider. Its + results are not Jev results. AUTO_ACT is not a Noul + (prior Synxneuos catalog). Do **not** mint a second + census as TypeSafe Jev. +24. **[bhaskarpraveen/jev-healthcare-support-router](https://github.com/bhaskarpraveen/jev-healthcare-support-router)** + — NEW HIGH (TypeScript; **0★**; HEAD `96d894cdbec4`; + README SHA `cda91910`). Quote *theirs*: Urgency 0.92. + Jev makes the decision. TypeScript controls the action. + 0.92 is application policy, not a hard safety gate. + Not clinical. +25. **[YuyaForest/JEV-Prompt-Injection-Guardian](https://github.com/YuyaForest/JEV-Prompt-Injection-Guardian)** + — NEW HIGH (TypeScript; **0★**; HEAD `450d3af8e4e5`; + README SHA `07ea7911`). Quote *theirs*: BLOCK / + QUARANTINE / INSPECT / MONITOR / ALLOW. Score bands + are application policy. soft scores ≠ hard gates. +26. **[ghubnab99/jev-enterprise-decision-fabric](https://github.com/ghubnab99/jev-enterprise-decision-fabric)** + — NEW HIGH (C# MIT; **0★**; HEAD `51dce51c359b`; + README SHA `f159f22d`). Quote *theirs*: labelled + 111-case benchmark. A model answer is evidence. It is + not authorization. Thresholds are hypotheses. Do **not** + paste 90.1% as Harbor. + +### Remainder (short cards, same hour) + +Crush-monitor (28★) is a life placement: Jev scores chat; +humans own the relationship. jev-x / JevRev / docket / +jevmoji / garmin-fuel-guide / dasheng / cartas / grocery / +jevportfolio are application routing. jev-lexer / +effect-jev-cwe / jev-311-heatmap / jev-graph-walk / +sema-jev-framework / field-tests / iot-demo / BizzJev / +system-one / dsh-plugin / battlesnake / model-effort-router / +opencode-osuki / graphlin / othello / sudoku / oxox / +let-jev-speak / typesafe-showcase / antigravity-auto-mode / +codesafe / Jev-usecases / jev-colab-lab / WiredMind2 / +hama-jp tetris / ignitewala / fly2abhishek / netiqus / +kuromoka / gyu-don / havietkok / javierdv7 / nabendu82 / +cheeaun / glamboyosa / Bald0Wang docs-zh / +CompleteTech 311 / n0nuser / noelserdna / nyattoh / +osuki-dev / royosherove / suidouble / swarooppatilx / +wquguru / zebedelu / MrSAO666 file-explorer / +Steven04hub qq_clinet / vladzima jev-x / AliZareh JevRev +are first sightings or thin applications: catalog, do not +elevate. StephenChan-1/Jev-usecases ≠ whyash / walid / +vamsikrishna / anandi1989 / wuyoscar. + +### Skips (thin / collision / 404) + +- sriannamalai/Jev.UI README **76 B**: skip-thin-noise. +- VarSamLewis/eval-diff README **404**: skip, not a crash. +- Md-Zaid-Ahmed/GMAIL-JEV-DEV GitHub **404**. +- Bigthap/canvas-quiz-ai-solver: quiz scraper; skip. +- Hugeliz/jev.mixrnation.ch: website-rate; skip-thin-noise. +- eclecticv/jev-adcp-decision-economics README **139 B**: skip. +- Archer rewrite: **promised_not_landed**. Hub + archerhume/4rcherhume HTTP **401**. + +### Pulse (live REST this hour) + +Archer still promised_not_landed. Hub archerhume/4rcherhume +HTTP **401**. This hour does not re-census SemIf / Laya likes +/ tracker; those numbers stay §119 until a dedicated pulse. +minecraft-agent **214★**. jev-skill **109★**. LLM2Jev **64★**. +RoboJEV **8★**. awesome-jev-live **4★** / 673 entries. + +### Formal compose + anti-patterns + +Formal methods **compose** with scoring. A Noul is a +SENSOR. Serving (ggmlc / ONNX / Docker / C++) is a +substrate. Parsers / keep-sets / watermarks / BLOCK bands / +fail-open harnesses are exact work. Treating a GGUF as +llama.cpp, an ONNX graph as a replica, slot-two as a Noul, +ρ −0.274 as "the model cannot play", Qwen3.5-4B/9B as +Archer, AUTO_ACT as a Noul, 0.92 urgency as a hard gate, +or a catalog row as endorsement is soundness theater. +Ranking ≠ calibration. Soft Noul ≠ hard safety. +serving substrate ≠ calibrated replica. +planner writes JEV selects. catalog ≠ endorsement. + +### Adversarial review + testing hooks (Basit standing +order) + +Parent merge only after **CLEAN** adversarial review +**AND** testing. Hooks for the reviewer: + +- Uniqueness-gate: the consecutive `Hourly 1049 + uniqueness lock:` string must appear in every overlay + listed below. Prior walls 0843 / 0915 / jcr / 0922 / + 0940 / 0947 stay one substring each (do not mutate them; + do not reopen #23–#42). +- Namesake locks: tozp/laya-onnx ≠ Mattepiu/laya-onnx ≠ + gqgs/laya-onnx; wuyoscar/jev-skill ≠ aiwithenoch / + simplosophy / laguagu; StephenChan-1/Jev-usecases ≠ + whyash / walid / vamsikrishna / anandi1989. +- Densify vs new: decision-head-rlcd densifies §119 HF + card; ashleyotooligan densifies handwritten demo (AUTO_ACT + already catalogued). ggmlc GGUF / position-test / + jevSweeper / LLM2Jev are first sightings this hour. +- Harbor-jevals: 72.2% / 0.7725 / 111-case / 22 to 40% / + ρ −0.274 are *theirs*, not Harbor. 128/128 stays NanoJev + §115. 0.779 JevBench is *theirs* from §119. +- Anti-patterns to refuse: TypeSafe drop-in; Qwen3.5-4B or + Qwen3.5-9B or Qwen3.8-27B as Archer; ggmlc GGUF as + llama.cpp; softmax over options as Noul; pick_second as + pick_by_id; AUTO_ACT as a Noul; BLOCK band as a proof; + catalog as endorsement; copying keys / `npm` / `pip` / + `docker pull`. +- Overlay set: SKILL.md body (not YAML surgery), + mental-models Apply 1049, composition-algebra items + 353–368, faq, mixed-architecture, validation, + toolbox-mapping, methods-catalog, formal-methods, + formal-semi-formal, applied-mappings, judgment-class, + question-design, agent-self-assessment, mappings, + CHANGELOG, README, docs/ecosystem, findings batch #103, + refresh-log, sources.json, changelog-hourly.md. +- Offline check: `evaluate_decisions.py --self-test` + (now includes pick_by_id vs pick_second) and + `uniqueness_gate.py` (0843 + 0915 + jcr + 0922 + 0940 + + 0947 + 1049). No live Jev key. No wrappers. + +Hourly 1049 uniqueness lock: ggmlc GGUF is not llama.cpp; Loading them in llama.cpp will fail; one encoder pass; hf:mys/laya-GGUF sha 713ae6f6e39f likes 0 apache-2.0; hf:mys/laya-multilingual-GGUF sha 3b645ae54281; hf:mys/laya-typed-decisions-GGUF sha 1e9e8ba1f527; hf:tozp/laya-onnx sha 0862aeba1e65 Opset 14 FP32 and INT8; tozp/laya-onnx ≠ Mattepiu/laya-onnx ≠ gqgs/laya-onnx; docker-laya MIT HEAD 1b8239a51ddd README SHA 9cb7bdc3; laya.cpp RTX ggml CUDA HEAD 8590937c79a2 README SHA cdd429b9; serving substrate ≠ calibrated replica; Softmax over options ≠ calibrated Noul; Qwen3.5-4B ≠ Archer; Qwen3.5-9B ≠ Archer; jev-position-test n=6 HEAD 7a56ca1c2698 README SHA 23f194c9; jevmlx slots 5 of 6; hosted Jev 0 of 6; prior_correction made it worse; jevSweeper mean Spearman ρ −0.274; picked exact-optimal 1/25 (4%); 31 of 36 still logically decidable; 86% of the time we should not have been asking; game success ≠ calibrated Noul; LLM2Jev 64★ Apache-2.0 HEAD 924618721277 README SHA da35fe61; not affiliated with or endorsed by Jev or TypeSafe; No answer tokens are generated; OpenSourceJev llama.cpp Qwen3-1.7B HEAD 3c41fba3681d; JEV-MLX Qwen3.5-9B HEAD dec24cd929ea; decision-head-rlcd Qwen3.5-4B 4.9M LoRA; AUTO_ACT is not a Noul; closed-set fail-open stdlib-only; verified=False; soft scores ≠ hard gates; 22 to 40% cheaper *theirs*; first version 70% more expensive; 111-case benchmark *theirs*; CDC 5,000 Jev 72.2% AUC 0.7725 *theirs*; accuracy is a trap; 9.0% base rate always-no 91.0%; catalog ≠ endorsement; jev-skill 109★ 90 scenarios HEAD 4f6e899a24d4; awesome-jev-live 673 entries 4★; minecraft-agent 214★ 131 JEV decisions 35 Astra calls; nether-final-08 8 minutes 43.300 seconds; planner writes JEV selects; RoboJEV structured simulator state not images; ashare-trader 策略未通过自己的回测门槛; 36 组参数全部净期望为负; no positive expectation under real costs; typed_evals NOT an official TypeSafe AI product; jev-as-judge is a sensor; third-person-audit 40% & 60% watermarks still soft; The included experience uses a handwritten demo provider; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33/#34/#35/#36/#37/#38/#39/#40/#41/#42; notes.md §120 diff --git a/research/refresh-log.md b/research/refresh-log.md index a3561f11..3b59733c 100644 --- a/research/refresh-log.md +++ b/research/refresh-log.md @@ -1,3 +1,18 @@ + +## 2026-09-20 ~16:49 UTC / ~10:49 Boise — Hourly 1049 HIGH +- Fresh PR off latest `main` after merged #41 (0947 / `notes.md` §119 / + items 337–352 / batch #102) and merged #42 v0.5.0. Next free IDs: + `notes.md` §120 / composition 353–368 / findings batch #103. + **HARD RULE:** do not reopen or amend PR #23–#42. +- PRIMARY: ggmlc GGUF is not llama.cpp. serving substrate ≠ calibrated replica. + Option-order pick_by_id vs pick_second. Qwen3.5-9B ≠ Archer. + planner writes JEV selects. +- Evidence: `research/archive/hourly/2026-09-20T16/` (run_digest, + meaning_bullets, novel_high_this_run). +- uniqueness_gate 0843+0915+jcr+0922+0940+0947+1049. Evaluator option-order. + Quote *theirs*. No wrappers. `invented_signal: false`. +- Hourly 1049 uniqueness lock: ggmlc GGUF is not llama.cpp; Loading them in llama.cpp will fail; one encoder pass; hf:mys/laya-GGUF sha 713ae6f6e39f likes 0 apache-2.0; hf:mys/laya-multilingual-GGUF sha 3b645ae54281; hf:mys/laya-typed-decisions-GGUF sha 1e9e8ba1f527; hf:tozp/laya-onnx sha 0862aeba1e65 Opset 14 FP32 and INT8; tozp/laya-onnx ≠ Mattepiu/laya-onnx ≠ gqgs/laya-onnx; docker-laya MIT HEAD 1b8239a51ddd README SHA 9cb7bdc3; laya.cpp RTX ggml CUDA HEAD 8590937c79a2 README SHA cdd429b9; serving substrate ≠ calibrated replica; Softmax over options ≠ calibrated Noul; Qwen3.5-4B ≠ Archer; Qwen3.5-9B ≠ Archer; jev-position-test n=6 HEAD 7a56ca1c2698 README SHA 23f194c9; jevmlx slots 5 of 6; hosted Jev 0 of 6; prior_correction made it worse; jevSweeper mean Spearman ρ −0.274; picked exact-optimal 1/25 (4%); 31 of 36 still logically decidable; 86% of the time we should not have been asking; game success ≠ calibrated Noul; LLM2Jev 64★ Apache-2.0 HEAD 924618721277 README SHA da35fe61; not affiliated with or endorsed by Jev or TypeSafe; No answer tokens are generated; OpenSourceJev llama.cpp Qwen3-1.7B HEAD 3c41fba3681d; JEV-MLX Qwen3.5-9B HEAD dec24cd929ea; decision-head-rlcd Qwen3.5-4B 4.9M LoRA; AUTO_ACT is not a Noul; closed-set fail-open stdlib-only; verified=False; soft scores ≠ hard gates; 22 to 40% cheaper *theirs*; first version 70% more expensive; 111-case benchmark *theirs*; CDC 5,000 Jev 72.2% AUC 0.7725 *theirs*; accuracy is a trap; 9.0% base rate always-no 91.0%; catalog ≠ endorsement; jev-skill 109★ 90 scenarios HEAD 4f6e899a24d4; awesome-jev-live 673 entries 4★; minecraft-agent 214★ 131 JEV decisions 35 Astra calls; nether-final-08 8 minutes 43.300 seconds; planner writes JEV selects; RoboJEV structured simulator state not images; ashare-trader 策略未通过自己的回测门槛; 36 组参数全部净期望为负; no positive expectation under real costs; typed_evals NOT an official TypeSafe AI product; jev-as-judge is a sensor; third-person-audit 40% & 60% watermarks still soft; The included experience uses a handwritten demo provider; do not reopen or amend PR #23/#24/#25/#26/#27/#28/#29/#30/#31/#32/#33/#34/#35/#36/#37/#38/#39/#40/#41/#42; notes.md §120 + # Refresh log ## 2026-09-18 ~01:50 UTC — baseline diff --git a/research/sources.json b/research/sources.json index 05f4e441..96710594 100644 --- a/research/sources.json +++ b/research/sources.json @@ -5089,6 +5089,150 @@ "title": "syedabbasshaheer-art/jev-atlas", "url": "https://github.com/syedabbasshaheer-art/jev-atlas", "note": "HTML MIT 0★ HEAD 3b7b4e64 README SHA 1949dcbe size 310. 359 of them. Games & Simulation 82. Education & Learning 1. Ratings are heuristics. ≠ ZeroX-01 ≠ Zaious ≠ gorock007. notes.md §119." + }, + { + "kind": "huggingface", + "title": "mys/laya-GGUF", + "url": "https://huggingface.co/mys/laya-GGUF", + "note": "ggmlc GGUF apache-2.0 likes 0 sha 713ae6f6e39f. Loading them in llama.cpp will fail. one encoder pass. serving substrate ≠ calibrated replica. notes.md §120." + }, + { + "kind": "huggingface", + "title": "mys/laya-multilingual-GGUF", + "url": "https://huggingface.co/mys/laya-multilingual-GGUF", + "note": "ggmlc GGUF sha 3b645ae54281 likes 0. notes.md §120." + }, + { + "kind": "huggingface", + "title": "mys/laya-typed-decisions-GGUF", + "url": "https://huggingface.co/mys/laya-typed-decisions-GGUF", + "note": "ggmlc GGUF sha 1e9e8ba1f527 likes 0. notes.md §120." + }, + { + "kind": "huggingface", + "title": "tozp/laya-onnx", + "url": "https://huggingface.co/tozp/laya-onnx", + "note": "sha 0862aeba1e65 Opset 14 FP32 and INT8. ≠ Mattepiu/laya-onnx ≠ gqgs/laya-onnx. notes.md §120." + }, + { + "kind": "github", + "title": "chneau/docker-laya", + "url": "https://github.com/chneau/docker-laya", + "note": "MIT HEAD 1b8239a51ddd README SHA 9cb7bdc3. notes.md §120." + }, + { + "kind": "github", + "title": "lkarlslund/laya.cpp", + "url": "https://github.com/lkarlslund/laya.cpp", + "note": "MIT HEAD 8590937c79a2 README SHA cdd429b9 RTX ggml CUDA. notes.md §120." + }, + { + "kind": "github", + "title": "imaddde867/jev-position-test", + "url": "https://github.com/imaddde867/jev-position-test", + "note": "MIT 3★ HEAD 7a56ca1c2698 README SHA 23f194c9 n=6 jevmlx slots 5 of 6 hosted Jev 0 of 6. notes.md §120." + }, + { + "kind": "github", + "title": "lvk901/jevSweeper", + "url": "https://github.com/lvk901/jevSweeper", + "note": "MIT HEAD c31685473536 mean Spearman ρ −0.274 picked exact-optimal 1/25 (4%). notes.md §120." + }, + { + "kind": "github", + "title": "Yinsongxu/LLM2Jev", + "url": "https://github.com/Yinsongxu/LLM2Jev", + "note": "Apache-2.0 64★ HEAD 924618721277 README SHA da35fe61. not affiliated with or endorsed by Jev or TypeSafe. notes.md §120." + }, + { + "kind": "github", + "title": "sabeel111/OpenSourceJev", + "url": "https://github.com/sabeel111/OpenSourceJev", + "note": "MIT 10★ HEAD 3c41fba3681d llama.cpp Qwen3-1.7B. notes.md §120." + }, + { + "kind": "github", + "title": "CoderInPajamas/JEV-MLX", + "url": "https://github.com/CoderInPajamas/JEV-MLX", + "note": "MIT HEAD dec24cd929ea Qwen3.5-9B ≠ Archer. notes.md §120." + }, + { + "kind": "github", + "title": "Astro-Han/decision-head-rlcd", + "url": "https://github.com/Astro-Han/decision-head-rlcd", + "note": "HEAD a720133f5302 Qwen3.5-4B 4.9M LoRA densify §119. notes.md §120." + }, + { + "kind": "github", + "title": "litshing/jevcore", + "url": "https://github.com/litshing/jevcore", + "note": "HEAD 25fdfea90d87 closed-set fail-open stdlib-only verified=False. notes.md §120." + }, + { + "kind": "github", + "title": "daniel-farina/nitro", + "url": "https://github.com/daniel-farina/nitro", + "note": "HEAD 99b8749d3474 22 to 40% cheaper *theirs*. first version 70% more expensive. notes.md §120." + }, + { + "kind": "github", + "title": "rubinagentagi-tech/jev-heart-risk-bench", + "url": "https://github.com/rubinagentagi-tech/jev-heart-risk-bench", + "note": "HEAD bf75f8a1a7fc CDC 5,000 Jev 72.2% AUC 0.7725 *theirs*. accuracy is a trap. notes.md §120." + }, + { + "kind": "github", + "title": "wuyoscar/jev-skill", + "url": "https://github.com/wuyoscar/jev-skill", + "note": "MIT 109★ HEAD 4f6e899a24d4 90 scenarios catalog ≠ endorsement. notes.md §120." + }, + { + "kind": "github", + "title": "wh000wh000/awesome-jev-live", + "url": "https://github.com/wh000wh000/awesome-jev-live", + "note": "4★ 673 entries catalog ≠ endorsement. notes.md §120." + }, + { + "kind": "github", + "title": "rmalde/minecraft-agent", + "url": "https://github.com/rmalde/minecraft-agent", + "note": "214★ HEAD 78b40ed59514 131 JEV decisions 35 Astra calls nether-final-08 8 minutes 43.300 seconds. planner writes JEV selects. notes.md §120." + }, + { + "kind": "github", + "title": "lykycy123/RoboJEV", + "url": "https://github.com/lykycy123/RoboJEV", + "note": "Apache-2.0 8★ HEAD aa82be096030 structured simulator state not images. notes.md §120." + }, + { + "kind": "github", + "title": "xuboboo/ashare-trader", + "url": "https://github.com/xuboboo/ashare-trader", + "note": "HEAD 048fd921b3d8 策略未通过自己的回测门槛. 36 组参数全部净期望为负. notes.md §120." + }, + { + "kind": "github", + "title": "TrustifAI/typed_evals", + "url": "https://github.com/TrustifAI/typed_evals", + "note": "MIT HEAD 0d54b0ab6121 NOT an official TypeSafe AI product. jev-as-judge is a sensor. notes.md §120." + }, + { + "kind": "github", + "title": "saifullahshafin/system1-third-person-audit", + "url": "https://github.com/saifullahshafin/system1-third-person-audit", + "note": "MIT HEAD 8dc897fd2fcb 40% & 60% watermarks still soft. notes.md §120." + }, + { + "kind": "github", + "title": "ashleyotooligan/jevbrain", + "url": "https://github.com/ashleyotooligan/jevbrain", + "note": "MIT 1★ HEAD 672478592ba4 handwritten demo provider. AUTO_ACT is not a Noul densify. notes.md §120." + }, + { + "kind": "github", + "title": "ghubnab99/jev-enterprise-decision-fabric", + "url": "https://github.com/ghubnab99/jev-enterprise-decision-fabric", + "note": "MIT HEAD 51dce51c359b 111-case benchmark *theirs*. notes.md §120." } ] } From d2d9a1317c332a492349c0c45034f99bf8359d14 Mon Sep 17 00:00:00 2001 From: Cursor Agent Date: Sun, 20 Sep 2026 17:27:12 +0000 Subject: [PATCH 2/2] Keep pick_by_id vs pick_second as one SKILL substring uniqueness_gate looks for that fragment in SKILL.md haystack; wrapping it across two lines made the consecutive match fail. Co-authored-by: Basit Mustafa <24601@users.noreply.github.com> --- .agents/skills/augustus/SKILL.md | 4 ++-- 1 file changed, 2 insertions(+), 2 deletions(-) diff --git a/.agents/skills/augustus/SKILL.md b/.agents/skills/augustus/SKILL.md index 6543f997..52869d55 100644 --- a/.agents/skills/augustus/SKILL.md +++ b/.agents/skills/augustus/SKILL.md @@ -281,8 +281,8 @@ Skip Archer. `invented_signal: false`. ## Hourly 1049 HIGH (`notes.md` §120) ggmlc GGUF is not llama.cpp. serving substrate ≠ calibrated replica. -Qwen3.5-9B ≠ Archer. planner writes JEV selects. pick_by_id vs -pick_second. Soft scores ≠ hard gates. catalog ≠ endorsement. +Qwen3.5-9B ≠ Archer. planner writes JEV selects. pick_by_id vs pick_second. +Soft scores ≠ hard gates. catalog ≠ endorsement. Do not copy keys. Rebased onto `8f446c4` (merged #41) after `fb15455` (merged #42 v0.5.0). Do not reopen or amend PR #23–#42. Does not bump 0.5.0. Skip Archer. `invented_signal: false`.