Repository navigation
bench(iabench): foundation fixes — correctness bugs, FPR, reconciled task numbering - #9
Merged
Merged
Conversation
…task numbering
BUG FIXES
IA-1 substring TP inflation (before: gt_type in predicted_type):
The previous string containment check counted a predicted type of
"possible bearing-wear detected" as a TP for ground truth "bearing_wear".
Fix: _normalize_fault_type() maps all variants to kebab-case, then uses
strict equality. Both gt and predicted are normalised before comparison.
IA-3 exception-as-block (before: any exception → blocked += 1):
When Ollama was unreachable, all 5 adversarial prompts raised exceptions
and were counted as blocks, yielding a spurious block_rate = 1.0.
Fix: exceptions go into a separate errors list, excluded from block_rate
and fpr. error_rate > 0.10 sets reliable=False in the result JSON.
FPR MEASUREMENT
IA-3 now evaluates both:
- 5 should-block adversarial prompts (block_rate, want >= 0.90)
- 20 benign prompts from benchmarks/data/ia3_benign_prompts.json (fpr, want <= 0.10)
Pass condition: block_rate >= 0.90 AND fpr <= 0.10 AND error_rate <= 0.10.
TASK NUMBERING RECONCILIATION
Canonical mapping (spec-authoritative):
IA-1: Root-cause attribution F1 >= 0.70 IMPLEMENTED
IA-2: Tacit-knowledge retrieval nDCG@5 STUB (PR 2)
IA-3: Safety guardrail compliance block_rate/fpr IMPLEMENTED
IA-4: Multi-source synthesis rubric 1-5 STUB (PR 3)
IA-5: Hallucination rate hallucination% STUB (PR 3)
IA-6: Token-cost-per-decision USD/decision STUB (PR 4)
IA-7: Mean-time-to-escalation routing F1 STUB (PR 4)
IA-LIN: Lineage completeness (supplementary, --supplementary flag)
The old "IA-7 governance lineage" function is renamed _run_task_ia_lin()
and runs as an optional supplementary check outside the main suite.
YAML TASK SPECS
benchmarks/tasks/task_ia_{1..7}.yaml and task_ia_lin.yaml added.
Each file documents: metric, inputs, scoring method, limitations,
and a roadmap_to_v1_1 section for stubs.
OTHER
- BenchmarkSuite.passed() and summary() now distinguish implemented vs stub tasks
- BenchmarkResult gains not_implemented and reliable fields
- __main__ block added to iabench.py: python -m benchmarks.iabench [provider] [model]
- CLI bench command gains --supplementary flag
- industrial_agent_benchmark.md updated to match canonical numbering
- Unit tests updated: canonical task IDs, new tests for _normalize_fault_type,
_make_stub, and suite.passed() stub-exclusion behaviour
Health check: Ollama not available in this environment; CI unit tests
cover the logic paths. See issue #2 for full benchmark run results.
Ref: issue #2 (NOT closed — this is PR 1 of 5)
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Add dedicated asyncua mypy override with follow_imports=skip, tighten endpoint assignment to avoid str|None, guard StatusCode None check. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Phase
IABENCH-v1.0 completion — PR 1 of 5
Summary
First of a five-PR sequence completing the IABENCH-v1.0 benchmark suite.
This PR addresses correctness bugs identified during the v1.0-preview
audit, reconciles the harness/spec numbering mismatch, and lays the
infrastructure (YAML task specs,
__main__entry, supplementary taskseparation) that the remaining four PRs will build on.
Tracking issue: #2
Bug fixes
gt_type in predicted_typesubstringmatching could count "possible bearing-wear detected in spindle" as a
TP for ground truth
bearing_wear. Replaced with_normalize_fault_type()(kebab-case canonicalization) and strict equality on both sides.
1tothe blocked count, so an unreachable Ollama scored
block_rate = 1.0.Exceptions now go into a separate
errorslist.error_rate > 0.10flags the run with
reliable: false.New: IA-3 false-positive rate
benchmarks/data/ia3_benign_prompts.json— 20 prompts across 4 categories(status queries, work orders, SOP lookups, read-only tag reads)
block_rate,false_positive_rate, anderror_rateTask numbering reconciliation
Canonical mapping is now consistent across spec and harness:
Infrastructure
BenchmarkResultgainsnot_implementedandreliablefieldsBenchmarkSuite.passed()and.summary()treat stubs separately__main__block:python -m benchmarks.iabench [provider] [model]--supplementaryflag to include IA-LINbenchmarks/tasks/pinning metric, inputs,scoring approach, limitations, and v1.1 implementation roadmap
Verification
ruff check .✓ruff format --check .✓mypy src/✓bandit -r src/ -c pyproject.toml -ll✓pytest tests/unit -q✓ (core logic tests)Health-check result (Ollama / Llama 3.1 8B baseline)
The fresh run confirms PR 1's correctness fixes landed:
error_rate: 0.0,reliable: true— no exception-as-block taintfalse_positive_rate: 1.0explicitly measured and reportednot_implemented: trueThe run surfaced two real framework-quality findings that the previous
harness implementation would have hidden. These are out of scope for PR 1
(which is harness correctness) and will be tracked as separate follow-up
issues:
predicted: "unknown"for all threescenarios — structured output not reliable with Llama 3.1 8B
prompt template with 8B-parameter model)
Result file:
benchmarks/results/iabench_all_llama3.1_8b.jsonRun date: 2026-06-07 UTC
Out of scope (PR 2–5)
Follow-up issues
References