feat(evaluation): bridge installed datasets into production runner - #258
Draft
daniele21 wants to merge 9 commits into
Draft
feat(evaluation): bridge installed datasets into production runner#258daniele21 wants to merge 9 commits into
daniele21 wants to merge 9 commits into
Conversation
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Scope
Adds the concrete production dataset-to-runner bridge required to move EVAL-2/EVAL-4 beyond fake case sources without coupling
evaluation:engineto filesystem or parser implementations.evaluation:dataset-adapteras a concrete integration module depending onevaluation:datasetsandevaluation:engine;EvaluationDatasetPreflightandEvaluationCaseDefinitionSourceports from registry-published installed packs;EvaluationDatasetRegistryand verifies the requested content digest;cases.jsonlthrough the canonical boundedEvaluationDatasetJsonlParserand verifies canonical content digest before exposing cases;Boundary
This PR does not add filesystem access to
evaluation:engine, does not create an alternate parser, and does not makeevaluation:datasetsdepend on the engine. The adapter is the composition seam between those two existing ownership boundaries.This slice also does not yet wire
EvaluationRunAggregatorinto the engine terminal path; the exposed category definitions are the input for that next composition slice after R-09 converges.Gate
Keep draft until module registration/navigation, scoped adapter tests/lint and repository validation are green. EVAL-2 should not be marked DONE until this bridge is integrated and used by production evaluation composition.