Repository navigation
Proposal: an evalport-openeval-adapter for LexBench-Browser tasks/results #118
adhabnr-ux
started this conversation in
Ideas
Replies: 0 comments
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
Proposal: an
evalport-openeval-adapterfor LexBench-Browser tasks/results (portable interchange, not a replacement for anything here)Hi — I maintain EvalPort, an Apache-2.0 open interchange spec (
TestCase/Grader/Result/ResultSet) for portable LLM eval data, with a Python/TS SDK and 37 framework adapters merged so far (DeepEval, Ragas, MLflow, LangSmith, Braintrust, plus aFinanceBenchadapter for a static benchmark dataset — same shape this would be). Not affiliated with this project; just read the actual schema/code before writing this.Why I think this is a genuinely close fit, not a generic pitch:
I pulled the real field names from
browseruse_bench/data/LexBench-Browser/task.jsonl,data_info.json,browseruse_bench/schemas/agent_result.py,eval_result.py, andEVALUATION_PROTOCOL.md— not the docs. A LexBench-Browser task already looks almost like an EvalPortTestCasewith a rubric-basedllm_judgegrader attached:task.jsonl)TestCaseid(int)id(str(id))queryinputtarget_website,domain,difficulty,task_type,reasoning_type,language,robustness_tags,login_required,risk_controltags+metadata["lexbench"]reference_answer.steps,.key_points,.common_mistakescontextexpected_outputreference_answer.scoring(total,items[].name/score/description) +score_thresholdllm_judgeGraderscore_threshold/scoring.totalbecomes the pass thresholdThat last row is the real design decision, and it mirrors what your own
EVALUATION_PROTOCOL.mdalready specifies: official runs usegpt-5.4with a stepwise judge strategy. So the adapter wouldn't invent a grading approach — it would encode the one this repo already runs, as onellm_judgegrader per task with thescoring.itemsrubric embedded inparams.promptandparams.threshold = score_threshold / scoring.total.On the results side,
AgentResult+EvalResult(frombrowseruse_bench/schemas/) map onto EvalPort'sResultSet.results[]almost field-for-field:ResultAgentResult.task_idtest_case_idAgentResult.answeractual_outputAgentResult.metrics.end_to_end_msduration_msAgentResult.metrics.steps,.usage,agent_metadata,action_historymetadataEvalResult.predicted_label == 1passedEvalResult.evaluation_details.score / scoring.totalgrader_results[0].score(clamped[0,1]— EvalPort requires it, yourscorefield doesn't cap it)EvalResult.evaluation_details.reasoninggrader_results[0].reasonEvalResult.failure_classificationgrader_results[0].metadata["lexbench_failure_classification"]AgentResult.error/agent_done != doneResult.error({"type": "runner_error", ...})I'd write this the same way
financebench-openeval-adapterhandles a static benchmark-plus-results shape (closest precedent — read that README, it documents exactly this kind of "dataset, not a live SDK" mapping honestly, including what does not round-trip), and the same package layoutdeepeval-openeval-adapteruses for the trajectory/tool-call side (tools_called/expected_toolsas name-only arrays is a real schema constraint on both sides, not a guess).Rough sketch, using the real field names above (untested pseudocode, not a claim this is finished):
What this would not claim to solve: it doesn't touch
bubench run/bubench evalinternals, the leaderboard, orcommunity/results/submission.json— those stay exactly as they are. This is purely an optional, standaloneto_openeval()/from_openeval()package (own repo, own tests against the real validator, zero footprint here) so atask.jsonl+experiments/.../result.jsonpair could round-trip into a format other eval tooling (Promptfoo, Inspect AI, etc.) can also read, the same way the FinanceBench and DeepEval adapters already do for their respective ecosystems.If this is useful, I'd rather build it as a real, tested PR against
adapters/in the EvalPort repo (linked above) than ask for anything to change here — just wanted to check first whether the mapping actually makes sense to people who know this dataset and its scoring model better than I do, and whether there's prior art or a reason this wouldn't be worth the effort (e.g. if the rubric-per-task shape is expected to change before v2.0 — I sawLexBench-Browser2.0already exists as a data directory).— Sahi, independent contributor (not affiliated with this project)
All reactions