Follows #21
What is being looked for, and why
The first L3 subset should be an environment set with three properties.
-
Frozen and self-hostable, so the content can live in fixtures/ and fall under the checksum that every run records over the fixture tree. Anything fetched at run time cannot be pinned, and a subset that cannot be pinned cannot be compared across runs.
-
Genuinely interactive, with multi-step browser operations, clicking, typing, selecting, navigating, and with state that carries across steps rather than a single mechanism probed once. This is the property L1 and L2 structurally cannot have.
-
Shaped like an RL environment, with discrete episodes and a defined success condition, so an episode can be a scoring unit and the subset speaks directly to whether an engine works as a rollout browser.
An externally defined set is much better than tasks written here. The neutrality argument in docs/RESULTS.md rests on the task set having been fixed independently of the results, and a subset whose selection could be described as "engine X does well on this" weakens it. External selection is the cheapest credibility available.
One candidate
MiniWoB++ from the Farama Foundation meets all three. It has over 100 web interaction environments under a Gymnasium API. It is MIT licensed, so it can be vendored under the existing THIRD_PARTY_NOTICES.md practice. It drives through element references rather than pixel coordinates, so it does not wait on #20. Upstream it runs through Selenium WebDriver, and this repository already has a Selenium adapter.
It is a candidate, not the definition. Any set meeting the three properties should be weighed alongside it before the subset is fixed.
Known hard parts
-
Where the environment grades itself in JavaScript, an engine's JavaScript bug can corrupt both the action and the grade. The design has to tell "the engine did not complete the task" apart from "it completed the task but computed the reward wrongly".
-
Instances generated from a seed need that seed frozen per task id, or the expected-answer model does not apply and reruns are not comparable.
-
The rest of the bench is binary pass or fail, so whether an episode or a task is the scoring unit changes what a headline number means.
-
100+ environments is not 100+ appropriate environments. Some depend on rendering or drag interactions that are out of reach today, and every exclusion needs a written skip reason, which the scenario checker already requires elsewhere.
Follows #21
What is being looked for, and why
The first L3 subset should be an environment set with three properties.
Frozen and self-hostable, so the content can live in
fixtures/and fall under the checksum that every run records over the fixture tree. Anything fetched at run time cannot be pinned, and a subset that cannot be pinned cannot be compared across runs.Genuinely interactive, with multi-step browser operations, clicking, typing, selecting, navigating, and with state that carries across steps rather than a single mechanism probed once. This is the property L1 and L2 structurally cannot have.
Shaped like an RL environment, with discrete episodes and a defined success condition, so an episode can be a scoring unit and the subset speaks directly to whether an engine works as a rollout browser.
An externally defined set is much better than tasks written here. The neutrality argument in
docs/RESULTS.mdrests on the task set having been fixed independently of the results, and a subset whose selection could be described as "engine X does well on this" weakens it. External selection is the cheapest credibility available.One candidate
MiniWoB++ from the Farama Foundation meets all three. It has over 100 web interaction environments under a Gymnasium API. It is MIT licensed, so it can be vendored under the existing
THIRD_PARTY_NOTICES.mdpractice. It drives through element references rather than pixel coordinates, so it does not wait on #20. Upstream it runs through Selenium WebDriver, and this repository already has a Selenium adapter.It is a candidate, not the definition. Any set meeting the three properties should be weighed alongside it before the subset is fixed.
Known hard parts
Where the environment grades itself in JavaScript, an engine's JavaScript bug can corrupt both the action and the grade. The design has to tell "the engine did not complete the task" apart from "it completed the task but computed the reward wrongly".
Instances generated from a seed need that seed frozen per task id, or the expected-answer model does not apply and reruns are not comparable.
The rest of the bench is binary pass or fail, so whether an episode or a task is the scoring unit changes what a headline number means.
100+ environments is not 100+ appropriate environments. Some depend on rendering or drag interactions that are out of reach today, and every exclusion needs a written skip reason, which the scenario checker already requires elsewhere.