Follows #21
L4 measures a pair, an agent together with an engine, rather than an engine on its own. An LLM agent solves tasks on live pages and the metric is task success. Because the object of measurement differs from L1 through L3, it is a separate layer, it is not formally scored, and it never appears in the headline table.
Recorded now so the layer boundary in #21 has something concrete on the other side of it. Points to settle before it is scheduled:
-
Task sets should come from their primary sources, with licensing and access gating checked first, since some publish answers only for a validation split.
-
The agent, the model and the scaffold have to be pinned in the run record with only the browser varying, or the comparison is not readable.
-
Two costs stay with the layer for as long as it exists: results expire when the model is upgraded, and the token bill scales with tasks times engines. Both should be estimated before the layer is scheduled.
Follows #21
L4 measures a pair, an agent together with an engine, rather than an engine on its own. An LLM agent solves tasks on live pages and the metric is task success. Because the object of measurement differs from L1 through L3, it is a separate layer, it is not formally scored, and it never appears in the headline table.
Recorded now so the layer boundary in #21 has something concrete on the other side of it. Points to settle before it is scheduled:
Task sets should come from their primary sources, with licensing and access gating checked first, since some publish answers only for a validation split.
The agent, the model and the scaffold have to be pinned in the run record with only the browser varying, or the comparison is not readable.
Two costs stay with the layer for as long as it exists: results expire when the model is upgraded, and the token bill scales with tasks times engines. Both should be estimated before the layer is scheduled.