Skip to content

L4: end-to-end agent evaluation on live pages #24

Description

@euyis1019

Follows #21

L4 measures a pair, an agent together with an engine, rather than an engine on its own. An LLM agent solves tasks on live pages and the metric is task success. Because the object of measurement differs from L1 through L3, it is a separate layer, it is not formally scored, and it never appears in the headline table.

Recorded now so the layer boundary in #21 has something concrete on the other side of it. Points to settle before it is scheduled:

  • Task sets should come from their primary sources, with licensing and access gating checked first, since some publish answers only for a validation split.

  • The agent, the model and the scaffold have to be pinned in the run record with only the browser varying, or the comparison is not readable.

  • Two costs stay with the layer for as long as it exists: results expire when the model is upgraded, and the token bill scales with tasks times engines. Both should be estimated before the layer is scheduled.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

L4End-to-end agent evaluation layeragentAgent-in-the-loop evaluationresearchOpen question, not yet scheduled

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions