Skip to content

Eight capabilities produce no score #19

Description

@euyis1019

The capability map holds 72 capabilities. 64 of them produce a score. The other eight are marked as evidence rather than as scored measurements, so they add nothing to the denominator and nothing they observe reaches a published number.

The eight are layout geometry, element visibility, the Clipboard API, Web Animations, CSS Typed OM, the Cache Storage API, View Transitions, and CSS.supports().

They run in every round. The fixture server grades them against checked-in expected answers, in the same way it grades the scored ones. Their results are then dropped.

Every capability should produce a score. A capability with no scored probe under it should fail validation instead of silently counting as nothing.

Why it is worth doing

Seven of the eight separate the engines, so this is not marginal evidence being withheld. Layout geometry is the sharpest result the layer has: Chrome passes it and no candidate engine does. On a 300 pixel box the three candidates answer 100, 5 and 100 pixels, so the declared CSS width never reached layout at all.

docs/RESULTS.md argues that layout and computed style are the substrate the driver frameworks stand on, since Playwright checks for a non-empty bounding box and computed visibility before every click or fill. The capabilities that measure exactly that currently count for zero.

The published output also cannot be reconciled with itself. docs/RESULTS.md says the L2 tasks map onto 72 capabilities. The report's axis row says 192 units, which is 64 capabilities across 3 attempts. Nothing names the eight that went missing between the two numbers.

What it costs

The axis grows from 192 units to 216, so reports are not forward compatible across this change. A run scored under the new rule cannot be read against a published one. Chrome stays at 100%. Moli moves from 95.31% to 94.44%, because that round was measured before the engine had layout support, which is a property of the round rather than of the change.

The scope is those eight capabilities. Every other evidence-only probe sits under a capability that already produces a score, and most of them repeat a claim a scored probe already covers, so they stay as they are.

One small fix belongs with this: r3_fv_checkvisibility declares the feature tag web.css.cascade, which does not describe Element.checkVisibility().

The figures above come from the published release archive evidence-four_engine_full_20260812.tar.gz (sha256 3052461b…f12f9). The aggregation reproduced the published 192 / 183 / 132 / 84 exactly before it was used to project anything.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

L2Web platform semantics layerscoringScoring rules and denominators

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions