Skip to content

release: 0.55.0 — the effort a run used is now a decision, and part of the evidence - #499

Merged
danielgwilson merged 2 commits into
mainfrom
release/0.55.0
Aug 20, 2026
Merged

release: 0.55.0 — the effort a run used is now a decision, and part of the evidence#499
danielgwilson merged 2 commits into
mainfrom
release/0.55.0

Conversation

@danielgwilson

Copy link
Copy Markdown
Owner

release: 0.55.0 — the effort a run used is now a decision, and part of the evidence

Reasoning effort was reachable in the provider and unreachable from a lab, so
every computer-use run humanish had ever done was the provider default — by
omission, not by decision. A participant who abandons a flow at medium and
finishes it at high did not find a usability problem; the harness did.

  • actors[].reasoningEffort, overridable per lane, validated against the
    documented vocabulary. An unsupported level fails on the first turn with the
    provider's own message rather than being silently downgraded into a trace
    that would then misreport what ran.
  • ActorTrace.modelSettings (humanish.model-settings.v1) records the effort
    the request ACTUALLY carried, default included.
  • The lab screen and the fan-out plan line show it, and the plan line is what
    you read before spending money.
  • labs/effort-contrast-demo.yaml — one persona, one mission, two efforts.

gpt-5.6-sol stays the default: Artificial Analysis, who ran OpenAI's pre-release
eval, reports Luna and Sol always on the Pareto frontier ahead of Terra.

…f the evidence

Reasoning effort was reachable in the provider and unreachable from a lab, so
every computer-use run humanish had ever done was the provider default — by
omission, not by decision. A participant who abandons a flow at medium and
finishes it at high did not find a usability problem; the harness did.

  - `actors[].reasoningEffort`, overridable per lane, validated against the
    documented vocabulary. An unsupported level fails on the first turn with the
    provider's own message rather than being silently downgraded into a trace
    that would then misreport what ran.
  - `ActorTrace.modelSettings` (humanish.model-settings.v1) records the effort
    the request ACTUALLY carried, default included.
  - The lab screen and the fan-out plan line show it, and the plan line is what
    you read before spending money.
  - labs/effort-contrast-demo.yaml — one persona, one mission, two efforts.

gpt-5.6-sol stays the default: Artificial Analysis, who ran OpenAI's pre-release
eval, reports Luna and Sol always on the Pareto frontier ahead of Terra.
@vercel

vercel Bot commented Aug 20, 2026

Copy link
Copy Markdown

The latest updates on your projects. Learn more about Vercel for GitHub.

1 Skipped Deployment
Project Deployment Actions Updated (UTC)
humanish Ignored Ignored Preview Aug 20, 2026 8:44pm

Request Review

… effort (#497)

A real run found the gap this closes: the trace recorded medium while the lab
looked like it declared high. That one was my test setup, not the code — but the
link it exercised is the one no unit test could see and only a paid run could
check. Two lanes, one persona, one mission, different efforts, real orchestration
on the fake substrate at $0.
@danielgwilson
danielgwilson merged commit c52395a into main Aug 20, 2026
13 of 14 checks passed
@danielgwilson
danielgwilson deleted the release/0.55.0 branch August 20, 2026 20:51
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant