Keep the loop moving when an action leaves no trace - #2
Conversation
Four faults, all of them the same shape: the action worked, the agent could not tell, and the step budget went to repeating it. Text values came back empty. mercury-2.5 pads JSON output, and an 800-token field value was hitting a 200-token ceiling -- finish_reason 'length', content '[]'. The agent read that as 'nothing to type', skipped the step, and chose the same field again. Budget raised, and a truncated or empty reply now retries in plain text instead of returning nothing. A JSON array from the text model crashed on .get(). Handles dicts, lists and bare strings now. Waiting had no floor. Three waits in a row means the change being waited for is not coming -- usually a click that took effect without moving the snapshot marker. The history now says so instead of spending the budget. Controls that repeatedly do nothing are withdrawn after two attempts, matching how inert scroll actions were already handled. A confirm button that updates state the snapshot does not read looks identical to a dead control from the loop's side; either way, offering it a third time cannot help. Also: the typing history now names the suggestions an autocomplete field opened. Measured on Google Flights, the field clears itself a second after the value lands -- the framework owns it and re-renders from state that does not have the text yet -- and the value is filled back in only when a suggestion is chosen. Verified: Wikipedia, Python docs and GitHub navigation all still complete in two steps and land on the expected URL. Signed-off-by: waple0820 <feng.wang@lexmount.com>
|
Claude finished @waple0820's task in 1m 50s —— View job PR Review
Two bugs in the Bug 1 —
|
There was a problem hiding this comment.
Automated Claude review completed. Review run: https://github.com/lexmount/jev-nolayout/actions/runs/35683348913. Please check Claude's inline feedback and summary comment for details.
Four faults, all the same shape: the action worked, the agent could not tell, and the step budget went to repeating it.
Text values came back empty
The one that mattered.
mercury-2.5pads JSON output, and a field value was hitting a 200-token ceiling:The agent read that as "nothing to type", skipped the step, and chose the same field again next round. Budget raised to 800, and a truncated or empty reply now retries in plain text rather than returning nothing.
A JSON array from the text model also crashed on
.get()— dicts, lists and bare strings are all handled now.Waiting had no floor
Three waits in a row means the change being waited for is not coming — usually a click that took effect without moving the snapshot marker. The history now says so instead of spending the remaining budget on it.
Controls that do nothing are withdrawn
After two attempts with no observable effect, a control stops being offered. This matches how inert scrolls were already handled. A confirm button that updates state the snapshot does not read is indistinguishable from a dead control from the loop's side; either way, a third offer cannot help.
Autocomplete history says what happened
Measured on Google Flights: the value reads
Zurichat the moment of writing and""a second later, while five suggestions appear. The framework owns the field and re-renders from state that does not have the typed text yet; the value is filled back in only once a suggestion is chosen.A history line saying
page_changed: truecannot express that, so the model looked for its text, did not find it, and concluded the typing had failed. The entry now names the suggestions and says what they mean.Verified
All land on the expected URL.