Skip to content

Keep the loop moving when an action leaves no trace - #2

Merged
waple0820 merged 1 commit into
mainfrom
fix/agent-loop-robustness
Sep 22, 2026
Merged

waple0820 merged 1 commit into
mainfrom
fix/agent-loop-robustness

Conversation

@waple0820

Copy link
Copy Markdown
Contributor

Four faults, all the same shape: the action worked, the agent could not tell, and the step budget went to repeating it.

Text values came back empty

The one that mattered. mercury-2.5 pads JSON output, and a field value was hitting a 200-token ceiling:

content:       '[]'
finish_reason: length

The agent read that as "nothing to type", skipped the step, and chose the same field again next round. Budget raised to 800, and a truncated or empty reply now retries in plain text rather than returning nothing.

A JSON array from the text model also crashed on .get() — dicts, lists and bare strings are all handled now.

Waiting had no floor

Three waits in a row means the change being waited for is not coming — usually a click that took effect without moving the snapshot marker. The history now says so instead of spending the remaining budget on it.

Controls that do nothing are withdrawn

After two attempts with no observable effect, a control stops being offered. This matches how inert scrolls were already handled. A confirm button that updates state the snapshot does not read is indistinguishable from a dead control from the loop's side; either way, a third offer cannot help.

Autocomplete history says what happened

Measured on Google Flights: the value reads Zurich at the moment of writing and "" a second later, while five suggestions appear. The framework owns the field and re-renders from state that does not have the typed text yet; the value is filled back in only once a suggestion is chosen.

A history line saying page_changed: true cannot express that, so the model looked for its text, did not find it, and concluded the typing had failed. The entry now names the suggestions and says what they mean.

Verified

Task Steps Time
Wikipedia article → linked article 2 6.7s
Python docs → Standard Library 2 7.9s
GitHub repo → Issues tab 2 13.5s

All land on the expected URL.

Four faults, all of them the same shape: the action worked, the agent could
not tell, and the step budget went to repeating it.

Text values came back empty. mercury-2.5 pads JSON output, and an
800-token field value was hitting a 200-token ceiling -- finish_reason
'length', content '[]'. The agent read that as 'nothing to type', skipped
the step, and chose the same field again. Budget raised, and a truncated or
empty reply now retries in plain text instead of returning nothing.

A JSON array from the text model crashed on .get(). Handles dicts, lists
and bare strings now.

Waiting had no floor. Three waits in a row means the change being waited
for is not coming -- usually a click that took effect without moving the
snapshot marker. The history now says so instead of spending the budget.

Controls that repeatedly do nothing are withdrawn after two attempts,
matching how inert scroll actions were already handled. A confirm button
that updates state the snapshot does not read looks identical to a dead
control from the loop's side; either way, offering it a third time cannot
help.

Also: the typing history now names the suggestions an autocomplete field
opened. Measured on Google Flights, the field clears itself a second after
the value lands -- the framework owns it and re-renders from state that
does not have the text yet -- and the value is filled back in only when a
suggestion is chosen.

Verified: Wikipedia, Python docs and GitHub navigation all still complete
in two steps and land on the expected URL.

Signed-off-by: waple0820 <feng.wang@lexmount.com>
@github-actions

github-actions Bot commented Sep 22, 2026 •

Copy link
Copy Markdown

Claude finished @waple0820's task in 1m 50s —— View job


PR Review

  • Read diff and full changed files
  • Identify bugs and security issues

Two bugs in the spent tracking in agent.py.

Bug 1 — spent persists stale node IDs across page navigations

agent.py:188-191

if changed:
    spent.pop(action.node, None)   # only removes the one interacted node
else:
    spent[action.node] = spent.get(action.node, 0) + 1

When a page navigation occurs (changed=True), only the node that was just clicked is evicted from spent. Every other stale entry from the previous page remains. Node IDs on the new page can collide with those old IDs (they're often sequential integers), so a fresh control on the new page could start life with spent[node] == 1 or == 2 and be suppressed immediately—before it's ever tried.

Fix: clear spent entirely whenever the URL changes, or conservatively on any changed=True.

Fix this →

Bug 2 — spent not cleared in the navigation-during-observe path

agent.py:139-153

except PageChanged:
    self.run_state.history.append(
        {"operation": operation, "label": action.label,
         "text": text or None, "page_changed": True})
    ...
    continue    # ← spent is never touched

A navigation during observe() (in-flight when settle timer expired) takes the same continue without updating spent. The old entries stay, same collision risk as Bug 1.


No security issues found. The field_text retry logic and JSON type-dispatch look correct.

@github-actions github-actions Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Automated Claude review completed. Review run: https://github.com/lexmount/jev-nolayout/actions/runs/35683348913. Please check Claude's inline feedback and summary comment for details.

@waple0820
waple0820 merged commit e50fd41 into main Sep 22, 2026
1 check passed
@waple0820
waple0820 deleted the fix/agent-loop-robustness branch September 24, 2026 13:14
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant