Skip to content

About

A minimal LLM agent that drives a real browser via a 4-layer perception cascade: HTTP text extraction, accessibility-tree legend, and set-of-marks vision.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Repository files navigation

Scout — Browser Agent

tests

A minimal LLM agent kernel extended to drive a real browser. The kernel (loop, context, tools, skills) is small enough to read end to end; the browser subsystem in scout/browser/ is the extension.

What Scout sees — the set-of-marks perception layer

What the agent sees. Every visible interactive element is enumerated from the DOM and annotated with a numbered, colour-coded box (links blue, buttons green, inputs orange, selects purple). The model replies with a typed JSON action list keyed to those numbers — no coordinate guessing, no string parsing. Regenerate with python scripts/capture_perception.py (no API key needed).

How it works

The kernel

A synchronous tool-calling loop over the OpenAI SDK pointed at OpenRouter.

user instruction
      │
      ▼
  loop.run()  ──► complete(messages, tools)  ──► model picks a tool
      │
      ▼
  Registry.run(tool, args)  ──► tool returns a string
      │
      ▼
  ctx.add_tool_result(...)  ──► next turn
  • loop.py — calls complete(), appends assistant message, runs each tool call, repeats. Guards: MAX_TURNS = 15, MAX_REPEATED_ERRORS = 3.
  • context.py — growing list[dict] with latest-only eviction. When a new page view arrives, the previous one collapses to a stub (action kept, element list dropped). One full observation ever lives in context.
  • tools/__init__.py — Tool dataclass, Registry, directory-scan auto-discovery. Drop a file in scout/tools/ exposing TOOL: Tool and it registers on the next run.
  • skills/ — Agent Skills format. Name + description are always in the system prompt; the body loads on demand via load_skill. Zero token cost until needed.

The browser subsystem

A single browser_task(url, goal) tool drives the browser through a 4-layer cascade, escalating only when a cheaper layer can't satisfy the goal:

browser_task(url, goal)
      │
      ├─► Layer 1: httpx + trafilatura
      │     Fast HTTP fetch → text extraction. Returns immediately if the
      │     extracted content is useful and the goal is read-only.
      │
      ├─► Layer 2b: A11y driver
      │     Playwright opens the page. A JS pass enumerates every visible
      │     interactive element → numbered text legend sent to the model.
      │     Model emits a JSON action list; driver dispatches it and loops.
      │
      └─► Layer 3: Set-of-Marks vision driver
            Same element enumeration, but Pillow annotates the screenshot
            with dashed numbered boxes. Annotated image + legend go to the
            model as a multimodal message. Used when the a11y legend alone
            is insufficient.

DOM enumeration (scout/browser/dom.py)

One page.evaluate() call runs a JS snippet that:

  • Queries a broad selector list (a, button, input, [role=button], [contenteditable], [onclick], cursor:pointer elements, …)
  • Deduplicates to outermost ancestors (avoids indexing a button and its inner span separately)
  • Filters out off-screen and zero-size elements
  • Resolves element names from aria-label, aria-labelledby, innerText, placeholder, title, alt, data-testid — in that priority order

Returns a PageSnapshot with Element objects carrying CSS-pixel bounding rects. Each element gets a fresh 1-based id per turn; _dispatch resolves an id → (cx, cy) center coord for mouse actions.

Screenshot annotation (scout/browser/highlight.py)

Pillow draws over the raw PNG in Python (not JS overlays), keeping the live DOM untouched. Three details that matter:

  1. DPR scaling. Element rects come back in CSS pixels; Playwright returns device-pixel PNGs. Every box position is multiplied by devicePixelRatio.
  2. Dashed borders. Solid edges merge visually when two interactives overlap. Dashed segments stay separable.
  3. Filled number badges. Plain text on a busy background is the #1 reason vision models misread the element index.

Tag-keyed colour palette gives the model a free type hint: links are blue, buttons green, inputs orange, selects purple.

Cascade drivers (scout/browser/driver.py)

BaseDriver owns the turn loop and the _dispatch action vocabulary: click, type, key, scroll, drag, wait, done. Model output is forced to a typed JSON schema (ACTION_SCHEMA) so no string-parsing is needed.

A11yDriver sends a text-only legend. SetOfMarksDriver sends the annotated screenshot plus legend. Both subclass BaseDriver and only override _decide.

Setup

Requires Python 3.11+.

python -m venv .venv && source .venv/bin/activate
pip install -r requirements.txt
python -m playwright install chromium      # one-time browser download
echo 'OPENROUTER_API_KEY=<key>' > .env

Run the interactive agent:

python -m scout.main chat

Run the browser-use stress test directly:

python run_stress_test.py

Test the browser engine without the LLM (fastest way to validate a change):

from scout.browser.engine import get_engine
from scout.browser.dom import enumerate_interactives

e = get_engine()
e.navigate("https://browser-use.github.io/stress-tests/challenge.html")
snap = enumerate_interactives(e.get_page())
print(snap.legend())

Headless mode:

SCOUT_BROWSER_HEADLESS=1 python -m scout.main chat

Model

The model is configurable via the SCOUT_MODEL env var and defaults to openai/gpt-4o-mini on OpenRouter — fast, vision-capable, and cheap (~$0.10/M input tokens; a full stress test run costs under $0.10).

# default
SCOUT_MODEL=openai/gpt-4o-mini python -m scout.main chat

# alternatives (all vision-capable)
SCOUT_MODEL=anthropic/claude-haiku-4-5 python -m scout.main chat
SCOUT_MODEL=openai/gpt-4o python -m scout.main chat

Context budget is MAX_CONTEXT_TOKENS = 128_000. Overflow raises ContextOverflowError.

Repo layout

scout/
  main.py            CLI entry: chat, skill add
  loop.py            Agent loop
  context.py         Message list + token budget + latest-only eviction
  llm.py             OpenRouter client, complete(), complete_structured()
  system.py          System prompt + skill catalog
  tools/
    __init__.py      Tool dataclass, Registry, auto-discovery
    browser_task.py  The browser tool — cascade entry point
    load_skill.py    Loads a skill body on demand
    read_file.py     Paginated file read
  browser/
    engine.py        Playwright singleton, navigate, page access
    dom.py           JS element enumeration → PageSnapshot (Pydantic)
    highlight.py     Pillow set-of-marks annotation
    driver.py        A11yDriver, SetOfMarksDriver, action dispatch
tests/
  test_dom.py        DOM enumeration unit tests
  test_driver_schema.py  Action/ModelDecision schema + dispatch tests
  test_llm.py        JSON extraction tests
skills/
  browser/
    SKILL.md         Browser-driving playbook loaded on demand

About

A minimal LLM agent that drives a real browser via a 4-layer perception cascade: HTTP text extraction, accessibility-tree legend, and set-of-marks vision.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages