A minimal LLM agent kernel extended to drive a real browser. The kernel (loop,
context, tools, skills) is small enough to read end to end; the browser
subsystem in scout/browser/ is the extension.
What the agent sees. Every visible interactive element is enumerated from the
DOM and annotated with a numbered, colour-coded box (links blue, buttons green,
inputs orange, selects purple). The model replies with a typed JSON action list
keyed to those numbers — no coordinate guessing, no string parsing. Regenerate
with python scripts/capture_perception.py (no API key needed).
A synchronous tool-calling loop over the OpenAI SDK pointed at OpenRouter.
user instruction
│
▼
loop.run() ──► complete(messages, tools) ──► model picks a tool
│
▼
Registry.run(tool, args) ──► tool returns a string
│
▼
ctx.add_tool_result(...) ──► next turn
loop.py— callscomplete(), appends assistant message, runs each tool call, repeats. Guards:MAX_TURNS = 15,MAX_REPEATED_ERRORS = 3.context.py— growinglist[dict]with latest-only eviction. When a new page view arrives, the previous one collapses to a stub (action kept, element list dropped). One full observation ever lives in context.tools/__init__.py—Tooldataclass,Registry, directory-scan auto-discovery. Drop a file inscout/tools/exposingTOOL: Tooland it registers on the next run.skills/— Agent Skills format. Name + description are always in the system prompt; the body loads on demand viaload_skill. Zero token cost until needed.
A single browser_task(url, goal) tool drives the browser through a
4-layer cascade, escalating only when a cheaper layer can't satisfy the
goal:
browser_task(url, goal)
│
├─► Layer 1: httpx + trafilatura
│ Fast HTTP fetch → text extraction. Returns immediately if the
│ extracted content is useful and the goal is read-only.
│
├─► Layer 2b: A11y driver
│ Playwright opens the page. A JS pass enumerates every visible
│ interactive element → numbered text legend sent to the model.
│ Model emits a JSON action list; driver dispatches it and loops.
│
└─► Layer 3: Set-of-Marks vision driver
Same element enumeration, but Pillow annotates the screenshot
with dashed numbered boxes. Annotated image + legend go to the
model as a multimodal message. Used when the a11y legend alone
is insufficient.
One page.evaluate() call runs a JS snippet that:
- Queries a broad selector list (
a,button,input,[role=button],[contenteditable],[onclick],cursor:pointerelements, …) - Deduplicates to outermost ancestors (avoids indexing a button and its inner span separately)
- Filters out off-screen and zero-size elements
- Resolves element names from
aria-label,aria-labelledby,innerText,placeholder,title,alt,data-testid— in that priority order
Returns a PageSnapshot with Element objects carrying CSS-pixel bounding
rects. Each element gets a fresh 1-based id per turn; _dispatch resolves
an id → (cx, cy) center coord for mouse actions.
Pillow draws over the raw PNG in Python (not JS overlays), keeping the live DOM untouched. Three details that matter:
- DPR scaling. Element rects come back in CSS pixels; Playwright returns
device-pixel PNGs. Every box position is multiplied by
devicePixelRatio. - Dashed borders. Solid edges merge visually when two interactives overlap. Dashed segments stay separable.
- Filled number badges. Plain text on a busy background is the #1 reason vision models misread the element index.
Tag-keyed colour palette gives the model a free type hint: links are blue, buttons green, inputs orange, selects purple.
BaseDriver owns the turn loop and the _dispatch action vocabulary:
click, type, key, scroll, drag, wait, done. Model output is
forced to a typed JSON schema (ACTION_SCHEMA) so no string-parsing is
needed.
A11yDriver sends a text-only legend. SetOfMarksDriver sends the annotated
screenshot plus legend. Both subclass BaseDriver and only override _decide.
Requires Python 3.11+.
python -m venv .venv && source .venv/bin/activate
pip install -r requirements.txt
python -m playwright install chromium # one-time browser download
echo 'OPENROUTER_API_KEY=<key>' > .envRun the interactive agent:
python -m scout.main chatRun the browser-use stress test directly:
python run_stress_test.pyTest the browser engine without the LLM (fastest way to validate a change):
from scout.browser.engine import get_engine
from scout.browser.dom import enumerate_interactives
e = get_engine()
e.navigate("https://browser-use.github.io/stress-tests/challenge.html")
snap = enumerate_interactives(e.get_page())
print(snap.legend())Headless mode:
SCOUT_BROWSER_HEADLESS=1 python -m scout.main chatThe model is configurable via the SCOUT_MODEL env var and defaults to
openai/gpt-4o-mini on OpenRouter — fast, vision-capable, and cheap
(~$0.10/M input tokens; a full stress test run costs under $0.10).
# default
SCOUT_MODEL=openai/gpt-4o-mini python -m scout.main chat
# alternatives (all vision-capable)
SCOUT_MODEL=anthropic/claude-haiku-4-5 python -m scout.main chat
SCOUT_MODEL=openai/gpt-4o python -m scout.main chatContext budget is MAX_CONTEXT_TOKENS = 128_000. Overflow raises
ContextOverflowError.
scout/
main.py CLI entry: chat, skill add
loop.py Agent loop
context.py Message list + token budget + latest-only eviction
llm.py OpenRouter client, complete(), complete_structured()
system.py System prompt + skill catalog
tools/
__init__.py Tool dataclass, Registry, auto-discovery
browser_task.py The browser tool — cascade entry point
load_skill.py Loads a skill body on demand
read_file.py Paginated file read
browser/
engine.py Playwright singleton, navigate, page access
dom.py JS element enumeration → PageSnapshot (Pydantic)
highlight.py Pillow set-of-marks annotation
driver.py A11yDriver, SetOfMarksDriver, action dispatch
tests/
test_dom.py DOM enumeration unit tests
test_driver_schema.py Action/ModelDecision schema + dispatch tests
test_llm.py JSON extraction tests
skills/
browser/
SKILL.md Browser-driving playbook loaded on demand
