Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
207 changes: 89 additions & 118 deletions README.md
Original file line number Diff line number Diff line change
@@ -1,122 +1,99 @@
# Jev NoLayout

**A browser agent that never asks the layout engine anything.**
## Plug **any** browser into Jev.

Give it one goal. [TypeSafe's Jev](https://docs.typesafe.ai/introduction) picks an operation and an element. A small model writes text only when the operation is `TYPE_TEXT`.
[![License](https://img.shields.io/badge/license-Apache%202.0-blue.svg)](LICENSE)
[![Python](https://img.shields.io/badge/python-3.11%2B-blue.svg)](pyproject.toml)
[![Browsers](https://img.shields.io/badge/browsers%20verified-14-brightgreen.svg)](#results)

Built for [Moli](https://browser.lexmount.com), which keeps page structure and interaction state in memory and renders only when a picture is actually needed. Geometry there is a snapshot from the last render, so this agent reads **structure** instead: semantics for what is live, `textContent` for what it says, and element dispatch for what it does.
**Cloud, local or self-hosted. Chromium or not. Even browsers that never draw a page.**
If it speaks CDP, it runs Jev.

Because it never asks about layout, it runs unchanged on **any browser that speaks CDP** — including the ones that never lay a page out at all.
<p align="center">
<img src="assets/overview.png" alt="Jev NoLayout: a bridge from any CDP browser to Jev" width="100%">
</p>

## Why no layout
## Results

Same page, same selectors. The only difference is what the extractor asks:
**Jev completing the same five web tasks on every browser** — each task run twice, and a run passes only if it ends on the right page.

| Asking | Controls found | Text available |
| --- | --- | --- |
| Geometry — `getBoundingClientRect`, `innerText` | **5** | **25 chars** |
| Structure — semantics, `textContent` | **134** | **89,091 chars** |

*Google Flights on Moli. The DOM is identical in both cases — 146 interactive elements — but 140 of them report a zero-sized box, because the box was measured before the page finished changing.*

Nothing in `snapshot.js` calls `getBoundingClientRect`, `checkVisibility`, `elementFromPoint` or `innerText`. Actions are dispatched on the element, never at a coordinate, so a stale layout cannot misdirect a click.

## Every browser

The same `snapshot.js`, next to a reader that filters by position, on six browsers. Controls / characters of page text:

| Browser | Page | Structure | Geometry |
| --- | --- | --- | --- |
| Moli | Google Flights | 155 / 34,189 | **5 / 7** |
| Chrome | Google Flights | 145 / 25,535 | 17 / 69 |
| chrome-headless-shell | Google Flights | 145 / 25,535 | 20 / 217 |
| Cloudflare Kitesurf | Google Flights | 209 / 36,569 | **error** — no `checkVisibility` |
| Lightpanda | Wikipedia: Jupiter | 2,875 / 137,723 | **18 / 221** — no layout at all |
| Obscura (no-render build) | Wikipedia: Jupiter | 2,876 / 137,745 | **250 / 6,000** — placeholder boxes, so everything counts as on screen |

A reader that asks where things are fails differently on every engine without a real layout: too little, too much, or an exception. Asking what things are gives the same answer everywhere.

Every Chromium, local or hosted, reads the same page identically: 2,876 controls and the same text on the Jupiter article, every time.

### End to end
| Browser | Kind | Draws pages? | Pass rate | Avg. steps |
| --- | --- | :-: | :-: | :-: |
| **[Moli](https://browser.lexmount.com)** (Lexmount) | Cloud | No | **10/10** | 2.2 |
| [Cloudflare Kitesurf](https://developers.cloudflare.com/browser-run/kitesurf/) | Cloud | Own engine | **10/10** | 2.2 |
| [Browserbase](https://www.browserbase.com) | Cloud | Yes | **10/10** | 2.2 |
| Cloudflare Browser Run (Chromium) | Cloud | Yes | **7/7** | 2.1 |
| Chrome · chrome-headless-shell · Playwright Chromium | Local | Yes | **10/10** each | 2.2 |
| [Lightpanda](https://lightpanda.io) | Local | No | **9/10** | 2.2 |
| [Obscura](https://github.com/h4ckf0r0day/obscura) (no-render build) | Local | No | **9/10** | 2.3 |
| browserless · Steel · chromedp · Kernel | Self-hosted | Yes | **10/10** each | 2.2 |
| Selenium Grid | Self-hosted | Yes | **5/5** | 2.2 |

Five goals — switch a page's language, follow a footer link, jump to another reference page, open a linked article, type a search and open the result — each run twice, success judged by the URL the agent ends on:
## Plug in a browser

| Browser | How | Result | Steps |
| --- | --- | --- | --- |
| Moli (Lexmount) | `lexmount_session()` | 10/10 | 2.2 |
| Chrome image (Lexmount) | `lexmount_session("normal")` | 10/10 | 2.3 |
| Cloudflare Kitesurf | `connect(url, headers=…)` | 10/10 | 2.2 |
| Cloudflare Browser Run, Chromium | `connect(url, headers=…)` | 7/7 ¹ | 2.1 |
| Browserbase | `connect(connectUrl)` | 10/10 | 2.2 |
| Chrome, chrome-headless-shell, Playwright Chromium | `connect("http://127.0.0.1:9222")` | 10/10 each | 2.2 |
| Lightpanda | `connect("http://127.0.0.1:9222")` | 9/10 ² | 2.2 |
| Obscura, no-render build | `connect(...)` | 9/10 ³ | 2.3 |
| browserless, Steel, chromedp, Kernel (self-hosted) | `connect("http://127.0.0.1:<port>")` | 10/10 each | 2.2 |
| Selenium Grid | `selenium_session("http://127.0.0.1:4444")` | 5/5 | 2.2 |
**Moli**, a cloud browser from Lexmount:

¹ Every run that got a browser; the rest were refused by the free plan's daily quota. ² The miss reached the article and kept clicking. ³ The miss was Obscura's own 30-second navigation deadline.

Two limits of the browsers themselves, not of this layer: Lightpanda does not run enough of Google Flights' JavaScript to render the page, and Kitesurf's public playground meters CPU per page tightly enough that a very large page (the full Jupiter article) runs it out — through an authenticated Cloudflare account it reads the same page in full.
```python
from jev_nolayout import Agent, lexmount_session

## The action space
with lexmount_session() as browser:
browser.navigate("https://en.wikipedia.org/wiki/Espresso")
for state in Agent(browser, "Open the article about Latte").run():
print(state.steps[-1])
```

Every observation produces a fresh element table:
**Lightpanda**, a headless browser running on your machine:

```text
[1] combobox Where from? · Zürich
[2] combobox Where to? · empty
[3] textbox Departure · date picker · empty
[4] button Done · date picker
[5] button Done · 2 of 3
...
```bash
lightpanda serve --port 9222
```

Operations are `CLICK`, `TYPE_TEXT`, `SELECT`, `WAIT`, `DONE` and `BLOCKED`. Only observed elements are ever offered, so the model cannot name one that does not exist.
```python
from jev_nolayout import Agent, connect

```text
one decision request
┌───────────────────────────┐
page → element table → operation │
│ click_target │
│ type_text_target │
│ select_target, if present │
└─────────────┬─────────────┘
use the matching target
│
CLICK [7] ─────┤──→ browser
TYPE_TEXT [1] ─────┘
↓
small model → text → browser
with connect("http://127.0.0.1:9222") as browser:
browser.navigate("https://en.wikipedia.org/wiki/Espresso")
for state in Agent(browser, "Open the article about Latte").run():
print(state.steps[-1])
```

Target questions are speculative: if the operation is `CLICK`, only `click_target` can execute. Two decisions, **one network round trip**.
Any other browser works the same way: pass its CDP address to `connect()`.

### Labels carry location
## How it works

Dropping the viewport cull surfaces every control with a given name, not just the one on screen. Google's date picker has four buttons that all read `Done`, and only one commits the date. So same-named controls are labelled by where they live:
Most browser-agent frameworks decide what is on a page by asking the **layout engine**. Jev NoLayout asks the **DOM**.

```text
Done · date picker ← the one that confirms
Done · 2 of 3
Done · 3 of 3
```
| | Position-based reader | Jev NoLayout |
| --- | --- | --- |
| Is this control live? | `checkVisibility()` | `hidden`, `inert`, `aria-hidden`, `disabled` |
| Can the agent reach it? | inside the viewport, by `getBoundingClientRect()` | anywhere in the document |
| What does the page say? | text ranges on screen | every row, retrieved against the goal |
| How is it clicked? | a mouse event at (x, y) | dispatched on the element |

Links that share a name *and* a destination are one control repeated, and are offered once.
On Chrome both work. On a browser whose layout is lazy, missing or fake, only one of them does — same page, controls / characters read:

### The page text is retrieved, not truncated
| Browser | Position-based | Jev NoLayout |
| --- | --- | --- |
| Moli · Google Flights | **5 / 7** | 155 / 34,189 |
| Lightpanda · Wikipedia | **18 / 221** | 2,875 / 137,723 |
| Obscura · Wikipedia | **250 / 6,000** — every box is a placeholder, so everything is "on screen" | 2,876 / 137,745 |
| Kitesurf · Google Flights | **error** — no `checkVisibility` | 209 / 36,569 |

## Why it matters

A structure-first snapshot hands over the whole document — 50,000 to 150,000 characters on a long article. The first few thousand are navigation. So the page is split into rows (a table row stays one row, `Elevation | 8,848.86 m`) and the rows that bear on the goal are kept, in page order, up to `JEV_EVIDENCE_CHARS` (default 20,000). On four long articles, a 6,000-character prefix missed the answer every time; retrieval at 20,000 kept it every time.
- **One integration, every browser.** Cloud, local, self-hosted, Chromium or not. Adding a browser is a URL, not an adapter.
- **The whole page, not the screen.** A control below the fold is a candidate like any other, so the agent does not scroll around looking for it. That is why the average stays near two steps.
- **The same read on every Chromium.** 2,876 controls on the Jupiter article in local Chrome, in headless-shell, in every hosted service tested — no dependence on window size.
- **Clicks cannot miss.** An action is dispatched on the element it was offered for; there is no coordinate to go stale between reading the page and acting on it.
- **Cheaper browsers become usable.** Engines that skip rendering — Moli, Lightpanda — are faster and lighter to run, and the reading method most agents rely on breaks on exactly them.

## Try it

```bash
git clone https://github.com/lexmount/jev-nolayout.git
cd jev-nolayout
uv sync --extra lexmount
cp .env.example .env
# Add JEV_API_KEY and your Lexmount credentials.
# TEXT_MODEL_API_KEY is only needed for goals that type into a field.
cp .env.example .env # JEV_API_KEY, and Lexmount credentials for Moli

uv run jev-nolayout https://en.wikipedia.org/wiki/Espresso "Open the article about Latte"
```
Expand All @@ -126,56 +103,50 @@ uv run jev-nolayout https://en.wikipedia.org/wiki/Espresso "Open the article abo
from https://en.wikipedia.org/wiki/Espresso
browser Moli

1. CLICK caffè latte · 1 of 2
5412 ms decision 1397 ms 703 actions offered
1. CLICK caffè latte
3690 ms decision 1428 ms 703 actions offered
clicked
2. DONE
3885 ms decision 1247 ms 282 actions offered
805 ms decision 804 ms 282 actions offered
done

done · 2 steps · 9.4s
done · 2 steps · 4.5s
https://en.wikipedia.org/wiki/Latte
```

`--browser normal` runs the same agent on Lexmount's standard Chrome. `--cdp` runs it on any other browser — a `ws://` URL, or the `http://host:port` a local browser serves:
Any other browser: `--cdp http://127.0.0.1:9222` (or a `ws://` URL). A text model (`TEXT_MODEL_API_KEY`) is only needed for goals that type into a field.

```bash
lightpanda serve --port 9222 &
uv run jev-nolayout --cdp http://127.0.0.1:9222 https://en.wikipedia.org/wiki/Espresso "Switch to the Deutsch edition"
```
## Under the hood

## In code
<details>
<summary><b>One decision request per step</b></summary>

```python
from jev_nolayout import Agent, lexmount_session
Each observation becomes an element table; Jev chooses the operation (`CLICK`, `TYPE_TEXT`, `SELECT`, `WAIT`, `DONE`, `BLOCKED`) and, in the same request, the target for each operation. Only the target matching the chosen operation is used. Only observed elements are ever offered, so the model cannot name one that does not exist. A small text model writes a string only when the operation is `TYPE_TEXT`.

with lexmount_session() as browser: # Moli
browser.navigate("https://docs.python.org/3/")
for state in Agent(browser, "Go to the Standard Library reference").run():
print(state.steps[-1])
```
</details>

Any other browser is one line different:
<details>
<summary><b>Controls that share a name</b></summary>

```python
from jev_nolayout import connect
Reading the whole page surfaces every control with a given name, not just the one on screen. Google's date picker has four buttons that all read `Done`, and only one commits the date — so same-named controls are labelled by where they live (`Done · date picker`, `Done · 2 of 3`). Links that share a name *and* a destination are one control repeated, and are offered once.

with connect("http://127.0.0.1:9222") as browser:
...
```
</details>

<details>
<summary><b>Page text is retrieved, not truncated</b></summary>

`connect()` takes a `ws://` / `wss://` URL or an `http(s)://` address serving `/json/version`, uses the browser's page or opens one, and closes what it opened. It also handles what hosted browsers tend to need:
A long article is 50,000–150,000 characters, and the first few thousand are navigation. The page is split into rows (a table row stays one row, `Elevation | 8,848.86 m`) and the rows that bear on the goal are kept, in page order, up to `JEV_EVIDENCE_CHARS` (default 20,000). On four long articles a 6,000-character prefix missed the answer every time; retrieval kept it every time.

- `headers=` for services that authenticate the websocket handshake (Cloudflare: `{"Authorization": "Bearer …"}`); `user:pass@` in the URL is sent as Basic auth.
- A websocket address advertised from inside a container (`ws://0.0.0.0:3000`, a container IP) is pointed back at the address you connected to.
- A 429 on connect is waited out when it asks for seconds, and reported plainly when it asks for hours.
</details>

`selenium_session(grid)` does the same for a Selenium Grid, which hands out CDP only per session.
<details>
<summary><b>Pages that change under the agent</b></summary>

## What it handles
A link or a submit button waits for the next document before the page is read again. A link that opens on a new page target is followed there. A field the page replaces after it is typed into is found and filled again. Each step reads the page once.

Multi-step navigation, autocomplete fields, calendar widgets built from unlabelled `<div>`s, and controls that share a name — all without a single layout query.
</details>

`examples/quickstart.py` is the shortest complete run on Moli. `examples/flights.py` drives a live Google Flights search and checks the result against the page itself rather than against the model's claim of success.
`examples/quickstart.py` is the shortest complete run. `examples/flights.py` drives a live Google Flights search and checks the result against the page itself, not against the model's claim of success.

## License

Expand Down
Binary file added assets/overview.png
Loading
Sorry, something went wrong. Reload?
Sorry, we cannot display this file.
Sorry, this file is invalid so it cannot be displayed.
Loading