Skip to content

Latest commit

 

History

24 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

SidePilot

SidePilot provides local Chrome automation through an MV3 extension and Windows desktop automation through Python and a C# runtime. An authenticated loopback coordinator manages task generations, authority, operation IDs, cancellation, durable receipts, and target leases.

Browser actions use DOM elements and an independent in-page cursor. Native UIA and window-message actions can work without moving the physical pointer. Python click and drag currently use hardware input, move the physical pointer temporarily, and may alter window stacking or focus. Their receipts report physicalCursorMoved: true.

Purpose and use

SidePilot gives AI agents tools to carry out tasks in Chrome and Windows applications on your own computer. Use it for website navigation, searches, forms, scrolling, media playback, and desktop workflows involving application controls or dialogs. The agent decides what to do; SidePilot observes the target, executes actions, and returns results. It does not include an AI model.

Task ownership, target leases, cancellation, and durable action receipts help agents coordinate work and track what actually happened. Semantic browser refs and Windows UIA controls provide precise targets when available; image annotations support visual targeting when semantic controls are unavailable.

Which skill to use

The bundled skills tell an agent how to operate SidePilot, establish a task, select targets, and check results.

Skill Use it for
sp-web-use Website tasks in Chrome: search, navigate, click, type, scroll, and play media. Use this first for web content.
sp-computer-use Windows applications, browser chrome, file pickers, dialogs, and visual interfaces that browser semantic tools cannot target. Prefer UIA controls when available.

Ask your agent to use sp-web-use or sp-computer-use with your task. To install both skills into supported agent environments, run npm run install:agent-skills:all; npm run install:agent-skills installs them for Antigravity only.

Recognition

  • Browser snapshots provide semantic refs across scriptable frames and open shadow roots. Browser omniparse provides up to 96 viewport-clipped marks, configurable grid cells, and 3×3 subcells. Marks validate document, task generation, and control semantics before use.
  • Desktop omniparse prepares an image with an outer 0..1000 ruler. It annotates supplied LLM boxes or optional UIA controls; it does not run a vision model or local OCR. Annotation footers adapt to image width; badges are omitted when no free position exists.
  • Normalized desktop points bind to the explicit HWND. Cached marks reject changed window identity or geometry. Ruler coordinates refer to the inner frame, excluding border and legend.
  • PrintWindow support depends on the app. Failed/black captures report errors instead of substituting covering-window pixels. Use a visible-window screenshot when background capture is unavailable.
  • Pixel differences indicate observed change, not proof of action success. Capture failure returns uiChanged: null and verified: false.

Setup

Requires Node.js 20+, Windows, Python 3.11+, and the .NET 10 Windows desktop runtime. Ruler rendering requires NumPy and OpenCV. Native builds require the .NET 10 SDK.

powershell -ExecutionPolicy Bypass -File scripts/setup.ps1
sidepilot server start
sidepilot status

Load extension/ as an unpacked extension in Chrome. See bundled agent skills for task setup and authority. See AUDIT.md for findings, measurements, and remaining bottlenecks.

Verification and performance

npm test
npm run build:release
python scripts/benchmark-vision.py --baseline HEAD
node scripts/benchmark-browser.mjs HEAD

Regression checks mock browser/Win32 boundaries and do not operate applications. Benchmarks compare against a Git revision and exclude end-to-end automation costs; use a fixed commit after committing changes.

Prefer semantic refs/UIA when available. Browser --fast reduces presentation waits. Desktop click/scroll setup skips animated movement; explicit moves retain it. Frame observations run in groups of eight. Desktop batches amortize interpreter startup and check cancellation during waits. Restricted native tasks must use individual coordinator actions because batches lack per-step authority validation.

About

Give AI agents their own smooth cursor and semantic control for Chrome & Windows — browse, scroll, and automate apps at native speed without hijacking your physical mouse or keyboard.

Topics

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages