Skip to content

feat(sandbox): run agents in pydantic ai workspaces - #2011

Open
DEENUU1 wants to merge 8 commits into
mainfrom
feat/2003-pydantic-ai-workspaces
Open

DEENUU1 wants to merge 8 commits into
mainfrom
feat/2003-pydantic-ai-workspaces

Conversation

@DEENUU1

@DEENUU1 DEENUU1 commented Oct 5, 2026

Copy link
Copy Markdown
Member

Problem

pydantic-ai-backend 0.2.30 replaced its backend protocol with Pydantic AI workspaces and removed what the sandbox capability was built on: StateBackend as a backend, RemoteSandbox, DaytonaSandbox, ConsoleCapability(backend=...). AgenticOS was pinned to the old API, so it could not take any library fix or Pydantic AI release after 2.45.

What changes

Every place that handed a backend object around now works in the run's workspace (ctx.workspace, a pydantic_ai.workspaces.Workspace):

  • The sandbox capability is StateWorkspace() (the in-memory fallback for a preview or a test) combined with the console. The runner opens the agent's workspace and passes it to agent.run/agent.iter as workspace=. That way it also takes precedence over a ref in a stored history: a conversation continued on another host, or restored from another environment, works in the workspace it has now instead of failing with UserError.
  • state workspaces are a CappedStateBackend document served through StateWorkspaceBackend under the scope key. A write past SANDBOX_STATE_MAX_BYTES is refused at the call with ENOSPC and rolled back, files and directories both.
  • Container and Daytona workspaces are SandboxdWorkspace(session_name=key) / DaytonaWorkspace(sandbox_name=key) (decision 1). The scope key names the session, so the first run of a scope opens it by name, concurrent ones included, without recording a provider id first. When that run ends, the row records the session (session_id), and every later run attaches by that ref. So a session whose files were purged on the host surfaces as WorkspaceUnavailableError: the console tells the model its files are gone instead of the run carrying on in a new, empty session. The row then forgets the lost session, so the next turn starts fresh. A run-scoped sandbox is destroyed by its ref at close; one that was never touched isn't asked for. Deleting a conversation destroys the recorded session.
  • RecordingWorkspace replaces RecordingBackend and records at the workspace layer (decision 2). The operations are read, write, ls_info, mkdir, remove and execute. An edit_file shows as a read and a write, and a glob or grep as the command it ran. Rows hold a path and never a payload, as before.
  • A directories column (migration 0104_workspace_directories.py) stores a state document's directories beside its files, so the document round-trips exactly as StateBackend holds it (decision 3). They count against SANDBOX_STATE_MAX_BYTES. No tool in the product creates an empty directory today, so for now this is about fidelity, not a feature.
  • Attachments, skills, artifacts, image generation and tool-output spills write and read through the workspace. The spill store is the one consumer that's built before the run, so it gets the workspace as the WORKSPACE_RESOURCE build resource (renamed from WORKSPACE_BACKEND_RESOURCE).
  • A delegate works in the parent's workspace only when sandbox is shared. subagents-pydantic-ai hands every delegation the parent's workspace, so the delegate proxy now applies share_with_delegates: a shared delegate keeps it, one binding its own sandbox gets a fresh one (workspace='new'), and one binding none gets none. This keeps main's behaviour. Without it, an unshared delegate would have run its own execute in the parent's container.
  • A container's shell failures are handled like a full document. A sandbox with no native filesystem moves bytes through its shell, which fails with WorkspaceError, not OSError. Attachments, generated images, artifact reads and spills handle both. Skill proposals stat a file the listing didn't measure, so the size ceiling also applies on containers.
  • The Builder's tool contracts list a capability's tools with a probe context that registers the capability tree, as a run does. Without it, the combined sandbox binding raised and the Builder lost the sandbox tools' full descriptions.

The licence review records bidict (MPL-2.0) and obstore (MIT, no licence file), both new through the Daytona SDK, and THIRD_PARTY_NOTICES.md is regenerated.

Dependencies: pydantic-ai-slim 2.45 → 2.54, pydantic-ai-harness 0.35 → 0.54 (now released from the pydantic-ai monorepo with an exact slim pin), pydantic-ai-backend>=0.2.32, subagents-pydantic-ai>=0.2.25. Docs: docs/sandbox.md, docs/reference/capabilities.md and docs/architecture.md, in all four languages. CHANGELOG [Unreleased] is updated.

Verification

  • make test equivalent against pgvector (pgvector/pgvector:pg16): 11,260 passed, 27 skipped, 100% platform coverage, integration tests included.
  • Backend lint steps: ruff check, ruff format --check, ty check, vulture, deptry, and the guard scripts (backticks, routes, comments, docs paragraphs, docs i18n), all clean. codespell ran through pre-commit on every commit.
  • Migrations: alembic upgrade head, then downgrade base, then upgrade head on a fresh database, and alembic check with no drift. Head is 0104_workspace_directories.
  • scripts/license_inventory.py check: reviewed, with the one open finding already on main (PyMuPDF is AGPL-3.0: decide whether the backend image may ship it #1602).
  • The container paths run against a stand-in provider backed by a real directory and shell. Spills are pruned by a real rm/rmdir, and a purged session is simulated as the provider reports it. The run-level tests use real agents (FunctionModel): one with a history naming another organization's document or another connection's session writes only into the workspace the runner opened, and a delegate's file lands in the parent's workspace only when sandbox is shared.
  • A reviewer pass over the diff found the unshared-delegate leak, the WorkspaceError/OSError gap, the unmeasured skill files, and an overclaim about directories in the docs and CHANGELOG. All four are fixed here with tests.
  • Not run: make check as one target (no frontend change), Playwright, and a live sandboxd or Daytona conversation (see Limitations).

Limitations

  • Anthropic max_tokens default. Pydantic AI 2.52 defaults AnthropicModel's max_tokens to the model's maximum output (previously 4096). An agent or profile that sets no max_tokens can now answer at length. The budget pre-check doesn't read max_tokens, so nothing breaks, but the cost per answer can rise. It's in the CHANGELOG.
  • Harness compaction now measures the run's history (ctx.messages), not the request it edits. In a run these are the same. The unit tests built a context without history and were updated.
  • Older activity-log rows keep their old operation names (edit, glob_info, grep_raw, read_bytes) until the 30-day retention removes them.
  • Container file I/O is chattier. Neither sandboxd nor Daytona has a native filesystem here, so Pydantic AI moves files through the shell: a few /run calls per write plus one per 48 KiB, and one per 64 KiB read. A large attachment written into a container takes many sequential calls. The results are correct, but nothing here measured the latency.
  • Not run live: no sandboxd or Daytona account was exercised end to end. The container paths are tested against a stand-in provider backed by a real local directory and shell (so pruning runs a real rm/rmdir). The library's own sandboxd and Daytona backends are tested in pydantic-ai-backend.
  • The sandboxd:0.2 image tag floats and already serves 0.2.32's wire, so the compose files are unchanged.

Closes #2003

Work in progress: the sandbox capability, recording, the capped state
document and every caller move from the backend protocol to
ctx.workspace. Workspace tests and docs follow.
The delegation library hands every delegation the parent's workspace
whenever one is attached. The share list now decides: a delegate shared
sandbox works in the parent's files, one binding its own works in a
fresh one, one binding none gets none.

Also requires pydantic-ai-backend 0.2.32 from PyPI and records the
licence review for the Daytona SDK's bidict and obstore.
A sandbox with no native filesystem moves bytes through its shell, which
fails with WorkspaceError rather than OSError: attachments, generated
images, artifact reads and spills now treat both alike. A container's
listing carries no sizes, so the skill proposal ceiling asks stat.
Directories count against the state document's ceiling.
A workspace row records the host session once a run opened it; later
runs attach by that ref, so a session whose files were purged surfaces
as WorkspaceUnavailableError instead of an empty one under the same
name. The row forgets a lost session so the next turn starts afresh.

Refs #2003
@chatgpt-codex-connector

chatgpt-codex-connector Bot commented Oct 5, 2026 •

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

Review Status Commit Review trigger
📝 Code Review ✅ Completed 2026-10-05T22:20:13.211086Z fe21cea PR opened
🔒 Security Review ✅ Completed 2026-10-05T22:20:21.818490Z fe21cea PR opened
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

async def test_the_handle_is_the_path_the_workspace_resolved(self):
"""What a later read, and the prune at close, have to name exactly."""
store = WorkspaceOverflowStore(document_workspace())
assert await store.write("run-1/call-1.0", b"x") == "/tool_output/run-1/call-1.0"

from alembic import op

revision: str = "0104_workspace_directories"
from alembic import op

revision: str = "0104_workspace_directories"
down_revision: str | None = "0103_skill_library_fingerprint"

revision: str = "0104_workspace_directories"
down_revision: str | None = "0103_skill_library_fingerprint"
branch_labels: str | Sequence[str] | None = None
revision: str = "0104_workspace_directories"
down_revision: str | None = "0103_skill_library_fingerprint"
branch_labels: str | Sequence[str] | None = None
depends_on: str | Sequence[str] | None = None

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: fe21ceacd3

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment on lines +101 to +104
# Whether an operation found the session gone. Noted here because every
# operation passes through, and the close needs it: a lost session is
# forgotten, so the next run opens a fresh one instead of failing again.
self.lost = False

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Track session loss from unrecorded workspace operations

Mark the workspace as lost when inherited operations such as exists, stat, resolve, or working_dir raise WorkspaceUnavailableError, not only when one of the six recorded methods fails. For example, if a sandbox host is rebuilt after the initial snapshot but before AttachmentRouter calls workspace.exists(), that exception bypasses _recorded, so lost remains false and _settle_session preserves the stale session_id; the following turn attaches to the missing session and reports the same loss again instead of starting fresh as intended.

Useful? React with 👍 / 👎.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Migrate the sandbox onto Pydantic AI workspaces (pydantic-ai-backend breaking release)

1 participant