Repository navigation
Conversation
Work in progress: the sandbox capability, recording, the capped state document and every caller move from the backend protocol to ctx.workspace. Workspace tests and docs follow.
The delegation library hands every delegation the parent's workspace whenever one is attached. The share list now decides: a delegate shared sandbox works in the parent's files, one binding its own works in a fresh one, one binding none gets none. Also requires pydantic-ai-backend 0.2.32 from PyPI and records the licence review for the Daytona SDK's bidict and obstore.
A sandbox with no native filesystem moves bytes through its shell, which fails with WorkspaceError rather than OSError: attachments, generated images, artifact reads and spills now treat both alike. A container's listing carries no sizes, so the skill proposal ceiling asks stat. Directories count against the state document's ceiling.
A workspace row records the host session once a run opened it; later runs attach by that ref, so a session whose files were purged surfaces as WorkspaceUnavailableError instead of an empty one under the same name. The row forgets a lost session so the next turn starts afresh. Refs #2003
Codex Review SummaryThis comment shows the latest Codex review activity on this pull request.
ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings. |
| async def test_the_handle_is_the_path_the_workspace_resolved(self): | ||
| """What a later read, and the prune at close, have to name exactly.""" | ||
| store = WorkspaceOverflowStore(document_workspace()) | ||
| assert await store.write("run-1/call-1.0", b"x") == "/tool_output/run-1/call-1.0" |
|
|
||
| from alembic import op | ||
|
|
||
| revision: str = "0104_workspace_directories" |
| from alembic import op | ||
|
|
||
| revision: str = "0104_workspace_directories" | ||
| down_revision: str | None = "0103_skill_library_fingerprint" |
|
|
||
| revision: str = "0104_workspace_directories" | ||
| down_revision: str | None = "0103_skill_library_fingerprint" | ||
| branch_labels: str | Sequence[str] | None = None |
| revision: str = "0104_workspace_directories" | ||
| down_revision: str | None = "0103_skill_library_fingerprint" | ||
| branch_labels: str | Sequence[str] | None = None | ||
| depends_on: str | Sequence[str] | None = None |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: fe21ceacd3
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| # Whether an operation found the session gone. Noted here because every | ||
| # operation passes through, and the close needs it: a lost session is | ||
| # forgotten, so the next run opens a fresh one instead of failing again. | ||
| self.lost = False |
There was a problem hiding this comment.
Track session loss from unrecorded workspace operations
Mark the workspace as lost when inherited operations such as exists, stat, resolve, or working_dir raise WorkspaceUnavailableError, not only when one of the six recorded methods fails. For example, if a sandbox host is rebuilt after the initial snapshot but before AttachmentRouter calls workspace.exists(), that exception bypasses _recorded, so lost remains false and _settle_session preserves the stale session_id; the following turn attaches to the missing session and reports the same loss again instead of starting fresh as intended.
Useful? React with 👍 / 👎.
Problem
pydantic-ai-backend 0.2.30 replaced its backend protocol with Pydantic AI workspaces and removed what the sandbox capability was built on:
StateBackendas a backend,RemoteSandbox,DaytonaSandbox,ConsoleCapability(backend=...). AgenticOS was pinned to the old API, so it could not take any library fix or Pydantic AI release after 2.45.What changes
Every place that handed a backend object around now works in the run's workspace (
ctx.workspace, apydantic_ai.workspaces.Workspace):StateWorkspace()(the in-memory fallback for a preview or a test) combined with the console. The runner opens the agent's workspace and passes it toagent.run/agent.iterasworkspace=. That way it also takes precedence over a ref in a stored history: a conversation continued on another host, or restored from another environment, works in the workspace it has now instead of failing withUserError.stateworkspaces are aCappedStateBackenddocument served throughStateWorkspaceBackendunder the scope key. A write pastSANDBOX_STATE_MAX_BYTESis refused at the call withENOSPCand rolled back, files and directories both.SandboxdWorkspace(session_name=key)/DaytonaWorkspace(sandbox_name=key)(decision 1). The scope key names the session, so the first run of a scope opens it by name, concurrent ones included, without recording a provider id first. When that run ends, the row records the session (session_id), and every later run attaches by that ref. So a session whose files were purged on the host surfaces asWorkspaceUnavailableError: the console tells the model its files are gone instead of the run carrying on in a new, empty session. The row then forgets the lost session, so the next turn starts fresh. A run-scoped sandbox is destroyed by its ref at close; one that was never touched isn't asked for. Deleting a conversation destroys the recorded session.RecordingWorkspacereplacesRecordingBackendand records at the workspace layer (decision 2). The operations areread,write,ls_info,mkdir,removeandexecute. Anedit_fileshows as areadand awrite, and agloborgrepas the command it ran. Rows hold a path and never a payload, as before.directoriescolumn (migration0104_workspace_directories.py) stores astatedocument's directories beside its files, so the document round-trips exactly asStateBackendholds it (decision 3). They count againstSANDBOX_STATE_MAX_BYTES. No tool in the product creates an empty directory today, so for now this is about fidelity, not a feature.WORKSPACE_RESOURCEbuild resource (renamed fromWORKSPACE_BACKEND_RESOURCE).sandboxis shared. subagents-pydantic-ai hands every delegation the parent's workspace, so the delegate proxy now appliesshare_with_delegates: a shared delegate keeps it, one binding its ownsandboxgets a fresh one (workspace='new'), and one binding none gets none. This keeps main's behaviour. Without it, an unshared delegate would have run its ownexecutein the parent's container.WorkspaceError, notOSError. Attachments, generated images, artifact reads and spills handle both. Skill proposalsstata file the listing didn't measure, so the size ceiling also applies on containers.The licence review records
bidict(MPL-2.0) andobstore(MIT, no licence file), both new through the Daytona SDK, andTHIRD_PARTY_NOTICES.mdis regenerated.Dependencies:
pydantic-ai-slim2.45 → 2.54,pydantic-ai-harness0.35 → 0.54 (now released from the pydantic-ai monorepo with an exact slim pin),pydantic-ai-backend>=0.2.32,subagents-pydantic-ai>=0.2.25. Docs:docs/sandbox.md,docs/reference/capabilities.mdanddocs/architecture.md, in all four languages. CHANGELOG[Unreleased]is updated.Verification
make testequivalent against pgvector (pgvector/pgvector:pg16): 11,260 passed, 27 skipped, 100% platform coverage, integration tests included.ruff check,ruff format --check,ty check,vulture,deptry, and the guard scripts (backticks, routes, comments, docs paragraphs, docs i18n), all clean.codespellran through pre-commit on every commit.alembic upgrade head, thendowngrade base, thenupgrade headon a fresh database, andalembic checkwith no drift. Head is0104_workspace_directories.scripts/license_inventory.py check: reviewed, with the one open finding already onmain(PyMuPDF is AGPL-3.0: decide whether the backend image may ship it #1602).rm/rmdir, and a purged session is simulated as the provider reports it. The run-level tests use real agents (FunctionModel): one with a history naming another organization's document or another connection's session writes only into the workspace the runner opened, and a delegate's file lands in the parent's workspace only whensandboxis shared.WorkspaceError/OSErrorgap, the unmeasured skill files, and an overclaim about directories in the docs and CHANGELOG. All four are fixed here with tests.make checkas one target (no frontend change), Playwright, and a live sandboxd or Daytona conversation (see Limitations).Limitations
max_tokensdefault. Pydantic AI 2.52 defaultsAnthropicModel'smax_tokensto the model's maximum output (previously 4096). An agent or profile that sets nomax_tokenscan now answer at length. The budget pre-check doesn't readmax_tokens, so nothing breaks, but the cost per answer can rise. It's in the CHANGELOG.ctx.messages), not the request it edits. In a run these are the same. The unit tests built a context without history and were updated.edit,glob_info,grep_raw,read_bytes) until the 30-day retention removes them./runcalls per write plus one per 48 KiB, and one per 64 KiB read. A large attachment written into a container takes many sequential calls. The results are correct, but nothing here measured the latency.sandboxdor Daytona account was exercised end to end. The container paths are tested against a stand-in provider backed by a real local directory and shell (so pruning runs a realrm/rmdir). The library's own sandboxd and Daytona backends are tested in pydantic-ai-backend.sandboxd:0.2image tag floats and already serves 0.2.32's wire, so the compose files are unchanged.Closes #2003