Agent Sandboxing

A cloud coding-agent system that isolates repository work in Docker sandboxes and opens pull requests from agent-authored changes.

  • TypeScript
  • Node.js
  • Docker
  • PostgreSQL
  • Prisma
  • Oauth

Agent Sandboxing is a cloud coding-agent system for making controlled changes to GitHub repositories. A user connects an installation, selects a repository and base branch, then works with an agent in a repo-scoped chat session. The platform creates an isolated workspace, gives the agent tools against that workspace, records the resulting diff and artifacts, and publishes approved work to a pull request.

The project is built around a security and ownership constraint: agent reasoning is part of the control plane, while repository commands execute in an untrusted execution plane. The sandbox is deliberately useful enough to build and inspect a project, but never receives model-provider keys, GitHub App private keys, database access, or long-lived user credentials.

Technical deep dive: Sandboxing Coding Agents.


Product boundary

A chat session is the unit of work. It owns exactly one sandbox and one working branch, named agent/<sessionId> for GitHub-backed sessions. Every message in that session operates on the same workspace, so follow-up instructions can inspect or amend earlier work without cloning again or reconstructing state from the model transcript.

There is intentionally no separate run resource. A user message is queued, processed once, and reaches one terminal state:

queued -> working -> completed
                 -> failed
                 -> cancelled

Only one message may be active per session. This serializes access to the shared filesystem and branch, preventing two agent turns from concurrently changing the same working tree. Completing, failing, or cancelling a message releases the lock; it does not discard the sandbox. The next message can continue from the current repository state.


Control plane and execution plane

The control plane contains the model integration, versioned prompts, context construction, tool policy, tracing, and application secrets. It stays outside Docker. The execution plane is the session-owned Docker container where repository commands run.

This split is the primary security boundary:

  • The sandbox receives only a short-lived, repository-scoped GitHub installation token during provisioning.
  • The clone remote is changed to a token-free URL immediately after cloning.
  • OPENROUTER_API_KEY, GitHub App private keys, database credentials, and backend configuration do not cross into the container.
  • The agent never receives a Docker client, a Prisma client, raw repository credentials, or a provider token.
  • The platform, not the agent, owns commits, pushes, and pull request creation.

The Docker runtime is an MVP implementation detail behind the Sandbox Service contract. The service exposes a constrained runtime seam rather than Docker control, so the execution plane can later move to a VM, microVM, Kubernetes pod, or hosted sandbox without changing the chat and agent boundaries.


Repository provisioning and branch ownership

Provisioning is platform-managed because clone, checkout, and branch selection establish the trusted starting state for agent work. They are not tasks for the model.

For a GitHub repository, the Sandbox Service creates the container, clones into /workspace/repo with an installation token, removes that token from the remote URL, checks out the selected base branch, and verifies the selected base SHA when one was supplied. If the branch moved and the checkout does not match the expected SHA, provisioning fails rather than silently running against a different revision.

Once the workspace is ready, the service prepares the session branch. It creates agent/<sessionId> from the selected base or checks out the existing session branch if the workspace is clean. A dirty workspace on the wrong branch is rejected. In particular, an existing session workspace is never reset to the base branch: that would destroy earlier session work.

Fixture repositories follow the same workspace contract but are copied from a locally configured directory. They are restricted to explicitly enabled test, evaluation, and acceptance environments.


Agent harness

Each user message gets one primary agent invocation. The message processor assembles a bounded request from five inputs:

  1. A compacted session summary.
  2. Recent conversation history.
  3. Recent tool activity.
  4. Workspace state, including branch and diff status.
  5. The new user request.

The agent receives this composed context with a versioned system prompt and uses the main tool profile. Tools are serialized because they share the same working tree. A tool call cannot race a second call that reads or rewrites the same file.

The workspace profile includes read, write, edit, bash, grep, find, and ls. Filesystem tools validate paths under /workspace/repo or /tmp; command output and execution time are bounded, and output truncation preserves valid UTF-8. The bash tool reports non-zero exit status and output to the model instead of converting every failing command into an opaque tool error.

The main agent can delegate bounded investigation to a nested subagent. That subagent runs a separate restricted profile with only read, grep, find, and ls; it cannot write files, execute shell commands, or use pull request tools. Its report is capped at 20,000 characters and remains internal to the parent turn. It does not create another user-visible message.

A failed primary invocation is terminal for that message. There is no hidden retry or recovery loop. After a completed turn, summary compaction updates the persisted session summary with the objective, current state, result, blockers, and a capped file context so later messages retain useful state without repeatedly sending the entire transcript.


Pull request publishing

GitHub operations are brokered capabilities, not shell access. The agent may ask to publish, but cannot supply a remote, token, or arbitrary provider command. Direct git commit and git push commands are rejected from the sandbox shell.

The backend verifies the configured remote and session branch, refuses an empty workspace, commits the changes, and pushes HEAD to refs/heads/agent/<sessionId>. The first publication creates a pull request targeting the session base branch. Later publications update that same pull request.

After publishing, the local commit is reset so the message result can still retain the workspace diff. Before a later publication, the service synchronizes the local branch from the pull request branch. This preserves the invariant that the session’s branch and pull request are platform-owned while retaining a useful uncommitted diff for the chat experience.


Durable events and live progress

The Event Service makes Postgres the canonical record of session activity. Sandbox lifecycle, messages, commands, agent tool calls, diffs, artifacts, and pull request transitions are persisted as append-only session events with a strictly increasing numeric sequence.

A state mutation and its lifecycle event are committed in the same transaction. Only after that transaction commits does the application publish the event to the in-process SSE hub. This prevents clients from receiving progress for state that later rolls back.

Clients consume the stream at:

GET /chat-sessions/:sessionId/events

They can resume with after=<sequence> or Last-Event-ID. The hub subscribes before replaying durable events, buffers events published during that replay, sorts them, and drops already-replayed sequences. This closes the replay/live race: a reconnecting client can recover from an SSE disconnect without receiving gaps or duplicates. A process restart only drops live connections; the event history remains replayable from Postgres.

The agent’s prose response is not the source of truth for completed work. The backend derives changed files, artifacts, pull request state, and terminal status from persisted records. This prevents a model claim from being mistaken for an execution result.


Artifacts, traces, and observability

The system keeps user-facing session data separate from operator-facing traces.

ArtifactStore retains bounded, redacted diffs and oversized tool output outside the main chat context. Artifact retrieval is scoped to the owning session. The final assistant message gives the user a concise report, while event records expose live operational progress.

The trace recorder captures the detailed execution hierarchy for operators: identity, composed context, session-agent timing and usage, redacted tool input and output, nested subagent activity, outcome, and safe errors. Traces can be written to JSONL and exported to Langfuse. They are intentionally not stored as database artifacts: traces are diagnostic records with a different size, retention, and audience than chat results.

Tool lifecycle events are also persisted around every call. Result snippets are capped at 500 bytes, known service errors expose only their public code and message, and raw provider errors, command environments, and secret values are excluded from persisted data.


Evaluation and verification strategy

The repository has conventional unit and integration coverage for service boundaries, lifecycle behavior, events, GitHub integration, and HTTP routes. The agent itself also has policy evaluations that run against a real configured model using an in-memory runtime. This runtime implements only the command forms emitted by the tools, so agent behavior can be measured without provisioning Docker, GitHub, Prisma, or HTTP.

The evaluation catalogue currently covers 15 policy categories, including read-only investigation, no-op behavior when work is already complete, minimal edits, destructive requests, and subagent investigation. Cases describe a prompt, fixture files, and expected observable behavior such as changed files, tool usage, and final response. The intended expansion is ten independently reported variants per category.

At the service level, cancellation propagates through the model and runtime with an AbortSignal; the processor attempts diff capture, records cancellation events, and releases the session lock. Runtime failures, clone failures, checkout mismatches, command timeouts, unexpected container exits, and GitHub publication failures become structured safe states rather than unhandled backend errors.


Deliberate limits

The system does not attempt to solve every agent platform problem. It excludes multi-repository tasks, persistent per-user VMs, parallel terminal sessions, deployment automation, browser previews, arbitrary chatbot mode, and autonomous production monitoring. Network access remains enabled because dependency installation and test execution often require it. The design keeps these limits explicit so that additional capability does not quietly erode the execution boundary.

The result is a narrow loop: a user asks for a repository change, one agent works in an isolated and observable workspace, and the platform produces a reviewable pull request without giving the model direct control over credentials or infrastructure.

Site search and portfolio assistant