All posts
Engineering16 min read

Sandboxing on Brimble

Ilerioluwa David
Ilerioluwa David@pipe_dev
Sandboxing on Brimble

AI agents eventually ask for a shell.

Give an agent a CSV and it wants Python. Ask it to fix a repository and it wants Git, Node.js, a package manager, and permission to run the test suite. Ask it to process an image and it will probably install Pillow before explaining why your JPEG is haunted.

Generating code is useful. Executing it is where the trouble starts.

At Brimble, we run a bare-metal platform on Hetzner and Scaleway hardware. Customer workloads share physical machines, with HashiCorp Nomad scheduling containers across the fleet.

Running arbitrary, AI-generated code directly on those machines would be fast, cheap, and a terrible idea.

The code may be malicious. More commonly, it is simply wrong in ways nobody predicted. Either way, it should not be able to inspect the host, reach internal services, interfere with another workload, or leave a machine full of processes after the user has gone home.

So we built Sandboxes: short-lived compute environments that developers and AI agents can create, use, snapshot, pause, and destroy through an API.

The user-facing model is intentionally small:

const sandbox = await client.sandboxes.create({
  template: "python-3.12",
  egress: {
    mode: "deny_all",
  },
});

const result = await sandbox.exec({
  command: "python -c 'print(\"still contained\")'",
});

console.log(result.stdout);

await sandbox.destroy();

Underneath that API is Nomad, gVisor (https://gvisor.dev/), custom CNI networks, RabbitMQ workers, CSI volumes, several reapers, and a Go service called heracle.

This post is about what is underneath.

The shape of the system

The architecture is easiest to understand as two planes.

Our main API, which we call core, owns the control-plane concerns: authentication, quotas, billing, the sandbox record, regions, templates, and the public REST API.

heracle owns the machinery. It is a Go gRPC service with access to Nomad and the systems around it. It resolves images, creates jobs, waits for allocations, executes commands, applies network policies, manages persistent volumes, creates snapshots, and destroys whatever it created.

The path from API request to running code looks like this:

core never constructs a Nomad job or talks directly to Docker, CNI, or CSI. It sends a provisioning request to heracle and expects either a running allocation or an error it can unwind.

That boundary has been useful. Billing code cannot accidentally learn how to attach a CNI network, while the sandbox orchestrator does not need to understand subscription plans. Both systems have enough problems already.

Provisioning is synchronous from the caller’s point of view. The create request waits until the sandbox allocation is running, then returns a handle that can immediately execute commands and transfer files.

With a warm image cache, creation currently takes approximately 1.5–2 seconds from API request to runnable allocation. Since images are already available on the nodes, there is no additional latency from cold pulls.

Lifecycle operations that may take longer, like pause, resume, snapshot completion, and expiry, also use asynchronous events through RabbitMQ.

Containers were too weak. VMs were more than we needed.

The central decision was the isolation boundary.

A normal container uses namespaces, cgroups, capabilities, and seccomp to isolate a process while still letting that process talk to the host’s Linux kernel. That is a perfectly good model when you have some trust in the software.

We have no such trust.

The sandbox may be running a package from a public registry, a shell command produced by a model, or a repository supplied by an end user. Plain runc would put all of that code one kernel vulnerability away from the host.

At the other end of the spectrum, we could give every sandbox a VM.

Firecracker was the obvious candidate. It is designed for lightweight microVMs and gives a stronger boundary than a shared-kernel container. We considered it seriously.

We did not choose it.

A Firecracker platform would have meant operating another virtualization layer: kernel images, root filesystems, VM networking, guest lifecycle, snapshot compatibility, and a tighter integration between the scheduler and the virtual machine monitor.

None of that is impossible. It is also not free.

We already had a mature container path through Nomad and Docker. We wanted stronger isolation without replacing the execution model beneath the rest of our platform.

gVisor fit that gap.

Instead of allowing the sandbox process to send most syscalls directly to the host kernel, gVisor provides a userspace kernel that implements much of the Linux syscall surface itself.

Conceptually:

untrusted process
        ↓
 gVisor userspace kernel
        ↓
    host kernel

Each sandbox runs through runsc, gVisor’s OCI runtime. That additional boundary reduces how much of the host kernel is exposed to untrusted code while preserving the parts of containers we wanted: OCI images, Docker tooling, Nomad’s Docker driver, high workload density, and startup behavior much closer to a container than a conventional VM.

This does not make gVisor identical to a VM.

A hardware-backed microVM still provides a stronger and better-understood isolation boundary. gVisor also has compatibility gaps. Some syscalls behave differently, some kernel features are unavailable, and anything expecting to observe a normal Linux network stack may be surprised.

We chose gVisor’s compatibility and observability problems over building a microVM platform.

What Nomad actually runs

gVisor is not our only isolation layer.

Sandbox jobs run in a dedicated Nomad node pool. Those nodes do not run ordinary Brimble customer deployments, so a sandbox never shares a machine with someone’s production API or database.

If one of our assumptions fails, the next boundary is a machine reserved for hostile workloads.

The generated job is larger than this, but the important parts look roughly like the following:

job "sandbox-<sandbox-id>" {
  type      = "service"
  namespace = "sandboxes"
  node_pool = "sandbox"

  group "sbx" {
    count = 1

    restart {
      attempts = 0
      mode     = "fail"
    }

    reschedule {
      attempts = 0
    }

    network {
      mode = "cni/sandbox-no-egress"
    }

    task "run" {
      driver = "docker"

      config {
        image      = "brimble/sandbox-python:<version>"
        runtime    = "runsc"
        command    = "/bin/sh"
        args       = ["-c", "tail -f /dev/null"]
        force_pull = false
        pids_limit = 4096
        cap_drop   = ["ALL"]
        cap_add    = ["ALL"]
      }

      resources {
        cpu        = <cpu-mhz>
        memory     = <memory-mb>
        memory_max = <memory-mb-plus-headroom>
      }
    }
  }
}

Notice cap_drop = ["ALL"] sitting right next to cap_add = ["ALL"]. On a normal container that pairing would be pointless: you'd be dropping every Linux capability and then handing all of them straight back. Under runsc, it means something different: the syscalls those capabilities gate never reach the host kernel to begin with, because gVisor's userspace kernel intercepts and re-implements them first. Isolation lives in the boundary between the sandbox and gVisor, not in which capabilities the container happens to hold, so a sandbox can run tools that expect a normal, fully-privileged Linux process (Docker-in-Docker, for instance) without that privilege ever meaning what it would mean on a bare runc container. It's the detail that makes the rest of this section's argument concrete rather than theoretical.

We disable automatic restart and rescheduling deliberately.

If a sandbox process dies, Nomad does not quietly recreate it and pretend nothing happened. That would be especially confusing for an interactive environment whose local filesystem may be ephemeral.

A fresh allocation with missing state is not recovery. It is a new computer wearing the old computer’s name tag.

Heracle waits until the Nomad allocation and its run task both enter the running state. The container stays alive with tail -f /dev/null; actual work is invoked later through Nomad’s allocation exec API, proxied by Heracle’s gRPC methods.

There is no HTTP server inside the sandbox to health-check. “Ready” means the isolated execution environment exists and can accept commands.

Prebuilt templates

An empty Ubuntu container is technically a development environment. It is also an excellent way to spend the first minute of every agent session installing the same packages.

We provide prebuilt templates for common workloads:

python-3.12
node-24
bun-1
ubuntu-24
claude-code
codex

Templates are ordinary OCI images with the expected runtime and common tools already present. The base images stay lean, the language runtimes bundle their toolchains, and the agent images ship an entire coding assistant, trading size for capability.

The agent templates are the interesting ones. Images like claude-code and codex come with the coding agent already installed and wired up to run unattended, so you can start a sandbox and have Claude Code or Codex working inside it by default. There is no install step and no bespoke Dockerfile to maintain. You create the sandbox, hand it a task or a repository, and the agent runs.

That shape is already useful for people building agent workflows on top of Brimble. Here is a Codex sandbox running on Brimble:

Paul Adams is using Brimble sandboxes inside TestSnag, a product worth checking out if you build or test with AI agents:

You are not limited to what a template ships, either. Because runsc implements enough of the Linux surface for a fully-privileged process, a sandbox can run Docker itself. Docker-in-Docker works inside gVisor (see gVisor's own guide), so an agent that needs to build and run its own containers can do so without ever touching the host Docker daemon.

Every template image is baked onto the sandbox node pool ahead of time, so creating a sandbox never pulls an image. The container starts from something already on disk, which is what keeps creation in the 1.5–2s range regardless of how large the template is.

The safest network is no network

Filesystem isolation gets most of the attention in sandbox products. Networking deserves at least as much.

Code that cannot escape the runtime may still be able to scan private services, probe metadata endpoints, call arbitrary third-party systems, or exfiltrate whatever data the user gave it.

For untrusted execution, our safe network profile is deny_all.

Nomad attaches those sandboxes to a custom CNI network:

cni/sandbox-no-egress

The corresponding sandbox-no-egress conflist gives the allocation no outbound route and no DNS configuration. The Brimble API can still reach the sandbox through Nomad’s exec and file APIs, but processes inside the sandbox cannot call the internet or our internal network.

We also rewrite well-known cloud metadata endpoints back to localhost inside the container. Defense in depth is mostly refusing to trust your previous defense.

Not every workload can live without a network. Agents may need to install packages, clone repositories, or call approved APIs, so the platform supports three egress modes:

open
restricted
deny_all

restricted uses a separate CNI bridge and a per-node egress worker.

The flow looks like this:

Once Nomad places the allocation, Heracle publishes an egress:apply message through RabbitMQ, routed using the Nomad node ID. Only the worker on that physical node receives it.

The worker resolves the sandbox IP and installs nftables rules for the approved hosts, addresses, or CIDRs.

The critical behavior is fail-closed.

If a sandbox requests restricted egress and Heracle cannot deliver the enforcement message, provisioning fails and the allocation is cleaned up. We do not return a sandbox marked “restricted” while quietly leaving it open.

That would be less of a security feature and more of a UI theme.

Executing code without exposing the sandbox

Sandboxes do not join our service mesh. They do not register a public service in Consul, and Caddy does not create a route for them. There is no inbound address that an end user connects to directly.

Heracle exposes operations for commands, streaming output, file upload, file download, lifecycle changes, snapshots, and egress updates.

The same SDK surface used during creation is used for execution:

const sandbox = await client.sandboxes.create({
  template: "node-24",
  egress: {
    mode: "restricted",
    allow: ["registry.npmjs.org"],
  },
});

const execution = await sandbox.exec({
  command: "npm install && npm test",
  stream: true,
});

for await (const frame of execution.output) {
  if (frame.stream === "stdout") {
    process.stdout.write(frame.data);
  } else {
    process.stderr.write(frame.data);
  }
}

const result = await execution.result();

console.log(result.exitCode);

await sandbox.destroy();

The SDK presents a sandbox handle rather than making developers repeatedly pass allocation IDs or understand Nomad. That abstraction also keeps Nomad credentials far away from the caller.

Heracle is the only component allowed to perform the dangerous operations, and its internal gRPC endpoint authenticates callers before touching the scheduler.

A sandbox API should feel like renting a shell, not like joining the platform engineering team.

Ephemeral until it shouldn't be

By default, a sandbox is disposable. Its local allocation disk is not sticky and does not migrate. Destroy the sandbox and the state disappears.

That is the correct default for one-shot execution. It is less useful for an agent that installs dependencies, checks out a repository, spends an hour modifying files, and needs to continue tomorrow.

We support two persistence models.

Persistent workspaces use JuiceFS-backed CSI volumes. The volume is independent of the sandbox allocation, so it can survive pause, destruction, and placement on another physical node.

Snapshots preserve the sandbox filesystem as a restorable image.

Snapshot work happens asynchronously. Heracle sends work to a RabbitMQ-backed snapshot worker running on the sandbox node. The worker commits the container filesystem, pushes the resulting image to our registry, reports completion, and removes the local committed image so snapshotting does not slowly fill every host disk.

For a typical sandbox, snapshot creation completes asynchronously in the background. Restoring from that snapshot is quick with a warm image cache, and longer when the image must be pulled.

Restoring creates a new sandbox using the committed snapshot image as an image override.

This is not process checkpointing. We do not preserve RAM, running processes, sockets, or a Python interpreter halfway through a function.

A snapshot captures filesystem state. Pause and resume similarly create a fresh allocation; persistent files return, process state does not. “Resume” is one of those words that can promise more than an implementation actually does.

For agent workflows, filesystem persistence gets us most of the value without operating CRIU across arbitrary runtimes and kernel versions.

We know where that road leads. We have chosen not to drive down it yet.

Failure spans multiple systems

Provisioning crosses an HTTP request, a database write, a gRPC call, a Nomad job registration, an allocation, and sometimes a RabbitMQ message.

There is no transaction around all of that.

Core writes the sandbox record as STARTING before it asks Heracle to provision it. On success, it stores the returned Nomad job and allocation IDs and marks the sandbox ready.

On failure, it runs a compensating cleanup: destroy any partially created job, detach an attached volume, cancel scheduled snapshot work, notify the client, and remove the incomplete sandbox record.

But requests can die between steps, so both sides run reapers.

Core cleans up sandboxes stuck in STARTING and expires old snapshots. Heracle independently finds expired sandboxes, failed allocations, completed one-shot jobs, and environments paused for too long.

This is redundant on purpose. When three systems each have a different opinion about whether a computer exists, a second cleanup loop is cheaper than optimism.

gVisor made observability weird

Our most useful lesson was that isolation changes what the host can observe.

We use Grafana Beyla elsewhere in the platform for application instrumentation. Beyla relies on eBPF visibility from the host kernel, including syscall and network activity. That works well for ordinary containers.

gVisor changes that assumption.

One of the reasons we use gVisor is that it intercepts a large part of the Linux syscall surface and handles it in runsc instead of letting the sandboxed process talk directly to the host kernel. That is the isolation boundary doing its job.

It also means host-level eBPF tooling does not see the workload the same way it would see a normal container. Some of the syscall and network activity Beyla expects to trace is now behind runsc, or appears as runtime activity rather than application activity.

Beyla could see the runtime process, but it could not provide the workload-level visibility we expected inside the sandbox.

Nothing was broken. Our mental model was.

We had treated observability as something the host could always recover from the kernel. gVisor deliberately reduces the workload’s direct interaction with that kernel. The isolation feature and the instrumentation technique were pulling in opposite directions.

We now treat sandbox monitoring differently: runtime and allocation health from Nomad, host and runsc metrics from the nodes, API-level command timings from Heracle, and explicit output from the exec path.

We may add in-sandbox instrumentation for workloads that need deeper telemetry. We do not currently have an eBPF dashboard, and while we may reintroduce Beyla or similar tooling in the future, for untrusted code we would rather accept a visibility gap than weaken the boundary.

Bare metal makes this cheaper and harder

Owning the machines changes the economics.

We are not paying the per-VM premium of a hyperscaler every time an agent needs to run a three-second Python script. Nomad can bin-pack many isolated sandboxes onto dedicated hosts, and gVisor gives us better density than one full VM per execution environment.

On our current sandbox nodes, we can schedule many isolated sandboxes at the standard 0.5 vCPU / 512 MB profile before memory becomes the limiting resource. The equivalent per-workload VM model would reserve substantially more baseline memory, although we are deliberately not publishing a synthetic VM-versus-gVisor benchmark until we have one we trust.

The less photogenic part is that we own the rest: image caching, CNI and nftables policy, scheduler capacity, runsc upgrades, CSI volumes, snapshot workers, and the cleanup paths when those systems disagree.

There is no cloud support ticket that ends with somebody else replacing the host.

Bare metal is cheaper because the provider hands you more of the problem. For Brimble, that tradeoff makes sense. We already operate the underlying platform, and sandboxes benefit directly from predictable capacity and high workload density.

But “we run this cheaply on bare metal” should never be confused with “the machinery is free.”

The work is worth it because the product surface stays small: create a sandbox, run code, keep it contained, and throw it away when you are done.

That simplicity is the point. The user should not have to care about Nomad allocations, CNI bridges, RabbitMQ routing keys, nftables rules, CSI volumes, or runsc. They should get an environment that starts quickly, behaves like a useful machine, and fails closed when the platform cannot enforce the promises it made.

Most of the system exists to make that promise boring.

If you want to get started, the Brimble Sandboxes documentation walks through the API, templates, egress modes, persistence, and lifecycle operations.

Written by

Ilerioluwa David
Ilerioluwa David@pipe_dev