plimsoll

module
v0.21.0 Latest Latest
Warning

This package is not in the latest version of its module.

Go to latest
Published: Oct 9, 2026 License: Apache-2.0

README

plimsoll

plimsoll runs untrusted code, such as code an AI agent wrote, inside a sandbox, and every result says how strong that sandbox was. A request can demand a minimum strength; if the sandbox is weaker, nothing runs.

audit gvisor Go Reference Go version release license

The name comes from the Plimsoll line, the load limit painted on the outside of a ship's hull where anyone can check it. plimsoll does the same for the isolation tier: how strong the wall around a run is, stated as one of four levels, weakest first: process, container, kernel, vm. The tier is a value on every result, not a sentence in a datasheet. A request can set a floor, the weakest tier it will accept, and the official Go client checks the tier again when the result comes back.

The same machinery lets plimsoll run an exam on code an agent wrote, not only contain it. A trial runner built into the simulation image runs the agent's program against scenarios that program cannot read: each one reaches the runner in a file deleted before the program starts. Every run returns SHA-256 hashes of the request as sent and the result as returned, which the official clients recompute (how). A separate program outside the daemon, plimsoll-attest, signs the checked results, refuses a chain of calls with one missing, and replays a run. The exam itself (the simulated system, the scenarios, the pass mark) stays yours: plimsoll supplies the boundary and the evidence.

See one real run

An AI agent wrote a controller: a program that reads a system's state at every tick (here, one 10 ms step of the simulation) and decides how to push it. plimsoll ran it against a simulated cart-pole, a cart on a rail with a pole hinged on top, which the controller keeps upright by pushing the cart left and right. The simulator is compiled to WebAssembly, a portable bytecode that runs inside a host program, and built into the sandbox image. The controller was the only file the caller sent, through the ordinary project API, and it ran at the container tier.

The cart-pole is the standard teaching problem in control engineering. The runner, a program built into the image that runs the controller against the simulator, records the run's trajectory, every state at every tick, and hashes it into a fingerprint: a SHA-256 hash, so equal fingerprints mean identical numbers. Run the accepted controller twice and the two fingerprints match. The agent's first draft had two gains (multipliers in its formula) with the wrong sign: it drops the pole at 5.96 seconds, and its fingerprint differs.

Pole angle and cart position over twenty seconds, for the accepted controller and for the draft that fell

Open the live run report ↗ to replay both runs in the browser, read the controller the agent wrote, and see what the page deliberately does not claim. One execution of the example makes three sandbox runs: the accepted controller twice and the draft once. Reproduce them with make docker-images && go run ./examples/oracle.

The two things it does

It reports the wall. Every result states the isolation tier the run actually got. Every request can set a floor, and a request whose floor the daemon cannot meet is refused before any code runs: ErrInsufficientIsolation means nothing ran.

res, err := provider.Sandbox.RunJavaScript(ctx, sandbox.Request{
    Code:             "console.log(6 * 7)",
    MinimumIsolation: sandbox.IsolationVM, // ErrInsufficientIsolation means nothing ran
})
// res.Isolation reports the boundary that actually ran.

Every refusal that ran nothing is marked as one, with a reason (request, permission, protocol, unsupported, isolation, environment, capacity). sandbox.NotDispatchedReason(err) reads the mark, whether the provider (the backend that runs the code) is in your own process or behind the daemon. An error without the mark may have come after the code started, so it is never safe to retry automatically. A mark counts only on an answer that carries the client's request ID back, so a proxy that serves one request's answer for another cannot make a call that ran read "nothing ran" (what comes back, and what it means).

The tier is evidence the daemon collected: its configuration and its own startup checks. It is not attestation, cryptographic proof from the hardware of what software is running, and no provider here offers that. docs/isolation-tiers.md lists what each tier rests on.

It lets agent-written code call your API without ever holding your credential, your API's address, or open network access. A grant gives one run permission to call listed routes of one API. The code calls through a small client plimsoll puts in the sandbox, and the broker, the part of plimsoll outside the sandbox that makes the real call, checks the caller and the exact route, then attaches the credential itself. The credential is minted per run: plimsoll asks for it once per run, so it can be a fresh short-lived token, and it never enters the sandbox. Code that skips the client and calls out by hand gets no further, because the broker, not the client, enforces the rules. See docs/capability-grants.md.

Use it from an agent framework

These agent frameworks can hand their model's code to plimsoll through a tool or code executor of their own (Trigger.dev, one of them, is a hosted TypeScript job runner). Each row names what to install, the guide, and a page that replays recorded calls from that framework's tool to a sandbox and back, refusals included.

Framework Package or extra Guide See a run
Vercel AI SDK @plimsollmark/client/ai-sdk Vercel AI SDK guide (INTERNAL · documentation →) Vercel AI SDK page (INTERNAL · plimsoll site →)
Google ADK plimsoll-client[adk] Google ADK guide (INTERNAL · documentation →) Google ADK page (INTERNAL · plimsoll site →)
Agno plimsoll-client[agno] Python client guide, Agno and CrewAI tools (INTERNAL · documentation →) Agno page (INTERNAL · plimsoll site →)
CrewAI plimsoll-client[crewai] Python client guide, Agno and CrewAI tools (INTERNAL · documentation →) CrewAI page (INTERNAL · plimsoll site →)
Trigger.dev @plimsollmark/client/trigger Trigger.dev guide (INTERNAL · documentation →) Trigger.dev page (INTERNAL · plimsoll site →)
Mastra @plimsollmark/client/mastra TypeScript client guide, Mastra (INTERNAL · documentation →) Mastra page (INTERNAL · plimsoll site →)

Give a Trigger.dev agent a code tool

The TypeScript add-on gives an executeCode tool to a chat agent on Trigger.dev. On a daemon with a session, one sandbox kept for several calls, each call is a cell, code run inside the interpreter kept for that conversation: variables and files can survive the next call. Each tool result says whether they can survive, whether this call started fresh, the isolation tier it ran behind, and the SHA-256 of the run record, a receipt for what was sent and returned that the client checked before returning. The caller sets a minimum tier, so a weaker provider refuses before running code.

See the feature and deployment guide for the exact guarantees and production floor. The deployable Trigger.dev starter has a chat agent and a task that proves two cells share one interpreter through a real daemon, without making an AI model call.

Run something in one minute

The example needs no daemon, no docker, no credentials and no network: it uses the wasm provider, which runs JavaScript on QuickJS, a small JavaScript engine compiled to WebAssembly, inside your own process. Cloning and fetching Go modules do use the network.

git clone https://github.com/plimsollmark/plimsoll && cd plimsoll
go run ./examples/minimal
provider   wasm
isolation  process
exit code  0
timed out  false
truncated  stdout=false stderr=false
duration   304ms
stdout     {"Engineering":59000000,"Sales":20300000,"Operations":9800000}

The isolation line comes from the run's result, not from the example's own text. process means there is no operating-system wall at all: do not use it for hostile code. The same snippet runs at the kernel tier once gVisor is installed. gVisor is a layer between a container and your machine's kernel that handles the container's requests to the operating system itself, so the code never talks to your kernel directly. sudo ./docker/install-gvisor.sh installs a pinned gVisor release (one exact version, changed only by editing this repository) and registers its runtime, runsc, with docker. Nothing else about the program changes:

make docker-images   # builds the project image (plimsoll/sandbox) the docker provider checks for
SANDBOX_PROVIDER=docker SANDBOX_DOCKER_RUNTIME=runsc go run ./examples/minimal

With SANDBOX_PROVIDER unset, plimsoll uses the Disabled provider, which runs nothing, so code execution is never on by accident.

Next: docs/getting-started.md walks through the same pieces on one machine: start the daemon, create a caller credential, write your own client, watch a floor refuse a run, then move the daemon from process to kernel without changing the client. It also covers embedding the Go package instead of running the daemon.

Where the pieces sit

flowchart LR
  A["Agent or MCP gateway (MCP: how an AI app offers tools to a model)"] -->|"Run: one payload, a protocol number, minimum_isolation"| AU

  subgraph D["plimsolld"]
    AU["authenticate the caller"] --> FL["compare the floor with current provider evidence"]
    FL --> DI["dispatch"]
    BR["broker: holds the minted credential, matches the exact route, counts every call"]
  end

  DI --> W["wasm: QuickJS on wazero (process tier)"]
  DI --> K["docker with runsc (kernel tier)"]
  DI --> V["e2b Firecracker (VM tier)"]

  W -. "host.get / host.post" .-> BR
  K -. "per-run Unix socket" .-> BR
  V -. "authenticated guard" .-> BR
  BR ==> |"your credential, attached host-side"| API["Your API"]

The code inside the sandbox, the guest, never holds the credential and never reaches the network directly. A run with no grant has no network at all.

Isolation tiers

plimsolld is the plimsoll server. Each one runs exactly one provider, chosen by SANDBOX_PROVIDER:

Provider Boundary Tier reported Use
wasm QuickJS on wazero (a WebAssembly runtime written in Go), inside plimsolld itself process Fast local development. An escape (a bug that lets code out of its sandbox) in the engine lands inside your daemon.
docker with runc, docker's default runtime a container sharing your machine's kernel container Self-hosting when a kernel bug is not one of the attacks you plan for. With a project image it keeps a session: one container for many calls, with a JavaScript interpreter (and a Python one with plimsoll/sandbox-python as the project image) whose variables survive between calls.
docker with runsc gVisor, once the startup checks confirm docker has runsc registered kernel Hostile code on your own machines. Keeps sessions too.
e2b a microVM (a small virtual machine made for one run, then destroyed) from E2B, a hosted service that runs them on Firecracker, AWS's open-source VM monitor vm Hostile code, on E2B's machines rather than yours; billed per run. Keeps sessions too, billed while the microVM runs; a suspend pauses it, which keeps the files but not an interpreter's variables.
dockercloud a microVM from Docker Cloud Sandboxes, Docker's hosted sandbox service vm Hostile code, on Docker's machines; billed per run. It speaks two APIs, chosen by SANDBOX_DOCKERCLOUD_API, and both passed the live suite on 2026-10-04: the one Docker released before launch, no longer documented, which is the default because only it has grants and image evidence, and the REST API Docker documents, kept as a backup until it covers those (no grants, refused under PLIMSOLL_HARDENED=1, runs of at most 270 s). The account's network policy must be deny-all (no connection unless a rule allows it), and every run checks that. Grants work through the same guard as E2B (an address on the plimsoll server, the only place the microVM may connect to) when SANDBOX_DOCKERCLOUD_GUARD_URL is set. Unlike E2B, the guest holds its own run's short-lived credential for the guard, and the one network rule is set through a Docker call outside its published API.
openshell a sandbox from OpenShell, NVIDIA's agent sandbox runtime, created by its gateway server on docker container Agent platforms that already run an OpenShell gateway. Each run gets its own sandbox with no network; plimsoll reads its settings back, refuses to run on any difference, and deletes it afterwards. Tested against a v0.1.2 gateway on 2026-09-28. Grants reach the broker through a relay inside the sandbox that plimsoll connects to from outside, so the sandbox needs no network rules. Keeps sessions too.
unset nothing runs n/a The default.

Each tier rests on the daemon's configuration, what the provider reports, and a real test run at startup; none is attestation. docs/isolation-tiers.md lists the evidence for each tier and shows how a request sets its floor. A session trades the fresh sandbox of every call for speed and kept state: code an earlier call ran can change what later calls see, and a call with API access needs a grant that allows sessions (docs/sessions.md). A session is one trust domain, and plimsoll ties it to the authenticated caller, not to that caller's customers: a service that runs many customers' code through one credential must keep each customer in a session of their own (who may share a session). The docker provider can also apply the shipped seccomp profile, a list of the only system calls (requests to the kernel, such as opening a file) its containers may make (docs/seccomp.md).

More real runs

Page What it shows
Simulation replay pages, all eight simulators ↗ The eight simulators built into the simulation image for controller runs: shower, buck converter (a power supply that steps a voltage down by switching it on and off), ship heading, black hole orbit, relativistic rocket, satellite clock, double slit, and cart-pole swing-up (the image also carries the models the module-run tests use). Each page runs a failing and a passing hand-written controller through the sandbox and replays both trajectories with their fingerprints.
A controller in C, compiled in the sandbox ↗ The same cart-pole simulator, with a swing-up controller (it swings the pole up from hanging, then balances it) written in C and compiled to WebAssembly by the run's own first step, so the simulator and the controller are both WebAssembly. The run report compares it with the JavaScript version tick by tick. The two languages' cos functions disagree in the last bit on about one input in a hundred, so in three of the four scenarios one or two force values differ, by a few representable doubles (at most 14). The fingerprint catches that difference, and the motion itself is identical to the bit. Source and caveats in its README; reproduce with make docker-images && go run ./examples/wasm-controller.
A buck converter controller in C ↗ A power supply's control law (the formula its controller applies at each tick) written in C, the language converter firmware is written in. It is compiled to WebAssembly in the sandbox and run against a simulated converter, whose output it must hold at 5 V while the load changes suddenly. It calls no library function, so its trajectory equals the JavaScript version's byte for byte in all four scenarios; the page charts one of them tick by tick. Source and caveats in its README; reproduce with make docker-images && go run ./examples/wasm-buck.
Same run, different sandboxes ↗ The cart-pole run from the top of this page, on local runc and gVisor, an OpenShell gateway, E2B and Docker Cloud Sandboxes: three isolation tiers, two Node versions, and one fingerprint, the same one the first report published. Reproduce the configured rows with go run ./examples/providers; the cloud rows are billed.
One sandbox, five calls ↗ A session on an OpenShell sandbox: a failing test, a patch, the test passing without the files being sent again, and a leftover process that is gone by the next call. Each call's run record (the daemon's statement of what was sent, what came back and where it ran) is signed and linked to the previous one. The verifier accepts the signed set, and refuses it when one call is dropped or one byte is changed. Reproduce with go run ./examples/sessions and a gateway.
The efficiency advisor's report ↗ The advisor reads the API calls a run made and points out wasteful patterns. One measured run: the same question asked of an API as 13 calls, then as 1, and the advisor's finding that names the route that answers it in one call.
Twelve interactive lessons ↗ How a run is executed, the providers, the API broker, and connecting an AI agent. Static pages: no network calls, no analytics, no third-party scripts.

Status, plainly

  • Pre-1.0, single author, no external users yet. Breaking changes land without a deprecation path, on purpose.
  • No third-party security audit has ever been performed, and the author's own self-review ledgers are not published either. SECURITY.md says what exists, what it is worth, and what you can check yourself instead.
  • CI runs the gate, in the open, and each check means what it ran. The audit workflow has three jobs. audit runs plain make audit on every push to main and every pull request: build, vet, race tests, lint, buf lint, a generated-code drift check, and govulncheck. After audit, clients runs make clients-suite: npm ci in clients/typescript and examples/trigger-chat, then the Python and TypeScript Go test packages in required mode. A missing runtime, dependency or skipped Go test fails that job. audit-docker also follows audit and runs make docker-suite with the images prepared, in required mode: a missing daemon, a missing image or a skipped test fails the job. A green audit-docker therefore means the docker suite ran under runc with the shipped seccomp profile and proved a read-only root, sized noexec writable mounts, and the broker's refusals. The gvisor workflow runs the same suite under runsc, the kernel tier, from the pinned installer. The race detector runs in the plain audit job and both docker-suite jobs; the client job runs without it. No check exercises E2B or Docker Cloud Sandboxes: those suites drive live paid services and are deliberately never wired to a runner. See CONTRIBUTING.md.

Learn how it works

The interactive lessons are the fastest way in if you would rather read than clone. Start with Plain English for the idea before the API, or Quick start to run something. Every term these docs use has a one-sentence definition in the glossary. The lessons live in docs/trainers/ and work offline: open any file from a clone in a browser. GitHub shows .html files as source instead of rendering them, which is why the links above point at the published copy.

Documentation

Each of these answers one question, end to end.

Document Answers
docs/getting-started.md How do I build it, embed it, start it as a service with authentication, and watch a floor refuse a run?
docs/example-programs.md What do the runnable examples prove, and which should I read first?
docs/isolation-tiers.md What does each tier rest on, and how do I demand one per request?
docs/capability-grants.md How does agent code call my API without ever holding my credential?
docs/run-results.md What comes back, and when is a failure an error rather than a result?
docs/run-records.md What does each run's record state, how do I recompute it in another language, and how do I sign, verify and replay records outside the daemon?
docs/sessions.md How do I keep one sandbox for many calls, and what holds between the calls, an interpreter's variables included?
clients/python How do I call plimsolld from Python?
clients/typescript How do I call it from TypeScript, and give a Trigger.dev or Mastra agent a code tool that keeps its state?
docs/trigger-dev.md How do calls that reuse one interpreter, an isolation floor and checked run records appear together in a Trigger.dev tool, and how do I deploy a task?
docs/ai-sdk.md How do I give a Vercel AI SDK agent an executeCode tool, fresh per call or kept for one user's conversation?
docs/google-adk.md How do I make plimsoll the code executor of a Google ADK agent, with input files, output artifacts and a refusal below the floor?
docs/placement.md I run several daemons: how do I pick one per request, and when is a refusal safe to retry elsewhere?
docs/inner-loop-workflow.md How do I iterate fast locally without shipping a weak sandbox to production?
docs/efficiency-advisor.md What does the advisor see, why can what it records never include the data the code sent or received, and how do I choose what it emits?
docs/hardened-mode.md How do I make the daemon refuse to start unless every production safeguard is set?
docs/dependencies.md Which dependencies must be trusted for the sandbox to hold, and how are the tools that check them pinned?
docs/limitations.md What does this deliberately not do?
docs/dockercloud.md What does the Docker Cloud Sandboxes provider need from the operator, and what does each run check?
docs/openshell.md What does the OpenShell provider need from the operator, and what does each run check?
docs/releasing.md Why is the module path public, why do releases start at v0.2.0, and why is there no checksum exemption?
docs/callers.md How do I create, rotate and revoke caller credentials?
docs/seccomp.md and docs/gvisor.md What do the seccomp filter and gVisor enforce?
docs/guest-dependencies.md How do npm packages get into a sandbox that has no network?
docs/architecture/credential-minting.md Where does a per-run credential come from, and where does it stay?
docs/seams.md Where are the deliberate extension points?
Glossary ↗ What does this word mean? One plain sentence per term.
AGENTS.md The architecture reference: providers, the rules no change may break, and every environment variable.

License

Apache-2.0.

Directories

Path Synopsis
Package attest is the harness's half of plimsoll's run records: it signs a checked record as a DSSE envelope around an in-toto Statement v1, verifies envelopes and whole bundles (signatures, digests, and each session's chain), and replays stored requests to compare results.
Package attest is the harness's half of plimsoll's run records: it signs a checked record as a DSSE envelope around an in-toto Statement v1, verifies envelopes and whole bundles (signatures, digests, and each session's chain), and replays stored requests to compare results.
Package client dials a remote plimsolld and adapts it to sandbox.Sandbox, so isolated calls can swap providers behind the same interface.
Package client dials a remote plimsolld and adapts it to sandbox.Sandbox, so isolated calls can swap providers behind the same interface.
cmd
plimsoll-attest command
Command plimsoll-attest is a harness for plimsoll's run records, run outside the daemon: it makes a signing key, runs a request and signs its checked record into a bundle, verifies a bundle (signatures, digests, the bundle's chain of links, session chains, and with -expect the caller's own list of its calls), and replays a bundle's single runs to compare results.
Command plimsoll-attest is a harness for plimsoll's run records, run outside the daemon: it makes a signing key, runs a request and signs its checked record into a bundle, verifies a bundle (signatures, digests, the bundle's chain of links, session chains, and with -expect the caller's own list of its calls), and replays a bundle's single runs to compare results.
plimsoll-clients command
Command plimsoll-clients manages caller credentials in an offline registry.
Command plimsoll-clients manages caller credentials in an offline registry.
plimsoll-specgen command
Command plimsoll-specgen derives a host-API grant description from an OpenAPI 3.x document (Sluice "spec-import").
Command plimsoll-specgen derives a host-API grant description from an OpenAPI 3.x document (Sluice "spec-import").
plimsolld command
Command plimsolld serves the plimsoll SandboxService over Connect (h2c, so Connect, gRPC, and gRPC-Web clients all work).
Command plimsolld serves the plimsoll SandboxService over Connect (h2c, so Connect, gRPC, and gRPC-Web clients all work).
prospector-report command
Command prospector-report renders plimsoll's exported audit stream into a single self-contained HTML report (Prospector Phase 4) — the no-Grafana companion to the dashboard.
Command prospector-report renders plimsoll's exported audit stream into a single self-contained HTML report (Prospector Phase 4) — the no-Grafana companion to the dashboard.
examples
advisor command
Command advisor shows the efficiency advisor end to end: the same question asked of a small inventory API twice, first the way an agent tends to write it (list the ids, then one call per item, then add the column up in its own code) and then the way the advice that came back suggests.
Command advisor shows the efficiency advisor end to end: the same question asked of a small inventory API twice, first the way an agent tends to write it (list the ids, then one call per item, then add the column up in its own code) and then the way the advice that came back suggests.
daemon command
Command daemon runs the whole service path end to end: it starts plimsolld with a real multi-client auth file, calls it with the official Go client, and shows what the server refuses.
Command daemon runs the whole service path end to end: it starts plimsolld with a real multi-client auth file, calls it with the official Go client, and shows what the server refuses.
grant command
Command grant demonstrates the capability model: agent-authored code reaching a real HTTP API through an injected client, over a route allowlist the code cannot widen, with a credential it never sees.
Command grant demonstrates the capability model: agent-authored code reaching a real HTTP API through an injected client, over a route allowlist the code cannot widen, with a credential it never sees.
internal/daemonproc
Package daemonproc builds plimsolld from source and runs it on a loopback port for the example programs, so an example exercises the real daemon, not a copy of its wiring.
Package daemonproc builds plimsolld from source and runs it on a loopback port for the example programs, so an example exercises the real daemon, not a copy of its wiring.
internal/wasmshim
Package wasmshim holds the Node shim that runs a controller compiled to WebAssembly under the runner: controller.js, sent as a project file beside controller.wasm.
Package wasmshim holds the Node shim that runs a controller compiled to WebAssembly under the runner: controller.js, sent as a project file beside controller.wasm.
minimal command
Command minimal runs one JavaScript snippet in a sandbox and prints what ran, where, and behind which isolation boundary.
Command minimal runs one JavaScript snippet in a sandbox and prints what ran, where, and behind which isolation boundary.
oracle command
Command oracle is the physics oracle demo: an agent writes a controller, plimsoll runs it against a compiled cart-pole plant, and the trajectory's fingerprint is the acceptance test.
Command oracle is the physics oracle demo: an agent writes a controller, plimsoll runs it against a compiled cart-pole plant, and the trajectory's fingerprint is the acceptance test.
providers command
Command providers runs the physics oracle's run on every provider this machine can reach and writes a page comparing them: the same controller, the same runner, the same plant, and the fingerprint of the trajectory from each provider.
Command providers runs the physics oracle's run on every provider this machine can reach and writes a page comparing them: the same controller, the same runner, the same plant, and the fingerprint of the trajectory from each provider.
sessions command
Command sessions runs one session on an NVIDIA OpenShell sandbox through the real daemon, with the harness signing every call, and writes a page that shows the calls, their chained records, and the verifier accepting the signed bundle and refusing it with one call dropped or one byte changed.
Command sessions runs one session on an NVIDIA OpenShell sandbox through the real daemon, with the harness signing every call, and writes a page that shows the calls, their chained records, and the verifier accepting the signed bundle and refusing it with one call dropped or one byte changed.
wasm-buck command
Command wasm-buck runs and scores a buck converter controller written in C: the project run compiles it to WebAssembly inside the sandbox, then runs it against the buck plant, which is WebAssembly too, under the same shim as examples/wasm-controller.
Command wasm-buck runs and scores a buck converter controller written in C: the project run compiles it to WebAssembly inside the sandbox, then runs it against the buck plant, which is WebAssembly too, under the same shim as examples/wasm-controller.
wasm-controller command
Command wasm-controller runs and scores a controller written in C: the project run compiles it to WebAssembly inside the sandbox, then runs it against the cart-pole plant, which is WebAssembly too.
Command wasm-controller runs and scores a controller written in C: the project run compiles it to WebAssembly inside the sandbox, then runs it against the cart-pole plant, which is WebAssembly too.
gen
go/plimsoll/v1/plimsollv1connect
Package plimsoll.v1 is the RPC surface of the standalone code-execution sandbox.
Package plimsoll.v1 is the RPC surface of the standalone code-execution sandbox.
internal
answertest
Package answertest serves the intermediary the official clients' answer binding is tested against (failure-injection review, finding 1): a proxy that forwards every request to the daemon, so the daemon runs it, and answers every request after the first with the first one's answer, headers and all.
Package answertest serves the intermediary the official clients' answer binding is tested against (failure-injection review, finding 1): a proxy that forwards every request to the daemon, so the daemon runs it, and answers every request after the first with the first one's answer, headers and all.
clientconfig
Package clientconfig defines the caller registry shared by the daemon and the offline operator CLI.
Package clientconfig defines the caller registry shared by the daemon and the offline operator CLI.
grants
Package grants resolves named host-API capability profiles for the RPC server.
Package grants resolves named host-API capability profiles for the RPC server.
insights
Package insights derives deterministic, metadata-only efficiency findings from a run's brokered host.* call trace (Prospector Phase 1).
Package insights derives deterministic, metadata-only efficiency findings from a run's brokered host.* call trace (Prospector Phase 1).
quantise
Package quantise reports the resolution at which a module run's swept parameter was actually resolved, which is not always the resolution the caller asked for.
Package quantise reports the resolution at which a module run's swept parameter was actually resolved, which is not always the resolution the caller asked for.
report
Package report renders plimsoll's exported audit stream into a single self-contained, dependency-free HTML report (Prospector Phase 4).
Package report renders plimsoll's exported audit stream into a single self-contained, dependency-free HTML report (Prospector Phase 4).
rpc
Package rpc serves the plimsoll SandboxService over Connect, wrapping the transport-agnostic sandbox package with auth, rate limiting, and request validation.
Package rpc serves the plimsoll SandboxService over Connect, wrapping the transport-agnostic sandbox package with auth, rate limiting, and request validation.
softwarewire
Package softwarewire maps the transport-neutral software admission rule to its wire form without changing its meaning.
Package softwarewire maps the transport-neutral software admission rule to its wire form without changing its meaning.
specgen
Package specgen derives a plimsoll host-API grant description from an OpenAPI 3.x document.
Package specgen derives a plimsoll host-API grant description from an OpenAPI 3.x document.
toolchain
Package toolchain exposes the single source of truth for the baked TypeScript toolchain versions (the repo-root toolchain.versions file), shared by the docker provider image and the e2b template so they cannot drift.
Package toolchain exposes the single source of truth for the baked TypeScript toolchain versions (the repo-root toolchain.versions file), shared by the docker provider image and the e2b template so they cannot drift.
unreadbody
Package unreadbody keeps an HTTP/1.x server from waiting on a request body its handler answered without reading.
Package unreadbody keeps an HTTP/1.x server from waiting on a request body its handler answered without reading.
measurements
session-latency command
Command session-latency times what an agent's code tool pays per call on each provider with sessions: a fresh run, a session's snippet call, a cell in a warm interpreter (JavaScript and Python), the first cell (which starts the interpreter), a call after an idle suspend, and the data-reload case (a fresh run that rebuilds an array every call against a cell that keeps it).
Command session-latency times what an agent's code tool pays per call on each provider with sessions: a fresh run, a session's snippet call, a cell in a warm interpreter (JavaScript and Python), the first cell (which starts the interpreter), a call after an idle suspend, and the data-reload case (a fresh run that rebuilds an array every call against a cell that keeps it).
session-latency/report command
Command report renders docs/measurements/session-latency/*.json, written by the session-latency command, as one self-contained HTML page:
Command report renders docs/measurements/session-latency/*.json, written by the session-latency command, as one self-contained HTML page:
warm-pool/hints command
Command hints measures the warm pool's language hints: how the pool's split across language sets follows the hints sessions open with, what a member of each set holds in memory, and the time from a hinted open to its first cell's answer.
Command hints measures the warm pool's language hints: how the pool's split across language sets follows the hints sessions open with, what a member of each set holds in memory, and the time from a hinted open to its first cell's answer.
warm-pool/pooled command
Command pooled measures the docker session pool (SANDBOX_SESSION_POOL): the time from asking for a session to the first cell's answer, with and without a pool member waiting, for each language; and what a waiting member holds in memory.
Command pooled measures the docker session pool (SANDBOX_SESSION_POOL): the time from asking for a session to the first cell's answer, with and without a pool member waiting, for each language; and what a waiting member holds in memory.
Package placement chooses which plimsoll daemon runs a request, and sends it there.
Package placement chooses which plimsoll daemon runs a request, and sends it there.
Package protocol holds the one number the wire protocol is versioned by.
Package protocol holds the one number the wire protocol is versioned by.
Package record computes the digests in a plimsoll run record: what a caller sent, what came back, and the record that states both with the evidence the run executed under.
Package record computes the digests in a plimsoll run record: what a caller sent, what came back, and the record that states both with the evidence the run executed under.
Package sandbox runs untrusted, agent-authored code through an explicitly tiered execution provider.
Package sandbox runs untrusted, agent-authored code through an explicitly tiered execution provider.
internal/deadline
Package deadline holds the one rule every provider uses to decide that a run timed out.
Package deadline holds the one rule every provider uses to decide that a run timed out.
internal/e2bfake/envd command
Command envd is a stand-in for E2B's envd, the agent inside an E2B sandbox, for the e2b provider's session tests: run as root inside a local container, it really starts processes, so the session conformance suite can run against the provider without E2B. It speaks the part of envd's API the provider uses, as envd does:
Command envd is a stand-in for E2B's envd, the agent inside an E2B sandbox, for the e2b provider's session tests: run as root inside a local container, it really starts processes, so the session conformance suite can run against the provider without E2B. It speaks the part of envd's API the provider uses, as envd does:
internal/lease
Package lease records what a provider asked a remote service to create and has not yet seen deleted, so orphan reconciliation never races a create.
Package lease records what a provider asked a remote service to create and has not yet seen deleted, so orphan reconciliation never races a create.
internal/runnerwire
Package runnerwire is the host side of the in-sandbox project runner protocol (docker/runner.mjs): the JSON plan a provider writes to the runner's stdin, the authenticated frame that carries the runner's report on its stdout, and the decoding and classification of that report.
Package runnerwire is the host side of the in-sandbox project runner protocol (docker/runner.mjs): the JSON plan a provider writes to the runner's stdin, the authenticated frame that carries the runner's report on its stdout, and the decoding and classification of that report.
internal/sessionkit
Package sessionkit holds what every session provider runs inside its sandbox to keep the boundary between calls: the process lister that records a fresh sandbox's own processes, the sweep that runs after every call, and the interpreters a session may keep alive between calls (interp.go).
Package sessionkit holds what every session provider runs inside its sandbox to keep the boundary between calls: the process lister that records a fresh sandbox's own processes, the sweep that runs after every call, and the interpreters a session may keep alive between calls (interp.go).
openshell
Package openshell runs plimsoll snippets and projects in NVIDIA OpenShell sandboxes, through an OpenShell gateway's gRPC API (openshell.v1, vendored at v0.1.2 under third_party/openshell and generated into gen/go/openshell).
Package openshell runs plimsoll snippets and projects in NVIDIA OpenShell sandboxes, through an OpenShell gateway's gRPC API (openshell.v1, vendored at v0.1.2 under third_party/openshell and generated into gen/go/openshell).
sessiontest
Package sessiontest is the session conformance suite: the behavior every sandbox.SessionProvider promises, checked with node snippets and project steps through the provider itself.
Package sessiontest is the session conformance suite: the behavior every sandbox.SessionProvider promises, checked with node snippets and project steps through the provider itself.
Package sandboxtest provides shared test helpers for projects that embed the plimsoll sandbox, so the cross-cutting guarantees are asserted by ONE implementation here instead of being copy-pasted into every consumer's tests.
Package sandboxtest provides shared test helpers for projects that embed the plimsoll sandbox, so the cross-cutting guarantees are asserted by ONE implementation here instead of being copy-pasted into every consumer's tests.

Jump to

Keyboard shortcuts

? : This menu
/ : Search site
f or F : Jump to
y or Y : Canonical URL