An AI agent analysing data often needs to run code: calculate a total, inspect a file, or test an idea. That generated code needs a sandbox.
I built plimsoll, an Apache-2.0 code execution sandbox, with an integration for Trigger.dev. It runs Python and JavaScript separately from your task, and supports sessions that keep variables and files available between calls.
Plimsoll is pre-1.0. There is a working starter you can deploy, including a fixed Python task that checks the connection without calling an AI model.
What runs where?
Your Trigger.dev task sends code to plimsoll. Plimsoll executes it in a sandbox and returns the output, along with the isolation tier used.
For self-hosted Trigger.dev, a Docker Compose recipe runs the sandbox service beside your worker on your infrastructure. Task containers reach it over a private Docker network.
The self-hosted recipe was tested on one host with Trigger.dev v4.7.2. Its README describes the setup and limitations:
https://github.com/plimsollmark/plimsoll-trigger-starter/tree/main/self-hosted
Start with the working example
Clone the starter and install its dependencies using Node.js 22.18 or later:
git clone https://github.com/plimsollmark/plimsoll-trigger-starter.git
cd plimsoll-trigger-starter
npm ci
npm run typecheck
Before deploying, follow the starter’s instructions to configure a plimsoll service your worker can reach:
https://github.com/plimsollmark/plimsoll-trigger-starter#prepare-plimsolld
For self-hosted workers, use the Compose instructions linked above instead.
Set PLIMSOLL_URL and the secret PLIMSOLL_TOKEN in your Trigger.dev project’s production environment. Set TRIGGER_PROJECT_REF for the deployment CLI.
Then run:
npm run deploy -- --dry-run
npm run deploy
Your project should list two tasks: code-chat and deployed-cell-trial.
Check that Python remembers its variables
The verification task sends two separate Python calls to one sandbox session.
The first creates a list and returns its length:
numbers = [2, 3, 5]
len(numbers)
The second uses the same list:
sum(numbers)
The expected results are 3 and 10.
The second call depends on a variable created by the first. Like two cells in a notebook, they share a running Python interpreter.
Supply your Trigger.dev production API key as TRIGGER_SECRET_KEY through your shell or secret manager, then run:
npm run trial
For self-hosted Trigger.dev, also set TRIGGER_API_URL to your own Trigger.dev address before running the command.
The task checks the outputs, interpreter reuse, and minimum isolation level. It closes the sandbox when finished, including when a check fails.
A successful run reports interpreterReused: true and, with the default isolation requirement, kernel for both calls. A broken session or insufficient isolation makes the task fail.
This trial makes no AI call. Infrastructure charges can still apply.
Give the agent a sandbox tool
The included chat agent exposes an executeCode tool and follows Trigger.dev’s documented sandbox lifecycle:
- Warm the sandbox in
onTurnStart. - Reuse it for tool calls while the run remains active.
- Close it in
onChatSuspendoronComplete.
Variables and files remain available while the session exists. Once it closes, a later session starts fresh.
This lets an agent load data, inspect it, and calculate results across several calls without rebuilding its workspace every time.
Running the chat agent requires its model credentials and incurs model usage charges. The fixed verification task does not.
Trigger.dev’s sandbox pattern is documented here:
https://trigger.dev/docs/ai-chat/patterns/code-sandbox
Know what the sandbox protects
The self-hosted Compose setup selects gVisor when available and runc otherwise.
gVisor adds a separate kernel boundary for sandboxed code. Ordinary runc containers share the host kernel and provide a weaker boundary.
The starter tasks require the kernel isolation tier by default. If the service only provides the container tier, the tasks refuse to execute unless you explicitly lower that requirement. Keep the stronger requirement for hostile code.
Plimsoll’s daemon is trusted infrastructure: it holds the Docker socket, which gives it powerful access to the host. The self-hosted README also documents credential logging observed with Trigger.dev v4.7.2 and other deployment limits.
Try it
The starter includes the chat integration, a deployed verification task, and the self-hosted Compose recipe:
https://github.com/plimsollmark/plimsoll-trigger-starter
I maintain plimsoll, and I’d like feedback from people using it with Trigger.dev. If you try it, tell me what broke, especially during setup or when connecting it to an existing worker.
Top comments (0)