Caged runs coding agents in isolated Firecracker VMs. Every file edit, terminal command, and LLM call is recorded, costed, and scored — automatically.
No credit card · 5 sandboxes/day free · Works with Claude Code, Cursor, Aider, any MCP agent
3 active
sandboxes
$1.42
today's spend
87
trust score
Other platforms give agents a computer. Caged gives you full visibility and control over what they do with it.
Every sandbox gets its own Linux kernel via Firecracker microVMs. Hardware-enforced isolation — an agent can't escape to your host or other sandboxes.
File edits, terminal commands, LLM calls, network requests — all captured with sub-millisecond timestamps. Replay any session from start to finish.
Set a dollar limit. When the agent hits it, the sandbox dies — no overruns. Track real-time spend per model, per session, per day.
Quantified risk per session. Attempted sudo? Score drops. Accessed .env? Score drops. Ran tests? Score goes up. Transparent and explainable.
Workflow
Pick a template, set CPU/RAM limits, define a budget. One API call, one CLI command, or one click.
$ caged run --template node-20 \
--repo github.com/acme/app \
--budget 5.00
⚡ Sandbox sbx_7f3k ready (247ms)
🔗 ssh [email protected]Point Claude Code, Cursor, Aider, or any MCP-compatible agent at the sandbox. Full Linux environment with your code pre-loaded.
Watch live or replay later. Costs tick up in real-time. Trust score drops? Kill it instantly. Budget exceeded? Auto-killed.
Capabilities
Dedicated kernel per sandbox. Hardware-enforced boundaries — real VMs, not container namespaces.
Track every token and compute second as it happens. Hard caps auto-kill sessions that overspend.
Scrub through file edits, terminal commands, and LLM calls. Structured playback of entire sessions.
Flags credential access, destructive commands, and network anomalies. Quantified risk per session.
Claude Code, Cursor, Aider, OpenHands, custom MCP servers. If it speaks SSH or MCP, it works.
Read-only keys for dashboards, full-access for CI. Create, rotate, revoke — zero downtime.
Per-model breakdowns, daily trends, top spenders. Understand spend before the invoice arrives.
Allowlist domains your agent needs. Block everything else. Reach npm — not your production DB.
Define rules per tool, file, command, or network call. SOC2 and HIPAA templates built in. Policies inherit org → team → project.
Pause execution when an agent hits a guardrail. Approve or reject from dashboard, Slack, or email. SLA timers auto-escalate.
Define test scenarios in YAML. Run agents against assertions. Detect flakiness and regressions across runs.
Sandboxes launch in <300ms via Firecracker. No cold starts. Your agent is working before you look away.
Kick off agent tasks from the tools your team already uses. Results flow back automatically.
Slack
/caged run from any channel
GitHub
Label issue → agent starts
Linear
Assign to agent → sandbox
Jira
Move to Agent Review status
Discord
/caged run bot command
Scheduled Runs
Run agents on a cron schedule. Nightly test suites, weekly dependency audits, daily code reviews — all automated.
Webhook Triggers
POST to a webhook URL and Caged spins up a sandbox. Wire it to anything — Zapier, n8n, custom backends.
Results Everywhere
Agent posts results back to the originating thread, issue, or PR. Status, cost, trust score — all reported automatically.
Other tools give agents a place to run. Caged tells you what they did, what they spent, and whether to trust them.
| Caged | boxes.dev | E2B | Codespaces | Local | |
|---|---|---|---|---|---|
| VM-level isolation | |||||
| Real-time cost tracking | |||||
| Session replay | |||||
| Trust scoring | |||||
| Budget auto-kill | |||||
| Network policy | |||||
| Policy engine (RBAC guardrails) | |||||
| Human-in-the-loop approvals | |||||
| Agent eval & testing | |||||
| Any agent (Claude, Cursor, Aider) | |||||
| Sub-second boot | |||||
| Snapshot & fork | |||||
| CLI + API + Dashboard |
Use your own LLM API keys — we don't sell tokens. You pay only for compute and platform features.
Free
$0
Experiment freely
Pro
$29/mo
Ship with AI daily
Team
$99/mo
Agents at scale
Uses your existing Codex / Claude Code / OpenAI subscription — we don't sell tokens.
Your team is already using Claude Code and Cursor. Caged gives them isolated environments and gives you the controls to roll it out responsibly.
Still have a question? Get in touch.
Locally, agents run with your credentials, your filesystem, and no guardrails. Caged gives each agent its own VM with network isolation, budget limits, and full observability. If something goes wrong, it's contained.
Yes. Caged doesn't proxy or resell LLM tokens. Your agent connects to Anthropic, OpenAI, or any provider using your own API keys. We track the estimated cost by observing token counts and model usage.
Any agent that works over SSH or MCP: Claude Code, Cursor, Aider, OpenHands, Codex, or custom MCP servers. If it can connect to a Linux machine, it works with Caged.
The observability layer inside each sandbox intercepts LLM request metadata (model, token counts) — never your API keys. We estimate cost using a maintained pricing table of model rates.
Risk signals: sudo attempts (-20), unknown network calls (-15), .env access (-25). Positive signals: running tests (+3), git operations (+5). The score starts at 100 and adjusts in real-time. You set the threshold for alerts or auto-kill.
Your code runs inside a Firecracker microVM with its own Linux kernel. It never touches our host filesystem. Sandboxes are destroyed after use. We don't store your code beyond the session.
Under 300 milliseconds. We use Firecracker microVMs with pre-warmed pools — your agent gets a running VM almost instantly.
Self-hosted deployment is on our roadmap. The sandbox agent and MCP server will be open-sourced. Contact us if you need on-prem before the public release.
Your agents get their own environment. You get full control.
Free. No credit card. 30 seconds to start.