{"licence":{"name":"CC BY-SA 4.0","spdx":"CC-BY-SA-4.0","url":"https://creativecommons.org/licenses/by-sa/4.0/","attribution":"Atlas, a bilingual technical dictionary (https://atlas.maintz.dev/)"},"id":"ai/agent-sandbox","url":{"en":"https://atlas.maintz.dev/en/terms/ai/agent-sandbox/","da":"https://atlas.maintz.dev/da/terms/ai/agent-sandbox/"},"term":{"en":"Agent sandbox","da":"Agent-sandkasse"},"aka":{"en":["sandboxing"],"da":["sandboxing"]},"domain":["ai","security"],"cluster":"agents","layer":"agent","status":"emerging","summary":{"en":"A walled-off space where an AI agent runs code and uses tools, so a mistake or trick cannot reach the rest of your systems.","da":"Et lukket rum, hvor en AI-agent kører kode og bruger værktøjer, så en fejl eller et trick ikke kan nå resten af ens systemer."},"body":{"formal":{"en":"A closed-off place to run code - often a container or small virtual machine - in which an agent's commands execute with only chosen folders, network destinations and rights, and which is thrown away after the task.","da":"Et lukket sted at køre kode - ofte en container eller en lille virtuel maskine - hvor en agents kommandoer udføres med kun udvalgte mapper, netværksmål og rettigheder, og som smides væk efter opgaven."},"plain":{"en":"Like letting a new cook practise in a test kitchen instead of the restaurant - if they set the pan on fire, only the test kitchen gets burnt.","da":"Som at lade en ny kok øve sig i et prøvekøkken i stedet for i restauranten - hvis panden bryder i brand, er det kun prøvekøkkenet, der brænder."},"inPractice":{"en":"A developer in a region's IT department lets a coding agent run tests inside a fresh container that sees only the booking app's folder and cannot reach the internet; when the task ends, the container is deleted.","da":"En udvikler i en regions IT-afdeling lader en kodeagent køre tests i en frisk container, der kun kan se bookingappens mappe og ikke kan nå internettet; når opgaven er løst, slettes containeren."},"whyItMatters":{"en":"No one can promise an agent will never be tricked or wrong, so limiting what it can touch is what turns a possible disaster into a small failure that goes no further.","da":"Ingen kan love, at en agent aldrig bliver narret eller tager fejl, så det at begrænse, hvad den kan røre, er det, der gør en mulig katastrofe til et lille uheld, der ikke breder sig."}},"deepDive":{"en":"An agent sandbox is a containment boundary for code whose behaviour is chosen at runtime by a model that can be wrong or manipulated. The threat model therefore assumes that any command inside the boundary may be hostile, and the design question is what that command can read, write, execute and reach over the network. Isolation strength forms a spectrum. Process-level sandboxes apply kernel policy to an ordinary process: seccomp-bpf syscall filters and Landlock on Linux, unprivileged namespace tools such as bubblewrap, and the Seatbelt framework on macOS. Containers add namespaces and cgroups but still share the host kernel, so a kernel vulnerability can become an escape, a limitation NIST SP 800-190 discusses at length. gVisor interposes a user-space kernel that services most syscalls itself, and microVMs such as Firecracker give each workload its own KVM guest kernel while booting in around 125 ms with a few MiB of memory overhead, which is why many hosted code-execution services use them.\n\nThe filesystem policy usually mounts only the working directory read-write, keeps system paths read-only, and deliberately omits credential locations such as ~/.ssh, ~/.aws or browser profiles. Network policy is typically default-deny egress with an allowlist enforced by a proxy, since blocking only by IP misses DNS-based exfiltration and CDN-hosted endpoints. Credentials that the task genuinely needs should be short-lived and narrowly scoped, and some designs never expose them inside the sandbox at all, letting the egress proxy inject them into requests to approved hosts. Ephemerality matters as much as walls: a fresh sandbox per task prevents persistence, and snapshots make it cheap to reset.\n\nClaude Code's sandboxed Bash tool illustrates a local implementation: it uses Seatbelt on macOS and bubblewrap plus a socat-relayed proxy on Linux and WSL2, allows writes to the working and session temp directories by default, and restricts egress to allowed domains. Notably, if dependencies are missing it warns and runs unsandboxed unless configured to fail closed, a detail auditors should check.\n\nCommon misconceptions: a sandbox does not stop prompt injection, it bounds its blast radius; and it does not neutralise misuse of channels it allows. An agent permitted to push to GitHub can exfiltrate through a commit, and code written inside the sandbox may plant a malicious git hook or CI step that executes later outside it. The sandbox is thus the enforcement mechanism for least privilege, complementary to human-in-the-loop approval, and it reduces the approval fatigue that arises when every command needs a click.","da":"En agent-sandkasse er en indkapslingsgrænse for kode, hvis adfærd vælges under kørsel af en model, der kan tage fejl eller blive manipuleret. Trusselsmodellen antager derfor, at enhver kommando inden for grænsen kan være fjendtlig, og designspørgsmålet er, hvad kommandoen kan læse, skrive, afvikle og nå over netværket. Isolationsstyrken ligger på en skala. Sandkasser på procesniveau lægger kernepolitik på en almindelig proces: seccomp-bpf-filtre på systemkald og Landlock på Linux, uprivilegerede namespace-værktøjer som bubblewrap og Seatbelt-frameworket på macOS. Containere tilføjer namespaces og cgroups, men deler stadig værtens kerne, så en sårbarhed i kernen kan blive til et udbrud, en begrænsning NIST SP 800-190 gennemgår grundigt. gVisor indskyder en kerne i brugerrummet, der selv håndterer de fleste systemkald, og microVM'er som Firecracker giver hver arbejdsbyrde sin egen KVM-gæstekerne og starter på omkring 125 ms med få MiB hukommelsesoverhead, hvilket er grunden til, at mange hostede tjenester til kodeafvikling bruger dem.\n\nFilsystempolitikken monterer typisk kun arbejdsmappen med skriveadgang, holder systemstier skrivebeskyttede og udelader bevidst steder med legitimationsoplysninger som ~/.ssh, ~/.aws eller browserprofiler. Netværkspolitikken er normalt udgående trafik afvist som standard med en tilladelsesliste håndhævet af en proxy, fordi blokering alene på IP-adresser overser exfiltration via DNS og endpoints bag CDN'er. Legitimationsoplysninger, opgaven reelt har brug for, bør være kortlivede og snævert afgrænsede, og nogle design eksponerer dem slet ikke inde i sandkassen, men lader udgangsproxyen indsætte dem i forespørgsler til godkendte værter. Kortlivethed betyder lige så meget som mure: en frisk sandkasse pr. opgave forhindrer vedvarende fodfæste, og snapshots gør det billigt at nulstille.\n\nClaude Codes sandboxede Bash-værktøj viser en lokal implementering: det bruger Seatbelt på macOS og bubblewrap plus en proxy via socat på Linux og WSL2, tillader som standard skrivning i arbejdsmappen og sessionens temp-mappe og begrænser udgående trafik til tilladte domæner. Bemærk, at det ved manglende afhængigheder advarer og kører uden sandkasse, medmindre det er sat op til at fejle lukket, en detalje, som en revision bør tjekke.\n\nUdbredte misforståelser: En sandkasse stopper ikke prompt injection, den begrænser skadens omfang, og den neutraliserer ikke misbrug af de kanaler, den tillader. En agent, der må pushe til GitHub, kan exfiltrere via et commit, og kode skrevet inde i sandkassen kan plante en ondsindet git-hook eller et CI-trin, der senere kører uden for den. Sandkassen er altså håndhævelsesmekanismen for mindste privilegium, supplerer menneske i løkken-godkendelse og mindsker den godkendelsestræthed, der opstår, når hver kommando kræver et klik."},"howTo":{"steps":{"en":["List every place an agent runs code or uses tools (developer laptops, CI runners, hosted agents), and decide per use case how strong the wall must be, from a per-command sandbox for everyday work to a container or virtual machine for unattended runs and untrusted repositories.","Give each task a fresh sandbox that is deleted or reset from a snapshot afterwards, so nothing the agent leaves behind survives to the next task.","Mount only the project folder with write access, keep system paths read-only, and explicitly deny credential locations such as ~/.ssh, ~/.aws, .env files and browser profiles.","Block outgoing network traffic by default and allow only the domains the task needs, such as the package registry and the code host, enforced by a proxy rather than by IP address.","Hand the agent only short-lived, narrowly scoped tokens, or let the egress proxy inject credentials so the secret itself never enters the sandbox.","Make the sandbox fail closed, so commands stop instead of running unsandboxed when the sandbox cannot start (in Claude Code, sandbox.failIfUnavailable set to true and allowUnsandboxedCommands set to false), and enforce the settings centrally through managed settings.","Check what sits outside the wall, such as MCP servers, hooks and built-in file tools, and move the whole agent process into a container or VM if those must be contained too.","Review the channels you still allow, such as git push and CI, for exfiltration and planted git hooks or pipeline steps, and look at sandbox violation logs and the allowlist every quarter."],"da":["Lav en liste over alle steder, hvor en agent kører kode eller bruger værktøjer (udvikleres bærbare, CI-runners, hostede agenter), og beslut for hver use case, hvor stærk muren skal være, fra en sandkasse pr. kommando i hverdagen til en container eller virtuel maskine ved uovervågede kørsler og repositories, I ikke stoler på.","Giv hver opgave en frisk sandkasse, der slettes eller nulstilles fra et snapshot bagefter, så intet, agenten efterlader, overlever til næste opgave.","Montér kun projektmappen med skriveadgang, hold systemstier skrivebeskyttede, og afvis udtrykkeligt adgang til steder med loginoplysninger som ~/.ssh, ~/.aws, .env-filer og browserprofiler.","Bloker udgående netværkstrafik som standard, og tillad kun de domæner, opgaven har brug for, fx pakkeregistret og kodehosten, håndhævet af en proxy og ikke via IP-adresser.","Giv kun agenten kortlivede, snævert afgrænsede tokens, eller lad udgangsproxyen indsætte loginoplysningerne, så selve hemmeligheden aldrig kommer ind i sandkassen.","Sørg for, at sandkassen fejler lukket, så kommandoer stopper i stedet for at køre uden sandkasse, når den ikke kan starte (i Claude Code sandbox.failIfUnavailable sat til true og allowUnsandboxedCommands sat til false), og håndhæv indstillingerne centralt via managed settings.","Tjek, hvad der ligger uden for muren, fx MCP-servere, hooks og indbyggede filværktøjer, og flyt hele agentprocessen ind i en container eller VM, hvis de også skal holdes inde.","Gennemgå de kanaler, I stadig tillader, fx git push og CI, for datalæk og plantede git-hooks eller pipeline-trin, og kig på logs over sandkasseovertrædelser og tilladelseslisten hvert kvartal."]},"pitfalls":{"en":["Believing the sandbox stops prompt injection, when it only limits what a tricked agent can reach through the channels you left open.","Leaving the default read access to the whole home folder, so the agent can still read SSH keys and cloud credentials even though it cannot write there.","Letting the sandbox silently fall back to running commands unsandboxed when a dependency is missing, so the wall disappears without anyone noticing.","Sandboxing only shell commands while MCP servers and hooks run with full access on the host."],"da":["At tro, at sandkassen stopper prompt injection, selv om den kun begrænser, hvad en narret agent kan nå via de kanaler, I har ladet stå åbne.","At beholde standardlæseadgangen til hele hjemmemappen, så agenten stadig kan læse SSH-nøgler og loginoplysninger til skyen, selv om den ikke kan skrive der.","At lade sandkassen stille og roligt køre kommandoer uden sandkasse, når en afhængighed mangler, så muren forsvinder, uden at nogen opdager det.","Kun at sandboxe shell-kommandoer, mens MCP-servere og hooks kører med fuld adgang på værten."]},"guides":[{"title":"Claude Code documentation - Configure the sandboxed Bash tool","url":"https://code.claude.com/docs/en/sandboxing","publisher":"Anthropic","tier":"official-doc"},{"title":"Claude Code documentation - Choose a sandbox environment","url":"https://code.claude.com/docs/en/sandbox-environments","publisher":"Anthropic","tier":"official-doc"},{"title":"NIST SP 800-190 - Application Container Security Guide","url":"https://csrc.nist.gov/pubs/sp/800/190/final","publisher":"NIST","tier":"standard"},{"title":"AI Agent Security Cheat Sheet","url":"https://cheatsheetseries.owasp.org/cheatsheets/AI_Agent_Security_Cheat_Sheet.html","publisher":"OWASP","tier":"reference"}]},"edges":[{"type":"requires","to":"ai/ai-agent","confidence":"high","strength":"normal"},{"type":"mitigates","to":"ai/excessive-agency","why":{"en":"Even if an agent has broad tools, the walls around it cap what those tools can actually reach.","da":"Selv hvis en agent har brede værktøjer, sætter væggene omkring den en grænse for, hvad værktøjerne faktisk kan nå."},"confidence":"high","strength":"primary"},{"type":"mitigates","to":"ai/prompt-injection","why":{"en":"It does not stop the trick, but it limits the damage - a tricked agent can only harm what is inside the walls.","da":"Det stopper ikke tricket, men begrænser skaden - en narret agent kan kun skade det, der er inden for væggene."},"confidence":"high","strength":"normal"},{"type":"used-with","to":"platform/container","why":{"en":"Containers are the most common way to build an agent sandbox quickly and throw it away after each task.","da":"Containere er den mest udbredte måde at bygge en agent-sandkasse hurtigt og smide den væk efter hver opgave."},"confidence":"high","strength":"primary"},{"type":"used-with","to":"cs/least-privilege","confidence":"high","strength":"normal"},{"type":"used-with","to":"ai/computer-use","confidence":"high","strength":"normal"},{"type":"used-with","to":"ai/coding-agent","why":{"en":"Run the agent inside a sandbox so its commands, file access and network reach stop at the project; a tricked agent then cannot touch the rest of the machine.","da":"Kør agenten i en sandkasse, så dens kommandoer, filadgang og netværksadgang stopper ved projektet; en narret agent kan så ikke røre resten af maskinen."},"confidence":"medium","strength":"normal"}],"depth":5,"sources":[{"title":"OWASP Top 10 for LLM Applications 2025 (LLM06 Excessive Agency)","tier":"reference","publisher":"OWASP"},{"title":"NIST SP 800-190 - Application Container Security Guide","tier":"standard","publisher":"NIST"},{"title":"Claude Code documentation - Configure the sandboxed Bash tool","url":"https://code.claude.com/docs/en/sandboxing","tier":"official-doc","publisher":"Anthropic"}],"draft":true}