A talk by Mahika Khanna
Sandboxing
What actually stops an AI coding agent from wrecking your machine.
An agent doesn't need to be hacked. It only needs to read the wrong thing. Everything from the talk is here: the slides, the demos I run on stage, a page where you can try to break a pretend agent yourself, and the sources behind it all.
If you remember three things
The short version
Assume the agent will be fooled
Prompts, approvals and guardrails are all requests. Sooner or later a web page, a package or a pull request will talk the agent into something it shouldn't do.
Walls beat words
A sandbox doesn't depend on the agent behaving. It takes away the files and the network, so a fooled agent has nothing to steal and nowhere to send it.
Lock the exits first
If you only do one thing, lock the exits. Controlling what can leave the machine is what stops a fooled agent from robbing you.
What you'll see
Six parts in fifty minutes
- 0 to 9 minIntroA hook, a short history, and agents explained from scratch
- 9 to 18 minPractical A: no sandboxTwo attacks with nothing in the way
- 18 to 30 minExplanation and fixesWhy it happened, and what your options are
- 30 to 41 minPractical B: with a sandboxThe same attacks again, and this time they fail
- 41 to 45 minHonest limits and takeawaysWhat a sandbox can't do, and what to do on Monday
- 45 to 50 minOne last thing, then questionsA closing surprise, and your questions
Try it yourself
Can you talk an agent into leaking a secret?
I built a pretend agent that reads whatever page you give it and does what the page says. Hide an instruction in the page and see what leaves the machine. Then switch the sandbox on and try again.
A short history
Docker spent a decade saying VMs were too heavy. Now it sells a VM per chatbot.
- 1979/82chroot
- 2000FreeBSD jails
- 2008cgroups and LXC
- 2013Docker's first demo
- 2014“Containers do not contain”
- 2018Firecracker, built for Lambda
- TodayDocker ships microVM sandboxes
Everything in one place