Firecracker vs gVisor: why every app and serverless function on Synteq runs in its own microVM, and what this means for your AI Agent sandbox

The last week has been a good one for anyone who thinks about sandboxes for a living, or just frequently.
On October 2, Google announced it is donating gVisor to the CNCF, the same foundation that looks after Kubernetes. The application was accepted on September 28, and gVisor enters as a Sandbox project, with Incubation as the next step. The announcement names OpenAI, Anthropic, Modal and Tines among its users, with Tines engineer Shayon Mukherjee saying that gVisor "powers every sandbox execution today" in their product.
A day later, Vercel CEO Guillermo Rauch confirmed that a researcher, Paulos Yibelo, had escaped a Vercel Sandbox VM to root on the host through a KVM zero-day. Vercel Sandbox runs on Firecracker microVMs. The write-up isn't out yet, so we don't know whether the bug sits in KVM itself or in the code around it.
All of this lands in the middle of the AI agent sandboxing boom, with hundreds of startups building strictly for this environment, and virtually everyone spinning up their own. In the Kimi K3 technical report, Moonshot describes what happened when they ran agent training on container sandboxes:
"In our early experiments with traditional container-based sandbox runtimes, we observed several kernel panics and deadlocks caused by unintended agent operations."
They moved to a Firecracker-based runtime called AgentENV, which they say provides "a level of isolation and fidelity that container-based runtimes cannot match," and created more than 51 million sandboxes on it during training and evaluation. Rauch's take on this move - which we share - was blunt: "container-level isolation is not enough."
We made the same call when we built Synteq Apps and Functions. Every app instance and every function runs in its own Firecracker microVM, with its own Linux kernel. That choice costs us density, and in turn costs us money.
Three ways to isolate code you didn't write
When designing environments that will contain code other then yours, it's important to be paranoid, and treat all execution as if it's adversarial.
There are three common answers to "how do I run untrusted code next to other untrusted code?"
Containers (Docker, containerd, runc) run every workload on one shared host kernel. Namespaces decide what each workload can see, cgroups decide how much it can use, and seccomp filters decide which system calls it may make. Containers are fast and dense, and they are the right tool for packaging software. As a security boundary, though, the boundary is the host kernel's system-call surface. A single kernel bug reachable from inside a container can hand an attacker the host, and with it every other tenant on that host. Kimi's agents didn't even need a bug: ordinary agent behavior was enough to panic the shared kernel and take down every neighbor with it.
gVisor puts a second kernel in between. Its core, the Sentry, is a reimplementation of the Linux system-call interface, written in Go and running in user space. Your application's system calls are intercepted and served by the Sentry, which in turn makes a much smaller, tightly filtered set of calls to the real host kernel. File access goes through a separate helper process called the Gofer. The idea is to shrink the host kernel attack surface without paying for full hardware virtualization.
microVMs (Firecracker, Kata Containers, and others) give every workload a real virtual machine with its own guest kernel, isolated by the CPU's hardware virtualization features through KVM. AWS built Firecracker for Lambda and Fargate, and its NSDI '20 paper is worth reading in full. Firecracker drops everything a function doesn't need: BIOS, PCI, USB, graphics, and live migration don't exist. The authors report about 50,000 lines of Rust, 96% fewer than QEMU, less than 5 MB of memory overhead per VM, boot to application code in under 125 ms, and up to 150 new microVMs per second per host.
Why we chose microVMs
The boundary is hard(ware)
With a microVM, code inside the guest can do anything a Linux kernel allows. It can load modules, poke at /proc, or hit whatever obscure system call it likes, and the worst outcome is that it breaks its own kernel. To reach the host or another customer, an attack has to get through the hardware virtualization boundary and the small virtual machine monitor behind it
The Firecracker paper puts the trade-off plainly: virtualization "moves the security-critical interface from the OS boundary to a boundary supported in hardware and comparatively simpler software." With containers, and to a lesser degree with gVisor, you are trying to make a very large interface safe. With a microVM, you only expose a very small one.
Linux. Just plain Linux
gVisor implements the Linux system-call interface, but it is a reimplementation, and some system calls, flags, and /proc and /sys behaviors are missing or behave differently. Most web apps never notice, but others do: native extensions, database engines, tools that expect a real filesystem, and anything that wants to run containers or mount disks of its own.
Performance is close to native
Inside a Firecracker guest, ordinary CPU and memory work runs at native speed. Only device I/O crosses the virtual machine monitor. gVisor has to intercept every system call. The best-known measurement, The True Cost of Containing: A gVisor Case Study (University of Wisconsin, HotCloud '19), found simple system calls at least 2.2× slower, opening and closing files on an external tmpfs 216× slower, small-file reads 11× slower, and large downloads 2.8× slower.
To be fair to gVisor, those numbers are from 2019. Since then gVisor has shipped Systrap, which replaced the slow ptrace platform in 2023, and directfs, which removed most of the Gofer round trips for file access. Today's gaps are much smaller. But the cost problem follows the same structure: every system call is handled twice. Google's own donation post mentions "out-of-the-box performance" as one of the things holding back adoption.
"Every run is a fresh kernel" is easy to reason about
When an instance is torn down, its kernel, page cache, and memory go with it. Nothing from your workload is left on a kernel that the next customer will also use.
What it costs us
We want to be honest about this, because it is the reason most platforms don't do it.
Density. On a container platform, thousands of workloads share one kernel and one page cache. Each Firecracker microVM carries its own guest kernel and its own memory, so a host fits far fewer of them. Firecracker keeps the per-VM overhead tiny (the paper measured around 3 MB for the monitor itself), but the guest kernel and its caches are real memory that we pay for and you don't see. The Lambda team aimed for 10% overhead at most, and the per-VM cost is still exactly why most PaaS products run containers.
Cold starts. A VM boot is slower than starting a process in an existing container. We keep a pool of microVMs that are already booted and ready to take a request. In our current measurements, a cold request to a Node.js or Python function usually completes end to end server side in ~50-100 ms.
So is gVisor the wrong choice?
No. gVisor is a serious, well-engineered project, and its CNCF move is good news for everyone who runs untrusted code. It runs on any Linux machine without hardware virtualization, which matters a lot if you're already inside a cloud that doesn't expose nested virtualization. It starts quickly, packs densely, and plugs straight into Docker and Kubernetes. Google says it offers "empirically-equivalent security" to virtualization, and teams like Tines, Modal, and the big AI labs bet on it every day. Cloudflare's CTO Dane Knecht has said Cloudflare Containers uses a mix of microVM hypervisors "plus gvisor for certain workloads."
Neither boundary is magic
The Vercel escape is a useful reminder that microVMs are not invincible. KVM is a large, complex part of the Linux kernel, and Google runs kvmCTF precisely because KVM bugs exist. As security researcher s1r1us put it, microVMs are "everywhere, frontier labs use too," yet only a handful of companies, like Google and Vercel, pay researchers to break their production sandboxes. We think that should change, and we're grateful to the people who do their research in the open.
What this means if you're choosing a platform
If you're comparing Synteq with Render, Railway, Heroku, ask each one the same question: what is the isolation boundary between my workload and another customer's? "A container" means a shared kernel. "gVisor" means a user-space kernel in front of a shared kernel. "A microVM" means your own kernel, behind the CPU's virtualization boundary.
On Synteq:
- App: push code or connect a Git repo and get a URL. Every replica runs in its own microVM.
- Functions: serverless Node.js 22 and Python 3.12 (or anything else through a Dockerfile), triggered by HTTP, cron, or a webhook. Functions scale to zero, cost $1.40 per million invocations, and include 125,000 free invocations a month.
- AI agents and untrusted code: because every app instance is a real Linux machine of its own, it's a natural home for your agents to live off device.








