The Boundary Everyone Agreed On: A $50,000 KVM Escape at Vercel, and the Agent That Already Took Google’s kvmCTF Flag
Every serious answer to “how do we run agent-generated code safely” converged on the same primitive: give each workload its own kernel inside a hardware-virtualised microVM, and treat that microVM — not the container inside it — as the security boundary. On 3 October 2026, that primitive took a public hit. Researcher Paulos Yibelo announced a “full VM escape zeroday” achieving guest-to-host root, and Vercel CEO Guillermo Rauch confirmed a KVM zero-day had come through the company’s Sandbox bounty programme. Vercel paid $50,000 — the programme’s maximum for a single report.
Read on its own, that is a confirmed report with almost nothing attached: no CVE, no affected kernel versions, no patch, no exploit chain, no published root cause. What makes it worth writing about today is what landed five days earlier. On 28 September, the pwn.ai team published a detailed account of an AI agent building and operating a KVM guest-to-host exploit that captured a flag from Google’s live kvmCTF host — a result that maps onto a real, published kernel vulnerability, CVE-2026-46113. One story is a claim awaiting disclosure. The other is a documented end-to-end exploitation of the exact boundary the first story threatens.
What Vercel actually confirmed — and what it did not
Vercel’s architecture is published and specific: each sandbox runs in its own Firecracker microVM with a dedicated guest kernel on a bare-metal EC2 host, with a Linux container inside the microVM running the operator’s code. Vercel is explicit that the microVM is the security boundary and the container is not — its bounty rules place container namespace escapes that only reach the Firecracker guest OS out of scope, calling namespaces “a developer-experience feature, not the security boundary.”
That scoping is what gives the report its weight. Yibelo’s claim is not a container break; it is the crossing Vercel defined as the thing that must never happen. The company’s own bounty table reserves its top band — $25,000–$50,000 for Critical — for a vulnerability letting an attacker read or modify another tenant’s data, with microVM escape to an EC2 host named in that category. The payout therefore signals Vercel’s triage assessment of maximum demonstrable impact.
Several caveats deserve to survive the headline, because the available statements do not support more:
- No evidence of customer data access has been published. A bounty band described as covering cross-tenant impact is a severity classification, not a demonstration that real tenant data was read.
- “KVM zero-day” does not mean every KVM deployment is exploitable. Without a root cause, affected versions, processor requirements, or configuration preconditions, no operator can assess exposure — and patches for unrelated KVM issues should not be assumed to cover it.
- Whether guest privilege is required is undisclosed. In a sandbox threat model this matters less than usual — Vercel already assumes hostile code with root in the container and full kernel access in the microVM — but it still governs who else is affected.
Context on the programme itself: Vercel’s Sandbox challenge opened 18 August 2026 as a public HackerOne programme scheduled to run two weeks to 1 September, with up to $1,000,000 in total payouts and a documented requirement for a live proof of concept rather than a code-review finding. A full technical write-up has been promised but not published.
The documented case: an agent, Google’s host, and CVE-2026-46113
Where the Vercel report is a promise of detail, the pwn.ai write-up is the detail — and it is anchored to a CVE anyone can verify. CVE-2026-46113 is recorded in NVD as published 28 May 2026 with a CVSS v3.1 base score of 8.8 (High), titled in the kernel changelog as “KVM: x86: Fix shadow paging use-after-free due to unexpected GFN.” The NVD description matches the attack shape precisely: KVM’s shadow MMU computes GFNs for direct shadow pages from sp->gfn plus the SPTE index, an assumption that breaks when guest page tables are modified between VM entries — a new leaf SPTE and rmap entry get installed, and the rmap lands outside the range KVM later searches when zapping that shadow page. The entry remains NVD status Undergoing Analysis, and references seven stable-kernel commits including 0cb2af2ea66a, the commit Google attributed the submission to.
Google’s kvmCTF is about as unforgiving a proof mechanism as exists in this space: a bare-metal host running Linux 6.1.74, an ordinary Debian guest, a one-hour reserved slot, and a KASAN report path modified so that a qualifying host memory-safety violation — and only that — unlocks a random 64-bit token retrievable through a challenge hypercall. There is no model-graded answer and no human judging whether a crash “looks promising.” Either the host releases a secret that existed only for that boot, or it does not.
By pwn.ai’s account, the agent reached that bar on 17 June 2026, returning the token cdc63bb4b015eb8c on the fourth attempt of the session after three zero results.
The part that is genuinely interesting: the writer
The exploit engineering is where this stops being a bounty anecdote. The stale-rmap bug needs something to mutate a nested EPT entry without triggering KVM’s page-tracking cleanup. The obvious candidates — host userspace, or DMA from an emulated device — drag in a virtual device, a reachable backend, and a race window. Per the write-up, the agent discarded that path and found a cleaner one: make KVM perform the untracked write itself.
When an L2 guest executes a memory-destination VMREAD that KVM must emulate, KVM copies the value on the guest’s behalf via kvm_write_guest_virt_system() — a path that does not issue the kvm_page_track_write() notification the ordinary emulator write path uses. Pointing that write at an L1 virtual alias of a nested EPT leaf, using VMCS field 0x2006 (VM_EXIT_MSR_STORE_ADDR) purely as an eight-byte value carrier because its MSR-store count was zero, turns a three-byte instruction into an EPT entry rewrite from read-only RAM to MMIO/read-write, with no page-track event. The next L2 store faults against the stale read-only shadow SPTE, KVM re-walks the mutated leaf, resolves MMIO, overwrites the SPTE in place — and the rmap entry for the original page is left behind, pointing into a shadow page that can then be freed.
The second half is pure allocator work, and it is where the agent’s first twenty models were wrong. The initial theory — free the poisoned shadow page, then groom ordinary guest pages to reclaim it — produced nothing but zeros. Rather than adding sleeps or inflating the spray, the agent instrumented mmu_spte_clear_track_bits() on its lab host and generated 50,175 clear events, whose trace ended in a 1,001-entry rmap clear around KVM MMU zap activity. That was the correction: the poisoned object was not guest RAM but KVM-owned shadow memory, recycled by the host’s shadow MMU and unreachable by guest-page pressure. The final chain retired nested roots, rotated through eight prepared EPT roots, applied pressure with 49,152 sparse mappings and walked 1,002 aliases back to the same page, until KVM itself dereferenced the stale pointer. The retained harness is 14,338 lines and 539,206 bytes, SHA-256 b511156431b2b2e4d2695c845cf73f066fc520732b0fcfe5a47ebfc6b5649c94.
Credit where the write-up itself places it
Two pieces of restraint in pwn.ai’s own account are worth repeating, because they are the difference between a capability claim and a marketing claim. First, on attribution: the guest-visible hypercall transcript proves Google’s host recorded a qualifying invalid read, but it cannot contain the host-side KASAN stack — so the team uses Google’s CVE attribution without claiming Google supplied a host stack trace. Second, on provenance: a complete model conversation and token ledger were not retained, and the published trace is condensed from timestamped notes, source snapshots, result bundles and the official terminal log. The team states plainly that it has not manufactured a continuous agent transcript.
The honest reading of the result is the one the authors give: this was not a model with a terminal attached. Frontier models supplied reasoning; the harness supplied persistent state, specialised tooling, skeptical reviewer roles that attacked the exploit theory, evidence gates that refused to score intermediate markers as wins, and retry logic across a campaign measured in weeks. That distinction — the harness, not the model, as the operative system — is the same one that keeps showing up on the defensive side of this beat, and it cuts identically here.
It also did not earn a bounty. kvmCTF is a zero-day programme; by the time the flag was captured, a public report and patch for the same issue had already appeared, and Google’s eligibility review found the zero-day window had closed. The exploit ran on the official host and returned the official secret, but the clock had run out.
Why these two stories belong in the same briefing
The agent-infrastructure market spent 2026 standardising on microVM isolation precisely because containers stopped being defensible. We covered that collapse directly when CVE-2026-80521 escaped containers to host root through the AF_UNIX garbage collector. The industry’s answer was to move the boundary down a layer. These two reports are what it looks like when attention follows it there.
The pairing is instructive in both directions. Vercel’s report shows the economics working as intended: a vendor put up a seven-figure pool, invited the boundary to be tested on its own schedule, and got a confirmed critical finding instead of an incident. That is the system functioning. But it also establishes that the gold-standard isolation layer is reachable by a determined researcher — and the write-up is still pending, which means every operator running Firecracker or KVM for agent workloads currently has a confirmed threat and no exposure assessment.
The pwn.ai result supplies what the Vercel report withholds: proof that this class of boundary falls to sustained, instrumented, failure-driven exploitation, and that an agent harness can sustain it across weeks. The bug it exploited is patched. The method — audit the source, find an untracked write path, build the instrument that explains the failure, change the allocator model, tune the workload until the host cooperates — is not patchable, and it is the method that now has to be assumed against every new hypervisor CVE.
What to do with this now
- Patch CVE-2026-46113 rather than reasoning about it. It is a real, scored (CVSS 8.8), fixed kernel bug with seven stable commits in the NVD record. If you run KVM hosts for multi-tenant or agent workloads, confirm your kernels carry the shadow-MMU fix — this is the one item here that is actionable today.
- Do not treat the Vercel report as a patchable event yet. There is no CVE, no version range and no mitigation to apply. Track Vercel’s promised write-up and your Linux vendor’s advisories; resist the temptation to assume an unrelated KVM patch covers it.
- Know which boundary you are actually relying on. Vercel’s scoping is the model to copy: write down whether your isolation claim rests on the container, the guest kernel, or the hypervisor, and test the one you named. A namespace escape and a host-root escape are different incidents with different blast radii, and conflating them in a threat model produces false confidence.
- Assume the network boundary is tested alongside the compute boundary. Vercel enforces its sandbox firewall on the host, outside the microVM, specifically so sandboxed code cannot disable it, and brokers credentials at the boundary so they never enter the guest. If your sandbox has unrestricted egress, an attacker does not need a VM escape at all — a point already made by this year’s rogue-agent incidents.
- Calibrate on method, not on headline capability. The useful signal from the kvmCTF result is not “AI can write kernel exploits.” It is that a scaffolded agent sustained a multi-week campaign, correctly diagnosed why its own theory was wrong from 50,175 instrumented events, and converted a failed experiment into telemetry. Defensive programmes should be budgeting for adversaries with that persistence profile.
Sources
- NVD — CVE-2026-46113 (published 28 May 2026; CVSS v3.1 8.8 High; KVM x86 shadow-paging use-after-free; seven stable-kernel commit references)
- torvalds/linux — commit 0cb2af2ea66a (the fix Google attributed the kvmCTF submission to)
- pwn.ai — “How pwn Escaped Google kvmCTF (Before the Watershed)” (published 28 September 2026; exploit chain, timeline, reliability and provenance notes)
- Vercel — “A $1,000,000 hacker challenge for Vercel Sandbox” (programme window 18 August–1 September 2026; bounty table; microVM-as-boundary scoping)
- Google — kvmCTF rules (target configuration, reservation model, flag mechanism)
- Cyber Security News — “Vercel Confirms KVM Zero-Day VM Escape, Awards Researcher $50,000” (4 October 2026; reporting on the 3 October announcements by Paulos Yibelo and Guillermo Rauch)