Skip to contentNew accounts start with $100 in credits Get an API key →
← all posts

Building a VMM from scratch, #0: Why BoxLite is replacing libkrun

Part #0 of a build-in-public series: what BoxLite is, why we started on libkrun, why we're replacing it, and how a VMM works.

Dorian Zheng6 min read

Every BoxLite box runs in its own micro-VM, and today libkrun runs that VM. We're replacing it with a VMM we write ourselves. This series builds it in public, one merged PR at a time.

1. What BoxLite is

BoxLite is the compute substrate for AI agents — light enough to embed on your laptop, elastic enough to power an agentic cloud. Each box is a hardware-isolated micro-VM that runs any OCI image and keeps its state across turns.

BoxLite architecture on your machine and in the cloud

Figure 1. One engine in two places: the embedded BoxLite runtime starts one micro-VM per box, on your machine or on a cloud runner.

The part that runs each micro-VM is the VMM, the virtual machine monitor. That part is the subject of this series.

2. Why we started on libkrun

libkrun describes itself as "a dynamic library that allows programs to easily acquire the ability to run processes in a partially isolated environment", on KVM (Linux) and HVF (macOS on Apple silicon), behind "a simple C API". That fit BoxLite:

  • A library, not a daemon: BoxLite embeds in your app, with no root and no background service.
  • One API for KVM on Linux and Hypervisor.framework on macOS.
  • The guest kernel comes with it (libkrunfw), along with virtio-fs, virtio-blk and vsock.

libkrun is how BoxLite runs boxes on macOS and Linux today.

3. Why we're replacing it

libkrun is clear about what it isn't: its README lists "Become a generic VMM" and "Be compatible with all kinds of workloads" as non-goals. BoxLite's needs grew past those goals:

  • The VM's outcome. libkrun ends the process itself, so the caller gets no exit status, and every start failure comes back as -EINVAL.
  • The kernel as a shared library. libkrunfw is loaded with dlopen, and that keeps breaking on some hosts; #1512 lists six such issues.
  • Devices. Each directory volume gets its own device, and with three a box fails to start with IrqsExhausted (#935).
  • What comes next. Hot-plug mounts, live resizing, memory snapshots for AutoPause (#1003) and GPUs all need a VMM we control.

What we're not doing

  • Forking or patching libkrun. It ships unchanged until the new VMM is the default, stays one release as a fallback, and is then removed.
  • Building a general-purpose VMM. It runs one VM per process with BoxLite's devices only; confidential computing, live migration and Intel Macs are out of scope.

What it costs

  • The new VMM needs macOS 15 or later, for Hypervisor.framework's in-kernel interrupt controller. macOS 12–14 keep libkrun until cutover.
  • A long road: first boot (M1), then a real box (M2), volumes (M3), hardening (M4) and cutover (M5).

4. How a VMM works, in one screen

Inside one machine: host process, jailed boxlite-shim, micro-VM and state

Figure 2. Inside one machine. The runtime spawns a jailed boxlite-shim per box, and the VMM inside it runs the micro-VM. PROPOSED marks what this series replaces.

A hypervisor (KVM or Hypervisor.framework) runs guest instructions on the CPU. The VMM is the ordinary process around it: it gives the guest memory, creates vCPUs, loads the kernel and emulates every device. Its heart is one loop:

The VMM loop: run a vCPU, handle its exit, run it again

Figure 3. Run a vCPU until it exits, handle the exit, and run it again.

The first guest our new KVM backend ran was 7 bytes long:

mov dx, 0x3f8   ; COM1 serial port
mov al, 'K'
out dx, al      ; KVM exits to the VMM: port 0x3f8, byte 'K'
hlt

Everything else is detail, and each detail gets a part in this series:

PieceWhat it isPart
MemoryGuest RAM is host memory; an unmapped address is a device#1
ExitsOne exit contract for KVM, Hypervisor.framework and, later, WHP#2
StoppingPulling a vCPU out of the guest, safely#3
CPU stateRegisters, CPUID and MSRs on x86#4
BootThe VMM is the bootloader#6
DevicesSerial, RTC and reset first, then virtio#7, #10–12

5. The series

Status as of October 11, 2026. A part goes out only after its code merges, so every claim links to a merged PR.

ArcPartsShips when
1. First instructions#1 Your first VM is 7 bytes · #2 Design the exit contract first · #3 Stopping a vCPU is the hard part · #4 x86 CPU state, by hand · #5 A kernel we build ourselvesCode merged
2. First boot (M1)#6 You are the bootloader · #7 Boring devices first · #8 On Apple silicon, you decode the exits · #9 One kernel, three hostsAs M1 lands
3. First box (M2–M3)#10 virtio from scratch · #11 Disks, sockets, network · #12 One virtio-fs device for every volume · #13 Our agent as PID 1M2–M3
4. Ship it (M4–M5)#14 Guest input must never panic the host · #15 Measured against libkrun · #16 Deleting libkrunM4–M5
LaterMemory snapshots, hot-plug mounts, GPUs, our own network stackAfter M5

Where we are

  • ✅ Design and one hypervisor contract for KVM, Hypervisor.framework and, later, Windows
  • ✅ KVM on x86_64: VMs, memory, the run loop, I/O completion, kicks, boot registers, CPUID and MSRs
  • ✅ A reproducible build of our pinned Linux 6.12 kernel for x86_64
  • ⏳ M1: first boot on macOS arm64, Linux x86_64 and Linux arm64, checked in CI

Follow along on the RSS feed and in every PR on GitHub.

6. Come build it with us

BoxLite is open source, and so is every step of this VMM. We'd love your help:

  • Challenge the design. Read the VMM design doc and tell us where it's wrong.
  • Pick up M1 work. #1698 lists the open slices and their dependencies: devices, kernel loaders, the Hypervisor.framework backend, arm64 KVM and three-host CI.
  • Ask and discuss. Open a GitHub issue or find us on Discord.

Start with CONTRIBUTING.md. Your first PR asks you to sign our contributor license agreement.