sandkiln.

Real numbers fromreal hardware.

Every figure came from a criterion benchmark or a scripted load test against a live daemon on real KVM. One shared dev box, single node — directional, not authoritative.

measured
criterion, against the real Firecracker binary
Cold boot
10.5–10.9ms
Process spawn to InstanceStart acknowledged — was 32.3–33.1ms before fixing a 20ms fixed-sleep poll below
Exec round-trip
220–250µs
Over an already-open vsock connection
Resume from snapshot
6.5–7.7ms
Was 25.5–26.2ms — shares the same fixed socket wait
Snapshot take
309.6–335.0ms
Pause plus a full synchronous write of guest memory. The most load-sensitive figure here; re-runs on a busier box have read 550-900ms.
load test — 4 workers × 5 cycles, 0 errors, 5.38 cycles/sec
PhaseMinMeanP95
create137ms211ms577ms
exec—366ms—
delete—131ms—

A dash means the run didn't record that figure, not that it was zero. This table predates the fix below and hasn't been cleanly re-run yet (a live rerun this session was corrupted by unrelated load on the shared dev box).

correction

A measurement thatoverturned the plan.

create averaged 369ms against a ~33ms cold boot. Closing that gap took two attempts, and the second one proved the first explanation wrong.

Fixed. The ~300MB rootfs copy ran synchronously after the network lease. They are independent work, so they now run concurrently, and the copy switched to cp --reflink=auto. create's mean fell from 369ms to 211ms — real and repeatable. Throughput barely moved (5.59 → 5.38 cycles/sec), because create is one phase of a cycle and both runs are small samples.

Then re-measured, and wrong the first time. The plan said the rest needed a copy-on-write filesystem, or failing that a device-mapper layer. A real XFS loopback was set up and SANDKILN_BASE_ROOTFS pointed at it. The clone is genuinely CoW — four clones of a 300MiB rootfs added ~4MiB of real disk bydf, not ~1.2GiB, and a standalone copy dropped from ~110ms to ~0ms. But end-to-end POST /sandboxes stayed at ~160–170ms either way. The explanation offered at the time — the copy was already hidden behind the network lease — turned out to be wrong, caught by the next step below rather than by re-reading the reasoning.

A full per-phase profiling pass, not another guess. 20 isolated cold creates, one at a time, nothing else on the box, broke 167.65ms down completely — 0.32ms unaccounted for:

PhaseShare of a cold create
rootfs clone (cp --reflink=auto)74.0%
Vm::boot total20.3%
— wait for the API socket (fixed below; was 12.0%)0.5%
— InstanceStart7.1%
— configuration PUTs, all 7 combined1.0%
history-store write5.5%
network lease (fully concurrent with the clone)~0.1%

The network lease — three ip/bridge subprocess calls, the thing blamed above — measured ~4.36ms, not ~130ms. The rootfs clone measured 124.09ms, 74% of the whole create. The two were swapped: the clone was never hidden behind the lease, because the lease was never big enough to hide anything behind.

What that same pass found and fixed. Vm::boot's wait for Firecracker's freshly spawned API socket used a fixed sleep(20ms) before ever trying to connect — measured 20.11ms on every single one of the 20 boots (min 20.04, max 20.18), the signature of a sleep nobody needed to wait that long for. Replaced with a real connect-retry on a 200µs→5ms backoff, which also closes a genuine race the file-existence check had (the socket file appears at bind(), a moment before listen()). Re-measured: boot 34.00ms → 11.22/11.44ms, cold create 167.65ms → 144.59/143.56ms — a real ~14% cut to every cold create, and to snapshot resume too, since it shares the same code path.

The correction, stated plainly. The CoW filesystem test was real and its numbers stand — only the explanation attached to them was wrong. The leading replacement, not yet re-verified: the per-sandbox clone always lands in std::env::temp_dir() regardless of where the base image lives, so the earlier test's clone could never have actually reflinked even though the base image sat on XFS. A device-mapper layer would have hit the identical wall for the identical reason and remains not planned. See the engineering notebook for the full story, including this project's own mistake in reaching the wrong conclusion the first time.

pools

Pre-warmed pools win far biggerthan measured, and fail often.

POST /pools configures a pool, a replenisher keeps resumable snapshots ready, and a matching Sandbox.create() claims one instead of cold-booting. A claim that resumes cleanly measured 70–200ms against a cold create's 160–200ms — comparing how fast create() itself returns in each case.

That comparison was measuring the wrong finish line. create() returning 200 means Firecracker's InstanceStart succeeded — not that the guest agent is actually listening yet. Timed a cold sandbox's real first exec end to end: ~420-460ms, entirely inside the vsock client's own retry loop waiting for the agent to come up. Every exec after the first, on the same sandbox, measured ~3-5ms — confirming it's a one-time tax nothing before this had ever isolated, since every other benchmark on this page stops at boot or at create() returning.

A resumed sandbox skips almost all of it. Same test against a snapshot resumed from an already-warm agent: ~4-18ms to first exec — a real ~25-100x difference, repeated, not a one-off. Counted through to "the sandbox can actually run something" rather than just "create() returned," the real win a pre-warmed pool buys is closer to that 25-100x than the modest 2x the paragraph above suggests on its own — the feature was always this good, nothing had been able to see it yet. scripts/bench-report.sh now tracks this permanently as first_exec_client.

The clean case is not the common case. Building a pool means resuming far more often than any manual test had, and a resumed guest kernel can panic early in boot — an early-boot divide-by-zero trap in the console driver, caught through Firecracker's captured console log. Measured repeatedly across clean isolated runs, the failure rate ran from roughly 1 in 3 to 2 in 3 resumes. The existing integration check still passes every run, because one resume isn't enough attempts to hit it.

Handled, not hidden. Root cause is still open — plausibly TSC or clock-source drift between snapshot and restore, but that's a guess, not a diagnosis. Every claim runs a post-resume health check, and a failed check tears the sandbox down and falls back to a cold create. A caller never receives a dead sandbox id. The full investigation is in the engineering notebook.

next

Where the remaining gapactually lives, now measured.

The rootfs clone is 74% of a cold create — the genuine dominant cost, confirmed by direct measurement rather than assumption. Whether a CoW-capable filesystem actually closes that gap is still open (see the correction above): the clone always lands in std::env::temp_dir() regardless of where the base image lives, so today's config has no way to put both on the same filesystem to test cleanly. The real next step is a configurable clone destination, then a clean re-measurement — not another filesystem swap without fixing that first.

Also unexamined: a synchronous history-store write sits on the create critical path at 9.24ms, 5.5% of a create — for a write its own code already treats as best-effort. Looks like an easy, real win; not yet attempted or measured as a change.

Ideas, not results: whether jailer's chroot hard-linking adds measurable overhead is unbenchmarked, since jailer hasn't been proven on real hardware at all yet (see the Roadmap page).

Deeper write-up: Startup latency & the pre-warmed pool.