DEV Community

Chen Yuan
Chen Yuan

Posted on Originally published at dispatch-blog.hashnode.dev

MicroVMs at the Edge: What Netlify's Firecracker Migration Actually Changed

Netlify says its Edge Functions now handle about a billion invocations each day. That makes a small amount of per-request overhead expensive twice: users feel it as latency, and the platform pays it across a very large request volume.

The interesting part of Netlify's recent infrastructure migration is not simply that virtual machines became fast enough for edge compute. The deeper change is architectural. Netlify moved from a hosted V8-isolate execution service to Firecracker MicroVMs running inside its own edge network, while keeping the programming model for customers largely unchanged.

According to Netlify's engineering write-up, warm invocation overhead fell from roughly 25–40 ms on the previous infrastructure to about 5–6 ms at the median. Netlify also reports 47.4 percent faster p99 invocations, 99.998 percent availability, and log delivery that is about five times faster. Unikraft, which worked with Netlify on the execution layer, reports a 2 ms p99 for MicroVM startup in Netlify's fleet.

Those numbers matter, but the mechanisms behind them are more reusable than the benchmark headline.

The old boundary was outside Netlify's network

An Edge Function sits in the request path. If a request matches a configured route, the platform must select compute, start or reuse an execution environment, run customer code, and return the result before the request can continue.

On Netlify's previous architecture, a matching request left Netlify's own network for a hosted execution service based on V8 isolates. That provider handled execution, provisioning, and much of the runtime environment. The model offered operational simplicity, but it also created a network boundary Netlify did not control.

The new path keeps that work inside Netlify's network. The edge node that terminates TLS checks the deployment's Edge Function routes. If a route matches, it forwards the request to a regional compute node rather than sending it to an external execution provider.

The key lesson is that runtime startup time is only one term in request latency. Placement, network hops, image availability, service lookup, logging, and failure recovery all sit on the critical path too. Removing an external hop can therefore matter as much as making the runtime itself faster.

A request carries the machine specification with it

Before forwarding the request, the edge node writes a specification for the machine that should execute the function. Netlify describes three image layers in that specification: the runtime, a platform image, and the customer's Edge Function image. CPU, memory, and connection limits travel with the spec as well.

A simplified representation might look like this:

{
  "runtime": "deno-runtime-v42",
  "platform": "netlify-edge-v18",
  "function": "site-7f3a-image",
  "cpu": 1,
  "memory_mb": 512,
  "connection_limit": 64
}
Enter fullscreen mode Exit fullscreen mode

The exact object above is illustrative, not Netlify's production schema. The important design idea is that enough information travels with the request to identify the execution service deterministically.

Netlify hashes the machine specification together with site-specific information to derive a service ID. Different deploys or environment-variable sets therefore map to different services and do not share a MicroVM. That makes deployment identity part of the isolation boundary instead of a convention layered on top of a shared process.

Rendezvous hashing keeps the fast path warm

Each region contains a group of compute nodes. Netlify uses rendezvous hashing so the same service normally lands on the same node. That improves locality: once a node has fetched an Edge Function image, later requests can reuse the copy already on disk, and an existing MicroVM or snapshot can stay useful for that service.

A minimal version of the placement idea can be expressed as:

def choose_node(service_id, nodes):
    scored = [(score(service_id, node.id), node) for node in nodes]
    return max(scored, key=lambda item: item[0])[1]
Enter fullscreen mode Exit fullscreen mode

Real production placement has more inputs, but the property matters: stable input usually produces stable placement.

Pure stickiness creates another problem. A busy service can become a hot spot and saturate the node selected for it. Netlify therefore relaxes placement above a threshold and spreads a hot service across a slice of nodes. That is a useful pattern for state-light edge workloads: prefer locality when traffic is ordinary, then trade some cache warmth for capacity when one tenant becomes dominant.

Cold starts are mostly an image-distribution problem

A node that has never served a function needs the relevant images before it can run anything. Netlify reports that this cold path occurs on about 1.2 percent of invocations and takes about 9 ms on average.

The design avoids distributing every customer's code to every compute node in advance. A node fetches an image only when traffic for that service arrives in the region, stores it locally, and reuses it for later requests. Unikraft describes the same strategy as on-demand image resolution and notes that stale images can be pruned from the node cache automatically.

This changes the scaling equation. Pre-provisioning every deployment everywhere would consume storage and memory before demand exists. Fetch-on-first-use lets the fleet pay distribution cost where traffic actually appears.

The cache policy can be summarized as a small state machine:

request arrives
  -> image present: boot or restore MicroVM
  -> image missing: fetch image, cache image, then boot
  -> image idle long enough: prune local copy
Enter fullscreen mode Exit fullscreen mode

That mechanism also decouples deploy frequency from fleet-wide synchronization. A new function image does not need to be pushed to every node before it can receive traffic.

The MicroVM is fast because it avoids doing unnecessary work

Netlify says each function runs in its own Firecracker MicroVM. The platform creates these environments in under a millisecond and reports roughly 2 ms p99 startup. The VM boots a stripped-down Linux environment rather than a conventional general-purpose guest with a large boot sequence.

The function files are mounted as an uncompressed EROFS image and memory-mapped. That allows the VM to read the parts of the bundle it actually touches instead of copying the entire function into memory before execution can begin.

There is another optimization after boot. When the JavaScript server starts listening, the platform snapshots the MicroVM. Idle instances can scale to zero. A later request can restore a new MicroVM from the snapshot, and the snapshot itself is memory-mapped so execution does not have to wait for every page to be loaded first.

The practical takeaway is that "VM" is too broad a performance category. A large cloud VM that boots an ordinary operating system and a purpose-built MicroVM restored from a memory-mapped snapshot have very different startup paths.

Isolation became a property of the compute boundary

Netlify also frames the migration as a security change. A service with different code or environment variables receives a distinct service ID, and different services do not share a MicroVM. If customer code escapes the JavaScript runtime, it still faces a VM boundary before it can reach another tenant or the compute host.

That is a stronger boundary than placing mutually untrusted programs in separate V8 isolates inside a shared process. It does not make the system invulnerable, and no virtualization boundary should be described that way. It does change where a compromise has to cross next.

For edge platforms, that matters because performance and tenant isolation are usually in tension. The notable result here is not that Netlify chose stronger isolation at any cost. It reports lower request overhead after making that change.

The migration therefore challenges a common assumption: process-level isolation is not automatically the only practical option for very low-latency serverless work. If the VM lifecycle is reduced far enough, hardware virtualization can fit inside a single-digit-millisecond budget.

Owning the runtime changes the feature ceiling

The previous hosted provider also controlled the runtime. Netlify says that limited features such as filesystem access, WASM loading, and native npm modules. In the new system, Deno remains the JavaScript runtime, but Netlify ships and patches its own build inside the MicroVM.

That means infrastructure ownership affects product capability, not just cost or latency. A runtime provider can be an excellent abstraction until the platform needs a feature below that abstraction.

Netlify points to native npm modules, filesystem behavior, and future runtime choices as areas that become more tractable with a full VM. Unikraft separately notes that Netlify can deploy its own runtime changes without waiting for Unikraft to coordinate the rollout.

The distinction is useful for platform design. Outsourcing execution minimizes the amount of infrastructure a team must operate. Bringing execution in-house increases the operational surface, but it also moves the constraint boundary. The right choice depends on whether the managed boundary still matches the product roadmap.

Reliability comes from boring control-plane decisions

Fast boot is only useful if the fleet remains healthy under real traffic. Netlify describes several less glamorous choices that make the system operable: local DNS resolvers on compute nodes, circuit breakers for rerouting, health tracking in a control plane, and separate compute-node fleets that can be rolled out beside the current fleet before taking traffic.

The control plane gives edge nodes a list of healthy compute nodes. During a deployment, a new fleet comes up alongside the existing one, scales to the required size, proves healthy, and only then takes traffic. That keeps rollout and rollback separate from the per-request execution path.

Unikraft adds another scale detail: Netlify can create hundreds of millions of MicroVM invocations per day, and customer traffic can spike far above its normal level. At that scale, tiny-probability faults stop being theoretical. The architecture must assume rare failures will eventually happen.

That is why observability is part of the performance design. Netlify records boot time, time to first port open, and time until user code starts. Those measurements let engineers distinguish a slow image fetch from a slow VM restore or a slow customer function instead of collapsing everything into one latency number.

The migration is a lesson in moving work off the critical path

Several optimizations in the new design share the same shape. They do not merely make an operation faster; they avoid doing it for most requests.

Images are fetched only on a cache miss. The full fleet is not updated for every function deployment. Idle MicroVMs scale to zero instead of consuming capacity. A restored snapshot is memory-mapped rather than copied in full before execution. Rendezvous hashing keeps a service near already-cached code until load makes that placement harmful. Compute nodes use local DNS rather than turning every lookup into another remote dependency.

This is a useful way to review any serverless request path. List every operation between accepting a request and producing the first response byte. For each operation, ask whether it must happen on every request, whether its result can be cached, whether it can be prepared lazily, and whether failure can be isolated without blocking unrelated tenants.

A rough review template is:

for each critical-path operation:
    measure its p50 and p99
    identify reusable state
    move reusable state closer to the request
    remove remote dependencies where practical
    define a bounded fallback when locality fails
Enter fullscreen mode Exit fullscreen mode

The important word is bounded. A cache or sticky-placement strategy that has no escape route becomes a source of hot spots. A snapshot that cannot be invalidated becomes stale state. A local dependency without health checks becomes a silent single point of failure.

What to copy from this architecture

Most teams should not build an edge compute platform. The reusable ideas sit one level above the implementation details.

First, measure the entire request path, not just application execution time. A fast runtime can still sit behind slow routing, remote provisioning, or logging infrastructure.

Second, make locality intentional. Stable placement and local image caches reduce repeated work, but load-aware spreading is necessary when locality starts harming capacity.

Third, use identity to enforce isolation. Netlify derives a service ID from the machine specification and site-specific state, so a deployment change naturally creates a separate execution service. That is safer than hoping two logically different deployments remain separated by convention.

Fourth, treat cold starts as a distribution problem as well as a boot problem. If code or images must cross the network before execution, a very fast VM cannot hide an inefficient artifact-delivery path.

Fifth, keep rollout control outside the hot path. A control plane can decide which nodes are healthy and which fleet version should receive traffic without forcing every request to perform control-plane work.

Netlify's reported result is a warm Edge Function path of roughly 5–6 ms at the median, with cold image fetches affecting about 1.2 percent of invocations and averaging roughly 9 ms. The useful engineering story is how those numbers were obtained: fewer remote boundaries, on-demand artifacts, deterministic locality, fast snapshot restore, hardware-backed tenant separation, and explicit escape paths when the fast path becomes a hot spot.

That is a more durable lesson than "MicroVMs beat isolates." Architecture wins here because the whole request path changed.

Sources


Originally published on Dispatch.

Top comments (0)