DEV Community

Lightning Developer
Lightning Developer

Posted on

Beyond Isolates: How MicroVMs are Redefining Edge Function Performance

Moving Beyond the Isolate Paradigm

The serverless and edge computing landscape is currently undergoing a significant shift. For years, the industry standard for fast, ephemeral execution has been the V8 isolate. However, as infrastructure demands grow and security requirements tighten, developers are rethinking the balance between speed and isolation. Netlify recently made headlines by migrating their Edge Functions away from V8 isolates and onto Firecracker microVMs. This transition led to a dramatic reduction in median warm latency, dropping from the 25-40 ms range down to a blistering 5-6 ms. For those familiar with traditional virtualization, this sounds counterintuitive. We usually categorize isolates as the lightweight, high-performance option and virtual machines as heavy, slow, and resource-intensive.

To understand why this change makes sense, we need to peel back the layers of the architecture and look at what happens under the hood when a request hits the edge. The performance gains reported are not just about the runtime itself but about the elimination of unnecessary network hops and the clever orchestration of microVM lifecycles.

Blog Image

The Anatomy of an Isolate

A V8 isolate functions as a discrete heap and execution context contained within a single process. By utilizing a single runtime instance to manage thousands of these isolated contexts, platforms achieve remarkably fast startup times. Because there is no need to boot a kernel or manage a separate process for every tenant, isolates are effectively the speed-kings of serverless functions.

However, this speed comes at a cost regarding the isolation boundary. In a shared-process model, the security wall between your code and another tenant's code is maintained by the memory safety mechanisms of the engine itself. If that engine has a vulnerability, or if the runtime environment shares too many system-level resources, the isolation is significantly weaker than what hardware virtualization provides. Netlify’s move underscores a vital engineering principle: as the scale of deployments increases, the risk surface area grows, and hardware-backed isolation becomes a necessary trade-off.

Blog Image

Rethinking Latency at the Edge

Previously, Netlify handled high-volume traffic by routing requests out of their primary network to a separate execution service and back. This round-trip, combined with the overhead of the existing architecture, created the 25-40 ms latency floor. By redesigning the compute layer to reside directly inside the edge network, the system now functions in a more unified pipeline:

  1. Edge nodes handle TLS termination and request routing.
  2. Compute nodes manage the microVM lifecycle and cache function images locally.
  3. The control plane coordinates the entire orchestration layer.

By keeping the execution path contained within a single network environment, the system drastically reduces the latency previously lost to the network backhaul.

The MicroVM Performance Secret: Snapshots and Memory Mapping

If we assume a traditional VM boot process, we are talking about hundreds of milliseconds or even seconds. To achieve a single-digit millisecond response time, the system cannot perform a standard cold boot. Netlify achieves this by treating the microVM as an ephemeral entity that uses snapshots for instantaneous restoration.

Memory-Mapped Images

The function code is bundled as an uncompressed EROFS (Enhanced Read-Only File System) image. By memory-mapping this image, the guest VM does not need to load the entire bundle into memory. Instead, the guest kernel reads only the specific pages of code required to execute the function. Because this is done via memory mapping rather than standard file I/O, the performance hit is negligible, allowing the function to begin execution immediately.

The Lifecycle of a Snapshotted VM

When a microVM goes idle, the platform does not simply terminate it without a plan. It captures the state of the VM, including memory contents and emulated device states, into a snapshot. Subsequent requests can be served by restoring from these snapshots, effectively bypassing the entire OS boot sequence. This is a critical departure from traditional serverless cold starts.

Blog Image

However, there are technical challenges with snapshotting, specifically regarding guest state persistence. If a guest VM resumes from the same snapshot multiple times, it must ensure that unique values like random tokens or identifiers are handled correctly. Systems typically address this using technologies like VMGenID to trigger re-seeding of random number generators upon restoration, ensuring that every "clone" is unique enough to maintain security invariants.

Routing and Sticky Compute Nodes

Restoring a snapshot is only efficient if the necessary data is already present on the compute node. To solve this, the infrastructure utilizes rendezvous hashing. By assigning specific functions to specific nodes, the system ensures that the most frequently used functions have their images and snapshots already primed in memory.

This approach does come with a potential downside: hotspots. If a single function receives a massive influx of traffic, a single compute node could be overwhelmed. To mitigate this, the routing logic is designed to be elastic. If traffic patterns exceed predefined thresholds, the system relaxes the strict hash, spreading the load across a broader pool of nodes. While this might lead to a cold start for a small percentage of requests, it ensures the overall system remains resilient and available. Current benchmarks suggest that only about 1.2 percent of invocations face a cold start, which is a testament to the effectiveness of this routing strategy.

Practical Benchmarking

If you are skeptical about these latency numbers, the best way to verify them is to run your own experiments. You can measure the time-to-first-byte (TTFB) using a standard CLI command to see exactly how your functions are performing under load.

for i in $(seq 1 50); do
  curl -s -o /dev/null -w '%{time_starttransfer}\n' https://your-site.netlify.app/your-edge-path
done | sort -n | awk '{a[NR]=$1} END {print "p50:", a[int(NR*0.5)], "p99:", a[int(NR*0.99)]}'
Enter fullscreen mode Exit fullscreen mode

By comparing this result against a static asset on the same site, you can isolate the specific time added by the function execution itself. It is also worth noting that if you wish to experiment with Firecracker virtualization on your own hardware, you will need a machine capable of supporting KVM. The official Firecracker documentation provides an excellent roadmap for getting started with snapshotting and custom guest kernels.

Architectural Considerations and Future Limits

While the underlying infrastructure has changed, the developer experience remains unchanged. You continue to write, test, and deploy your code exactly as you did before. There is no manual migration or configuration update required to take advantage of this new compute model.

As of now, the current constraints on compute duration, memory, and package size remain in place. These limits were originally established to accommodate the constraints of the old isolate-based runtime, but they are not inherent to microVM technology. As the ecosystem matures, we can likely expect these limits to be revisited, potentially allowing for more complex applications and larger runtime environments at the edge.

Ultimately, the shift from isolates to microVMs demonstrates that performance at the edge is not solely determined by the "lightness" of the runtime. It is determined by the intersection of smart architectural design, efficient state management, and the ability to reduce overhead at every stage of the request path. By choosing hardware-level isolation, providers are opting for a more robust security model that doesn't have to sacrifice speed, provided they have the engineering rigor to implement snapshotting and intelligent routing effectively.

Reference

Top comments (0)