Introduction
Over three years, we migrated our latency-sensitive services from a mix of Python, Go, JVM, and C to Rust, achieving significant performance improvements. Load balancer latency dropped from 600ms to 101ms, publish API latency from ~350µs to ~50µs, and a Presence API redesign reduced peak memory by 6x. However, this migration also exposed a critical issue: unbounded in-flight work in the Tokio runtime led to excessive memory usage, causing a 100 MiB pod to balloon to 3.7 GiB under load.
The Rust Advantage
Rust's memory management, devoid of garbage collection, eliminated GC pauses, resulting in more predictable latency and reduced memory overhead. This was particularly beneficial for our high-volume message bus, where millions of concurrent occupants required efficient memory usage. For example, the replication layer achieved steady-state latency of 27-31µs, with a ~1.5µs gap between nodes becoming visible—a gap previously obscured by GC noise.
The Tokio Pitfall
While Rust's concurrency model provided safety through compile-time checks, Tokio's task scheduling allowed high concurrency with low CPU usage. However, each task retained state, and without an explicit cap on in-flight work, memory accumulated during bursts. This issue was exacerbated by the lack of backpressure mechanisms, allowing tasks to pile up unchecked. The causal chain was clear: bursts → unbounded tasks → memory accumulation → pod memory exhaustion.
Stakes and Timeliness
Without addressing the unbounded in-flight work issue, the system risked memory exhaustion, service instability, and potential downtime during traffic bursts, undermining the performance gains achieved through the migration. As more organizations adopt Rust for performance-critical systems, understanding both its advantages and the pitfalls of its ecosystem, such as Tokio's handling of asynchronous tasks, is crucial for building reliable and efficient services.
Technical Analysis
Performance Improvements
The migration to Rust yielded substantial performance gains. For instance, the load balancer latency dropped from 600ms to 101ms after replacing nginx with Pingora, Cloudflare's Rust proxy framework. Similarly, the publish API latency decreased from ~350µs to ~50µs, with variance reduced significantly. These improvements were not just in averages but also in predictable latency, as seen in the replication layer, where a 100µs excursion became noticeable and actionable.
Memory and CPU Efficiency
Rust's memory management eliminated GC pauses, making memory behavior easier to reason about. For example, the Presence API saw peak memory usage drop from 3.4 GiB to a 256-512 MiB band after a redesign. CPU usage also halved, with pods stabilizing at ~0.22 cores compared to 0.4-0.6 cores in the old implementation. This efficiency allowed us to run fewer pods with more headroom.
The Unbounded In-Flight Work Issue
Despite these gains, the unbounded in-flight work in Tokio became a critical issue. During bursts, tasks accumulated in the runtime, each holding buffers and state while waiting on I/O. This led to memory growth proportional to the arrival rate. Our initial response—tuning the Horizontal Pod Autoscaler (HPA)—only exacerbated the problem by providing more places for the backlog to accumulate. The solution was to introduce backpressure at the ingest edge, capping in-flight work and shedding excess requests. This change prevented memory from growing uncontrollably during bursts.
Lessons Learned
Instrument Before You Start
Every dashboard in this post existed before the migration, allowing us to quantify improvements like "600ms to 101ms". Instrumentation is critical for understanding baseline performance and measuring the impact of changes.
Budget for the Learning Curve
The first six months with Rust were challenging, with engineers frustrated by the borrow checker. By month eight, productivity improved as teams became more comfortable with Rust's concurrency model. This learning curve is a necessary investment for long-term gains.
Optimize for Predictable Latency
Customers notice spikes more than average latency. Rust's elimination of GC pauses and the introduction of backpressure mechanisms helped stabilize latency, making small regressions easier to detect and resolve.
Treat the First Rust Release as a Baseline
Comparing Rust implementations (e.g., panels 8 and 9) highlights that language choice is not the only factor. Architectural improvements, such as connection pooling and data layout optimizations, are equally critical for maximizing performance.
Work on Durability Alongside Performance
We replaced our custom TCP protocol with gRPC and a store-and-forward model to improve message durability and retry mechanisms. This change reduced the risk of message loss and simplified failure handling, even though it introduced the possibility of duplicate deliveries. The idempotency key ensured that duplicates were not observable at the application layer.
Conclusion
Migrating to Rust significantly improved performance and stability, but it also exposed the pitfalls of unbounded in-flight work in Tokio. By addressing this issue with backpressure mechanisms, we achieved a system that is both efficient and reliable. The migration justified its three-year investment through reduced memory and CPU usage, fewer warnings, and the retirement of a complex wire protocol. For latency-sensitive workloads, Rust's memory management and concurrency safety make it a compelling choice, but careful attention to task scheduling and backpressure is essential to avoid memory exhaustion.
Rule for Choosing a Solution
If you are migrating latency-sensitive services to Rust and using Tokio for asynchronous task handling, use explicit backpressure mechanisms to cap in-flight work. This prevents memory accumulation during bursts and ensures stable performance under load.
Scenarios and Analysis
Scenario 1: Load Balancer Latency Reduction
The migration of the load balancer from nginx to Pingora, a Rust-based proxy framework, resulted in a dramatic reduction in latency from 600ms to 101ms. This improvement is attributed to Rust's elimination of garbage collection (GC) pauses, which previously introduced unpredictable latency spikes. The causal chain is as follows: GC pauses → unpredictable latency → high average latency. By removing GC, Rust ensures more predictable memory management, allowing the load balancer to handle requests with consistent performance. However, this scenario also highlights the need for instrumentation to measure baseline performance and validate improvements, as evidenced by the detailed dashboards used to track the migration.
Scenario 2: Publish API Latency and Variance
The publish API latency dropped from ~350µs to ~50µs, with a significant reduction in variance. This improvement is due to Rust's fine-grained control over memory allocation and the absence of GC pauses. The causal mechanism is: GC pauses → latency spikes → high variance. Additionally, the use of connection pooling in the Rust implementation further stabilized latency by reusing established connections, reducing the overhead of connection setup. This scenario underscores the importance of architectural optimizations alongside language choice, as connection pooling was a critical factor in achieving low variance.
Scenario 3: Presence API Memory Reduction
The Presence API, which tracks millions of concurrent occupants, saw a 6x reduction in peak memory usage from 3.4 GiB to 256-512 MiB. This was achieved through a combination of Rust's memory management and a redesigned data layout. The causal chain is: inefficient data layout → high memory overhead → memory exhaustion risk. Rust's ownership model and lack of GC allowed for more efficient state management, while the redesign optimized per-occupant overhead. This scenario highlights the need to treat the first Rust release as a baseline and iteratively optimize, as the initial Rust implementation still required architectural improvements to maximize memory efficiency.
Scenario 4: Unbounded In-Flight Work in Tokio
The most critical issue emerged from unbounded in-flight work in the Tokio runtime, causing pod memory usage to spike from ~100 MiB to 3.7 GiB during bursts. The causal mechanism is: unbounded tasks → memory accumulation → pod memory exhaustion. Each task in Tokio holds state while waiting on I/O, and without an explicit cap, bursts of requests led to unchecked memory growth. The solution involved introducing backpressure at the ingest edge, capping in-flight work and shedding excess requests. This scenario demonstrates the risk of overlooking async runtime limitations and the importance of implementing explicit backpressure mechanisms. The rule for choosing a solution is: if using Tokio for high-concurrency services, always implement backpressure to prevent memory exhaustion during bursts.
Scenario 5: Replication Layer Optimization
The replication layer achieved a steady-state latency of 27-31µs, with a visible 1.5µs gap between nodes. This level of precision was only possible after stabilizing runtime noise, which previously masked small regressions. The causal chain is: runtime noise → obscured regressions → inability to optimize. Rust's elimination of GC pauses and fine-grained control over memory allowed for microsecond-level optimizations. This scenario emphasizes the need to optimize for predictable latency and the value of instrumentation in identifying and addressing subtle performance issues.
Root Cause Analysis and Implications
The root cause of the Tokio memory issue lies in the lack of explicit bounds on in-flight work, a common pitfall in async runtimes. The causal mechanism is: async task accumulation → state retention → memory growth. While Tokio's scheduler is highly efficient for low-CPU tasks, it does not inherently limit memory usage. This issue is exacerbated in latency-sensitive services, where bursts are common. The optimal solution is to implement backpressure, as it directly addresses the unbounded growth by capping the number of in-flight tasks. Without backpressure, even horizontal scaling (e.g., adding more pods) can worsen the problem by providing more places for tasks to accumulate. The rule for choosing a solution is: if handling bursts in an async runtime, use backpressure to prevent memory exhaustion.
Practical Insights
- Instrumentation is non-negotiable: Baseline metrics are essential for validating performance improvements and detecting regressions.
- Architectural optimizations matter: Language choice alone is insufficient; data layout, connection pooling, and state management are critical.
- Backpressure is mandatory for async runtimes: Without it, memory exhaustion during bursts is inevitable.
- Iterative optimization: Treat the first Rust release as a baseline and continuously refine both code and architecture.
Edge-Case Analysis
In edge cases where bursts are extremely large or unpredictable, even backpressure may not suffice. For such scenarios, consider shedding excess requests or implementing a priority queue to ensure critical tasks are processed first. However, shedding requests must be balanced against service availability, as it can lead to dropped messages. The optimal approach depends on the specific workload and SLOs. The rule for choosing a solution is: if bursts exceed backpressure capacity, combine backpressure with request shedding or prioritization.
Mitigation and Lessons Learned
The migration to Rust delivered significant performance gains, but the unbounded in-flight work issue in Tokio exposed a critical risk: memory exhaustion during traffic bursts. Here’s how we addressed it and the lessons learned from managing high-performance Rust applications.
Mitigation Strategy: Backpressure and Bounded Queues
The root cause was unbounded task accumulation in Tokio. Each task held state while awaiting I/O, and bursts overwhelmed the heap. Our solution:
- Backpressure at Ingest: We introduced a bounded queue at the ingest edge, capping in-flight work. Excess requests are either shed or queued, preventing memory growth with burst size.
- HPA Tuning: Lower scale-up thresholds and faster reaction times in the Horizontal Pod Autoscaler (HPA) were initially tried but failed, as scaling out merely distributed the unbounded backlog across more pods.
This change reduced pod memory from 3.7 GiB to ~100 MiB baseline during comparable bursts, stabilizing the system under load.
Key Lessons and Best Practices
- Explicit Backpressure is Mandatory in Async Runtimes
Tokio’s async model allows high concurrency with low CPU, but every task retains state. Without bounds, bursts consume memory linearly. Rule: Always implement backpressure in async runtimes handling bursts.
- Instrument Before Migrating
All dashboards existed pre-migration, enabling precise before/after comparisons (e.g., 600ms → 101ms for load balancer latency). Rule: Baseline metrics are non-negotiable for performance validation.
- Treat First Rust Release as a Baseline
Panels 8 and 9 show Rust vs. Rust comparisons: Presence memory dropped 6x post-redesign, and FCM variance vanished with connection pooling. Rule: Architectural optimizations are as critical as language choice.
- Budget for the Learning Curve
The first 6 months with Rust’s borrow checker were costly. By month 8, productivity improved. Rule: Allocate time for team adaptation to Rust’s strict ownership model.
Edge-Case Analysis: When Backpressure Fails
Backpressure works for predictable bursts, but may fail under:
- Extreme Burst Sizes: If bursts exceed queue capacity, requests are shed, potentially impacting availability.
- Unpredictable Traffic Patterns: Dynamic bounds are harder to manage than static ones.
Solution: Combine backpressure with priority queues or graceful shedding to balance load and availability. Rule: For extreme cases, prioritize critical requests over strict bounds.
Practical Insights
- Memory Predictability: Rust’s ownership model simplifies memory reasoning, but data layout still matters (e.g., Presence’s 6x memory reduction).
- Concurrency Safety: Rust’s compiler caught ownership and data-race errors, reducing post-deployment bugs.
-
Toolchain Efficiency:
cargounified build, test, and formatting, saving time compared to multi-language infrastructures.
In summary, Rust’s performance benefits are clear, but asynchronous runtime pitfalls like unbounded work require proactive mitigation. Backpressure is non-negotiable for high-concurrency systems.
Conclusion and Future Outlook
The migration to Rust delivered transformative performance gains, but the unbounded in-flight work issue in Tokio exposed a critical pitfall in asynchronous runtime management. This section distills the findings, emphasizes the importance of addressing runtime limitations, and outlines future improvements for Rust and Tokio.
Key Findings and Lessons
Rust’s memory management eliminated garbage collection (GC) pauses, reducing latency spikes and memory overhead. However, Tokio’s unbounded task scheduling led to memory accumulation during bursts, as each task retained state while awaiting I/O. This causal chain—bursts → unbounded tasks → memory accumulation → pod exhaustion—underscores the need for explicit backpressure mechanisms in async runtimes.
Architectural optimizations, such as connection pooling and data layout redesigns, were as critical as the language choice. For example, the Presence API’s 6x memory reduction was achieved by optimizing Rust’s data layout, not just by eliminating GC.
Future Improvements in Rust and Tokio
Rust’s ecosystem continues to mature, but Tokio’s handling of unbounded work remains a risk for high-concurrency systems. Future improvements should focus on:
- Built-in Backpressure Primitives: Tokio could introduce APIs for capping in-flight work, making backpressure implementation less error-prone.
- Memory Profiling Tools: Enhanced tooling to detect and diagnose memory accumulation patterns in async tasks.
- Concurrency Model Evolution: Rust’s compiler already enforces concurrency safety, but runtime-level safeguards for async workloads would further reduce risks.
Encouraging Community Research and Collaboration
The community must prioritize research into async runtime limitations, particularly in high-burst scenarios. Collaboration between Rust and Tokio developers could lead to:
- Standardized Backpressure Patterns: Establishing best practices for bounding in-flight work across async runtimes.
- Edge-Case Testing Frameworks: Tools to simulate extreme burst conditions and evaluate runtime resilience.
- Cross-Language Comparisons: Analyzing how other async runtimes (e.g., Go’s goroutines) handle similar challenges to inform Rust’s evolution.
Practical Insights for Performance-Critical Systems
When using async runtimes like Tokio, implement backpressure at the ingest edge to cap in-flight work. For example, a bounded queue with request shedding prevents memory growth during bursts. However, this approach may fail under extremely large or unpredictable bursts, requiring additional mechanisms like priority queues.
Rule for Choosing a Solution: If your system handles bursts with async tasks, use backpressure to bound in-flight work. Combine with request shedding or prioritization for edge cases.
In conclusion, Rust’s performance benefits are undeniable, but asynchronous runtime pitfalls like unbounded work require proactive mitigation. By addressing these limitations and fostering community collaboration, we can build more reliable and efficient systems.

Top comments (0)