DEV Community

Vincent Tran
Vincent Tran

Posted on Originally published at 0xgosu.dev on

How Rust Made Its Compiler Faster While Changing Its Core

Rust’s compiler is changing two of its hardest internal systems while its engineers are making ordinary builds faster. Nicholas Nethercote’s September performance report puts numbers on that work: across 629 benchmark measurements between July 29 and September 28, 555 improved and 74 regressed. Mean wall time fell 4.57%. A few crates saw much larger gains.

That average is useful, but it hides the real story. Some wins came from a newer LLVM or a better-trained Clippy binary. Others came from removing allocations in frequently visited compiler paths. The most striking case came from changing the order in which a dataflow analysis revisits blocks. Meanwhile, Rust enabled a more precise borrow checker and a new trait solver on nightly, each of which brought its own performance problems to solve.

These figures describe Rust’s benchmark suite over that interval. They are not a promise that every project will compile 4.57% faster, nor can every gain be added together as an independent percentage. Different changes affect different crates, build modes, and compiler phases.

A broad gain, assembled from many small changes

The compiler does more than translate source into machine code. It parses and lowers syntax, resolves types and traits, checks borrows, runs MIR analyses, and eventually asks LLVM to optimize and generate code. Clippy and rustdoc exercise related but distinct paths. A change in one layer can help a particular workload while barely moving another.

“Dark
There was no single speed switch. Different improvements reduced work in different compiler paths.

The LLVM 23 update delivered a 1.2% mean wall-time reduction across the compiler benchmarks on its own. That is a substantial result for a dependency upgrade, though it does not explain every September improvement. At the tooling edge, Rust’s build process began applying profile-guided optimization (PGO) to Clippy. PGO builds a program using profiles gathered from representative executions, allowing the optimizer to spend effort on paths the tool actually takes. Nethercote reports Clippy wall-time improvements across most of its benchmarks, reaching 18% in the best case. That figure is a best case for Clippy, not for all Rust builds.

The rustdoc gains discussed in an earlier report also had a follow-up explanation from Noah Lev. They belong to the same larger effort: measure a real workload, find a repeated cost, and remove it without weakening the tool’s behavior.

New semantics have a performance bill

Rust enabled Polonius Alpha , its next borrow checker, on nightly. It is more precise about when a borrow is live along a particular control-flow path. The Rust team’s announcement gives a useful example: a map lookup can return a mutable reference in one match arm while the other arm inserts a default value. The older analysis can treat the returned borrow as live across both arms and reject valid code. Polonius Alpha understands that the borrow from the first arm does not exist on the insertion path.

That precision takes work. Most measured crates have little change, but some see regressions, including serde. Jack Huey made liveness calculations lazy, reducing instruction counts for serde by 3–5% and producing smaller gains elsewhere. A data-structure adjustment and inlining changes trimmed more work across other benchmarks. The project still has slower outliers to investigate. These are nightly changes under evaluation, not a claim that the new checker is already faster for every crate.

The next-generation trait solver also became the nightly default. A trait solver proves obligations such as whether a type implements a trait and what an associated type resolves to. The Rust team’s release note calls this a major replacement of the compiler’s type reasoning machinery. It fixes long-standing behavior and prepares future language features, but its early implementations could perform much worse on certain trait-heavy crates.

Work by Jana Dönszelmann, Nethercote, and others attacked those outliers. Several targeted pull requests cut compile time for individual stress cases by 15%, 25%, or 50% , while Dönszelmann’s account explains the broader solver effort. Those numbers refer to particular cases; they should not be presented as suite-wide savings. They matter because a new default is only practical if pathological costs are brought under control.

Removing work from hot paths

Many compiler improvements are less dramatic to describe but easier to repeat. A specialization-graph change reduced mean cycle counts by 1.58% over the benchmark suite. Another change made loading incremental compilation data cheaper, with instruction-count wins of up to 6% in individual measurements. Avoiding allocations while processing obligations reached 2% in the best case. Replacing dynamic dispatch with static dispatch in the selection between old and new trait solvers removed another frequently observed allocation path, mostly yielding gains below 1%.

The pattern is important. A tiny allocation in cold initialization code barely matters. The same allocation inside a path visited for thousands of obligations can dominate a profile. A change that seems small in source code can therefore move a whole-suite metric when it sits at the right frequency. Conversely, a good-looking microbenchmark can fail to move actual compile time if its path is rarely used. That is why the report distinguishes instruction counts , CPU cycles , and wall time instead of treating them as interchangeable.

There were smaller wins in other corners. Rust increased the compiler’s default stack size so it could remove some manual stack-extension machinery from recursion-heavy paths; instruction counts fell across many benchmarks, by almost 3% in the best case. A cleanup in lowering the abstract syntax tree to the high-level intermediate representation unexpectedly cut instruction counts by up to 1.5%. The compiler team also sometimes merged several performance PRs as a rollup to manage CI capacity, then ran per-change performance measurements afterward so unexpected effects could still be found.

This work involved many contributors. Nethercote singled out new contributor xmakro for several of the hot-path changes. He also described using language models for private analysis on some PRs while writing his own code and prose, consistent with the Rust project’s LLM usage policy. The report closes with his move to work on the project’s compiler performance goal at Hexcat. Those details underline that the result is an ongoing engineering program, not a one-off optimization sprint.

One traversal change, 1.4 million fewer block visits

The clearest illustration of algorithmic cost is the control-flow graph traversal change. A dataflow analysis keeps facts about a program at each basic block. It repeatedly processes blocks until those facts stop changing: a fixpoint. The order of processing can decide how often the same block must be revisited.

Rust’s old worklist began in a sensible order, but when a loop sent new information backward it appended dirtied blocks to a first-in, first-out queue. Successors could then be processed before the earlier block had settled, causing another round of work later. The new algorithm selects the earliest dirty block in dataflow order. It aims to propagate changed input through the graph before spending time on downstream blocks whose inputs are likely to change again.

“Dark
Processing the earliest dirty block first avoids much repeated work in one unusually large control-flow graph.

For a cranelift-codegen function with more than 18,000 basic blocks , the EverInitializedPlaces analysis previously called apply_effects_in_block about 1.5 million times before reaching a fixpoint. The new traversal needed about 90,000 calls. Nethercote reports about a 30% wall-time reduction for a check build of that crate. The pull request’s measurements also show that the biggest gain was concentrated in that particular analysis. Most ordinary functions are small enough that the new ordering makes little difference. This is a case where one outlier exposed avoidable work in a general algorithm.

A second change stopped EverInitializedPlaces from tracking unnecessary data for projections. That reduced instruction counts on the match-stress benchmark by 17% , with smaller effects elsewhere. The two fixes illustrate different levers: do fewer iterations, and make each remaining iteration cheaper.

What the measurements actually tell us

Compiler performance is a distribution. The 4.57% mean improvement and the 555-to-74 split indicate broad progress, while the double-digit examples explain where engineering effort had unusually high payoff. The remaining regressions matter too, particularly for new compiler machinery that expands what Rust can accept. A fast average would be much less reassuring if a widely used crate became prohibitively slow.

The most useful lesson from this period is methodological. Rust’s team paired large semantic changes with benchmark runs, inspected specific slow crates, and kept cutting costs in the paths that profiles showed were hot. Some changes improved a whole suite; others rescued a single pathological case. Together they made the compiler faster during a period when it was also becoming more capable.

Sources

Top comments (0)