DEV Community

Masahiro Morino
Masahiro Morino

Posted on AI-assisted

Reading Rust's MIR: Following Control Flow and Values

I started looking into linkers, but a linker's job isn't specific to Rust. Before getting there, I thought it would be more interesting to see how rustc takes our Rust code toward code generation.

One of the representations we meet along the way is MIR. Apparently, it's used for both borrow checking and code generation. But what exactly is an “intermediate representation,” anyway?

In this article, I'll use a few small examples to explore how to read Rust operations in the compiler's MIR. We'll start with an if expression to get a feel for control flow, look closely at a single assignment to see how values and places are represented, and then move on to a call to Vec::push.

This article is based on the Rust Compiler Development Guide as consulted on September 29, 2026. The examples were checked with Rust 1.96.0. MIR's structure and textual output can change with the compiler version and optimization settings.

What is MIR for?

MIR stands for Mid-level Intermediate Representation. An intermediate representation is a way for a compiler to represent code while processing a program.

The Rust we write has constructs such as if, match, and nested expressions that make programs convenient to write and read. The compiler needs to answer questions like “Where is this value created, and where is it used?” or “After taking this branch, what happens next?” MIR organizes the program into a form that makes types, operations on values, and control flow explicit. That gives the compiler something it can analyze. The MIR — Rust Compiler Development Guide

For example, the expression if flag { x } else { 0 }, which we'll look at shortly, becomes a condition check, operations that set the return value, and a block where the paths meet. Complex expressions are broken down into smaller computations and temporary values. A more uniform representation lets the compiler follow values and control flow without having to handle every variation of the source syntax separately. MIR construction — Rust Compiler Development Guide

One major use of MIR is borrow checking. Remember borrowing? The compiler examines how values are used and borrowed on MIR, and also uses MIR for subsequent optimization and code generation. That doesn't mean one unchanged MIR body survives all the way through compilation. It passes through analysis and transformation stages on its way to the form used for code generation. Overview of the compiler

Let's pause and put these representations side by side.

Representation Its role in the path we're looking at
Rust source code Where we write the program using expressions and other language constructs
MIR Makes Rust types, operations on values, and branches explicit for borrow checking, optimization, and code generation
LLVM IR Carries the program from rustc into LLVM for further optimization and machine-code generation

There are other compiler representations before MIR, but I'm leaving those out here. I'd like to explore that part of the compiler sometime, too. For now, let's focus on MIR as a representation for analyzing Rust programs and look inside it. Overview of the compiler

In MIR, control flow becomes blocks

One characteristic of MIR is that control flow takes the form of a control-flow graph, or CFG. Operations are grouped into basic blocks, connected by branches and jumps. Each block contains statements, such as assignments, and ends with a terminator, such as a branch, a function call, or a return. The MIR — Rust Compiler Development Guide

Here's a small function:

pub fn choose(flag: bool, x: u32) -> u32 {
    if flag { x } else { 0 }
}
Enter fullscreen mode Exit fullscreen mode

It returns x when flag is true, and 0 when it's false. Compiling this function to MIR with Rust 1.96.0 gives us four basic blocks:

bb0 checks flag. The true branch enters bb1 and sets the return value to x; the false branch enters bb2 and sets it to 0. Both paths join at bb3, which returns.

Figure 1: The block structure of the generated MIR. The diagram uses flag, x, and return value in place of the internal local names.

In bb1, setting the return value to x is a statement, and jumping to bb3 is the terminator. bb0 has no statements at all, just the conditional-branch terminator. So a block doesn't have to do some computation before it branches.

A basic block is a unit that executes its statements in order and ends with a terminator. A terminator doesn't necessarily choose between multiple destinations. In this example, goto has one successor, switchInt selects a successor based on a value, and return has no successor block within this function. Thinking of a terminator as the block's exit helps explain why these operations belong to the same category.

Matching the MIR output to the diagram

In the output, arguments and the return value are represented by names such as _1, _2, and _0. Here, _1 is flag, _2 is x, and _0 holds the return value. If we omit the declarations and keep just the blocks, the correspondence with the diagram becomes clear:

bb0: {
    switchInt(copy _1) -> [0: bb2, otherwise: bb1];
}

bb1: {
    _0 = copy _2;
    goto -> bb3;
}

bb2: {
    _0 = const 0_u32;
    goto -> bb3;
}

bb3: {
    return;
}
Enter fullscreen mode Exit fullscreen mode

What was one if expression in the source has become a conditional branch, assignments, and a block where the paths meet. Making control flow explicit gives the compiler a way to analyze each path.

Reading one line of MIR: Places and values

Now that we know how the blocks connect, let's look more closely at the assignment in bb1:

_0 = copy _2;
Enter fullscreen mode Exit fullscreen mode

It simply puts x into the return value, but the pieces have different roles in MIR. The key terms are Local, Place, Operand, and Rvalue. We're starting to collect quite a few terms now.

A Local is storage; a Place identifies a location

A Local is a storage location for something like a function argument, a local variable, or a temporary value. MIR identifies locals by indices, written as _1, for example. These are the locals in choose:

Local Role in this function Type
_0 The special location for the return value u32
_1 The argument flag bool
_2 The argument x u32

Storage here is a compiler-level concept. It doesn't mean that every Local necessarily ends up with its own stack slot in the final machine code. And the same-looking local names can have different roles in different functions, as we'll see below.

A Place is an expression identifying a location to read from or write to. It can refer to an entire Local, a field within it, or a location reached through a reference. For example, if _1 were a tuple, _1 would identify the whole tuple and _1.0 its first field. Yes, this gets a little confusing.

A Local is a unit of storage; a Place specifies a particular location. In _0 = copy _2;, both the destination _0 and the source _2 are Places referring to entire Locals. Key MIR vocabulary

An Operand supplies a value; an Rvalue produces the value to assign

Identifying a location doesn't yet say how we'll use the value stored there. That's where an Operand comes in. Three forms we'll encounter are:

Operand Meaning
copy _2 Copy the value from the specified Place
move _3 Move the value from the specified Place
const 0_u32 Use a constant value

An Rvalue, meanwhile, is an expression that produces the value to assign. It might add values, create a reference to a Place, or simply use an Operand's value. The R comes from the right-hand side of an assignment.

Here's how the pieces fit together. This is an explanatory tree, not literal rustc output:

Statement: assignment  _0 = copy _2;
├─ Destination Place: _0
└─ Rvalue: use an Operand's value directly
   └─ Operand: copy _2
      └─ Source Place: _2
Enter fullscreen mode Exit fullscreen mode

The printed copy _2 can be a little puzzling: we're looking at an Rvalue when we consider the whole right-hand side, and an Operand when we focus on its input. The Rvalue in this example wraps one Operand and produces that Operand's value unchanged.

The assignment _0 = const 0_u32; in bb2 has the same shape. The input changes from a value stored in a Local to a constant, but we're still assigning an Operand's value to the return-value Place. Assignments, Rvalues, and Operands in MIR

Why is Vec::push a terminator?

Let's use those terms to read another example, also found in the official guide:

fn main() {
    let mut vec = Vec::new();
    vec.push(1);
    vec.push(2);
}
Enter fullscreen mode Exit fullscreen mode

Compiling it to MIR with Rust 1.96.0 produces this block for the first push. I explicitly used -C panic=unwind to inspect the unwinding path:

bb1: {
    _3 = &mut _1;
    _2 = Vec::<i32>::push(move _3, const 1_i32) -> [return: bb2, unwind: bb5];
}
Enter fullscreen mode Exit fullscreen mode

Here, _1 is the Local for vec. _3 is a temporary Local of type &mut Vec<i32>, and _2 receives the () returned by push. Local indices belong to their individual functions, so these _1 and _2 are different from the ones in choose. Watch out for that!

The first line, _3 = &mut _1;, is an assignment statement. On the left, _3 is the destination Place. On the right, &mut _1 is an Rvalue that creates a mutable reference to the Place _1. Our earlier Rvalue simply used an Operand; this one creates a reference value from a Place.

The next line passes the Operands move _3 and const 1_i32 to push. What's being moved is the mutable reference stored in _3, not the Vec itself.

And that whole call is a terminator. Despite its _2 = ... appearance, it isn't an assignment statement. Look at the destinations at the end:

Destination What happens in this example
return: bb2 If the call returns normally, continue to the block containing the next push
unwind: bb5 If the call unwinds, enter the cleanup block that drops vec

The return: bb2 label does not mean that main finishes. It tells us where execution resumes in this function after the called function returns.

In Rust source, vec.push(1); is one line. In MIR, it is split into preparing the reference argument and a call that explicitly includes the paths execution can take afterward. Treating a function call as a terminator lets us follow both normal execution and unwinding cleanup through the connections between blocks. The guide's Vec example

Generating MIR locally

For the first example, I saved the function in choose.rs and ran:

rustc choose.rs --crate-type lib --edition 2024 \
  -C opt-level=0 --emit=mir
Enter fullscreen mode Exit fullscreen mode

This writes the MIR to choose.mir. The environment was rustc 1.96.0 (ac68faa20 2026-05-25), targeting aarch64-apple-darwin, with LLVM 22.1.2. The output shown here comes from those conditions. Different optimization settings, for example, won't necessarily produce the same block structure.

For the Vec example, I saved the code in vec_push.rs and ran:

rustc vec_push.rs --edition 2024 \
  -C opt-level=0 -C panic=unwind --emit=mir
Enter fullscreen mode Exit fullscreen mode

Even with -C opt-level=0, the emitted MIR isn't necessarily the state before all transformations. The StorageLive statements shown in the official guide don't appear in this Vec output. When comparing dumps, we need to pay attention to both the Rust version and the stage of MIR we're looking at. Setting the MIR optimization level with -Z mir-opt-level=0 requires nightly; this article uses output from stable. MIR output and optimization

Reading control flow and values separately

In choose, a single if expression became a block that checks the condition, two blocks that set the return value, and a block that returns. Separating storage locations such as _0 from units of control flow such as bb0 makes the output easier to follow.

Inside each block, Local and Place tell us which location we're dealing with; Operand and Rvalue tell us how values are used and produced. In the Vec::push example, we also saw that a call is itself a terminator, with normal-return and unwinding paths.

MIR represents Rust operations in a form the compiler can analyze. Follow control flow through the connections between blocks, then distinguish places from values within each line. With those two ideas in mind, it becomes easier to find the original Rust operations in what initially looks like a wall of symbols.

So, when we move from MIR toward machine code, what happens to generic type arguments? And how is the code divided into units that can be compiled in parallel? In the follow-up, I plan to look at collecting the required code and dividing it into Codegen Units. That article is still in preparation.

There's a lot to wrap my head around. But that's what makes it interesting.

Top comments (0)