DEV Community

Masahiro Morino
Masahiro Morino

Posted on AI-assisted

Why Rust Object Files Still Need a Linker: Symbols and Relocations

When you compile Rust code, object files are produced along the way. They have the extension .o and contain machine code and data. The linker combines them to build an executable or another output.

But something bothers me here. If the machine code is already there, shouldn't the CPU be able to execute it? Why is linking still necessary?

Let's compile a pair of small functions separately and look at a state where the machine code exists, but the destination of a function call has not been determined yet. Symbols and relocations will help us read that state.

I tried this on an Apple Silicon Mac. The object format is Mach-O, and the instructions are ARM64. On Linux or Windows, the format and inspection commands differ. The addresses and instruction output below were obtained on October 10, 2026.

Can we generate machine code without the callee's implementation?

In everyday Rust, we define functions with fn and use Rust's usual mechanisms to call functions in other crates. We do not need to write unsafe or extern "C" just to use the linker.

For this experiment, I want to build the caller and callee as separate .o files and compare the state without the callee's implementation with the state after linking. That calls for a slightly unusual setup. extern "C" gives both sides the same calling convention, and #[unsafe(no_mangle)] makes the symbol names easy to follow. The "C" specifies a calling convention; it does not mean we are using C source code. Both implementations are written in Rust.

With an external function declared this way, however, the compiler alone cannot verify that the declaration matches the actual definition. That is why unsafe appears in this declaration and call. It is part of the setup chosen to observe the linker, rather than a requirement for ordinary Rust function calls.

First, let's write a function that simply calls answer. Save this as caller.rs:

#![no_std]

unsafe extern "C" {
    fn answer() -> u32;
}

#[unsafe(no_mangle)]
pub extern "C" fn call_answer() -> u32 {
    unsafe { answer() }
}
Enter fullscreen mode Exit fullscreen mode

We have the body of call_answer, but not the body of answer. All we have for answer is a declaration saying that it takes no arguments and returns a u32.

Since this looks a little different from everyday Rust, let's go through the settings used for the experiment:

  • The unsafe extern "C" block declares an external function. We specify its type and calling convention, but this declaration alone does not let the compiler verify that they match the actual definition.
  • extern "C" selects the C ABI calling convention. Here, it gives the caller and the definition the same convention. We will write the implementation in Rust later.
  • #[unsafe(no_mangle)] disables the usual symbol name mangling so that we can follow the function names easily. It is an unsafe attribute because the names can collide with other symbols.
  • #![no_std] prevents this small library from automatically using the standard library, std. The main function we write later will use std.

For more about these settings, see the Rust Reference sections on External blocks, The no_mangle attribute, and The no_std attribute.

The following command produces an object file before linking:

rustc caller.rs --crate-type lib --edition 2024 \
  -C opt-level=0 -C panic=abort --emit=obj -o caller.o
Enter fullscreen mode Exit fullscreen mode

--emit=obj requests a native object file. --crate-type lib tells rustc to compile code without a main function as a library. Here, we are taking the .o requested by --emit=obj, rather than an .rlib. We also specify -C panic=abort, so unwinding on panic is outside this experiment. Command-line Arguments: --emit

Even without the body of answer, caller.o was produced.

file caller.o
Enter fullscreen mode Exit fullscreen mode
caller.o: Mach-O 64-bit object arm64
Enter fullscreen mode Exit fullscreen mode

So this part works. But what happened to the call to answer?

Object files contain more than machine code

An object file is more than a sequence of machine instructions. Alongside code and data, it carries information needed to connect them later.

We will focus on three things:

Contents What we will examine
Code section What the machine code for call_answer looks like
Symbol table Which names this file defines, and which it needs from elsewhere
Relocation information Where references must be adjusted, and which targets they refer to

A section is a region that holds a particular kind of content, such as code or data. In my caller.o, the instructions were in the __text section of __TEXT. Here, text means code, rather than written text.

caller.o contains the instructions for call_answer, a symbol table with a definition of call_answer and an undefined reference to answer, and relocation information for the instruction at offset 8

Figure 1: The three kinds of information we examine in caller.o. This diagram does not show their actual positions or sizes in the file. Headers and other information are omitted.

Mach-O also has headers and load commands that let tools locate this information. We will read the parts relevant to the function call rather than decode the entire file format. On macOS, nm inspects the symbol table, and otool inspects sections and other contents. Building Mach-O Files — Apple

The symbol table records the missing implementation

First, let's use nm to see the externally visible symbols:

nm -g caller.o
Enter fullscreen mode Exit fullscreen mode
                 U _answer
0000000000000000 T _call_answer
Enter fullscreen mode Exit fullscreen mode

A symbol is information used to refer to a function or another entity by name. In this Mach-O output, the source names answer and call_answer appear with a leading underscore.

T denotes a definition in the code section; U denotes a symbol that is undefined in this file. In other words, call_answer is here, while the definition of answer must come from elsewhere. llvm-nm: symbol types

At this stage, U is not a failure. The object file records the intention to call something defined elsewhere.

The 0000000000000000 to the left of _call_answer does not mean the function will execute at address zero, either. In this .o, it is at the start of the code section. Its placement in the final executable is still to be determined.

So the name is still there. What about the machine code?

There is a call instruction. But where does it go?

Let's disassemble the code with otool. Disassembly turns machine code bytes back into a human-readable instruction listing.

otool -tvV caller.o
Enter fullscreen mode Exit fullscreen mode
caller.o:
(__TEXT,__text) section
_call_answer:
0000000000000000 stp x29, x30, [sp, #-0x10]!
0000000000000004 mov x29, sp
0000000000000008 bl  0x8
000000000000000c ldp x29, x30, [sp], #0x10
0000000000000010 ret
Enter fullscreen mode Exit fullscreen mode

The instruction to focus on is bl at 0000000000000008. ARM64 uses this instruction for function calls, and it also preserves information needed to return. Around it, we can see register saves and restores associated with the call. Arm: BL — Branch with Link

The 0000000000000008 on the left is the address where the bl instruction sits. The bl 0x8 on the right shows its branch destination as 0x8.

Wait. The instruction is at 0x8, and its destination is also 0x8? Read on its own, this looks like an instruction that branches to itself.

This bl instruction encodes the distance from the instruction to its destination. That distance is called a relative displacement. In this output before linking, it is still a placeholder value of zero. So 0x8 + 0 = 0x8, and the disassembler shows a branch to the same location. During linking, the linker writes the correct distance to answer into the instruction.

What we actually want to call is answer. To see that intention, we need to read the instruction together with the relocation information.

Relocation information says: adjust this reference to answer

Use otool -rv to display the relocations:

otool -rv caller.o
Enter fullscreen mode Exit fullscreen mode
caller.o:
Relocation information (__TEXT,__text) 1 entries
address  pcrel length extern type    scattered symbolnum/value
00000008 True  long   True   BR26    False     _answer
Enter fullscreen mode Exit fullscreen mode

Let's match this row to the instruction we just saw:

Field Meaning in this example
address: 00000008 Apply the relocation 8 bytes from the start of the __text section
pcrel: True Treat the reference as relative to the instruction's position
type: BR26 A relocation that adjusts the displacement of an ARM64 branch instruction
_answer The symbol used to determine the target

Here, address does not mean byte 8 from the start of the entire file. It is an offset within the code section. The file also contains headers and other information, so section offsets and file offsets must be read separately.

BR26 corresponds to Apple's ARM64_RELOC_BRANCH26, which handles the 26-bit displacement in B and BL instructions. Different instruction types need different adjustments, so a relocation needs both the referenced symbol and a relocation type. Apple: arm64 relocation types

The symbol table tells us which names exist. The relocation information tells us where to adjust references to those names.

That lets us produce the call instruction and the instructions for completing it later, even before we know where answer will be placed.

The word “relocation” might sound as though we are simply moving a function somewhere else. What we are examining here is information for adjusting a reference encoded in an instruction to match the final placement.

Put the body of answer in another .o

Now let's write the callee. Save this as provider.rs:

#![no_std]

#[unsafe(no_mangle)]
pub extern "C" fn answer() -> u32 {
    42
}
Enter fullscreen mode Exit fullscreen mode

Compile this into a separate object file as well:

rustc provider.rs --crate-type lib --edition 2024 \
  -C opt-level=0 -C panic=abort --emit=obj -o provider.o

nm -g provider.o
Enter fullscreen mode Exit fullscreen mode
0000000000000000 T _answer
Enter fullscreen mode Exit fullscreen mode

This time, _answer has a T. The definition that caller.o needs now exists in provider.o.

But this address is also zero. call_answer is at zero, and answer is at zero. Are they in the same place?

No: they are at the starts of their own sections in separate object files. During linking, they are placed in the output, and their references are connected using that placement.

The linker matches caller.o's undefined reference and relocation at offset 8 to the definition of answer in provider.o, then adjusts the call instruction using the output layout

Figure 2: A conceptual view of resolving the reference to answer. Other linker inputs, such as main and the standard library, are omitted.

Matching a name's reference to its definition is symbol resolution. Adjusting the necessary references according to the layout is the relocation processing we are looking at here.

Putting the files next to each other does not connect them. We need to establish which code calls which destination, and where that destination is.

Leave out the implementation, and linking fails

Let's add an entry point so we can run the program. Save this as main.rs:

unsafe extern "C" {
    fn call_answer() -> u32;
}

fn main() {
    println!("{}", unsafe { call_answer() });
}
Enter fullscreen mode Exit fullscreen mode

This main uses std. First, deliberately add only caller.o to the build:

rustc main.rs --edition 2024 -C opt-level=0 -C panic=abort \
  -C link-arg=caller.o -o demo-missing
Enter fullscreen mode Exit fullscreen mode

This time, linking failed. Here is the actual error, with only the long linker command and the note about its arguments omitted:

error: linking with `cc` failed: exit status: 1
  |
  … (linker command and argument note omitted)
  = note: Undefined symbols for architecture arm64:
            "_answer", referenced from:
                _call_answer in caller.o
          ld: symbol(s) not found for architecture arm64
          clang: error: linker command failed with exit code 1 (use -v to see invocation)

error: aborting due to 1 previous error
Enter fullscreen mode Exit fullscreen mode

Undefined symbols means that a symbol's definition could not be found. The missing name is _answer, and the error identifies _call_answer in caller.o as the code referencing it.

We have call_answer, but cannot find the answer it calls. A reference that could remain unresolved when producing the .o cannot be resolved when building this executable, so linking fails.

Now let's also provide provider.o:

rustc main.rs --edition 2024 -C opt-level=0 -C panic=abort \
  -C link-arg=caller.o -C link-arg=provider.o -o demo

./demo
Enter fullscreen mode Exit fullscreen mode
42
Enter fullscreen mode Exit fullscreen mode

It works.

-C link-arg adds an argument to the linker invocation. Here, we use it to add the two .o files. rustc also supplies inputs needed for main, the standard library, and so on. These two files are not the executable's only ingredients. Codegen Options: link-arg

The instruction really changed after linking

Finally, let's compare the machine code before and after linking. I used the LLVM disassembler available on my machine:

xcrun llvm-objdump --disassemble caller.o
xcrun llvm-objdump --disassemble demo
Enter fullscreen mode Exit fullscreen mode

Below are just the lines containing the call to answer. From left to right, they show the address, the machine instruction's value, and its human-readable form. Tabs and column spacing have been normalized.

Before linking, in caller.o:

8: 94000000  bl 0x8 <ltmp0+0x8>
Enter fullscreen mode Exit fullscreen mode

After linking, in demo:

100030a0c: 94000003  bl 0x100030a18 <_answer>
Enter fullscreen mode Exit fullscreen mode

94000000 became 94000003. This is more than a function name being added to the display: the value encoded in the call instruction itself changed. llvm-objdump

In my output, call_answer was placed at 0x100030a04, and answer at 0x100030a18. Given that layout, the call instruction at 0x100030a0c now branches to answer. We can see the relocation applied to the actual machine code.

The “3” encoded in the instruction does not mean 3 bytes. The 26-bit displacement in bl (imm26) is a signed value in units of 4 bytes. Here, it means 3 × 4 = 12 bytes ahead: 0x100030a0c + 12 = 0x100030a18, the address of answer. Arm instruction specification

The addresses shown here are virtual addresses recorded in the executable. They are distinct both from byte offsets within the file and from addresses necessarily used unchanged on every run. ASLR can shift the layout at load time, and references to shared libraries also involve runtime mechanisms. Here, we have focused on a call between two functions included in the same executable. Executing Mach-O Files — Apple, Apple Platform Security: ASLR

Why do we still need a linker when the machine code exists?

Our caller.o contained the machine code for call_answer. But answer remained an undefined symbol, accompanied by relocation information for adjusting the call instruction to reference it.

The compiler producing caller.o can generate an instruction saying “call answer here.” At that point, however, we have not supplied provider.o. We do not know where the body of answer will sit in the final executable, so we cannot determine the distance to it yet.

Instead, we leave behind the target's name and the location of the instruction that needs adjustment. Once the necessary inputs are available, the linker does the following:

  1. Match references to definitions. It connects the answer requested by caller.o to the definition in provider.o. When we left out the implementation, resolution failed and produced an error.
  2. Determine the layout of code and data. It places sections and other contents from separate .o files into the output executable. Positions local to each object file become part of the final layout.
  3. Adjust references to match that layout. It reads the relocation information and, in our example, writes the displacement to answer into the bl instruction. This is why 94000000 became 94000003.

Simply concatenating the files would leave the zero in the call instruction unchanged. Even if we included the function's body, the call would not be connected unless we adjusted the reference to reach it. Producing machine code and combining that code into a working program are separate jobs. Building Mach-O Files — Apple

Then why not compile everything together from the start? With a small example, that is tempting. But real programs use the standard library and other libraries alongside our own code. Compiling pieces separately and combining them later allows existing machine code to be reused. There is a qualification in Rust: generic functions, for example, may be monomorphized for concrete types and have machine code generated in the consuming crate. Not everything can be supplied solely as precompiled machine code. When producing the final executable, references between the generated pieces still need to be resolved. Monomorphization

The extern "C" in this experiment made that boundary easy to observe. Even without this syntax in everyday Rust, separately generated code still has to be connected. rustc can divide a single crate into multiple Codegen Units; with the LLVM backend, the object files produced from those units are linked. This normally happens as part of a build through rustc or Cargo, so we rarely need to operate the linker directly. Code generation

There was still work left between “the machine code exists” and “we have something executable.” The undefined symbol and relocation information we saw were how that work was handed to the next stage.

Top comments (0)