Every backend engineer learns early on that threads are not free. We are told to keep thread pools bounded, avoid context-switch storms, and prefer event loops or green threads for high concurrency.
Yet if you ask a developer what a context switch actually costs, the answer is usually vague: "It takes a microsecond or two to save some registers."
That answer hides a massive performance trap. Saving registers takes less than 50 nanoseconds. The real destruction happens in the CPU caches, the Translation Lookaside Buffer (TLB), and the branch predictor.
Here is the exact hardware and kernel mechanism that executes every time Linux stops one thread and runs another.
1. The Concurrency Illusion vs Physical Cores
A CPU core is a single instruction pipeline. It contains one Program Counter (rip), one stack pointer (rsp), and one set of architectural registers.
When your operating system runs 500 threads across 8 physical cores, those threads are not executing simultaneously. The Linux kernel rapidly multiplexes CPU execution time slices across runnable tasks.
A context switch is the mechanical procedure of:
- Freezing the execution state of the currently executing thread (
prev). - Storing its CPU register state and stack pointers.
- Optionally swapping the virtual address space (page tables).
- Restoring the saved registers and stack pointer of the next thread (
next). - Resuming execution at
next's saved instruction pointer.
This switch happens in two distinct flavors.
2. Voluntary vs Involuntary Switches
Linux categorizes every context switch into one of two counters inside task_struct:
// include/linux/sched.h
struct task_struct {
unsigned long nvcsw; /* voluntary context switches */
unsigned long nivcsw; /* involuntary context switches */
...
};
You can view these counters for any running process in /proc:
$ cat /proc/$$/status | grep ctxt
voluntary_ctxt_switches: 1420
nonvoluntary_ctxt_switches: 84
Voluntary Switches (nvcsw)
The thread actively yields the CPU because it cannot make progress. Common triggers include:
- Synchronous blocking I/O: waiting for a disk read or a socket
recv()buffer. - Lock contention: failing to acquire a
pthread_mutexand putting the thread to sleep viasys_futex(). - Explicit sleeping: calling
usleep()orsched_yield().
The thread calls into the kernel through a system call and willingly enters the TASK_INTERRUPTIBLE or TASK_UNRUNNABLE state.
Involuntary Switches (nivcsw)
The thread wanted to keep running, but the kernel forced it off the core. Triggers include:
- Time slice exhaustion: the scheduler (CFS or EEVDF) determined the task consumed its allocated runtime slice during the local APIC timer tick.
- Preemption: a higher-priority task (or a real-time thread) woke up and became runnable.
3. The Anatomy of a Switch: Step by Step
Let us trace what happens under the hood on an x86-64 system when thread prev is replaced by thread next.
[ User Space: Thread A ]
| (syscall / timer interrupt)
v
[ Ring 0 Kernel Mode ]
|
+---> 1. Save user registers on kernel stack
+---> 2. Enter scheduler (__schedule)
+---> 3. Pick next task_struct from runqueue
|
+-----+-----------------------------------+
| |
v v
[ Process Switch: prev->mm != next->mm ] [ Thread Switch: same mm ]
| |
+-> switch_mm_irqs_off() +-> Keep current CR3
+-> Reload CR3 (Page Tables) +-> Keep TLB warm
+-> PCID tag handling |
| |
+--------------------+--------------------+
|
v
[ 4. __switch_to_asm ]
+-> Push callee-saved registers (rbp, rbx, r12-r15)
+-> Swap kernel stack pointer (rsp)
+-> Update TSS.rsp0 (Ring 0 interrupt stack)
+-> Pop callee-saved registers of next
|
v
[ 5. Floating Point & Vector (FPU/AVX) ]
+-> Lazy restore / XSAVES state
|
v
[ 6. sysretq / iretq ]
|
v
[ User Space: Thread B ]
Step 1: Trapping to Kernel Space
The switch begins with a transition from Ring 3 (user space) to Ring 0 (kernel space).
When a system call or hardware timer interrupt fires, the CPU hardware automatically pushes five registers onto the thread's dedicated kernel stack:
-
%rsp(User stack pointer) -
%ss(User stack segment) -
%rflags(CPU status flags) -
%cs(User code segment) -
%rip(User instruction return pointer)
The CPU then switches the stack pointer %rsp to the kernel stack defined in the Task State Segment (TSS).
Step 2: Scheduler Core (__schedule)
Inside the kernel, __schedule() disables local preemption and inspects the CPU runqueue (rq). It invokes the scheduling class (such as EEVDF, the Early Eligible Virtual Deadline First scheduler in modern kernels) to choose the next task_struct.
Once next is chosen, the kernel invokes context_switch():
// kernel/sched/core.c (simplified)
static __always_inline struct rq *
context_switch(struct rq *rq, struct task_struct *prev,
struct task_struct *next, struct rq_flags *rf)
{
/* 1. Memory context switch */
if (!next->mm) {
/* Kernel thread: borrow previous active_mm (lazy TLB) */
next->active_mm = prev->active_mm;
} else {
/* User process: switch page tables if mm differs */
switch_mm_irqs_off(prev->active_mm, next->mm, next);
}
/* 2. Architecture register and stack switch */
switch_to(prev, next, prev);
return finish_task_switch(prev);
}
4. The Memory Context: Page Tables & TLB
This is where thread-to-thread switching and process-to-process switching diverge drastically.
Every user process has an mm_struct pointing to its Level 4 Page Table root (pgd). On x86-64, the CPU reads physical memory addresses through the hardware CR3 register.
Scenario A: Switching Threads Within the Same Process
If prev->mm == next->mm, both threads share the exact same virtual address space.
The kernel skips switch_mm_irqs_off() entirely. CR3 is never touched. Every virtual-to-physical address translation in the Translation Lookaside Buffer (TLB) remains valid and warm.
Scenario B: Switching to a Different Process
If prev->mm != next->mm, the incoming process has completely different memory mappings. The kernel must update the CR3 control register:
// arch/x86/mm/tlb.c
write_cr3(build_cr3(next->pgd, new_asid, prev_asid));
Writing to CR3 is an expensive serializing instruction. Historically, writing to CR3 instantly invalidated every non-global entry in the CPU's TLB. The next time the CPU executed user instructions, every single memory access incurred an MMU 4-level page table walk (PML4 -> PDPT -> PD -> PT -> RAM), stalling the execution pipeline for hundreds of cycles.
How Modern CPUs Tame This: PCID
Modern x86-64 processors feature PCID (Process Context Identifiers, enabled via CR4.PCIDE).
PCID adds a 12-bit tag (0 to 4095) to every TLB cache line representing the process ID. When switching page tables, Linux sets the NOFLUSH bit (bit 63) of CR3. The hardware preserves the cached translations for other processes. When the CPU switches back to prev later, its TLB entries are still warm.
However, if your system enables Kernel Page Table Isolation (KPTI, the mitigation for Meltdown), the CPU must flip between two distinct page tables (user space and kernel space) on every single system call and interrupt, compounding context switch overhead.
5. Register and Stack Switching in Assembly
Once the address space is ready, Linux calls switch_to(), which jumps to the architecture-specific assembly routine __switch_to_asm.
Here is the exact assembly code from the Linux kernel tree:
/* arch/x86/entry/entry_64.S */
SYM_FUNC_START(__switch_to_asm)
/* 1. Save callee-saved registers of prev on its kernel stack */
pushq %rbp
pushq %rbx
pushq %r12
pushq %r13
pushq %r14
pushq %r15
/* 2. Save current RSP to prev->thread.sp */
movq %rsp, TASK_threadsp(%rdi)
/* 3. Load next->thread.sp into RSP */
movq TASK_threadsp(%rsi), %rsp
#ifdef CONFIG_STACKPROTECTOR
movq TASK_stack_canary(%rsi), %rbx
movq %rbx, PER_CPU_VAR(__stack_chk_guard)
#endif
/* 4. Restore callee-saved registers from next's kernel stack */
popq %r15
popq %r14
popq %r13
popq %r12
popq %rbx
popq %rbp
/* 5. Jump to __switch_to C function to finish hardware state */
jmp __switch_to
SYM_FUNC_END(__switch_to_asm)
Look closely at lines 10 and 13:
-
%rdiholds the pointer toprev(task_struct). -
%rsiholds the pointer tonext(task_struct).
By writing %rsp into prev->thread.sp and reading next->thread.sp into %rsp, the kernel swaps the physical execution stack in two instructions.
When the subsequent popq instructions execute, the CPU is no longer popping prev's values; it is restoring the register state that next pushed onto its own stack days, hours, or milliseconds ago.
Updating the TSS (rsp0)
The kernel then updates the CPU's Task State Segment:
// arch/x86/kernel/process_64.c
this_cpu_write(cpu_tss_rw.x86_tss.sp0, next_top_of_stack);
This ensures that whenever the next hardware interrupt or syscall occurs, the CPU switches to next's kernel stack rather than corrupting prev's stack.
6. Vector Registers: The FPU and AVX-512 Penalty
General purpose registers (RAX, RBX, RCX, etc.) take only 128 bytes of storage. But modern CPUs feature SIMD vector extensions: AVX2 (256-bit) and AVX-512 (512-bit).
If a thread uses 32 AVX-512 ZMM registers and opmask registers, saving its state requires over 2,600 bytes of memory.
Doing a full vector register backup on every context switch would grind the CPU to a halt. Linux uses hardware-assisted state management (XSAVES / XRSTORS) and tracked thread flags (TIF_NEED_FPU_LOAD):
- If a thread never touched floating-point or vector math, the kernel skips saving FPU registers completely.
- When saving is necessary,
XSAVESonly serializes the specific register sub-components modified since the last switch.
7. The Direct vs Indirect Cost: Where Time Disappears
When benchmarked in isolation, the direct code path of a context switch (the instructions in __schedule and __switch_to_asm) is surprisingly fast:
- Direct Cost: ~1,000 to 2,500 CPU cycles (~0.3 to 1.2 microseconds on modern server hardware).
If direct execution is so cheap, why do high-concurrency systems collapse under thread contention?
The culprit is the Indirect Cost.
+-------------------------------------------------------------+
| Direct Cost (~1.2 µs): |
| - Scheduler decision logic (EEVDF runqueue lock) |
| - Pushing/popping 6 callee registers |
| - Swapping %rsp stack pointers |
| - TSS.sp0 reload & CR3 PCID tag write |
+-------------------------------------------------------------+
|
v
+-------------------------------------------------------------+
| Indirect Cost (~10 to 30+ µs): |
| - L1 Instruction Cache Misses (incoming code must fetch) |
| - L1/L2 Data Cache Eviction (working memory evicted) |
| - TLB Misses (multi-level page table walks) |
| - Branch Target Buffer (BTB) mispredictions |
| - CPU Core Migration (invalidating L1/L2 on another core) |
+-------------------------------------------------------------+
1. L1 and L2 Cache Pollution
An L1 cache lookup takes ~4 to 5 cycles (~1 nanosecond). An access to main memory (DRAM) takes ~200 cycles (~60 to 80 nanoseconds).
When Thread A runs for 1 millisecond, it fills the L1 and L2 caches with its active working set (hash tables, request buffers, stack frames).
When Thread B takes over, its instructions and memory references are not in L1/L2. Thread B stalls the execution pipeline waiting on DRAM while evicting Thread A's data. When Thread A gets scheduled back, all its data is gone. The CPU spends the majority of its time stalled on memory bus fetches rather than computing instructions.
2. Branch Target Buffer (BTB) Invalidation
Modern superscalar processors predict conditional branches and indirect function calls with high accuracy. When a new thread executes entirely different code paths, the hardware branch predictor suffers a wave of branch mispredictions, repeatedly flushing the out-of-order execution pipeline (a ~15-20 cycle penalty per misprediction).
8. Profiling Context Switches with perf
You can observe context switches and their microarchitectural fallout directly using Linux perf:
$ perf stat -e \
context-switches,cpu-migrations,page-faults,\
L1-dcache-load-misses,dTLB-load-misses,instructions,cycles \
./my_server_binary
Sample output from an overloaded thread-per-request architecture:
Performance counter stats for './my_server_binary':
1,248,391 context-switches # 0.124 M/sec
14,890 cpu-migrations # 1.482 K/sec
4,210 page-faults # 0.419 K/sec
489,120,441 L1-dcache-load-misses # 24.12% of all L1-dcache hits
18,940,110 dTLB-load-misses # 1.885 M/sec
3,410,294,012 instructions # 0.68 insn per cycle
5,015,133,892 cycles
10.051492104 seconds time elapsed
Notice the instructions per cycle (IPC) metric: 0.68.
A modern CPU core can easily achieve an IPC of 2.0 to 3.5 on compute workloads. An IPC below 1.0 is a screaming indicator that the CPU cores are spending over half their time stalled on memory cache misses and TLB walks caused by violent context switching.
9. Architectural Takeaways for Developers
Understanding the kernel mechanics leads directly to better system design choices:
1. Thread Pool Sizing: Match Physical Cores for CPU Work
For CPU-bound workloads (image processing, crypto, serialization), creating more threads than physical CPU cores adds zero throughput. It only introduces scheduler preemption, cache thrashing, and lock contention.
- Golden rule for CPU tasks:
Number of worker threads = Number of physical CPU cores.
2. Async I/O (epoll / io_uring) Over Thread-Per-Connection
A server handling 50,000 active connections with 50,000 OS threads will spend more time in __schedule and cache misses than processing business logic. Event-driven architectures (epoll, kqueue, Linux io_uring) keep a tiny pool of pinned worker threads running continuously on hot CPU caches without yielding.
3. Goroutines and Fibers: User-Space Multiplexing
Languages like Go (goroutines), Rust (Tokio green tasks), and Erlang (BEAM processes) bypass kernel context switches for concurrency:
- A goroutine switch only saves 3 general registers in user space.
- No kernel trap, no
CR3reload, no TSS update, no kernel stack allocation. - 100,000 goroutines multiplexed over 8 OS worker threads avoids kernel scheduling overhead entirely.
4. CPU Pinning and Thread Affinity
If you have high-throughput worker threads, bind them to dedicated physical cores using pthread_setaffinity_np or taskset:
$ taskset -c 0-3 ./high_frequency_worker
Pinning prevents the kernel scheduler from migrating threads across different CPU cores, keeping L1/L2 caches and NUMA memory nodes permanently hot.
Summary Mental Model
| Component | In-Process Thread Switch | Cross-Process Switch |
|---|---|---|
| Virtual Address Space | Kept intact (identical mm_struct) |
Swapped (CR3 reload) |
| TLB Cache | Remains 100% warm | Preserved via PCID or flushed |
| Kernel Stack | Swapped via TASK_threadsp(%rsp)
|
Swapped via TASK_threadsp(%rsp)
|
| Callee Registers | 6 registers saved (rbp, rbx, r12-r15) |
6 registers saved |
| FPU / Vector State | Saved lazily via XSAVES
|
Saved lazily via XSAVES
|
| Direct Hardware Cost | ~0.3 - 0.8 µs (~800 cycles) | ~1.0 - 2.0 µs (~2,500 cycles) |
| Indirect Cache Penalty | Moderate (L1/L2 data eviction) | Severe (L1/L2 + TLB eviction) |
The next time you design a backend service, remember: the cost of a thread switch is not the few instructions it takes to save %rsp. The cost is every cache line the CPU has to fetch from RAM all over again.
Top comments (0)