DEV Community

Cover image for What Happens to the CPU Cache During sched_yield
Sufyan bin Uzayr
Sufyan bin Uzayr

Posted on Originally published at code.zeba.academy

What Happens to the CPU Cache During sched_yield

Voluntary thread yielding via sched_yield() destroys hardware execution performance in production runtimes. Rather than executing an efficient wait, a yielded thread triggers a full Linux kernel context switch: general-purpose registers are spilled to struct pt_regs, execution stacks swap, and memory descriptors change. The resulting TLB flushes and L1/L2 cache pollution turn high-speed execution pipelines into stalled engines thrashing across the memory hierarchy. High-throughput architectures require deterministic hardware backoff or true kernel sleeping primitives, not uncontrolled scheduler concessions.

ielding CPU execution via sched_yield() introduces catastrophic tail latency by forcing kernel register spills, Translation Lookaside Buffer evictions, and L1/L2 cache pollution. Voluntary thread switches forfeit core affinity, turning deterministic sub-microsecond spin-waits into multi-microsecond memory stalls across NUMA domains.

Why Does sched_yield Degrade High-Throughput Engine Latency?

sched_yield latency spikes occur when a thread relinquishes its core execution context voluntarily, forcing the Completely Fair Scheduler (CFS) or EEVDF to trigger __schedule(). The yielding task is placed back onto the runqueue, incurring full hardware register evacuation and cache pollution overheads instead of progressing useful execution.

High-throughput concurrency runtimes—such as low-latency order matching engines or actor-framework thread pools—frequently substitute proper synchronization or bounded backoff with sched_yield(). Developers operate under the delusion that yielding CPU time is an innocuous, polite gesture to the OS scheduler. In reality, sched_yield() in Linux is an architectural disaster under load. Under the Completely Fair Scheduler (and the Earliest Eligible Virtual Deadline First scheduler), yielding does not guarantee an immediate or controlled return; rather, it relinquishes CPU hardware state and invites another thread to overwrite cached working sets.

THREAD CONTEXT SWITCH LIFECYCLE

+--------------------+        +---------------------+        +--------------------+
|  Running Thread A  |        |    Linux Kernel     |        |  Running Thread B  |
+--------------------+        +---------------------+        +--------------------+
          |                              |                              |
          |  sched_yield() [syscall]     |                              |
          |----------------------------->|                              |
          |                              | 1. Trap to Ring 0            |
          |                              | 2. Save GPRs to pt_regs      |
          |                              | 3. Invoke __schedule()       |
          |                              | 4. Select next task          |
          |                              | 5. switch_to() execution:    |
          |                              |    - Save RSP/RIP -> thread_struct
          |                              |    - Load Thread B thread_struct
          |                              | 6. switch_mm_irqs_off():     |
          |                              |    - Check mm_struct         |
          |                              |    - Reload CR3 (TLB Flush)  |
          |                              | 7. IRET / sysretq            |
          |                              |----------------------------->|
          |                              |                              | [Thread B Runs]
          |                              |                              | [L1d/L1i Polluted]
          |                              |                              | [TLB Overwritten]
          |                              |   Timer / Yield Interrupt    |
          |                              |<-----------------------------|
          |                              | 8. Reschedule Thread A       |
          |                              | 9. Restore CR3 + Registers   |
          |<-----------------------------|                              |
          | 10. Memory Stall Penalty:    |                              |
          |     - Cold L1/L2 Cache Misses|                              |
          |     - Page Table Walk Latency|
Enter fullscreen mode Exit fullscreen mode

The fundamental mechanics unravel at the silicon boundary. When an execution thread invokes system call 24 (sys_sched_yield), the processor transitions privilege domains from Ring 3 to Ring 0 via the syscall instruction. The user-space instruction pointer (RIP), stack pointer (RSP), and CPU flag register (RFLAGS) are written into architectural Model-Specific Registers (MSR_LSTAR, IA32_SYSENTER_EIP). At this juncture, the kernel must execute entry_SYSCALL_64, which pushes user General-Purpose Registers (GPRs: RAX, RBX, RCX, RDX, RSI, RDI, RBP, R8-R15) onto the process kernel stack, forming a struct pt_regs frame.

What Happens Inside the Kernel During a Task Context Switch?

Task switching occurs when the scheduler executes context_switch(), invoking switch_mm_irqs_off() to swap memory descriptors and switch_to() to swap register hardware states. The hardware stack pointer pivots, thread-local storage registers are loaded, and floating-point registers are swapped or marked unallocated via architectural control bits.

The kernel entry invokes the core scheduler via __schedule(). The scheduler identifies the next runnable entity via pick_next_task(). Once prev and next entities are verified, the execution path bifurcates into hardware memory mapping switch (switch_mm_irqs_off) and hardware architectural register switch (switch_to).

// Demonstrating the cost of sched_yield vs tight bounded pause in a spinloop
#define _GNU_SOURCE
#include <sched.h>
#include <x86intrin.h>
#include <stdint.h>
#include <stdio.h>
#include <stdatomic.h>

struct RingBuffer {
    atomic_uint_fast64_t head;
    atomic_uint_fast64_t tail;
    uint64_t buffer[1024];
};

// ANTI-PATTERN: Forcing CPU register eviction and pipeline stall via sched_yield
void broken_consumer_drain(struct RingBuffer *rb) {
    while (atomic_load_explicit(&rb->head, memory_order_acquire) == 
           atomic_load_explicit(&rb->tail, memory_order_relaxed)) {
        // TRAP: Incurs ring-transition, pt_regs frame store, 
        // __schedule() traversal, and CPU cache destruction.
        sched_yield(); 
    }
}

// CORRECT: Hardware-native instruction pipelines without context switch
void production_consumer_drain(struct RingBuffer *rb) {
    uint32_t backoff_counter = 0;
    while (atomic_load_explicit(&rb->head, memory_order_acquire) == 
           atomic_load_explicit(&rb->tail, memory_order_relaxed)) {
        if (backoff_counter < 32) {
            // Emits hardware PAUSE: de-pipelines speculative execution,
            // saves pipeline clear penalties without thread context swap.
            _mm_pause(); 
            backoff_counter++;
        } else {
            // Controlled sleep via futex/epoll instead of raw sched_yield thrashing
            struct timespec req = { .tv_sec = 0, .tv_nsec = 500 };
            nanosleep(&req, NULL);
            backoff_counter = 0;
        }
    }
}
Enter fullscreen mode Exit fullscreen mode

The register swap inside switch_to() is executed through inline assembly in arch/x86/entry/entry_64.S or __switch_to_asm. The active hardware stack pointer RSP is stored directly into prev->thread.sp, and the target thread stack pointer is dereferenced from next->thread.sp into the physical %rsp register. Execution then returns using a modified return frame: the hardware pulls the saved Instruction Pointer RIP off the new stack, immediately pivoting instruction execution to the alternate task's instruction stream.

Simultaneously, FS_BASE and GS_BASE registers which dictate Thread Local Storage (TLS) must be updated using wrfsbase or writing to MSR 0xC0000100 (MSR_FS_BASE), an operation with non-negligible cycle costs on microarchitectures prior to direct base manipulation instructions.

How Does a Context Switch Invalidate TLB and CPU Caches?

TLB and cache invalidation occurs when the memory descriptor changes between different address spaces, forcing a reload of the CR3 control register. The update flushes non-global page translations from the processor TLB and exposes L1 and L2 caches to eviction by foreign instructions and data access streams.

If the switch occurs between threads of disparate processes, the kernel must execute switch_mm_irqs_off(). This routine checks whether prev->mm == next->mm. When next->mm differs, the kernel loads the physical address of the next process's Page Global Directory (PGD) into control register CR3:

Writing to CR3 triggers an unconditional hardware flush of all Translation Lookaside Buffer (TLB) entries that do not carry the global flag (PGE bit in CR4). Even with Process-Context Identifiers (PCID) enabled via CR4.PCIDE which appends a 12-bit address space ID to TLB tags to prevent indiscriminate flushing the cache footprint suffers severe degradations.

Once a foreign task assumes execution on the physical core, it writes to L1 data (L1d) and instruction (L1i) caches. A modern Xeon or EPYC core maintains 32KB to 48KB of L1d cache, arranged in 64-byte cache lines. A foreign thread reading a 64KB array will completely evict the original thread's hot working set across all 8 associative cache ways within tens of microseconds.

When the original thread is eventually rescheduled, its first memory references generate immediate L1 misses (1.0–1.5 ns access time), L2 misses (3–5 ns access time), and L3/LLC misses (12–15 ns access time). Worse, if the scheduler migrated the task across core boundaries or NUMA nodes to balance queue depths, the thread must fetch lines over the UPI/Infinity Fabric interconnect from remote DDR memory, paying 80 to 120 nanoseconds per cache line fill.

The trade-off is categorical: calling sched_yield() under the belief that it resolves resource contention guarantees the destruction of CPU pipeline state and instruction locality, degrading tail-end determinism in exchange for zero scheduling control.

Technical Troubleshooting FAQ

How Do I Fix perf record Reporting High Overhead in __schedule and native_write_cr3?

This overhead occurs when excessive context switching causes the kernel to spend significant cycle budgets inside scheduler queues and page-directory register updates. Fix this by identifying the offending application using perf top -e sched:sched_switch, disabling application-level sched_yield() loops, pinning critical worker threads to isolated cores via taskset -c or pthread_setaffinity_np(), and enabling the isolcpus and nohz_full Linux boot parameters.

How Do I Fix sched_yield Not Yielding to Lower Priority Tasks Under CFS/EEVDF?

This issue occurs when an application calls sched_yield() expecting execution to drop down to normal-priority batch tasks while running under real-time policies (SCHED_FIFO or SCHED_RR). Under POSIX standards, sched_yield() only places the thread at the tail of its current priority level; if no other thread of identical priority exists, the calling thread resumes instantly, causing 100% CPU lockup. Fix this by replacing sched_yield() with true synchronizations such as pthread_cond_wait() or futex(FUTEX_WAIT) or by downgrading the thread scheduling class to SCHED_OTHER.

References

Top comments (1)

Collapse
 
respect17 profile image
Kudzai Murimi •

The _mm_pause() then nanosleep backoff pattern in your code sample shows the fix well instead of just telling people not to yield. Worth adding for readers that this cache thrashing cost is much smaller under SCHED_FIFO/SCHED_RR, where yield just drops to the tail of the same priority queue. The CFS/EEVDF case you're describing is really the worst case for using this as a spinlock replacement.