Have you ever launched 20–30 concurrent ffmpeg encoding jobs on a beefy multi-core server, only to watch the entire system grind to a halt?
Even on modern NVMe drives and 64-thread machines, running dozens of heavy video encoders simultaneously can trigger the dreaded I/O convoy effect. The standard Linux scheduler (EEVDF) tries to be fair by giving everyone CPU time at once. The result? All 25 processes decompress, encode, and flush buffers to disk at the exact same moment—saturating memory channels, dirtying page cache pages, and freezing interactive tasks.
With Linux 6.12, we finally have an elegant way to solve this at the kernel level without recompiling the kernel: sched_ext (SCX), the extensible BPF scheduler class.
In this article, we’ll build a custom eBPF scheduler that detects ffmpeg processes, isolates them into a dedicated queue, and staggers their execution to smooth out I/O bursts.
What is sched_ext?
Historically, the Linux CPU scheduler has been hardcoded into the kernel image (from O(1) to CFS and recently EEVDF). Modifying scheduling logic required compiling a custom kernel, testing out-of-tree patches, or maintaining custom kernel forks.
Merged into mainline in Linux 6.12, sched_ext lets developers write CPU scheduling policies in eBPF using struct_ops.
You can:
- Write scheduling logic in C or Rust.
- Load and unload schedulers at runtime via user space without restarting.
- Safely fall back to default kernel scheduling if your BPF program misbehaves.
Why Use a CPU Scheduler to Mitigate I/O Contention?
It’s important to understand the boundary:
-
sched_extgoverns CPU execution (which thread runs on which core and for how long). -
blk-mq/ cgroup v2 governs block storage queues (how bios are submitted to the disk controller).
So why fix I/O problems in a CPU scheduler?
Encoders operate in periodic cycles:
Read Input -> Heavy CPU Compute (Decode/Encode) -> Write/Flush Output
If 25 encoders run across all cores uncoordinated, their CPU phases align, meaning their I/O flush phases also align. By controlling CPU dispatching and pacing the number of concurrent encoders executing at any given millisecond, we naturally stagger the moments they hit the disk.
The Architecture: ffmpeg_pacer
Our scheduler enforces three rules:
-
Detection: Identify tasks belonging to
ffmpegvia task metadata. -
Isolation: Keep interactive / system tasks in the default global queue (
SCX_DSQ_GLOBAL), while routing allffmpegthreads into a dedicated Dispatch Queue (FFMPEG_DSQ). - Pacing / Staggering: Only allow up to N encoders (e.g., 6) to actively occupy CPU cores at any given time, staggering their time slices.
flowchart TD
Start([Incoming Tasks: enqueue]) --> Check{Is comm == 'ffmpeg'?}
Check -- Yes --> FFMPEG_Q["FFMPEG_DSQ (0x1001)<br/>(Paced & 20ms Slices)"]
Check -- No --> GLOBAL_Q["SCX_DSQ_GLOBAL<br/>(Standard Slices)"]
FFMPEG_Q --> Gatekeeper{Active Encoders <br/>< MAX_CONCURRENT?}
Gatekeeper -- Yes --> Dispatch([Dispatch to CPU Cores])
Gatekeeper -- No --> Wait["Hold in Queue<br/>(Staggered Execution)"]
GLOBAL_Q --> Dispatch
Wait -. When encoder finishes .- Gatekeeper
The BPF Code: ffmpeg_pacer.bpf.c
Here is a complete, working eBPF scheduler:
#include "vmlinux.h"
#include <bpf/bpf_helpers.h>
#include <bpf/bpf_tracing.h>
char _license[] SEC("license") = "GPL";
/* Custom DSQ ID for our encoder tasks */
#define FFMPEG_DSQ_ID 0x1001
#define MAX_CONCURRENT_FFMPEG 6
#define FFMPEG_TIME_SLICE_NS (20 * 1000 * 1000) /* 20ms slice */
/* Atomic counter tracking active encoders running on CPU */
volatile int active_encoders = 0;
/* Check if the task command starts with "ffmpeg" */
static inline bool is_ffmpeg(struct task_struct *p)
{
char comm[16];
bpf_get_current_comm(&comm, sizeof(comm));
return (comm[0] == 'f' && comm[1] == 'f' &&
comm[2] == 'm' && comm[3] == 'p' &&
comm[4] == 'e' && comm[5] == 'g');
}
/* 1. CPU Selection: Let the kernel find an idle core */
s32 BPF_STRUCT_OPS(ffmpeg_select_cpu, struct task_struct *p, s32 prev_cpu, u64 wake_flags)
{
return scx_bpf_select_cpu_dfl(p, prev_cpu, wake_flags, false);
}
/* 2. Enqueue: Segregate encoders from regular processes */
void BPF_STRUCT_OPS(ffmpeg_enqueue, struct task_struct *p, u64 enq_flags)
{
if (is_ffmpeg(p)) {
/* Route to our isolated DSQ with a fixed 20ms time slice */
scx_bpf_dispatch(p, FFMPEG_DSQ_ID, FFMPEG_TIME_SLICE_NS, enq_flags);
} else {
/* Standard tasks go straight to the global queue */
scx_bpf_dispatch(p, SCX_DSQ_GLOBAL, SCX_SLICE_DFL, enq_flags);
}
}
/* 3. Dispatch: Control when queued tasks are sent to CPU cores */
void BPF_STRUCT_OPS(ffmpeg_dispatch, s32 cpu, struct task_struct *prev)
{
/*
* Gatekeeper check: Only pull from FFMPEG_DSQ if the number of
* concurrently running encoders is below our threshold.
*/
if (__sync_val_compare_and_swap(&active_encoders, 0, 0) < MAX_CONCURRENT_FFMPEG) {
if (scx_bpf_consume(FFMPEG_DSQ_ID))
return;
}
/* Fallback: Keep the CPU busy with regular workloads */
scx_bpf_consume(SCX_DSQ_GLOBAL);
}
/* 4. Track when tasks start running */
void BPF_STRUCT_OPS(ffmpeg_running, struct task_struct *p)
{
if (is_ffmpeg(p)) {
__sync_fetch_and_add(&active_encoders, 1);
}
}
/* 5. Track when tasks yield, sleep, or exhaust slice */
void BPF_STRUCT_OPS(ffmpeg_stopping, struct task_struct *p, bool runnable)
{
if (is_ffmpeg(p)) {
__sync_fetch_and_sub(&active_encoders, 1);
}
}
/* 6. Initialization: Create our custom DSQ */
s32 BPF_STRUCT_OPS_SLEEPABLE(ffmpeg_init)
{
return scx_bpf_create_dsq(FFMPEG_DSQ_ID, -1);
}
void BPF_STRUCT_OPS(ffmpeg_exit, struct scx_exit_info *ei)
{
/* Cleanup hooks if needed */
}
/* Register with the kernel */
SEC(".struct_ops.link")
struct sched_ext_ops ffmpeg_pacer_ops = {
.select_cpu = (void *)ffmpeg_select_cpu,
.enqueue = (void *)ffmpeg_enqueue,
.dispatch = (void *)ffmpeg_dispatch,
.running = (void *)ffmpeg_running,
.stopping = (void *)ffmpeg_stopping,
.init = (void *)ffmpeg_init,
.exit = (void *)ffmpeg_exit,
.name = "ffmpeg_pacer",
};
Breaking Down the Code
-
scx_bpf_create_dsq: Inffmpeg_init, we instantiate a custom queue (0x1001). Custom DSQs allow you to build hierarchical, priority, or FIFO scheduling tiers without affecting the rest of the OS. -
is_ffmpeg(): We inspecttask_struct->comm. In a production setup, you would typically match by cgroup ID or Linux namespace rather than process name. -
Concurrency Limiter in
ffmpeg_dispatch: When a CPU core goes idle and asks for work,ffmpeg_dispatchinspectsactive_encoders. If 6 encoders are already working, it leaves the remaining 19 waiting inFFMPEG_DSQ_IDand executes other system tasks instead. -
Lifecycle Accounting: The
runningandstoppingcallbacks maintain an atomic counter so the scheduler always knows real-time concurrency.
Compiling & Loading the Scheduler
Prerequisites
- Kernel >= 6.12 compiled with
CONFIG_SCHED_CLASS_EXT=y. -
clang,llvm, andbpftool.
Check your kernel configuration:
zcat /proc/config.gz | grep CONFIG_SCHED_CLASS_EXT || grep CONFIG_SCHED_CLASS_EXT /boot/config-$(uname -r)
Compiling to BPF Object
clang -g -O2 -target bpf -D__TARGET_ARCH_x86 \
-I/usr/include/bpf \
-c ffmpeg_pacer.bpf.c -o ffmpeg_pacer.bpf.o
Activating the Scheduler
To load and switch the active Linux scheduler at runtime:
sudo bpftool struct_ops register ffmpeg_pacer.bpf.o
Verify that your scheduler is active (output should show enabled: ffmpeg_pacer):
cat /sys/kernel/sched_ext/state
To stop it and immediately revert back to the default kernel scheduler (EEVDF):
sudo bpftool struct_ops unregister name ffmpeg_pacer_ops
Taking It to Production
While writing custom C schedulers is great for specific needs, the open-source sched-ext/scx project already maintains several battle-tested schedulers written in Rust and C:
-
scx_lavd(Latency-Aware Virtual Deadline): An incredible scheduler for media pipelines and gaming. It dynamically calculates deadlines based on CPU memory footprints and thread interactivity, effectively preventing batch encoding workloads from causing UI stutter. -
scx_bpfland: A flexible, priority-aware general-purpose scheduler with vruntime tracking.
Combining with cgroup v2
For bulletproof enterprise stability, pair sched_ext with cgroup v2 I/O throttling:
Create a dedicated cgroup for your encoders:
sudo mkdir /sys/fs/cgroup/encoders
Limit aggregate write bandwidth to 100 MB/s to protect NVMe queues:
sudo echo "8:0 wbps=104857600" > /sys/fs/cgroup/encoders/io.max
Assign your FFmpeg worker processes into the cgroup:
echo $FFMPEG_PID > /sys/fs/cgroup/encoders/cgroup.procs
Conclusion
Linux kernel scheduling is no longer a monolithic black box reserved only for kernel maintainers. With sched_ext, you can design domain-specific scheduling algorithms that match your exact workload characteristics—whether that's pacing heavy video encoders, optimizing database thread locality, or improving game responsiveness.
Have you experimented with sched_ext yet? What custom scheduling policies would you build for your servers? Drop your thoughts in the comments below!
Top comments (0)