DEV Community

Cover image for Pacing Batch Workloads with Linux sched_ext: An Engineering PoC
Pawel Suchanecki
Pawel Suchanecki

Posted on

Pacing Batch Workloads with Linux sched_ext: An Engineering PoC

Pacing Batch Workloads with Linux sched_ext: An Engineering PoC

Running many heavy batch jobs in parallel can look deceptively simple:

for job in jobs/*; do
    ./worker "$job" &
done
wait
Enter fullscreen mode Exit fullscreen mode

On a large machine, this may work fine. Until it does not.

When enough CPU-heavy workers run at the same time, their phases can start aligning: compute, allocate, write, flush, repeat. Even if the workload is not purely I/O-bound, synchronized CPU progress can produce synchronized I/O pressure later.

This article describes an engineering PoC for controlling that pattern with Linux sched_ext.

The goal is not to claim a benchmark win upfront. The goal is narrower and more useful:

Can we use sched_ext to pace a selected class of batch tasks, reduce burstiness, and measure the trade-off against throughput?


What sched_ext Gives Us

Linux sched_ext, introduced in Linux 6.12, allows scheduling policies to be implemented with BPF struct_ops.

Instead of patching or rebuilding the kernel, we can load a BPF scheduler at runtime, route tasks into dispatch queues, and define custom policy decisions.

The official kernel documentation describes the core model: tasks are enqueued, dispatched into Dispatch Queues (DSQs), moved to local CPU queues, run, stop, and re-enter the lifecycle as needed.

Reference: Linux sched_ext documentation (v6.12)


Important Boundary: This Is Not an I/O Scheduler

sched_ext controls CPU scheduling.

It does not directly control block I/O queues. That remains the job of the block layer, writeback, page cache behavior, and cgroup v2 I/O controls.

So why use a CPU scheduler for an I/O-looking problem?

Because many batch workloads have coupled phases:

read -> compute -> allocate -> write -> flush
Enter fullscreen mode Exit fullscreen mode

If too many workers progress through those phases together, the system may see bursty writeback, higher tail latency, elevated Pressure Stall Information (PSI), or degraded interactive responsiveness.

CPU-side pacing cannot replace I/O control, but it can influence when workers reach their I/O-heavy phases.


Classification: Do Not Use bpf_get_current_comm() Here

A common trap is to classify the task inside enqueue(struct task_struct *p, ...) using:

/* INCORRECT PATTERN */
char comm[16];
bpf_get_current_comm(comm, sizeof(comm));
Enter fullscreen mode Exit fullscreen mode

That is not correct for this use case.

bpf_get_current_comm() reads the comm of the currently executing task on the CPU, not necessarily the arbitrary struct task_struct *p passed to the scheduler callback (which is often being woken up by another thread or interrupt handler).

Reference: Linux bpf.h helper documentation

For a solid PoC, the safer classification mechanism is cgroup-based:

sudo mkdir -p /sys/fs/cgroup/batch-pacer
echo "$PID" | sudo tee /sys/fs/cgroup/batch-pacer/cgroup.procs
Enter fullscreen mode Exit fullscreen mode

The scheduler identifies tasks belonging to that cgroup and routes only those tasks through the pacing policy.


Scheduler Policy

The PoC policy is intentionally simple:

  1. Tasks outside the target cgroup go to the normal/global DSQ.
  2. Tasks inside the target cgroup go to a dedicated DSQ (PACED_DSQ).
  3. Only up to $N$ target tasks may actively run at once.
  4. Other target tasks wait in the dedicated DSQ.
  5. The rest of the system continues to make unimpeded progress.

sched_ext CPU-side pacer architecture diagram


The Accounting Problem: Avoiding Dispatch Races

A loose counter updated across decoupled callbacks is not enough:

/* RACING PATTERN */
if (active_tasks < MAX_ACTIVE) {
    scx_bpf_consume(PACED_DSQ);
}
Enter fullscreen mode Exit fullscreen mode

Multiple CPUs can observe active_tasks == 5 at the same time and all decide to consume from the paced queue. The limit may be blown before running() has a chance to increment the counter.

The fix is to reserve a slot atomically before allowing dispatch:

static __always_inline bool reserve_slot(void)
{
    int old;

    #pragma unroll
    for (int i = 0; i < 8; i++) {
        old = __sync_val_compare_and_swap(&active_tasks, 0, 0);
        if (old >= MAX_ACTIVE)
            return false;
        if (__sync_val_compare_and_swap(&active_tasks, old, old + 1) == old)
            return true;
    }
    return false;
}
Enter fullscreen mode Exit fullscreen mode

A production-grade implementation should also track per-task state (for example, with BPF_MAP_TYPE_TASK_STORAGE), ensuring a slot is released only by a task that actually reserved one. Without per-task state, lifecycle callbacks (running(), stopping()) can overcount or undercount during preemption or requeue events.


Kernel API Detail: scx_bpf_select_cpu_dfl()

On Linux 6.12, scx_bpf_select_cpu_dfl() expects a bool *is_idle out parameter:

s32 BPF_STRUCT_OPS(simple_select_cpu, struct task_struct *p, s32 prev_cpu, u64 wake_flags)
{
    bool is_idle = false;
    s32 cpu;

    cpu = scx_bpf_select_cpu_dfl(p, prev_cpu, wake_flags, &is_idle);
    if (is_idle)
        scx_bpf_dispatch(p, SCX_DSQ_LOCAL, SCX_SLICE_DFL, 0);

    return cpu;
}
Enter fullscreen mode Exit fullscreen mode

Reference: Linux scx_simple.bpf.c (v6.12)

This matters because sched_ext APIs are still evolving across kernel versions. Any PoC should explicitly pin the kernel version and document the tested ABI.


What We Need to Measure

This PoC is incomplete without empirical measurements. At minimum, compare:

Variant Description
Baseline Default Linux scheduler (EEVDF)
Baseline + cgroup I/O Default scheduler with cgroup v2 I/O throttling (io.max / io.weight)
sched_ext pacer CPU-side pacing only
sched_ext pacer + cgroup I/O Combined CPU pacing and I/O controls

Key Metrics to Collect:

  • Total batch completion time & aggregate throughput
  • p95 / p99 latency of the batch workload
  • p95 / p99 latency of an interactive control workload
  • CPU utilization & %iowait
  • PSI (Pressure Stall Information) for CPU and I/O
  • Context switches & run-queue length
  • Dirty memory writeback behavior

Success Criteria

The PoC succeeds only if measurements show a justifiable engineering trade-off:

  • Lower p95/p99 tail latency,
  • Reduced PSI or burstiness,
  • Improved responsiveness for non-batch tasks,
  • No starvation,
  • Acceptable throughput loss.

A scheduler that cuts throughput by 40% just to make disk graphs look smoother is probably counterproductive. But a scheduler that reduces tail latency spikes and writeback stalls while maintaining 90–95% of baseline throughput is worth pursuing.


Conclusion

sched_ext is a powerful tool for workload-specific CPU scheduling experiments, but it should be treated as an engineering mechanism, not a magic performance wand.

For bursty batch workloads, the real question is not:

“Can we write a custom scheduler?”

The real question is:

“Can CPU-side pacing reduce harmful synchronization between workers, and can we prove it with measurements?”

Top comments (0)