DEV Community

mattleeee
mattleeee

Posted on Originally published at hkcode.dpdns.org

Taming BLAS Threads in Scheduled Python Jobs

Scheduled Python jobs on a shared box have a nasty habit of fighting each other. You write a script that runs in four seconds on your laptop, deploy it to a Windows Server with eight cores, and now it takes forty — but only when the nightly batch is also running. The culprit is usually not your code. It is the BLAS thread pool underneath numpy.

What Actually Spawns Threads

When you import numpy, the wheel you installed is linked against a BLAS implementation: OpenBLAS for the standard numpy wheels on PyPI, MKL for numpy built against Intel's math library (common with conda, or pip install mkl setups). Both ship multithreaded kernels for dgemm, sgemm, and friends. By default they read the machine's core count and spin up a pool of that size.

So import numpy on an 8-core server means eight OpenBLAS worker threads, plus the OpenMP runtime underneath, plus whatever your own concurrent.futures pool does. Matrix multiply is fast. Everything else pays for it.

The problem shows up when jobs overlap. Suppose Task Scheduler fires three jobs at 02:00:

  • Job A does a rolling regression over 200k rows.
  • Job B builds a covariance matrix.
  • Job C just reads a CSV with pandas and writes a parquet file.

Each one imports numpy, each one gets eight BLAS threads. That is 24 threads on 8 cores. The OS context-switches between them, cache lines bounce between cores, and the memory allocator thrashes. Job C, which does almost no math, is now slow because it is competing for scheduler time with two BLAS pools it never asked for.

The fix is boring and effective: pin every job to one BLAS thread and let the OS schedule your jobs in parallel instead of your math kernels.

The Three Variables

OPENBLAS_NUM_THREADS

Read by OpenBLAS at initialization. Controls the size of its internal thread pool. This is the one that matters for stock PyPI numpy wheels.

MKL_NUM_THREADS

The equivalent knob for Intel MKL. If your numpy is MKL-linked, this is the one that counts. MKL also honors MKL_DYNAMIC and MKL_THREADING_LAYER, but for our purposes MKL_NUM_THREADS=1 is enough.

OMP_NUM_THREADS

Read by the OpenMP runtime. Both OpenBLAS and MKL can use OpenMP as their threading backend, and so can other libraries (scikit-learn, some pandas paths, PyTorch). Setting this catches the cases where the BLAS-specific variable is ignored because the library was built with OpenMP threading.

Set all three. The cost of setting one that is unused is zero; the cost of missing the one that matters is a mystery slowdown you will spend an afternoon debugging.

There is also NUMEXPR_NUM_THREADS for numexpr, which pandas uses for some eval paths, and VECLIB_MAXIMUM_THREADS on macOS. On Windows servers the first three cover the realistic cases.

Where to Set Them

Two places, and they are not equivalent.

In the launcher

The safest place is the process environment before Python starts. In a .bat file:

@echo off
set OPENBLAS_NUM_THREADS=1
set MKL_NUM_THREADS=1
set OMP_NUM_THREADS=1
set NUMEXPR_NUM_THREADS=1
"C:\Python311\python.exe" "D:\jobs\build_features.py" >> "D:\logs\build_features.log" 2>&1
Enter fullscreen mode Exit fullscreen mode

This is the version I use in Task Scheduler actions. The variables are set before python.exe is even loaded, so every library that reads them at import time gets the right value. No ordering concerns, no surprises.

If you drive jobs from PowerShell instead:

$env:OPENBLAS_NUM_THREADS = "1"
$env:MKL_NUM_THREADS      = "1"
$env:OMP_NUM_THREADS      = "1"
& "C:\Python311\python.exe" "D:\jobs\build_features.py"
Enter fullscreen mode Exit fullscreen mode

Same effect. The environment block is inherited by the child process.

In Python, before importing numpy

Sometimes you cannot control the launcher — a CI runner, a wrapper someone else owns, a python -m invocation from a service. Then set the variables at the very top of your entry point:

# entry.py — must run before any numerical import
import os

os.environ.setdefault("OPENBLAS_NUM_THREADS", "1")
os.environ.setdefault("MKL_NUM_THREADS", "1")
os.environ.setdefault("OMP_NUM_THREADS", "1")
os.environ.setdefault("NUMEXPR_NUM_THREADS", "1")

# Now it is safe to import the numerical stack.
import numpy as np
import pandas as pd
Enter fullscreen mode Exit fullscreen mode

Two things to notice.

First, setdefault rather than =. If the launcher already set a value, respect it. This lets you override per-job without editing the script.

Second, the ordering is load-bearing. OpenBLAS reads OPENBLAS_NUM_THREADS inside its constructor, which runs when the shared library is dlopened — i.e. during import numpy. If you set the variable after that import, you are too late. The pool is already sized.

This bites people who put the os.environ lines in a config.py that gets imported after numpy somewhere up the chain. If you cannot guarantee import order, use the launcher.

A guard for libraries that import numpy for you

If your entry point imports a helper module that imports pandas at module scope, you have already lost. A defensive pattern:

# sitecustomize.py or a bootstrap module imported first
import os

for var in (
    "OPENBLAS_NUM_THREADS",
    "MKL_NUM_THREADS",
    "OMP_NUM_THREADS",
    "NUMEXPR_NUM_THREADS",
):
    os.environ.setdefault(var, "1")
Enter fullscreen mode Exit fullscreen mode

Drop this in a sitecustomize.py on PYTHONPATH and it runs before anything else in the interpreter, including numpy. It is heavy-handed for a laptop, but on a scheduled server it removes an entire class of ordering bugs.

When Single-Threading Wins, and When It Loses

Single-threading is not universally faster. It is faster when the machine is oversubscribed or the workload is memory-bound. It is slower when a single job has the box to itself and does real dense linear algebra.

Wins

  • Overlapping scheduled jobs. Three jobs each wanting eight threads on eight cores is worse than three jobs each wanting one.
  • Small matrices. A 200×200 dgemm does not have enough work to amortize thread wake-up. OpenBLAS's own heuristic often decides not to parallelize; forcing one thread removes the decision cost.
  • Memory-bound operations. Elementwise ops, np.where, pandas groupby aggregations — these are limited by RAM bandwidth, not FLOPs. Extra threads add cache-coherence traffic without adding throughput.
  • Predictable memory. Each BLAS thread gets its own scratch buffers. Eight threads means eight copies of the workspace. On a box with several jobs, that is the difference between fitting in RAM and swapping.
  • Container and CI environments. The container sees the host's core count but may be CPU-limited by cgroups. Threads get throttled and the pool spends time in futex waits.

Losses

  • A single large dense factorization. Cholesky on a 5000×5000 matrix, SVD on a big design matrix, a Monte Carlo sweep that is pure dgemm. Here BLAS threading is genuinely good, and forcing one thread can cost you 5–10×.
  • A job that owns the machine. If nothing else is scheduled and the box has 16 idle cores, pinning to one thread leaves them idle.

The practical rule: if jobs overlap, or if the workload is not dominated by large dense linear algebra, go single-threaded. If you have one job that does heavy dgemm and it runs alone, give it threads — but give it only that job, and set the thread count explicitly rather than letting it grab every core.

You can also split the difference per job. A feature builder that mostly does pandas joins gets OMP_NUM_THREADS=1. A nightly calibration that inverts a big matrix gets OPENBLAS_NUM_THREADS=4 and is scheduled in a window where nothing else runs.

A Benchmark Harness

Do not guess. Measure on the actual box with the actual workload. Here is a small harness that runs a representative operation at several thread counts and reports wall time and peak RSS.

"""bench_threads.py — compare BLAS thread counts on a real workload."""
import os
import subprocess
import sys
import time
import json

# The child process does the actual work so the environment is clean.
CHILD = r"""
import os, time, json, sys
import numpy as np

n = int(sys.argv[1])
reps = int(sys.argv[2])
rng = np.random.default_rng(0)

a = rng.standard_normal((n, n))
b = rng.standard_normal((n, n))

# warm up
_ = a @ b

t0 = time.perf_counter()
for _ in range(reps):
    c = a @ b
t1 = time.perf_counter()

# a memory-bound op for contrast
t2 = time.perf_counter()
for _ in range(reps):
    d = np.sqrt(np.abs(a))
t3 = time.perf_counter()

print(json.dumps({
    "matmul_s": (t1 - t0) / reps,
    "elementwise_s": (t3 - t2) / reps,
}))
"""


def run(n: int, reps: int, threads: int) -> dict:
    env = os.environ.copy()
    for var in ("OPENBLAS_NUM_THREADS", "MKL_NUM_THREADS",
                "OMP_NUM_THREADS", "NUMEXPR_NUM_THREADS"):
        env[var] = str(threads)

    proc = subprocess.run(
        [sys.executable, "-c", CHILD, str(n), str(reps)],
        env=env, capture_output=True, text=True, check=True,
    )
    return json.loads(proc.stdout.strip().splitlines()[-1])


if __name__ == "__main__":
    n = 2000
    reps = 3
    for threads in (1, 2, 4, 8):
        r = run(n, reps, threads)
        print(f"threads={threads:>2}  "
              f"matmul={r['matmul_s']*1000:8.1f} ms  "
              f"elementwise={r['elementwise_s']*1000:8.1f} ms")
Enter fullscreen mode Exit fullscreen mode

Run it on the target machine. On a typical 8-core VM with a 2000×2000 matmul you will see matmul drop steadily from 1 to 8 threads, while elementwise barely moves — it is memory-bound and the extra threads do nothing. That is the shape of the trade-off.

Now run it twice concurrently to simulate overlap:

import concurrent.futures as cf

with cf.ThreadPoolExecutor(max_workers=2) as pool:
    futures = [pool.submit(run, 2000, 3, t) for t in (1, 8)]
    for f in futures:
        print(f.result())
Enter fullscreen mode Exit fullscreen mode

With threads=8 on both, the combined time is usually worse than running them sequentially. With threads=1 on both, they finish in roughly the time of one job each, in parallel. That is the whole argument for pinning to one thread in a scheduled environment.

Reading the results

Two numbers matter: wall time and peak RSS. Wall time tells you whether you gained throughput. Peak RSS tells you whether you gained predictability. In my experience the memory win is the one that stops the 3 a.m. pages — a job that used to balloon to 2 GB under thread contention now sits flat at 600 MB.

If you want peak RSS in the harness, wrap the child with psutil:

import psutil, subprocess

p = subprocess.Popen([sys.executable, "-c", CHILD, "2000", "3"], env=env)
proc = psutil.Process(p.pid)
peak = 0
while p.poll() is None:
    try:
        peak = max(peak, proc.memory_info().rss)
    except psutil.NoSuchProcess:
        break
    time.sleep(0.05)
print(f"peak RSS: {peak / 1e6:.1f} MB")
Enter fullscreen mode Exit fullscreen mode

Checklist

For every scheduled numerical job on a shared Windows box:

  1. Set OPENBLAS_NUM_THREADS, MKL_NUM_THREADS, OMP_NUM_THREADS, and NUMEXPR_NUM_THREADS to 1 in the launcher, before python.exe.
  2. If you cannot control the launcher, set them at the top of the entry point with os.environ.setdefault, before any numerical import — or use a sitecustomize.py.
  3. Benchmark on the actual hardware with the actual matrix sizes. Raise the thread count only for a job that owns the machine and does large dense linear algebra.
  4. Watch peak RSS, not just wall time. Flat memory is the real prize.

The environment variables are not a micro-optimization. On a box running several jobs, they are the difference between a schedule that finishes and one that drifts into the morning.

More notes like this ship every week on this site.


Daily Picks

The following pairs are selected from the multi-timeframe trend scanner (Gate.io futures) and are for technical-analysis study only — not investment advice.
Data updated: 2026-10-06 12:36:33

Long

Pair Signal Price Take Profit Stop Loss R/R
SKYAI $0.0416 $0.0433 $0.0406 1:1.6
RE $0.4991 $0.5181 $0.4866 1:1.5

2 picks selected. Scanner runs every 15 minutes.

Top comments (0)