I had a twenty-file test suite that took just over four seconds to run. I added one flag, --parallel, and it took one second. Then I tried to push it further and found the exact point where the flag stops helping, which turned out to depend on what kind of work the tests were actually doing, not on how many files there were.
Bun 1.4 shipped on 20 August 2026 and brought bun test --parallel, which distributes test files across worker processes instead of running them one after another in a single process. I'm running Bun 1.4.2 on a 4-core Intel Xeon VM (KVM, 4 logical CPUs, no cgroup limit), and I wanted to know what the flag actually buys you, where it tops out, and what the "implies --isolate" line in the docs means in practice.
The headline number
I wrote twenty test files, two tests each, where every test awaits a 100ms setTimeout before asserting — a stand-in for the API calls and debounces that make a lot of real integration tests slow. Three runs each, fastest and slowest within 15ms of each other:
| Mode | Wall time |
|---|---|
bun test (serial) |
4.04s |
--parallel=1 |
4.06s |
--parallel=2 |
2.05s |
--parallel=4 (default, = CPU count) |
1.03s |
--parallel=8 |
0.63s |
With no argument, --parallel defaults to the number of CPU cores, which on this box is 4. That alone took the suite from 4.04s to 1.03s, a 3.9x speedup that lines up almost exactly with four workers splitting twenty files. What surprised me is that --parallel=8 kept improving, down to 0.63s, on a machine with only four logical cores. I'll get to why in a minute, because it's the most important part of this post.
The opposite reading: CPU-bound tests hit a wall at the core count
A test that awaits setTimeout isn't actually using the CPU while it waits. It's just occupying a slot. So I built a second suite, same shape, twenty files and two tests each, but this time each test runs a real computation: a prime sieve up to 20 million, which takes genuine CPU time with nothing to wait on.
| Mode | Wall time |
|---|---|
bun test (serial) |
6.1s |
--parallel=1 |
6.2s |
--parallel=2 |
3.3s |
--parallel=4 |
1.8s |
--parallel=6 |
1.9s |
--parallel=8 |
1.8s |
This is the number that argues against the headline. Scaling is close to linear up to four workers, matching the four physical cores, and then it stops. Going to six or eight workers bought nothing, and in a couple of runs was a touch slower than four, presumably from context-switch overhead as processes fight for the same cores. --parallel schedules OS processes; it doesn't make more CPU appear. For a suite that's actually CPU-bound, the number that matters isn't how many files you have, it's nproc.
Put the two tables together and the real rule is simpler than either one alone: --parallel can scale past your core count for tests that spend most of their time waiting, and it cannot for tests that spend most of their time computing. Most real suites are a mix of both, so the honest advice is to benchmark your own suite rather than assume either curve.
What "implies --isolate" actually isolates
The docs say --parallel implies --isolate, and describe it as giving "each file... a fresh global object even when two files land on the same worker." I wanted to know exactly what that resets, because the phrasing leaves open whether it's per-file or something coarser or finer.
I wrote a module with a plain top-level counter and four test files that import it and log what they see:
// shared.ts
let count = 0;
export function bump() { count++; return count; }
Under plain bun test, the whole run shares one process and one module cache, so the counter climbs across files in whatever order they happen to run:
file-4 sees counter = 1
file-1 sees counter = 2
file-2 sees counter = 3
file-3 sees counter = 4
Under --parallel, every file saw 1:
file-1 sees counter = 1 worker = 1
file-2 sees counter = 1 worker = 1
file-3 sees counter = 1 worker = 1
file-4 sees counter = 1 worker = 1
All four ran in the same worker process (same PID), so this isn't about separate OS processes — it's a fresh JS global per file inside one process, exactly as documented. Adding --no-isolate to --parallel brought the leak straight back: 1, 2, 3, 4 again, same as serial.
The part the one-line doc summary doesn't spell out: isolation is per file, not per test. Three tests inside a single file, run under --parallel, saw the counter climb 1, 2, 3 with no reset between them. If your tests rely on module-level setup running fresh for every single test, not just every file, --isolate won't give you that, parallel or not.
Bun also sets BUN_TEST_WORKER_ID and JEST_WORKER_ID to the worker's 1-based index, which I confirmed by printing both in the test above — worth knowing if you're porting a Jest setup that keys a database name or port off that variable.
--bail stops starting files, not running ones
The docs describe it precisely: "the coordinator handles --bail at file granularity: once the failure threshold is reached it starts no new files, but files already running finish." I tested this against twenty files where the first one fails a 300ms-long test.
Without --bail, under --parallel, all twenty files ran regardless of the failure:
19 pass
1 fail
Ran 20 tests across 20 files. [1.53s]
With --bail added:
Bailed out after 1 failure
3 pass
1 fail
Ran 4 tests across 4 files. [320.00ms]
Four files ran, not one. That matches four workers already mid-flight when the failure landed; the other sixteen files never started. This is a real time saving in CI (1.53s down to 0.32s here), but it also means --bail under --parallel won't tell you about any other failures in those sixteen untouched files, including ones with nothing to do with whatever broke the first.
Coverage merges correctly, refusals are clean, one bad file doesn't take down the run
Three more things I checked, more briefly:
Coverage merging. I split three functions of one module across three test files, each covering one function, and compared --coverage output. Serial and --parallel --coverage reported identical numbers — 75% functions, 60% lines, same uncovered line range — so the claimed coverage merge across workers held up exactly.
Invalid input is rejected before anything runs. --parallel=0, --parallel=-1, and --parallel=abc all produced the same clean message and exit code 1, with nothing executed:
error: --parallel expects a positive integer, received "0"
A crash in one file doesn't sink the run. I had one test file throw at module load time, independent of any test inside it. Under --parallel, the other three files in that batch still ran and passed, and the crash was reported as a distinct "error" rather than lumped in with assertion failures:
bun test v1.4.2 (744846f84) 4x PARALLEL
tests-crash/cr2.test.ts:
# Unhandled error between tests
-------------------------------
error: boom at module load time
-------------------------------
3 pass
1 fail
1 error
Ran 4 tests across 4 files. [11.00ms]
That 4x PARALLEL line in the banner is also the easiest way to confirm from a CI log how many workers actually ran, without adding anything to your own test code.
What I got wrong on the way
My first attempt at a "CPU-bound" suite used while (Date.now() - start < 100) {} as a stand-in for real work, expecting it to behave like the prime-sieve suite. It didn't: --parallel=8 kept getting faster on a 4-core box well past where real CPU work should have plateaued, the same shape as the IO-bound suite. I spent a while assuming Bun was doing something clever with scheduling before I worked out the actual cause: a spin loop checking the wall clock only needs to be scheduled once after its own deadline to notice time has passed and exit. It behaves like a sleep, not like computation, because the OS can preempt it for as long as it likes without the loop caring, so long as real time has moved on by the time it next gets to check. That's why I rebuilt the CPU-bound suite around an actual prime sieve with no clock check in it at all, which is the version in the table above.
Run it yourself
This is a trimmed version of what produced the first table. It needs Bun 1.4 or later.
curl -fsSL https://bun.sh/install | bash
mkdir bun-parallel-demo && cd bun-parallel-demo
for i in $(seq -w 1 20); do
cat > "t$i.test.ts" <<EOF
import { test, expect } from "bun:test";
test("io-$i", async () => {
await new Promise((r) => setTimeout(r, 100));
expect(1 + 1).toBe(2);
});
EOF
done
bun test # serial baseline
bun test --parallel # defaults to your CPU count
bun test --parallel=8 # push past it, since these tests just wait
Swap the body of the generated test for a real computation (a sieve, a sort, anything with no timer in it) to see the plateau instead of continued scaling, and compare against nproc on your own machine.
What to do with this
Before turning --parallel on for an existing suite, work out which kind of suite you actually have. If most of your tests wait on timers, sockets, or a database round trip, raising --parallel past your core count is free speed and worth trying immediately. If your tests are doing real computation in-process, there is no benefit past nproc, and you should set --parallel to that number rather than guessing higher. Either way, check what your tests actually share at module scope before you lean on --isolate for correctness: it resets state between files, not between tests in the same file, and that gap is exactly where flaky suites tend to live.
Top comments (0)