Most concurrency tutorials explain lock and Interlocked in a paragraph and move on. I wanted to see what they actually do, so I ran a deliberately tiny experiment on my laptop: increment a counter 100,000 times using Parallel.For, four different ways. The results contained one expected finding and one that changed how I think about parallel code.
The Experiment
The unsafe version looks harmless:
csharp
long counter = 0;
Parallel.For(0, 100_000, i => { counter++; });
Console.WriteLine(counter); // expected: 100000
I then wrote three corrected versions and a plain sequential loop for comparison:
csharp
// 1. lock
object gate = new object();
Parallel.For(0, N, i => { lock (gate) { counter++; } });
// 2. Interlocked
Parallel.For(0, N, i => { Interlocked.Increment(ref counter); });
// 3. Thread-local accumulation, merged once per task
long total = 0;
Parallel.For(0, N,
() => 0L, // each task starts with its own counter
(i, state, local) => local + 1, // count privately, no sharing
local => Interlocked.Add(ref total, local)); // merge once at the end
Why counter++ Isn't One Operation
The line looks atomic, but it compiles to three steps: read the current value, add one, write the result back. Two threads can interleave those steps:
Time Thread A Thread B Counter in memory
1 reads 5 5
2 reads 5 5
3 writes 6 6
4 writes 6 6
Two increments happened, but the counter moved by one. Across 100,000 attempts on several threads, most of those collisions go unnoticed, because nothing throws and nothing is logged.
Result 1: The Unsafe Version Was Wrong Every Time
Across 13 runs of the unsafe version, spread over three versions of the test program, the final value was never 100,000. The results ranged from 1,020 to 32,690, which means between roughly 67% and 99% of the increments were lost. The answer also changed on every run, which makes this kind of bug hard to catch: a test that passes once proves nothing.
Result 2: Timings (Release Build, .NET 8)
The first round of any benchmark includes compilation and thread-pool start-up, so the table below uses rounds 2 and 3 only. Times are in milliseconds.
Version Result Time (ms)
Sequential loop 100,000 0.033 – 0.034
Unsafe counter++ 17,825 and 22,577 (wrong) 0.751 – 1.093
lock 100,000 9.0 – 22.8
Interlocked 100,000 2.2 – 2.7
Thread-local + merge 100,000 0.277 – 0.287
These come from one machine and one process, and from Visual Studio, so the ordering and the rough size of the gaps are meaningful, but I wouldn't trust the decimals.
The Surprise: Parallel Made It Slower
The plain sequential loop beat every parallel version. Even the best correct parallel version, thread-local accumulation, was roughly eight to nine times slower than a simple for loop. Interlocked was around 65 to 80 times slower, and lock hundreds of times slower.
The sequential loop is so trivial that the compiler very likely optimized it heavily, so the number is best read as "this work costs almost nothing." That is exactly the point. Splitting work across threads has a fixed cost for scheduling, coordination, and merging results. When each item of work is tiny, that cost dominates, and the single-threaded version wins.
Parallelism only pays off when each unit of work is heavy enough to outweigh the overhead. Incrementing a number is not. Processing an employee record that involves database round trips might be, but that is something to measure, not assume.
Second Surprise: Unsafe Was Slower Than Correct
The unsafe version, which gives wrong answers, was also three to four times slower than the thread-local version, which gives right ones. A plausible explanation is contention: every thread repeatedly reads and writes the same memory location, so processor cores spend time passing it back and forth. I didn't measure the cause, so I'm presenting it as a likely explanation, not a finding. The measured fact is that sharing one variable was slower even before correctness was fixed.
lock vs. Interlocked
In these rounds, Interlocked was about 3 to 10 times faster than lock, and lock was the least consistent, ranging from 9 to 51 milliseconds across rounds. A reasonable reading is that Interlocked.Increment typically maps to a single atomic processor instruction, while lock has to coordinate threads that may end up waiting on each other.
An earlier, Debug-build version of the test showed a much larger gap, around 30 times. That used a different program, so the two aren't directly comparable, and I wouldn't conclude anything about Debug versus Release from them.
Interlocked has a limit that matters more than speed: it covers a single operation on a single variable. If an update involves several variables that must stay consistent with each other, Interlocked can't protect them, and lock (or another synchronization tool) is the right choice.
The Best Fix: Don't Share
The thread-local version was the fastest correct approach, and it works for a structural reason. Each task counts into its own private variable and only combines results once at the end. There is no contention because there is nothing shared to contend over. The same idea applies well beyond counters: summing totals, building per-thread lists, or collecting per-thread results and merging them afterward.
A Decision Rule
Situation Reach for
Many threads updating one number Interlocked
Several related values that must change together lock
Counting or totaling across a large loop Thread-local state, merged once
Very small work per item Check whether sequential is simply faster
What This Means for Batch Processing
The same reasoning applies to a payroll-style batch. Each employee's work involves database operations, which is far heavier than incrementing a number, so parallel processing may well help. But the earlier lesson holds: each thread needs its own connection and transaction, and the real speed-up depends on how much work each record carries and how the database handles concurrent writes. A four-thread split will not automatically finish in a quarter of the time. Measure sequential first, then measure parallel, and keep whichever is actually faster and correct.
How to Benchmark Without Fooling Yourself
Warm up first. The first round in my run was noticeably slower for several versions.
Report ranges, not single numbers. The lock timings varied by more than a factor of five.
Use Release builds. Debug builds change what the compiler does.
Include a sequential baseline. Without it, I would never have noticed that parallel was losing.
For precise numbers, use a proper tool such as BenchmarkDotNet, which handles warm-up and statistics for you.
Takeaway
An unsynchronized shared counter lost up to 99% of its updates, and a different answer every run is the signature of the bug. Of the fixes, Interlocked is the right tool for a single shared number, lock is for updates that span several values, and the fastest option was to avoid sharing at all. The most useful result, though, was that the parallel versions lost to a plain loop. Before parallelizing anything, measure the sequential version first.
Top comments (0)