DEV Community

Dhana
Dhana

Posted on

Multithreading in .NET for Batch Processing: What Actually Has to Change

A loop that processes 500 employee records one at a time is simple to reason about and easy to get right. Making that same loop run on multiple threads at once can cut processing time substantially, but it isn't a change you can make casually. This article walks through what multithreading actually means, how to apply it to a real batch-processing scenario, and the one structural change that makes it safe instead of dangerous.

What a Thread Actually Is

A thread is a single sequence of instructions executing in order, one line at a time. Code written and run normally executes on a single thread: step one happens, then step two, then step three, strictly in sequence. Multithreading means running more than one of these sequences at the same time, so independent pieces of work happen concurrently instead of waiting on each other.

For a loop processing 500 records sequentially, that's the equivalent of a single queue with one person serving every customer in turn. Splitting that work across four threads is closer to opening four separate queues, each handling its own share of the customers, finishing in roughly a quarter of the time.

Why This Matters Specifically for a Salary Batch

A payroll batch has a property that makes speed genuinely valuable: it's usually running against a real deadline, processing hundreds or thousands of employee records that each require their own database work. But payroll also has a property that makes carelessness genuinely costly: the result has to be correct, and a partial, inconsistent, or duplicated save is a real problem, not a minor inconvenience.

This combination is exactly where multithreading needs the most care. The speed benefit is real, but it has to be built on something that doesn't compromise the correctness guarantees a single-threaded, single-transaction approach already provides.

The Mistake: Sharing One Connection and Transaction Across Threads

A naive approach to speeding up the batch might try to process multiple employees at once while reusing the same database connection and transaction across all of them:

csharp
// Risky: one shared connection and transaction across multiple threads
using var connection = new OracleConnection(connectionString);
connection.Open();
using var transaction = connection.BeginTransaction();

Parallel.ForEach(employees, employee =>
{
ProcessSalary(employee, connection, transaction); // shared across threads
});

transaction.Commit();

This is the same category of problem covered in an earlier article about Dependency Injection lifetimes: a single shared resource, accessed by multiple concurrent operations at once, becomes a source of corruption rather than a convenience. Most database connection objects are not safe to use from multiple threads simultaneously. Two threads attempting to write through the same connection at the same moment can produce errors, corrupted data, or a transaction left in an undefined state. The underlying mistake is identical to a Singleton service holding shared mutable state across concurrent HTTP requests — just relocated from request-handling code to batch-processing code.

The Fix: One Connection and Transaction Per Thread

The structural change that makes this safe is straightforward to state, even though it requires rethinking the batch's structure: each thread needs its own connection and its own transaction, entirely separate from every other thread's.

csharp
var employeeBatches = SplitIntoBatches(employees, batchCount: 4);

Parallel.ForEach(employeeBatches, batch =>
{
using var connection = new OracleConnection(connectionString);
connection.Open();
using var transaction = connection.BeginTransaction();

try
{
    foreach (var employee in batch)
    {
        ProcessSalary(employee, connection, transaction);
    }

    transaction.Commit();
}
catch
{
    transaction.Rollback();
    throw;
}
Enter fullscreen mode Exit fullscreen mode

});

With this structure, each thread processes its own slice of employees against its own isolated connection and transaction. If one thread's batch encounters an error and rolls back, that rollback only affects the records that specific thread was handling. The employees already committed by other threads, running independently, are entirely unaffected. One thread's failure doesn't retroactively undo another thread's completed, correct work, and no two threads are ever contending for the same connection or transaction at once.

What Still Needs Attention Beyond Connections

Isolating connections and transactions per thread solves the most dangerous failure mode, but a few other details matter for a batch like this to behave correctly under concurrency:

Splitting the work evenly. How employees are divided across threads affects how balanced the workload is; an uneven split means some threads finish early while others remain the bottleneck.
Logging and reporting results. If each thread logs its own outcome independently, those results need to be collected and reconciled afterward into one overall batch report, rather than written to a single shared log object from multiple threads without coordination.
Degree of parallelism. Running too many threads at once against the database can overwhelm the connection pool or the database server itself, so the number of threads needs to be chosen deliberately rather than maximized blindly.
A Practical Question for Any Parallelized Batch

Before parallelizing any existing sequential batch process, it's worth asking directly: does every piece of shared state in this code, connections, transactions, in-memory counters, logging objects, genuinely belong to one unit of work, or is something being reused across what should be independent, isolated operations? If the answer reveals a connection or transaction being shared across threads, that's the one change that needs to happen before anything else, regardless of how much of a speed improvement the rest of the parallelization offers.

Takeaway

Multithreading a batch process like salary generation can deliver a genuine, substantial speed improvement, but the benefit depends entirely on isolating each thread's database work from every other thread's. Sharing a single connection and transaction across concurrent threads reproduces the same shared-state corruption risk found in other contexts, like a Singleton DI registration holding per-request data. Giving each thread its own connection and its own transaction removes that risk, keeping each thread's success or failure fully independent of every other thread's, which is exactly the guarantee a process like payroll generation can't afford to lose for the sake of speed.

Top comments (0)