The concepts that make concurrent code make sense, with enough real code that each one actually clicks.
Multithreading has a specific texture that's different from most programming topics. You can know exactly what synchronized does on paper and still get tripped up the moment you're asked to trace through what happens when two threads run the same three lines at the same time. A lot of the difficulty isn't the syntax, it's building the habit of imagining two (or ten) things happening at once, and noticing exactly where that stops being safe.
This post goes through multithreading as a concept, using Java to express it, since Java's concurrency vocabulary (synchronized, volatile, the java.util.concurrent package) is close to the vocabulary most languages end up reaching for anyway. It covers what a thread actually is, the lifecycle and the methods you'll actually use, race conditions, locks and monitors, the memory model, atomics, the executor framework, futures, concurrent collections, coordination utilities, deadlocks, virtual threads, and a way of reasoning through a stuck or misbehaving concurrent program when you're staring at one. The code targets Java 21, the current long-term support release, which matters here specifically because it's the release that made virtual threads a real, non-preview feature.
What multithreading actually is
Start simple. Multithreading means one program does more than one thing at the same time, by running several independent paths of execution instead of one. Each of those paths is a thread. Without threads, a program is a single line of instructions executing one after another; a slow file read blocks everything else, a slow network call blocks everything else, nothing overlaps. Threads let a program keep making progress on one piece of work while another piece is waiting on something slow.
Process vs thread. A process is a running program with its own private memory space, completely isolated from every other process. A thread is a path of execution inside a process, and every thread in the same process shares that process's memory. That sharing is the entire reason threads exist (they're cheap to create and can pass data to each other just by reading and writing the same variables) and it's also the entire reason multithreading is hard: two threads reading and writing the same variable at the same time is where nearly every concurrency bug traces back to.
Each thread gets its own private call stack (its own local variables and method-call history) but reads and writes the same heap as every other thread in the process. That's the picture in the cover image: shared counter, private cutting boards.
How the JVM actually runs a thread. For an ordinary Java Thread (called a "platform thread" from here on, to distinguish it from virtual threads, which work differently), the JVM maps it directly to a real operating system thread, one-to-one. The OS, not the JVM, decides when each thread actually gets CPU time and pays the cost of switching between them. That one-to-one mapping is simple and has worked for decades, but it means each thread carries real OS overhead (typically around a megabyte of stack memory, plus the cost of a context switch), which is why a server handling ten thousand concurrent connections by spinning up ten thousand platform threads runs into trouble. Virtual threads change exactly this mapping, not the language semantics around it, which is worth holding onto for now.
Creating a thread: Thread vs Runnable. Java gives you two ways to define what a thread should run.
// HelloThread.java
public class HelloThread extends Thread {
@Override
public void run() {
System.out.println("Hello from " + Thread.currentThread().getName());
}
}
// HelloTask.java
public class HelloTask implements Runnable {
@Override
public void run() {
System.out.println("Hello from " + Thread.currentThread().getName());
}
}
// Main.java
public class Main {
public static void main(String[] args) {
// Option 1: extending Thread directly
HelloThread helloThread = new HelloThread();
helloThread.start();
// Option 2: implementing Runnable, then handing it to a Thread
HelloTask task = new HelloTask();
Thread t = new Thread(task);
t.start();
// Runnable also works fine as a lambda, since it's a single-method interface
Thread lambdaThread = new Thread(() ->
System.out.println("Hello from " + Thread.currentThread().getName())
);
lambdaThread.start();
}
}
Runnable is the one to default to, and it's worth understanding why. Java doesn't allow multiple inheritance, so HelloThread, which extends Thread, can never extend anything else. HelloTask, which implements Runnable, stays free to extend whatever it needs to, and just as importantly, a Runnable is a plain unit of work, not tied to any particular thread, which means the exact same HelloTask object can be handed to a thread pool without any changes at all.
Thread lifecycle, states, and the methods that matter
A thread moves through a small set of states, and Java exposes them directly through Thread.getState(), which makes this an easy thing to check yourself if you're ever staring at a program that seems stuck.
-
NEW: the
Threadobject exists butstart()hasn't been called yet. - RUNNABLE: the thread is either actually running or ready and waiting for the OS scheduler to give it CPU time. Java doesn't distinguish these two as separate states, both count as RUNNABLE.
- BLOCKED: the thread wants a lock that another thread currently holds.
-
WAITING: the thread is waiting indefinitely for another thread to do something, typically via
wait(),join(), orLockSupport.park(). -
TIMED_WAITING: same as WAITING, but with a timeout, via
sleep(ms),wait(ms), orjoin(ms). -
TERMINATED:
run()has finished, normally or because of an uncaught exception.
start() vs run() is worth getting exactly right, because it's an easy thing to get quietly wrong. Calling start() asks the JVM to create a new OS thread and have it execute run() on that new thread. Calling run() directly just calls it like any ordinary method, on the current thread, synchronously, with no new thread created at all. The compiler will not stop you from writing myThread.run() instead of myThread.start(), your code will compile and even run correctly in terms of output, and it will silently not be multithreaded at all.
// StartVsRun.java
public class StartVsRun {
public static void main(String[] args) {
Thread t = new Thread(() ->
System.out.println("running on: " + Thread.currentThread().getName())
);
t.run(); // prints "main" — ran on the calling thread, no new thread created
t.start(); // prints "Thread-0" (or similar) — actually runs on a new thread
}
}
The methods worth knowing beyond start(). A handful of Thread methods come up constantly once you're actually writing concurrent code, and knowing their exact behaviour, not just their names, is what makes the difference between code that looks right and code that actually is right.
join() makes the calling thread wait until the target thread finishes.
// JoinExample.java
public class JoinExample {
public static void main(String[] args) throws InterruptedException {
Thread worker = new Thread(() -> {
try {
Thread.sleep(2000); // simulate slow work
} catch (InterruptedException e) {
Thread.currentThread().interrupt();
}
System.out.println("worker finished its work");
});
worker.start();
worker.join(); // main thread blocks here until worker finishes
System.out.println("worker is done, main can continue");
}
}
Without that join(), main has no guarantee it runs after the worker, it might print first, might print after, the order is undefined. join() also takes an optional timeout in milliseconds, after which the caller stops waiting regardless of whether the target thread finished.
sleep(ms) is a static method that pauses the current thread for at least the given duration. Two details are easy to miss here. First, it's static, so someOtherThread.sleep(1000) doesn't pause someOtherThread, it pauses whatever thread that line of code is actually running on, which trips people up more often than you'd expect. Second, and this one's genuinely important: sleep() does not release any lock the thread currently holds. A thread sleeping inside a synchronized block keeps that lock the entire time it's asleep, blocking every other thread that wants it, for the full duration.
interrupt() is Java's only real mechanism for asking a thread to stop, and it's cooperative, not forceful. Calling t.interrupt() doesn't stop the thread itself, it just sets an internal flag and, if the thread happens to be blocked in something like sleep(), wait(), or join() at that moment, wakes it up with an InterruptedException. If the thread is doing ordinary CPU work and never checks for the flag, interrupt() does nothing at all. That's why a well-behaved long-running task periodically checks whether it's been asked to stop and exits cleanly if so:
// InterruptibleTask.java
public class InterruptibleTask implements Runnable {
@Override
public void run() {
while (!Thread.currentThread().isInterrupted()) {
// do one unit of work
performOneStep();
}
System.out.println("cleanly stopped");
}
private void performOneStep() {
// stand-in for real work
}
}
// Main.java
public class Main {
public static void main(String[] args) throws InterruptedException {
Thread t = new Thread(new InterruptibleTask());
t.start();
Thread.sleep(1000); // let it run for a second
t.interrupt(); // ask it to stop
}
}
There is no safe way to force a thread to stop from the outside in modern Java. Thread.stop() exists but has been deprecated for a long time, because it can kill a thread mid-update and leave shared data in a half-modified, inconsistent state. Interruption is the sanctioned pattern precisely because it lets the thread finish its current unit of work and clean up before exiting, rather than being killed mid-sentence.
Daemon threads, set with t.setDaemon(true) before calling start(), are background threads the JVM won't wait for. A running JVM stays alive as long as at least one non-daemon thread is alive; the moment every remaining thread is a daemon, the JVM exits immediately, without letting those daemon threads finish whatever they were doing. Ordinary threads default to non-daemon, and this is a real, practical difference: use daemon threads for background housekeeping you're fine with being cut off abruptly, never for work that must complete. yield() is the least useful of the bunch and worth knowing mainly so you don't overestimate it: it's a hint to the scheduler that the current thread is willing to let other threads of the same priority run, but the scheduler is completely free to ignore it, and most real code never needs it.
Race conditions: the counter that doesn't count
Here's a small class with a counter, and a separate file that hammers it with ten threads, each incrementing it a thousand times:
// Counter.java
public class Counter {
private int count = 0;
public void increment() {
count++;
}
public int getCount() {
return count;
}
}
// Main.java
import java.util.ArrayList;
import java.util.List;
public class Main {
public static void main(String[] args) throws InterruptedException {
Counter counter = new Counter();
List<Thread> threads = new ArrayList<>();
for (int i = 0; i < 10; i++) {
Thread t = new Thread(() -> {
for (int j = 0; j < 1000; j++) {
counter.increment();
}
});
threads.add(t);
t.start();
}
for (Thread t : threads) {
t.join();
}
System.out.println("Final count: " + counter.getCount());
// Expected: 10000
// Actual: something smaller, and different almost every run
}
}
You would expect 10,000. You reliably get something smaller, and a different smaller number nearly every time you run it. This is the entire concept of a race condition, and the reason is that count++ isn't one operation, it's three: read the current value, add one to it, write the new value back. It's worth being able to narrate the exact interleaving that breaks it:
- Thread A reads
count, gets 5. - Thread A is paused by the OS scheduler before it writes anything back.
- Thread B reads
count, also gets 5 (A's increment hasn't been written yet). - Thread B computes 6 and writes it.
- Thread A resumes, computes 6 from the value it read earlier, and writes 6 too.
Two increments happened. The counter only went up by one. Nothing crashed, nothing threw an exception, the code just silently produced the wrong answer, which is exactly what makes race conditions dangerous in real systems: they don't announce themselves. The next few sections fix this same Counter class in three different ways.
synchronized, monitors, and reentrancy
Java's oldest fix for the counter problem is the synchronized keyword. Every object in Java has an associated monitor (an intrinsic lock), and synchronized makes a thread acquire that lock before entering a block of code, and release it on the way out, automatically, even if an exception is thrown inside.
// Counter.java
public class Counter {
private int count = 0;
public synchronized void increment() {
count++;
}
public synchronized int getCount() {
return count;
}
}
Run the exact same Main class from before, ten threads, a thousand increments each, against this version, and you get 10,000, every time. Only one thread can be inside a synchronized method (or block) on the same object at once; every other thread that wants in has to wait for the lock to be free.
You can synchronize a whole method, as above, which locks on this (the object instance), or synchronize a specific block, which lets you name exactly which object to lock on and keep the locked region as small as possible:
// Counter.java (block form instead of method form)
public class Counter {
private int count = 0;
private final Object lock = new Object();
public void increment() {
synchronized (lock) {
count++;
}
}
public int getCount() {
synchronized (lock) {
return count;
}
}
}
Smaller synchronized blocks are generally better, since anything inside the lock is a stretch of code where every other thread wanting that lock is stuck waiting, and keeping that stretch short keeps contention low.
Object-level vs class-level locking. synchronized on an instance method locks on this, so it only blocks other threads calling synchronized instance methods on that same object. Two different Counter instances don't block each other at all. If you instead want to lock across every instance of a class, for a static counter shared globally, you synchronize on the class object itself, or mark a static method synchronized, which does the same thing implicitly:
// GlobalCounter.java
public class GlobalCounter {
private static int count = 0;
// synchronized on GlobalCounter.class, shared by every call, from any instance
public static synchronized void increment() {
count++;
}
public static synchronized int getCount() {
return count;
}
}
Reentrancy is the property that lets a thread that already holds a lock re-acquire the same lock without deadlocking itself, which matters constantly in practice because one synchronized method commonly calls another synchronized method on the same object.
// ReentrantExample.java
public class ReentrantExample {
private int count = 0;
public synchronized void outer() {
System.out.println("in outer(), about to call inner()");
inner(); // the current thread already holds the lock, and re-enters safely
}
public synchronized void inner() {
count++;
System.out.println("in inner(), count is now " + count);
}
}
If synchronized locks weren't reentrant, that call to inner() from inside outer() would deadlock the thread against itself, waiting forever for a lock it's already holding. Java's intrinsic locks are reentrant by design, and that's a deliberate design choice, not an accident, one that ReentrantLock (covered next) takes its name directly from.
wait(), notify(), and coordinating threads
synchronized stops two threads from corrupting shared data at the same time. It doesn't help when one thread needs to wait for another thread to do something first, that's a different problem, and Java's answer is wait() and notify()/notifyAll(), defined on every Object.
The rules are strict and worth memorizing exactly, because getting any one of them wrong produces a program that hangs instead of one that crashes, which is far harder to debug. wait(), notify(), and notifyAll() can only be called from inside a synchronized block on the object you're calling them on, or Java throws IllegalMonitorStateException. Calling wait() releases the lock and puts the thread to sleep until another thread calls notify() or notifyAll() on the same object. notify() wakes up one waiting thread, chosen arbitrarily; notifyAll() wakes every thread waiting on that object, and they all then compete for the lock again.
There's one rule that's very easy to get wrong: always call wait() in a loop that rechecks the condition, never in a plain if. Java allows spurious wakeups, where a thread can wake from wait() even though nobody called notify(). If your code assumes an if check before wait() is still true the moment it wakes up, you've introduced a bug that shows up rarely, unpredictably, and is miserable to reproduce.
// SafeWaitPattern.java — the shape every wait() call should follow
public class SafeWaitPattern {
private boolean conditionIsTrue = false;
private final Object lock = new Object();
public void waitUntilReady() throws InterruptedException {
synchronized (lock) {
while (!conditionIsTrue) { // loop, never if
lock.wait();
}
// safe to proceed here, condition is genuinely true
}
}
}
A worked example: print odd and even alternately using two threads. This is a genuinely useful example to build once by hand, since it forces you to actually use wait()/notify() to coordinate two threads, rather than just define what they do.
// OddEvenPrinter.java
public class OddEvenPrinter {
private int number = 1;
private final int max;
private boolean oddTurn = true;
public OddEvenPrinter(int max) {
this.max = max;
}
public synchronized void printOdd() {
while (number <= max) {
while (!oddTurn) {
try {
wait();
} catch (InterruptedException e) {
Thread.currentThread().interrupt();
return;
}
}
System.out.println("Odd: " + number++);
oddTurn = false;
notify();
}
}
public synchronized void printEven() {
while (number <= max) {
while (oddTurn) {
try {
wait();
} catch (InterruptedException e) {
Thread.currentThread().interrupt();
return;
}
}
System.out.println("Even: " + number++);
oddTurn = true;
notify();
}
}
}
// Main.java
public class Main {
public static void main(String[] args) {
OddEvenPrinter printer = new OddEvenPrinter(10);
Thread oddThread = new Thread(printer::printOdd, "odd-thread");
Thread evenThread = new Thread(printer::printEven, "even-thread");
oddThread.start();
evenThread.start();
}
}
Walk through what's actually happening: both threads call into synchronized methods on the same printer object, so only one runs at a time. The odd-thread checks a shared boolean flag; if it's not its turn, it calls wait(), releasing the lock and going to sleep. The even-thread does the same in reverse. Whichever thread does get to print flips the flag and calls notify() to wake the other one up. Neither thread ever busy-waits, burning CPU checking a condition in a loop, they genuinely sleep until told to wake up, which is the entire point of wait()/notify() over a naive spin-loop.
The Java Memory Model and volatile
This is one of the places where understanding what's actually happening at the hardware level pays off, because the core distinction to hold onto is that visibility and atomicity are two separate problems, and synchronized happens to solve both at once, but volatile only solves one of them.
Visibility is about whether a change one thread makes to a variable is actually seen by another thread, at all, ever. Modern CPUs have per-core caches, and a thread running on one core can update a variable in its own cache without that update being immediately, or ever, flushed to main memory where another thread on another core would see it. Without any coordination, one thread can update a boolean running = true to false, and another thread checking running in a tight loop can keep reading a stale cached true forever, spinning in a loop that should have exited.
volatile fixes exactly this, and only this. A volatile field is never cached privately; every read goes to main memory and every write goes straight back to it, so every thread always sees the latest value. It also establishes a happens-before relationship: everything a thread wrote before writing a volatile variable becomes visible to any thread that reads that same volatile variable afterward, not just the volatile field itself.
// Worker.java
public class Worker implements Runnable {
private volatile boolean running = true;
public void stop() {
running = false;
}
@Override
public void run() {
int cycles = 0;
while (running) {
cycles++; // stand-in for real work
}
System.out.println("stopped after " + cycles + " cycles");
}
}
// Main.java
public class Main {
public static void main(String[] args) throws InterruptedException {
Worker worker = new Worker();
Thread t = new Thread(worker);
t.start();
Thread.sleep(1000);
worker.stop(); // without volatile, this update might never be seen by t
t.join();
}
}
But here's the part worth being careful about: volatile gives you visibility, it does not give you atomicity. Go back to count++ and make count volatile instead of synchronizing it: the race condition is completely unchanged, because count++ is still three separate operations (read, add, write), and volatile doesn't make those three operations happen as one atomic unit. Two threads can still both read the same value before either writes, and one increment still vanishes. volatile is the right tool for a single flag one thread sets and others read, like the running boolean above. It is the wrong tool for anything involving a read-modify-write sequence, like a counter, that needs synchronized, an explicit lock, or an atomic variable instead.
Happens-before, briefly, since it's the deeper idea underneath all of this: it's the JVM's guarantee about which memory effects are visible in which order across threads. A synchronized block establishes happens-before between the thread releasing a lock and the next thread acquiring that same lock, meaning everything the first thread did before releasing is guaranteed visible to the second thread after it acquires. volatile writes and reads establish the same guarantee for that specific field. Thread.start() happens-before anything the started thread does, and everything a thread does happens-before another thread's join() on it returns. Without one of these explicit happens-before relationships in place, the JVM and the CPU are both free to reorder instructions and cache values in ways that can genuinely surprise you, which is the deeper reason unsynchronized shared-memory code is unsafe even when it "looks" correct on paper.
Explicit locks: ReentrantLock, ReadWriteLock, StampedLock
synchronized is easy to use correctly but inflexible. java.util.concurrent.locks gives you explicit lock objects with more control, at the cost of having to remember to release them yourself.
// Counter.java
import java.util.concurrent.locks.ReentrantLock;
public class Counter {
private final ReentrantLock lock = new ReentrantLock();
private int count = 0;
public void increment() {
lock.lock();
try {
count++;
} finally {
lock.unlock(); // must be in finally, or an exception leaves the lock held forever
}
}
public int getCount() {
lock.lock();
try {
return count;
} finally {
lock.unlock();
}
}
}
That finally block is not optional. With synchronized, the JVM releases the lock automatically even if an exception is thrown inside. With ReentrantLock, nothing releases it for you, so a forgotten unlock() on an exception path leaves the lock held forever, and every other thread waiting for it waits forever too.
Here's what ReentrantLock gets you that synchronized doesn't, laid out directly:
synchronized |
ReentrantLock |
|
|---|---|---|
| Acquire/release | Automatic (block-scoped) | Manual (lock() / unlock()) |
| Try without blocking | Not possible |
tryLock(), returns immediately |
| Timeout on waiting | Not possible | tryLock(time, unit) |
| Interruptible while waiting | Not possible | lockInterruptibly() |
| Fairness (FIFO ordering) | Not possible | new ReentrantLock(true) |
| Multiple wait conditions | One implicit condition (wait/notify) |
Several, via newCondition()
|
tryLock() is how you'd avoid one specific flavour of deadlock rather than just waiting forever:
// TryLockExample.java
import java.util.concurrent.locks.ReentrantLock;
import java.util.concurrent.TimeUnit;
public class TryLockExample {
private final ReentrantLock lock = new ReentrantLock();
public void attemptWork() {
try {
if (lock.tryLock(2, TimeUnit.SECONDS)) {
try {
doProtectedWork();
} finally {
lock.unlock();
}
} else {
System.out.println("couldn't get the lock in time, doing something else instead");
}
} catch (InterruptedException e) {
Thread.currentThread().interrupt();
}
}
private void doProtectedWork() {
// the actual work that needs the lock
}
}
Fairness is worth a one-line caveat: a fair lock grants access in the order threads requested it, which sounds strictly better, but it comes with a real throughput cost, so it's an opt-in, not the default.
ReadWriteLock is the answer to a common, specific inefficiency: data that's read constantly and written rarely. A plain lock makes every reader wait for every other reader, even though two threads only reading the same data was never actually unsafe. ReadWriteLock splits this into a read lock, which any number of threads can hold at once as long as nobody holds the write lock, and a write lock, which is exclusive against everyone, readers included.
// SharedData.java
import java.util.concurrent.locks.ReadWriteLock;
import java.util.concurrent.locks.ReentrantReadWriteLock;
public class SharedData {
private final ReadWriteLock rwLock = new ReentrantReadWriteLock();
private String data = "initial value";
public String read() {
rwLock.readLock().lock();
try {
return data;
} finally {
rwLock.readLock().unlock();
}
}
public void write(String value) {
rwLock.writeLock().lock();
try {
data = value;
} finally {
rwLock.writeLock().unlock();
}
}
}
For a read-heavy cache with occasional writes, this can meaningfully beat a plain ReentrantLock, since readers stop blocking each other entirely. StampedLock takes this further with an optimistic read mode that doesn't block writers at all, a reader just validates afterward that no write happened in the meantime and retries if one did. It's a genuinely more advanced tool, worth being able to name and describe in one sentence, but you'd typically only reach for it once you've confirmed a ReadWriteLock isn't fast enough for what you're building.
Atomic variables and compare-and-swap
There's a third way to fix the counter, and it's usually the best one for a simple case like this: skip locking entirely and use a class built for lock-free atomic updates.
// Counter.java
import java.util.concurrent.atomic.AtomicInteger;
public class Counter {
private final AtomicInteger count = new AtomicInteger(0);
public void increment() {
count.incrementAndGet();
}
public int getCount() {
return count.get();
}
}
Run the same ten-threads Main class from the race condition section against this version and, again, you get exactly 10,000, with no synchronized keyword anywhere. AtomicInteger, AtomicLong, and AtomicReference work using a CPU-level instruction called compare-and-swap (CAS): "update this memory location to a new value, but only if it still holds the value I expect it to; tell me whether that succeeded."
incrementAndGet() is implemented, roughly, as a loop: read the current value, compute the new value, attempt a CAS from old to new, and if the CAS fails because another thread changed the value in between, loop around and try again with the fresh value. Here's that same loop using the lower-level compareAndSet() method directly, which shows what's happening underneath, though you'd never actually write it this way in real code, incrementAndGet() already does it for you:
// ManualCasIncrement.java — for illustration, showing what incrementAndGet() does internally
import java.util.concurrent.atomic.AtomicInteger;
public class ManualCasIncrement {
public static void incrementManually(AtomicInteger counter) {
while (true) {
int current = counter.get();
int updated = current + 1;
if (counter.compareAndSet(current, updated)) {
// succeeded: nothing else changed the value between get() and here
return;
}
// else: another thread updated it first, loop around and retry
}
}
}
No thread ever blocks waiting for a lock, they just occasionally retry, which tends to perform noticeably better than locking under high contention on simple operations like this, since there's no thread ever sitting idle waiting to be woken up.
The tradeoff is scope: atomics are excellent for a single variable, a counter, a flag, a reference that gets swapped. They don't help when you need to update several related variables together as one atomic unit, like moving money between two account balances, that's still a job for a lock.
Deadlock, livelock, and starvation: a bank transfer gone wrong
A deadlock happens when two or more threads are each waiting for a lock the other one holds, and neither will ever let go. Here's the version that shows up in real code, not a toy two-lock example: two bank accounts, two threads transferring money in opposite directions at the same time.
// Account.java
public class Account {
private final int id;
private double balance;
private final Object lock = new Object();
public Account(int id, double balance) {
this.id = id;
this.balance = balance;
}
public int getId() {
return id;
}
public Object getLock() {
return lock;
}
public void withdraw(double amount) {
balance -= amount;
}
public void deposit(double amount) {
balance += amount;
}
}
// BankTransfer.java — the broken version
public class BankTransfer {
public static void transfer(Account from, Account to, double amount) {
synchronized (from.getLock()) {
System.out.println(Thread.currentThread().getName() + " locked account " + from.getId());
synchronized (to.getLock()) {
System.out.println(Thread.currentThread().getName() + " locked account " + to.getId());
from.withdraw(amount);
to.deposit(amount);
}
}
}
}
// Main.java
public class Main {
public static void main(String[] args) {
Account accountA = new Account(1, 1000);
Account accountB = new Account(2, 1000);
Thread t1 = new Thread(() ->
BankTransfer.transfer(accountA, accountB, 100), "Thread-1"
);
Thread t2 = new Thread(() ->
BankTransfer.transfer(accountB, accountA, 50), "Thread-2"
);
t1.start();
t2.start();
// this can hang forever, depending on timing
}
}
Walk through why this locks up: Thread 1 acquires accountA's lock first, then tries to acquire accountB's lock. At almost the same moment, Thread 2 acquires accountB's lock first, then tries to acquire accountA's lock. If the timing lines up so each thread grabs its first lock before either reaches for its second, Thread 1 is now stuck waiting for a lock Thread 2 holds, and Thread 2 is stuck waiting for a lock Thread 1 holds. Neither thread will ever release what it's holding, because releasing only happens after the synchronized block completes, and neither block can complete. The program doesn't crash, it just stops, silently, forever, which is exactly what makes deadlocks nasty in production: nothing looks wrong until it does.
The fix is lock ordering: always acquire locks in the same global order, no matter which direction the transfer is going.
// BankTransfer.java — the fixed version
public class BankTransfer {
public static void transfer(Account from, Account to, double amount) {
Account first = from.getId() < to.getId() ? from : to;
Account second = from.getId() < to.getId() ? to : from;
synchronized (first.getLock()) {
synchronized (second.getLock()) {
from.withdraw(amount);
to.deposit(amount);
}
}
}
}
Now every thread, regardless of which direction it's transferring, locks the lower-ID account first. Thread 1 and Thread 2 can never both be holding one lock while wanting the other, because they always reach for the same first lock, so one of them simply waits at the very first synchronized block instead of deadlocking deep inside two nested ones. The general pattern worth taking away: assign a fixed, consistent order to every lock in the system, and always acquire in that order. The other standard tool is a timeout: tryLock(time, unit) on ReentrantLock lets a thread give up and back off instead of waiting forever, which avoids the deadlock at the cost of needing retry logic.
Two related terms are worth being able to define crisply, since they're easy to confuse with deadlock and with each other. Livelock is when threads aren't blocked, they're actively running, but they keep responding to each other in a way that prevents any of them from making real progress, like two people repeatedly stepping aside for each other in a hallway and never actually passing. Starvation is when one specific thread never gets access to a resource because other threads keep getting priority over it, it's not stuck waiting on a cycle, it's just perpetually unlucky (or deprioritized). Both come up less often than deadlock, but being able to name the difference correctly reflects real understanding of the distinction rather than a memorized list of terms.
The Executor framework and thread pools
Creating a new Thread() for every unit of work doesn't scale. Each thread carries real OS overhead, and a burst of ten thousand incoming requests each spinning up its own thread will exhaust memory and grind the machine to a halt on context switching long before it exhausts CPU. The Executor framework separates submitting work from how it actually gets run, so you write task-submission code once and swap the execution strategy underneath it.
// Main.java
import java.util.concurrent.ExecutorService;
import java.util.concurrent.Executors;
public class Main {
public static void main(String[] args) {
ExecutorService executor = Executors.newFixedThreadPool(4);
for (int i = 0; i < 10; i++) {
int taskNumber = i;
executor.submit(() -> {
System.out.println("running task " + taskNumber + " on " + Thread.currentThread().getName());
});
}
executor.shutdown(); // stop accepting new tasks, let submitted ones finish
}
}
Executors offers several convenience factory methods, and it's worth knowing them, along with why most production codebases have moved away from several of them in favour of building a ThreadPoolExecutor directly:
| Factory method | What it creates | The catch |
|---|---|---|
newFixedThreadPool(n) |
A pool of exactly n threads, an unbounded task queue | The unbounded queue can grow without limit if tasks arrive faster than they're processed, hiding a backlog until memory runs out |
newCachedThreadPool() |
Creates threads as needed, reuses idle ones, no upper bound | No upper bound means a burst of load can create an unbounded number of threads, exactly the problem pooling was meant to prevent |
newSingleThreadExecutor() |
One thread, tasks run strictly in order | Fine for its specific purpose, just not a general-purpose pool |
newScheduledThreadPool(n) |
Like fixed, plus delayed and repeating tasks | Same unbounded-queue concern as the fixed pool |
The common thread through the catches: several of these factories quietly use an unbounded queue or an unbounded thread count, which trades one failure mode (an OutOfMemoryError from too many raw threads) for a subtler one (silent, unbounded queue growth, or the same OutOfMemoryError from an unbounded cached pool under sustained load). This is exactly why it's worth reaching for ThreadPoolExecutor directly once you actually care about these bounds, where every one of them is explicit:
// Main.java
import java.util.concurrent.*;
public class Main {
public static void main(String[] args) {
ExecutorService executor = new ThreadPoolExecutor(
4, // core pool size: threads kept alive even when idle
8, // maximum pool size: the ceiling under load
60L, TimeUnit.SECONDS, // how long extra threads beyond core sit idle before dying
new LinkedBlockingQueue<>(100), // bounded queue: caps how much backlog can build up
new ThreadPoolExecutor.CallerRunsPolicy() // what to do when the queue is also full
);
for (int i = 0; i < 20; i++) {
int taskNumber = i;
executor.submit(() ->
System.out.println("task " + taskNumber + " on " + Thread.currentThread().getName())
);
}
executor.shutdown();
}
}
Walk through what each parameter is actually deciding. corePoolSize threads stay alive even with nothing to do. Once the queue fills up, the pool grows beyond core, up to maximumPoolSize, to absorb the extra load. keepAliveTime controls how long those extra threads beyond core linger idle before being shut down and reclaimed. The queue itself is bounded on purpose, so backlog has a hard ceiling instead of growing invisibly. And the rejection policy decides what happens once both the pool and the queue are completely full and a new task still arrives, options include throwing an exception (AbortPolicy, the default), silently dropping the task (DiscardPolicy), or, as used above, running the task on the calling thread itself as a form of backpressure (CallerRunsPolicy), which has the effect of slowing down whoever's submitting work until the pool catches up.
Sizing a pool sensibly depends on what the work actually is. For CPU-bound work, a pool size around the number of available CPU cores is typically a reasonable starting point, since more threads than cores just means more context switching with no extra throughput. For I/O-bound work, where threads spend most of their time waiting on a network call or a disk read rather than actually using the CPU, a pool considerably larger than the core count usually performs better, since idle-waiting threads aren't competing for CPU anyway. This exact tension, threads mostly sitting idle waiting on I/O, is precisely the problem virtual threads are built to solve.
Callable, Future, and CompletableFuture
Runnable doesn't return anything and can't throw a checked exception. Callable<V> fixes both: it returns a value and is allowed to throw.
// Main.java
import java.util.concurrent.*;
public class Main {
public static void main(String[] args) throws Exception {
ExecutorService executor = Executors.newFixedThreadPool(2);
Callable<Integer> task = () -> {
Thread.sleep(1000);
return 42;
};
Future<Integer> future = executor.submit(task);
System.out.println("doing other work while the task runs...");
Integer result = future.get(); // blocks until the result is ready, or throws
System.out.println("result: " + result);
executor.shutdown();
}
}
future.get() blocks the calling thread until the result is available, which is useful but limited: you can't easily chain what happens next, and you can't combine the results of several futures without writing a fair amount of blocking, coordinating code by hand.
CompletableFuture solves that by letting you describe a pipeline of what should happen when a result becomes available, without ever explicitly blocking:
// Main.java
import java.util.concurrent.CompletableFuture;
public class Main {
public static void main(String[] args) {
CompletableFuture<Integer> future = CompletableFuture
.supplyAsync(Main::fetchUserId)
.thenApply(id -> id * 2)
.thenApply(doubled -> doubled + 1);
future.thenAccept(result -> System.out.println("Final: " + result));
// give the async pipeline a moment to finish before the program exits
future.join();
}
private static int fetchUserId() {
// stand-in for a slow call, like hitting a database
return 21;
}
}
thenApply transforms a result once it's ready and hands you a new CompletableFuture wrapping the transformed value, thenAccept consumes the final result without producing a new value, and thenCompose is what you reach for when the next step is itself another asynchronous operation (chaining two CompletableFutures together rather than nesting one inside a plain function). CompletableFuture.allOf(f1, f2, f3) waits for several independent futures to all complete, which is the standard way to fan out several concurrent calls, say, three separate API calls, and combine their results once every one of them is done.
Concurrent collections, and building a blocking queue by hand
Wrapping a plain HashMap or ArrayList in synchronized on every call works, but it serializes every single access, one thread at a time, even for operations that could safely happen concurrently. java.util.concurrent gives you collections designed for concurrent access from the ground up, and they're almost always the better choice.
ConcurrentHashMap allows concurrent reads and writes without locking the entire map for every operation, internally splitting its locking so unrelated keys don't contend with each other. It's the default choice for a shared map accessed by multiple threads, essentially never Collections.synchronizedMap(new HashMap<>()) in new code.
// Main.java
import java.util.concurrent.ConcurrentHashMap;
import java.util.Map;
public class Main {
public static void main(String[] args) throws InterruptedException {
Map<String, Integer> visitCounts = new ConcurrentHashMap<>();
Runnable visitor = () -> {
for (int i = 0; i < 1000; i++) {
visitCounts.merge("home-page", 1, Integer::sum);
}
};
Thread t1 = new Thread(visitor);
Thread t2 = new Thread(visitor);
t1.start();
t2.start();
t1.join();
t2.join();
System.out.println(visitCounts.get("home-page")); // reliably 2000
}
}
CopyOnWriteArrayList takes a different approach entirely: every write creates a brand-new copy of the underlying array, while reads always work against a stable, never-changing snapshot and need no locking at all. That makes it excellent for read-heavy, write-rare situations, like a list of event listeners that's iterated constantly and modified rarely, and a poor choice for anything write-heavy, since copying the entire array on every single write gets expensive fast.
BlockingQueue implementations (ArrayBlockingQueue, LinkedBlockingQueue) add blocking behaviour on top of a queue: put() blocks the calling thread if the queue is full instead of failing, and take() blocks if the queue is empty instead of returning null. That blocking behaviour is exactly what makes the producer-consumer pattern simple to write correctly instead of needing manual wait()/notify() coordination.
It's worth actually building a small bounded blocking queue yourself once, using nothing but wait()/notify(), because it's a great way to see exactly what BlockingQueue is doing underneath rather than just knowing it exists:
// SimpleBlockingQueue.java
import java.util.LinkedList;
import java.util.Queue;
public class SimpleBlockingQueue<T> {
private final Queue<T> queue = new LinkedList<>();
private final int capacity;
public SimpleBlockingQueue(int capacity) {
this.capacity = capacity;
}
public synchronized void put(T item) throws InterruptedException {
while (queue.size() == capacity) {
wait(); // full, wait for a consumer to make room
}
queue.add(item);
notifyAll(); // wake any consumer waiting on take()
}
public synchronized T take() throws InterruptedException {
while (queue.isEmpty()) {
wait(); // empty, wait for a producer to add something
}
T item = queue.poll();
notifyAll(); // wake any producer waiting on put()
return item;
}
}
// Main.java
public class Main {
public static void main(String[] args) {
SimpleBlockingQueue<Integer> queue = new SimpleBlockingQueue<>(5);
Thread producer = new Thread(() -> {
try {
for (int i = 0; i < 20; i++) {
queue.put(i);
System.out.println("produced " + i);
}
} catch (InterruptedException e) {
Thread.currentThread().interrupt();
}
});
Thread consumer = new Thread(() -> {
try {
for (int i = 0; i < 20; i++) {
int item = queue.take();
System.out.println("consumed " + item);
}
} catch (InterruptedException e) {
Thread.currentThread().interrupt();
}
});
producer.start();
consumer.start();
}
}
Both put() and take() wait in a loop, not an if, for exactly the spurious-wakeup reason covered earlier. notifyAll() rather than notify() matters here specifically because both producers and consumers might be waiting on the same object at once, notify() would only wake one arbitrary thread, which might not even be the kind of thread that can currently make progress; notifyAll() wakes everyone and lets each one recheck its own condition. This is, in miniature, exactly what ArrayBlockingQueue does internally, just with more tuning and edge-case handling than is worth reproducing by hand in real code.
Coordination utilities: CountDownLatch, CyclicBarrier, Semaphore
These three exist for a specific, recurring shape of problem: several threads need to coordinate their timing with each other, not just their access to shared data.
CountDownLatch lets one or more threads wait until a fixed number of events have happened. It's initialized with a count, other threads call countDown() as they finish their part, and any thread that called await() unblocks the moment the count hits zero. It cannot be reset or reused once it reaches zero, it's a strictly one-time gate.
// Main.java
import java.util.concurrent.CountDownLatch;
public class Main {
public static void main(String[] args) throws InterruptedException {
CountDownLatch latch = new CountDownLatch(3);
for (int i = 0; i < 3; i++) {
int workerId = i;
new Thread(() -> {
System.out.println("worker " + workerId + " starting work");
doWork();
System.out.println("worker " + workerId + " finished");
latch.countDown();
}).start();
}
latch.await(); // blocks until all 3 threads have called countDown()
System.out.println("all workers finished, main can continue");
}
private static void doWork() {
try {
Thread.sleep(500);
} catch (InterruptedException e) {
Thread.currentThread().interrupt();
}
}
}
CyclicBarrier looks similar but solves a different shape of problem: instead of one thread waiting for several others, a group of threads all wait for each other to reach the same point before any of them proceeds, and once released, the barrier resets and can be used again for the next round.
// Main.java
import java.util.concurrent.CyclicBarrier;
public class Main {
public static void main(String[] args) {
CyclicBarrier barrier = new CyclicBarrier(3,
() -> System.out.println("all 3 threads reached the barrier, proceeding together")
);
for (int i = 0; i < 3; i++) {
int workerId = i;
new Thread(() -> {
doWorkPhaseOne(workerId);
try {
barrier.await();
} catch (Exception e) {
Thread.currentThread().interrupt();
}
doWorkPhaseTwo(workerId); // only starts once all 3 threads reached the barrier
}).start();
}
}
private static void doWorkPhaseOne(int workerId) {
System.out.println("worker " + workerId + " doing phase one");
}
private static void doWorkPhaseTwo(int workerId) {
System.out.println("worker " + workerId + " doing phase two");
}
}
The concrete difference worth taking away: CountDownLatch is for one thread (or several) waiting on a fixed number of other threads to finish, used once; CyclicBarrier is for a fixed group of threads waiting on each other to all reach the same checkpoint together, and it's reusable round after round, useful for something like a simulation that proceeds in synchronized phases.
Semaphore controls access to a resource with a limited number of permits, rather than the single all-or-nothing lock that synchronized and ReentrantLock give you. acquire() takes a permit, blocking if none are available; release() gives one back. A semaphore initialized with 3 permits allows exactly three threads through at once, a straightforward way to cap concurrent access to something like a pool of database connections or a rate-limited external API.
// Main.java
import java.util.concurrent.Semaphore;
public class Main {
public static void main(String[] args) {
Semaphore semaphore = new Semaphore(3); // at most 3 concurrent
for (int i = 0; i < 10; i++) {
int callId = i;
new Thread(() -> {
try {
semaphore.acquire();
System.out.println("call " + callId + " got a permit");
callLimitedResource();
} catch (InterruptedException e) {
Thread.currentThread().interrupt();
} finally {
semaphore.release();
}
}).start();
}
}
private static void callLimitedResource() {
try {
Thread.sleep(500);
} catch (InterruptedException e) {
Thread.currentThread().interrupt();
}
}
}
Virtual threads: what changed in Java 21
This is worth understanding properly, because virtual threads became a standard, non-preview feature in Java 21 (they'd been previewed in 19 and 20 first), and they've genuinely changed how high-throughput Java servers get built.
A traditional web server handles each incoming request on its own platform thread, and each platform thread is a real OS thread, carrying real memory overhead and real scheduling cost. Most of a typical request's time isn't spent computing, it's spent waiting: on a database query, on a downstream API call, on disk I/O. A platform thread sits there fully allocated, doing nothing, for the entire wait. Scale that to tens of thousands of concurrent requests and you run out of threads, and the machine, long before you run out of actual CPU work to do.
A virtual thread is a thread implemented by the JVM itself rather than mapped one-to-one to an OS thread. Many thousands of virtual threads run on top of a small pool of ordinary platform threads, called carrier threads. When a virtual thread blocks on I/O, the JVM unmounts it from its carrier thread entirely, freeing that carrier to run a different virtual thread, and remounts the original one onto some carrier once its I/O actually completes. The blocking, from the code's point of view, looks exactly like ordinary blocking code always has, Thread.sleep(), a blocking database call, a synchronous HTTP client, none of it needs to be rewritten in a reactive or callback style. That's the actual headline: you get the throughput characteristics of async code while writing plain, ordinary, sequential-looking blocking code.
// Main.java
import java.util.concurrent.ExecutorService;
import java.util.concurrent.Executors;
public class Main {
public static void main(String[] args) {
try (ExecutorService executor = Executors.newVirtualThreadPerTaskExecutor()) {
for (int i = 0; i < 10_000; i++) {
int taskId = i;
executor.submit(() -> {
makeSlowNetworkCall(); // blocks this virtual thread, not a whole OS thread
return null;
});
}
} // executor.close() is called automatically here, waiting for all tasks to finish
}
private static void makeSlowNetworkCall() {
try {
Thread.sleep(100); // stand-in for a real I/O-bound call
} catch (InterruptedException e) {
Thread.currentThread().interrupt();
}
}
}
Ten thousand platform threads would very plausibly crash a typical server. Ten thousand virtual threads, each cheap to create and cheap to sit idle, generally won't, because the JVM never had to allocate ten thousand OS threads' worth of stack memory and scheduling overhead to begin with.
The one caveat worth holding onto clearly: virtual threads help I/O-bound workloads, where threads spend most of their time waiting. They do essentially nothing for CPU-bound work, tight loops doing real computation with no blocking at all, because in that case there's no idle waiting time to reclaim in the first place; the bottleneck is genuinely the CPU, and virtual threads don't create more CPU cores. It's also worth knowing that two closely related Loom features, structured concurrency and scoped values, are still preview features even as of Java 25, not yet stable APIs, so it's fine to know what they're for in one sentence each (structured concurrency treats a group of related subtasks as a single unit that succeeds or fails together; scoped values are a safer, immutable alternative to ThreadLocal for sharing context across threads) without needing to write their exact preview syntax from memory.
Producer-consumer, done properly
Here's producer-consumer done once, cleanly, with the real BlockingQueue rather than the hand-rolled SimpleBlockingQueue from before.
// Main.java
import java.util.concurrent.ArrayBlockingQueue;
import java.util.concurrent.BlockingQueue;
public class Main {
public static void main(String[] args) {
BlockingQueue<Integer> queue = new ArrayBlockingQueue<>(10);
Thread producer = new Thread(() -> {
try {
for (int i = 0; i < 100; i++) {
queue.put(i); // blocks automatically if the queue is full
System.out.println("produced " + i);
}
} catch (InterruptedException e) {
Thread.currentThread().interrupt();
}
});
Thread consumer = new Thread(() -> {
try {
while (true) {
Integer item = queue.take(); // blocks automatically if the queue is empty
System.out.println("consumed " + item);
}
} catch (InterruptedException e) {
Thread.currentThread().interrupt();
}
});
producer.start();
consumer.start();
}
}
Notice how much of the earlier complexity has simply disappeared. No synchronized, no manual wait()/notify(), no explicit condition-checking loop, BlockingQueue has absorbed all of it internally, using essentially the same wait-loop-and-notify mechanics behind SimpleBlockingQueue. That's the real lesson here: the primitives (synchronized, wait, notify) are what you need to understand deeply, so you can reason about what's happening underneath and debug it when it breaks, but the higher-level tools (BlockingQueue, ExecutorService, CompletableFuture) are almost always what you should actually reach for in real code, since they've already handled the edge cases you'd otherwise get wrong.
Reasoning through a stuck or misbehaving program
There's one more skill worth building: being able to look at a concurrent program that's misbehaving and reason through why, methodically, rather than guessing. Here's a way to think about it that covers most real situations.
A program appears to hang. The first, fastest diagnostic is a thread dump, triggerable with jstack <pid> from the command line, or by sending SIGQUIT to the process, or by capturing one straight from your IDE's debugger. A thread dump lists every thread in the JVM along with its current state and, critically, its full stack trace at that exact instant. Scan specifically for threads in the BLOCKED state, that tells you they're waiting on a lock, and the dump names exactly which lock and which thread is currently holding it. If thread A is BLOCKED waiting for a lock thread B holds, and thread B's stack shows it BLOCKED waiting for a lock thread A holds, you've found your deadlock directly in the dump, no guessing required. Modern JVMs actually detect this specific cycle for you and print "Found one Java-level deadlock" right at the top of the dump, naming the exact threads and locks involved.
A program produces wrong results, but doesn't hang. This is the race-condition case, and it's harder precisely because nothing crashes and nothing blocks; the program simply computes something subtly incorrect, and often only under real concurrent load, never in a quick single-threaded test. The systematic approach is to look for exactly the pattern from the counter example: any place where a shared, mutable variable is read, modified, and written back without a lock, an atomic type, or a volatile covering the whole read-modify-write sequence. count++, list.add() on a plain ArrayList, if (map.get(k) == null) map.put(k, v), that check-then-act pattern in particular is a very common, very real source of race conditions, since the check and the act aren't atomic together even if each individually looks safe.
A program is slow, but not stuck. Here a thread dump is still the first move, but you're looking for something different: threads sitting RUNNABLE for a long stretch doing real work, one very large synchronized block that most other threads are BLOCKED waiting on, or, in older code, contention around a single coarse lock that could be split into a ReadWriteLock, a ConcurrentHashMap, or several finer-grained locks instead of one lock guarding far more than it needs to.
The consistent thread through all three: pin down the symptom precisely first, hung completely, wrong output, or just slow, because each symptom points toward a genuinely different class of cause, and starting from that distinction is almost always faster than diving straight into a specific fix and hoping.
Further reading
- Oracle's official Java Concurrency tutorial
- JEP 444: Virtual Threads
- Java Language Specification, Chapter 17: Threads and Locks
- Brian Goetz et al., Java Concurrency in Practice, dated in places, still the deepest treatment of the memory model and locking available
- Inside Java's ongoing coverage of Project Loom











Top comments (0)