Table of Contents
- 0. What Does This Article Answer Next?
- 1. What Guarantees Does Java volatile Provide at the Language Level?
- 2. How Does HotSpot Implement volatile?
- 3. Why Don't Go and CPython Have Java-style volatile?
- 4. How volatile, Atomics, and Mutexes Differ
- 5. Next: From Mutexes to Read-Write Locks
0. What Does This Article Answer Next?
The previous two articles traced atomic operations from language-level semantics down to the CPU, showing how Atomicity, Visibility, and Ordering are implemented.
This article follows the same path for volatile.
volatile does not make a compound update such as counter++ atomic. It is better suited to publication and observation:
Thread A Thread B
counter = 1 if (ready) {
ready = true print(counter)
}
The required guarantee is:
If Thread B observes
ready == true, a subsequent read ofcountermust observe1.
We first establish this guarantee in the Java Memory Model, then trace how HotSpot preserves it down to the CPU.
1. What Guarantees Does Java volatile Provide at the Language Level?
1.1 Visibility and Ordering
Declare ready as volatile:
int counter = 0;
volatile boolean ready = false;
// Thread A
counter = 1;
ready = true;
// Thread B
if (ready) {
System.out.println(counter);
}
counter remains an ordinary field.
JLS §17.4.4 states:
“A write to a volatile variable v synchronizes-with all subsequent reads of v by any thread.”
So if Thread B's volatile read observes the value true written by Thread A, the relationship is:
Thread A Thread B
counter = 1
│
│ program order
▼
volatile ready = true ── synchronizes-with ──► read volatile ready == true
│
│ program order
▼
read counter
By transitivity of happens-before:
write counter
↓ happens-before
write volatile ready
↓ synchronizes-with
read volatile ready
↓ happens-before
read counter
Once B observes ready == true, the earlier counter = 1 write must be visible to B. The following outcome is therefore forbidden:
ready == true
counter == 0
That gives us two guarantees:
- Visibility: writes completed by A before the volatile write become visible to B after B observes that write;
- Ordering: the operations must be observed in a way that is consistent with the happens-before relationship.
1.2 volatile Does Not Make Compound Operations Atomic
Consider:
volatile int counter = 0;
counter++;
counter++ is still a read-modify-write operation:
volatile read
↓
add
↓
volatile write
volatile constrains the visibility and ordering of the individual read and write, but it does not make the entire read → add → write sequence indivisible.
With multiple threads incrementing concurrently, updates can still be lost.
If you need an atomic RMW, use AtomicInteger, CAS, or a mutex.
2. How Does HotSpot Implement volatile?
2.1 Layers
As in the Mutex and Atomic articles, first establish the implementation stack:
Java Source Code
Example: volatile boolean ready
Role: declare a field with volatile memory semantics
│
▼
JVM Bytecode
Example: getfield / putfield (or getstatic / putstatic)
Field metadata: ACC_VOLATILE
Role: perform ordinary field access; volatility is carried by field metadata
│
▼
JVM Implementation (HotSpot)
Example: field->is_volatile() / MO_SEQ_CST / MemBar
Role: preserve and implement volatile memory-ordering semantics
│
▼
x86-64 Hardware
Example: Load / Store / Cache Coherence / Fence
Role: provide the hardware basis for Visibility and Ordering
This differs from both synchronized and AtomicInteger:
synchronized
→ dedicated monitorenter / monitorexit bytecodes
AtomicInteger
→ ordinary invokevirtual
→ no Atomic-specific bytecode
volatile
→ still uses getfield / putfield
→ no dedicated volatile-load / volatile-store bytecode
→ the volatile property is recorded in ACC_VOLATILE field metadata
2.2 End-to-End Implementation Path for a volatile Write and Read
Continue with the same counter / ready example:
The key point is not a particular machine instruction. HotSpot must recognize that ready is volatile and preserve its memory semantics all the way to the target architecture.
2.3 Guarantees at the Runtime / Language Implementation Layer
2.3.1 Visibility
HotSpot distinguishes ordinary fields from volatile fields while parsing field accesses. In the current C2 path, volatile accesses carry memory-order information such as MO_SEQ_CST, while ordinary fields use unordered access semantics.
Conceptually:
ordinary field
→ ordinary load / store
volatile field
→ volatile load / store
→ preserve cross-thread observable memory semantics
HotSpot therefore cannot optimize a volatile access as if it were an ordinary field access—for example, by eliminating, merging, or moving it across synchronization boundaries that matter.
2.3.2 Ordering
HotSpot must also constrain reordering around volatile accesses.
In C2's IR, the typical structure includes:
volatile store
← ordering constraints such as MemBarRelease
volatile load
→ ordering constraints such as MemBarAcquire
Cases that require stronger ordering—such as volatile store → volatile load—also require the corresponding StoreLoad / volatile-barrier constraint.
These are IR nodes and compiler-ordering constraints, not a one-to-one promise that every barrier becomes a separate CPU fence instruction.
Relevant OpenJDK sources include parse3.cpp, memnode.hpp, and templateTable_x86.cpp.
2.4 Guarantees at the Hardware Layer
At the hardware layer, these guarantees reduce to the same two capabilities introduced earlier in the series:
Visibility
→ Cache Coherence
→ other CPU cores cannot keep using an invalidated old cache line indefinitely
Ordering
→ CPU Memory Ordering + Fence / equivalent ordering constraints
→ preserve the memory-access order required at volatile boundaries
x86-64 already provides relatively strong memory ordering, so many Acquire / Release constraints do not require a standalone hardware fence instruction.
When a StoreLoad fence is required, HotSpot's Linux x86 path can use:
lock; addl $0, 0(%rsp)
to provide a full-fence effect.
See OpenJDK orderAccess_linux_x86.hpp.
3. Why Don't Go and CPython Have Java-style volatile?
Go and Python do not expose Java's field-level volatile mechanism. The reasons are different, so consider them separately.
3.1 Go: A Language Design Choice
Effective Go summarizes one of Go's central concurrency design principles with a well-known rule:
“Do not communicate by sharing memory; instead, share memory by communicating.”
That reflects a language-design choice: synchronization should be expressed through explicit concurrency primitives rather than hidden in field modifiers.
Accordingly, Go does not provide a Java-style field-level volatile modifier. It expresses the same kinds of synchronization through channels, mutexes, and atomic operations.
For the use case closest to Java volatile, Go uses sync/atomic. The Go Memory Model states:
“This definition provides the same semantics as C++'s sequentially consistent atomics and Java's volatile variables.”
Use the same counter / ready example:
var counter int
var ready atomic.Bool
// Goroutine A
counter = 1
ready.Store(true)
// Goroutine B
if ready.Load() {
fmt.Println(counter)
}
If ready.Load() observes the value written by ready.Store(true), the two atomic operations establish a synchronized-before relationship:
Goroutine A Goroutine B
counter = 1
│
│ sequenced-before
▼
ready.Store(true) ── synchronized-before ──► ready.Load() == true
│
│ sequenced-before
▼
read counter
So once B observes ready == true, the earlier counter = 1 write must be visible to B.
3.2 CPython: Synchronization Remains a Runtime Responsibility
Python does not expose Java-style volatile. The key reason is not syntax; synchronization has historically been a responsibility of the CPython runtime.
There are two main points.
1. Traditional CPython uses the GIL to cover many Visibility / Ordering concerns
The CPython C API documentation states:
“only a thread that holds the GIL may operate on Python objects or invoke Python’s C API.”
In traditional CPython, the GIL serializes many operations on Python objects.
That serialization also provides many of the practical Visibility and Ordering effects that Java code might otherwise obtain from volatile.
The abstractions are still different:
Java volatile
→ field-level language semantics
CPython GIL
→ runtime-level synchronization mechanism
2. Free-threaded CPython still uses runtime locks and atomics
Making the GIL optional did not introduce field-level volatile semantics.
PEP 703 instead says:
“This PEP proposes using per-object locks...”
The runtime also uses atomic operations internally.
So the implementation shifts from:
GIL
to:
Per-object lock + atomic
Synchronization therefore remains a runtime responsibility rather than becoming part of Python's field-level language semantics.
At the application level, publication and observation are expressed through synchronization objects. The documentation for threading.Event says:
“one thread signals an event and other threads wait for it.”
Use the same counter / ready example:
counter = 0
ready = threading.Event()
# Thread A
counter = 1
ready.set()
# Thread B
ready.wait()
print(counter)
Here, Event plays the synchronization role that ready played in the Java example:
Thread A Thread B
counter = 1
│
▼
ready.set() ───────────────► ready.wait() returns
│
▼
read counter
So Python application code uses Event / Lock for synchronization, while lower-level atomics and memory ordering remain implementation details of the CPython runtime.
4. How volatile, Atomics, and Mutexes Differ
The three abstractions operate at different granularities:
volatile |
Atomic | Mutex | |
|---|---|---|---|
| Core problem | publish / observe shared state | one shared-state operation | a critical section |
| Atomicity | does not make compound RMW atomic | one atomic operation is indivisible | critical section is mutually exclusive |
| Visibility | Yes | Yes | Yes |
| Ordering | Yes | Yes | Yes |
| Typical use | initialization / configuration / stop flags; state publication | counters / Add / CAS / Swap | multi-step logic / multi-field updates |
A compact way to view the relationship is:
volatile
→ establish Visibility / Ordering between a publication and a later observation
Atomic
→ make one shared-state operation indivisible
→ also provide the corresponding Visibility / Ordering
Mutex
→ use lower-level primitives such as atomics to arbitrate lock state
→ then protect an entire critical section
These are not simply stronger and weaker versions of the same tool. They solve different problems at different granularities.
5. Next: From Mutexes to Read-Write Locks
A Mutex allows only one execution unit to enter a critical section at a time.
If most operations only read shared state, making readers exclude one another is unnecessary.
The next article looks at read-write locks: multiple readers can proceed concurrently, while writers still retain exclusive access.
This article was first published on ThinkerQAQ's personal blog and syndicated here by the author. The original article may be revised over time; please refer to the personal blog for the latest version.

Top comments (0)