DEV Community

Cover image for Concurrency Programming (7): volatile — From Language Semantics to the CPU
ThinkerQAQ
ThinkerQAQ

Posted on Originally published at thinkerqaq.github.io

Concurrency Programming (7): volatile — From Language Semantics to the CPU

Table of Contents


0. What Does This Article Answer Next?

The previous two articles traced atomic operations from language-level semantics down to the CPU, showing how Atomicity, Visibility, and Ordering are implemented.

This article follows the same path for volatile.

volatile does not make a compound update such as counter++ atomic. It is better suited to publication and observation:

Thread A                         Thread B

counter = 1                     if (ready) {
ready = true                        print(counter)
                                }
Enter fullscreen mode Exit fullscreen mode

The required guarantee is:

If Thread B observes ready == true, a subsequent read of counter must observe 1.

We first establish this guarantee in the Java Memory Model, then trace how HotSpot preserves it down to the CPU.


1. What Guarantees Does Java volatile Provide at the Language Level?

1.1 Visibility and Ordering

Declare ready as volatile:

int counter = 0;
volatile boolean ready = false;

// Thread A
counter = 1;
ready = true;

// Thread B
if (ready) {
    System.out.println(counter);
}
Enter fullscreen mode Exit fullscreen mode

counter remains an ordinary field.

JLS §17.4.4 states:

“A write to a volatile variable v synchronizes-with all subsequent reads of v by any thread.”

So if Thread B's volatile read observes the value true written by Thread A, the relationship is:

Thread A                              Thread B

counter = 1
    │
    │ program order
    ▼
volatile ready = true ── synchronizes-with ──► read volatile ready == true
                                                   │
                                                   │ program order
                                                   ▼
                                               read counter
Enter fullscreen mode Exit fullscreen mode

By transitivity of happens-before:

write counter
    ↓ happens-before
write volatile ready
    ↓ synchronizes-with
read volatile ready
    ↓ happens-before
read counter
Enter fullscreen mode Exit fullscreen mode

Once B observes ready == true, the earlier counter = 1 write must be visible to B. The following outcome is therefore forbidden:

ready == true
counter == 0
Enter fullscreen mode Exit fullscreen mode

That gives us two guarantees:

  • Visibility: writes completed by A before the volatile write become visible to B after B observes that write;
  • Ordering: the operations must be observed in a way that is consistent with the happens-before relationship.

1.2 volatile Does Not Make Compound Operations Atomic

Consider:

volatile int counter = 0;

counter++;
Enter fullscreen mode Exit fullscreen mode

counter++ is still a read-modify-write operation:

volatile read
      ↓
     add
      ↓
volatile write
Enter fullscreen mode Exit fullscreen mode

volatile constrains the visibility and ordering of the individual read and write, but it does not make the entire read → add → write sequence indivisible.

With multiple threads incrementing concurrently, updates can still be lost.

If you need an atomic RMW, use AtomicInteger, CAS, or a mutex.


2. How Does HotSpot Implement volatile?

2.1 Layers

As in the Mutex and Atomic articles, first establish the implementation stack:

Java Source Code
Example: volatile boolean ready
Role: declare a field with volatile memory semantics
        │
        ▼
JVM Bytecode
Example: getfield / putfield (or getstatic / putstatic)
Field metadata: ACC_VOLATILE
Role: perform ordinary field access; volatility is carried by field metadata
        │
        ▼
JVM Implementation (HotSpot)
Example: field->is_volatile() / MO_SEQ_CST / MemBar
Role: preserve and implement volatile memory-ordering semantics
        │
        ▼
x86-64 Hardware
Example: Load / Store / Cache Coherence / Fence
Role: provide the hardware basis for Visibility and Ordering
Enter fullscreen mode Exit fullscreen mode

This differs from both synchronized and AtomicInteger:

synchronized
  → dedicated monitorenter / monitorexit bytecodes

AtomicInteger
  → ordinary invokevirtual
  → no Atomic-specific bytecode

volatile
  → still uses getfield / putfield
  → no dedicated volatile-load / volatile-store bytecode
  → the volatile property is recorded in ACC_VOLATILE field metadata
Enter fullscreen mode Exit fullscreen mode

2.2 End-to-End Implementation Path for a volatile Write and Read

Continue with the same counter / ready example:

How the ordinary write to counter and the volatile write/read of ready pass through JVM Bytecode and HotSpot before reaching x86-64 Hardware.

The key point is not a particular machine instruction. HotSpot must recognize that ready is volatile and preserve its memory semantics all the way to the target architecture.

2.3 Guarantees at the Runtime / Language Implementation Layer

2.3.1 Visibility

HotSpot distinguishes ordinary fields from volatile fields while parsing field accesses. In the current C2 path, volatile accesses carry memory-order information such as MO_SEQ_CST, while ordinary fields use unordered access semantics.

Conceptually:

ordinary field
  → ordinary load / store

volatile field
  → volatile load / store
  → preserve cross-thread observable memory semantics
Enter fullscreen mode Exit fullscreen mode

HotSpot therefore cannot optimize a volatile access as if it were an ordinary field access—for example, by eliminating, merging, or moving it across synchronization boundaries that matter.

2.3.2 Ordering

HotSpot must also constrain reordering around volatile accesses.

In C2's IR, the typical structure includes:

volatile store
  ← ordering constraints such as MemBarRelease

volatile load
  → ordering constraints such as MemBarAcquire
Enter fullscreen mode Exit fullscreen mode

Cases that require stronger ordering—such as volatile store → volatile load—also require the corresponding StoreLoad / volatile-barrier constraint.

These are IR nodes and compiler-ordering constraints, not a one-to-one promise that every barrier becomes a separate CPU fence instruction.

Relevant OpenJDK sources include parse3.cpp, memnode.hpp, and templateTable_x86.cpp.

2.4 Guarantees at the Hardware Layer

At the hardware layer, these guarantees reduce to the same two capabilities introduced earlier in the series:

Visibility
  → Cache Coherence
  → other CPU cores cannot keep using an invalidated old cache line indefinitely

Ordering
  → CPU Memory Ordering + Fence / equivalent ordering constraints
  → preserve the memory-access order required at volatile boundaries
Enter fullscreen mode Exit fullscreen mode

x86-64 already provides relatively strong memory ordering, so many Acquire / Release constraints do not require a standalone hardware fence instruction.

When a StoreLoad fence is required, HotSpot's Linux x86 path can use:

lock; addl $0, 0(%rsp)
Enter fullscreen mode Exit fullscreen mode

to provide a full-fence effect.

See OpenJDK orderAccess_linux_x86.hpp.


3. Why Don't Go and CPython Have Java-style volatile?

Go and Python do not expose Java's field-level volatile mechanism. The reasons are different, so consider them separately.

3.1 Go: A Language Design Choice

Effective Go summarizes one of Go's central concurrency design principles with a well-known rule:

“Do not communicate by sharing memory; instead, share memory by communicating.”

That reflects a language-design choice: synchronization should be expressed through explicit concurrency primitives rather than hidden in field modifiers.

Accordingly, Go does not provide a Java-style field-level volatile modifier. It expresses the same kinds of synchronization through channels, mutexes, and atomic operations.

For the use case closest to Java volatile, Go uses sync/atomic. The Go Memory Model states:

“This definition provides the same semantics as C++'s sequentially consistent atomics and Java's volatile variables.”

Use the same counter / ready example:

var counter int
var ready atomic.Bool

// Goroutine A
counter = 1
ready.Store(true)

// Goroutine B
if ready.Load() {
    fmt.Println(counter)
}
Enter fullscreen mode Exit fullscreen mode

If ready.Load() observes the value written by ready.Store(true), the two atomic operations establish a synchronized-before relationship:

Goroutine A                          Goroutine B

counter = 1
    │
    │ sequenced-before
    ▼
ready.Store(true) ── synchronized-before ──► ready.Load() == true
                                                 │
                                                 │ sequenced-before
                                                 ▼
                                             read counter
Enter fullscreen mode Exit fullscreen mode

So once B observes ready == true, the earlier counter = 1 write must be visible to B.

3.2 CPython: Synchronization Remains a Runtime Responsibility

Python does not expose Java-style volatile. The key reason is not syntax; synchronization has historically been a responsibility of the CPython runtime.

There are two main points.

1. Traditional CPython uses the GIL to cover many Visibility / Ordering concerns

The CPython C API documentation states:

“only a thread that holds the GIL may operate on Python objects or invoke Python’s C API.”

In traditional CPython, the GIL serializes many operations on Python objects.

That serialization also provides many of the practical Visibility and Ordering effects that Java code might otherwise obtain from volatile.

The abstractions are still different:

Java volatile
  → field-level language semantics

CPython GIL
  → runtime-level synchronization mechanism
Enter fullscreen mode Exit fullscreen mode

2. Free-threaded CPython still uses runtime locks and atomics

Making the GIL optional did not introduce field-level volatile semantics.

PEP 703 instead says:

“This PEP proposes using per-object locks...”

The runtime also uses atomic operations internally.

So the implementation shifts from:

GIL
Enter fullscreen mode Exit fullscreen mode

to:

Per-object lock + atomic
Enter fullscreen mode Exit fullscreen mode

Synchronization therefore remains a runtime responsibility rather than becoming part of Python's field-level language semantics.

At the application level, publication and observation are expressed through synchronization objects. The documentation for threading.Event says:

“one thread signals an event and other threads wait for it.”

Use the same counter / ready example:

counter = 0
ready = threading.Event()

# Thread A
counter = 1
ready.set()

# Thread B
ready.wait()
print(counter)
Enter fullscreen mode Exit fullscreen mode

Here, Event plays the synchronization role that ready played in the Java example:

Thread A                         Thread B

counter = 1
    │
    ▼
ready.set()  ───────────────►  ready.wait() returns
                                  │
                                  ▼
                              read counter
Enter fullscreen mode Exit fullscreen mode

So Python application code uses Event / Lock for synchronization, while lower-level atomics and memory ordering remain implementation details of the CPython runtime.


4. How volatile, Atomics, and Mutexes Differ

The three abstractions operate at different granularities:

volatile Atomic Mutex
Core problem publish / observe shared state one shared-state operation a critical section
Atomicity does not make compound RMW atomic one atomic operation is indivisible critical section is mutually exclusive
Visibility Yes Yes Yes
Ordering Yes Yes Yes
Typical use initialization / configuration / stop flags; state publication counters / Add / CAS / Swap multi-step logic / multi-field updates

A compact way to view the relationship is:

volatile
  → establish Visibility / Ordering between a publication and a later observation

Atomic
  → make one shared-state operation indivisible
  → also provide the corresponding Visibility / Ordering

Mutex
  → use lower-level primitives such as atomics to arbitrate lock state
  → then protect an entire critical section
Enter fullscreen mode Exit fullscreen mode

These are not simply stronger and weaker versions of the same tool. They solve different problems at different granularities.


5. Next: From Mutexes to Read-Write Locks

A Mutex allows only one execution unit to enter a critical section at a time.

If most operations only read shared state, making readers exclude one another is unnecessary.

The next article looks at read-write locks: multiple readers can proceed concurrently, while writers still retain exclusive access.


This article was first published on ThinkerQAQ's personal blog and syndicated here by the author. The original article may be revised over time; please refer to the personal blog for the latest version.

Top comments (0)