Table of Contents
- 0. What Does This Article Continue to Answer?
- 1. How Does HotSpot Implement AtomicInteger?
- 2. How Does Go Implement sync/atomic?
- 3. How Does CPython Implement Internal Atomics?
- 4. The Common Pattern Behind All Three Implementations
- 5. Next: volatile
0. What Does This Article Continue to Answer?
The previous article described the atomicity, visibility, and ordering guarantees exposed by Atomic APIs at the language level. This article follows those guarantees down through the compiler and runtime to the CPU.
We continue to use the same two examples:
counter++
and:
counter = 1
ready = true
The implementation path can be summarized as:
Language API and memory model
↓
Compiler / Runtime
↓
CPU Atomic Instruction / Memory Ordering
1. How Does HotSpot Implement AtomicInteger?
1.1 Layers
Java API
Example: AtomicInteger.incrementAndGet() / compareAndSet() / get()
Role: declare single-variable Atomic operations
│
▼
JDK Implementation
Example: AtomicInteger / Unsafe / VarHandle Memory Effects
Role: express Atomic RMW, CAS, and memory semantics
│
▼
JVM Implementation (HotSpot)
Example: Intrinsic / Atomic Node
Role: map Atomic operations to the target architecture
│
▼
x86-64 Hardware
Example: LOCK XADD / LOCK CMPXCHG / Load
Role: provide the hardware basis for Atomicity, Visibility, and Ordering
1.2 A Complete AtomicInteger Implementation Path
Common Atomic operations can be expanded as:
1.3 Atomicity
On x86-64, an atomic addition can be simplified as:
mov eax, 1
lock xadd dword ptr [counter], eax
XADD reads the old value and writes old + 1; the LOCK prefix makes the whole RMW indivisible with respect to competing cores.
CAS can be simplified as:
; EAX contains expected
; ECX contains newValue
lock cmpxchg dword ptr [counter], ecx
sete al
These snippets illustrate equivalent machine-level operations; they are not a claim about the exact full instruction sequence emitted by every JVM version.
For counter + ready, HotSpot must also constrain compile-time reordering according to the JMM and the memory effects of the specific Atomic method, then map those constraints onto x86-64 memory ordering. Looking only at a LOCK instruction and ignoring the compiler layer would miss part of the implementation.
1.4 Visibility and Ordering
Java Atomic methods carry explicit memory effects in addition to indivisible RMW. incrementAndGet() / getAndAdd() use the volatile read/write effects of VarHandle.getAndAdd, while get() has volatile-read semantics.
HotSpot must preserve those compiler-ordering constraints before x86-64 memory ordering, cache coherence, and locked RMW provide the hardware behavior.
Atomicity
→ Atomic Instruction
→ HotSpot / x86-64: LOCK XADD / LOCK CMPXCHG
Visibility
→ Cache Coherence
→ Typical x86-64 CPU: MESI-family (for example MESIF / MOESI)
Ordering
→ Compiler Ordering + Hardware Memory Ordering
→ HotSpot Memory Effects + x86-64 Locked RMW / Load rules
LOCK XADD / LOCK CMPXCHG are typical x86-64 mappings, not a promise that every JVM version or architecture emits the same instruction sequence.
See OpenJDK AtomicInteger.java, Unsafe.java, and HotSpot atomicAccess_linux_x86.hpp.
2. How Does Go Implement sync/atomic?
2.1 Layers
Go API
Example: atomic.Int64.Add() / CompareAndSwap() / Load()
Role: declare single-variable Atomic operations
│
▼
Go Implementation (sync/atomic)
Example: atomic.Int64 / AddInt64 / CompareAndSwapInt64
Role: express Add, CAS, Load, and related Atomic semantics
│
▼
Go Runtime Atomic
Example: internal/runtime/atomic
Role: implement target-architecture Atomic primitives
│
▼
x86-64 Hardware
Example: LOCK XADDQ / LOCK CMPXCHGQ / Load
Role: provide the hardware basis for Atomicity, Visibility, and Ordering
2.2 A Complete atomic.Int64 Implementation Path
Common Atomic operations can be expanded as:
2.3 Atomicity
On amd64, atomic addition can reach an XADD with the LOCK prefix:
LOCK
XADDQ AX, 0(BX)
CAS can reach:
LOCK
CMPXCHGQ CX, 0(BX)
So counter.Add(1) can use a hardware atomic-add operation directly rather than first being rewritten as a CAS loop.
For counter + ready, the compiler must also preserve the ordering required by the Go Memory Model. The amd64 implementation uses strongly ordered atomic instructions; other architectures can use different instruction and barrier combinations while preserving the same observable semantics.
2.4 Visibility and Ordering
The Go Memory Model defines synchronized-before for Atomic operations and requires all Atomic operations to behave as though they execute in some sequentially consistent order. The compiler and Runtime must preserve that observable behavior.
In the current non-race sync/atomic path, primitives enter internal/runtime/atomic; on amd64, Add and CAS then use locked RMW instructions.
Atomicity
→ Atomic Instruction
→ Go / amd64: LOCK XADDQ / LOCK CMPXCHGQ
Visibility
→ Cache Coherence
→ Typical x86-64 CPU: MESI-family (for example MESIF / MOESI)
Ordering
→ Compiler Ordering + Hardware Memory Ordering
→ Go Atomic semantics + amd64 Locked RMW / Load rules
See sync/atomic, internal/runtime/atomic, and atomic_amd64.s.
3. How Does CPython Implement Internal Atomics?
This section discusses only the Atomics used inside the current CPython Runtime. It does not turn them into a public application-level Python API.
3.1 Layers
CPython Runtime
Example: Reference Count / Runtime State / Lock State
Role: create internal shared-state update requirements
│
▼
CPython Atomic Abstraction
Example: _Py_atomic_add_* / _Py_atomic_compare_exchange_* / _Py_atomic_load_*
Role: express Atomic operations and Memory Order
│
▼
C Compiler Atomic Builtin
Example: GCC / Clang __atomic_*
Role: lower Atomic semantics to the target architecture
│
▼
x86-64 Hardware
Example: LOCK XADD / LOCK CMPXCHG / Load
Role: provide the hardware basis for Atomicity, Visibility, and Ordering
3.2 A Complete CPython Atomic Implementation Path
Common internal Atomic operations can be expanded as:
3.3 Atomicity
On Linux x86-64 with GCC or Clang, an atomic addition can typically compile to LOCK XADD, while CAS can typically compile to LOCK CMPXCHG:
_Py_atomic_add_*
↓
__atomic_fetch_add
↓
LOCK XADD
_Py_atomic_compare_exchange_*
↓
__atomic_compare_exchange_n
↓
LOCK CMPXCHG
The exact instruction depends on compiler version, target architecture, operand width, and requested memory order, so these are representative paths rather than universal instruction listings.
3.4 Visibility and Ordering
CPython's _Py_atomic_* layer supports more than one memory order. Current implementations include default sequentially consistent operations as well as variants such as _relaxed, _acquire, and _release.
Visibility and Ordering therefore depend on the Memory Order selected at the call site:
CPython Runtime
→ choose _Py_atomic_* Memory Order
→ GCC / Clang __atomic_* order
→ x86-64 Load / Store / Locked RMW and architecture ordering
Mapped back to the hardware model:
Atomicity
→ Atomic Instruction
→ x86-64: LOCK XADD / LOCK CMPXCHG
Visibility
→ Cache Coherence
→ Typical x86-64 CPU: MESI-family (for example MESIF / MOESI)
Ordering
→ Compiler Atomic Memory Order + Hardware Memory Ordering
→ __atomic_* order + x86-64 ordering rules
This is an implementation contract of the CPython Runtime, not a Python-language Memory Model.
See pyatomic.h, pyatomic_gcc.h, and a real use site in Python/lock.c.
4. The Common Pattern Behind All Three Implementations
Java, Go, and CPython expose different APIs, but their Atomic implementations converge on the same three problems:
1. How can one update become indivisible?
2. If a conditional update fails, who decides whether to retry?
3. How do ordinary reads and writes around the atomic operation gain the required visibility and ordering?
4.1 How Does a Single-variable Update Become Indivisible? — Atomic RMW
An ordinary read → modify → write contains several interleavable steps. Atomic Add, Exchange, CAS, and other RMW operations hand that shared-state update to a CPU Atomic Instruction:
Language / Runtime Atomic
↓
Compiler / Intrinsic
↓
Atomic RMW Instruction
↓
Cache Coherence
On x86-64, typical examples are LOCK XADD and LOCK CMPXCHG.
4.2 What Happens When a Conditional Update Fails? — CAS / Retry
CAS performs one atomic "compare and conditionally write":
expected still matches
→ update succeeds
expected is stale
→ update fails
CAS itself does not promise unlimited retries. compareAndSet() / CompareAndSwap() may simply return failure. updateAndGet(), some Runtime algorithms, or caller-written CAS loops are the layers that decide to reload, recompute, and retry.
CAS
= one atomic conditional update
CAS Loop
= CAS + algorithm-level retry after failure
4.3 Why Can Atomics Also Establish Visibility and Ordering?
Indivisible RMW solves only Atomicity.
Java VarHandle Memory Effects, the Go Memory Model's Atomic rules, and the Memory Order chosen by CPython _Py_atomic_* also constrain:
Compiler Reordering
+
CPU Memory Ordering
+
Cache Coherence
So the full path remains:
Language / Runtime defines Memory Semantics
↓
Compiler preserves ordering constraints
↓
CPU uses Atomic Instruction / Load / Store / Fence mechanisms
↓
Cache Coherence propagates shared state
↓
Program receives the corresponding Visibility and Ordering
This is also where Atomics and the previous Mutex article converge at the bottom: both depend on Atomic Instructions, Cache Coherence, and Memory Ordering. Mutexes additionally need waiting and wakeup paths under serious contention; the Atomic operation itself has no Park / Wakeup flow.
5. Next: volatile
Atomic operations address the atomicity of RMW updates such as counter++. If no atomic update is needed and the goal is only to use ready to publish an already-written counter, Java provides volatile. The next article explains how it provides visibility and ordering, while also explaining why Go and Python do not have a corresponding volatile keyword.
This article was first published on ThinkerQAQ's personal blog and syndicated here by the author. The original article may be revised over time; please refer to the personal blog for the latest version.






Top comments (0)