DEV Community

Cover image for Concurrency Programming (6): Atomic Implementation — From Runtime to CPU
ThinkerQAQ
ThinkerQAQ

Posted on Originally published at thinkerqaq.github.io

Concurrency Programming (6): Atomic Implementation — From Runtime to CPU

Table of Contents


0. What Does This Article Continue to Answer?

The previous article described the atomicity, visibility, and ordering guarantees exposed by Atomic APIs at the language level. This article follows those guarantees down through the compiler and runtime to the CPU.

We continue to use the same two examples:

counter++
Enter fullscreen mode Exit fullscreen mode

and:

counter = 1
ready = true
Enter fullscreen mode Exit fullscreen mode

The implementation path can be summarized as:

Language API and memory model
        ↓
Compiler / Runtime
        ↓
CPU Atomic Instruction / Memory Ordering
Enter fullscreen mode Exit fullscreen mode

1. How Does HotSpot Implement AtomicInteger?

1.1 Layers

Java API
Example: AtomicInteger.incrementAndGet() / compareAndSet() / get()
Role: declare single-variable Atomic operations
        │
        ▼
JDK Implementation
Example: AtomicInteger / Unsafe / VarHandle Memory Effects
Role: express Atomic RMW, CAS, and memory semantics
        │
        ▼
JVM Implementation (HotSpot)
Example: Intrinsic / Atomic Node
Role: map Atomic operations to the target architecture
        │
        ▼
x86-64 Hardware
Example: LOCK XADD / LOCK CMPXCHG / Load
Role: provide the hardware basis for Atomicity, Visibility, and Ordering
Enter fullscreen mode Exit fullscreen mode

1.2 A Complete AtomicInteger Implementation Path

How incrementAndGet travels through AtomicInteger, Unsafe, HotSpot, and x86-64 hardware to complete one atomic RMW.

Common Atomic operations can be expanded as:

Add, CAS, and Load pass through Unsafe and HotSpot intrinsics, then map to atomic RMW, compare-and-swap, and an appropriately ordered load.

1.3 Atomicity

On x86-64, an atomic addition can be simplified as:

mov  eax, 1
lock xadd dword ptr [counter], eax
Enter fullscreen mode Exit fullscreen mode

XADD reads the old value and writes old + 1; the LOCK prefix makes the whole RMW indivisible with respect to competing cores.

CAS can be simplified as:

; EAX contains expected
; ECX contains newValue
lock cmpxchg dword ptr [counter], ecx
sete al
Enter fullscreen mode Exit fullscreen mode

These snippets illustrate equivalent machine-level operations; they are not a claim about the exact full instruction sequence emitted by every JVM version.

For counter + ready, HotSpot must also constrain compile-time reordering according to the JMM and the memory effects of the specific Atomic method, then map those constraints onto x86-64 memory ordering. Looking only at a LOCK instruction and ignoring the compiler layer would miss part of the implementation.

1.4 Visibility and Ordering

Java Atomic methods carry explicit memory effects in addition to indivisible RMW. incrementAndGet() / getAndAdd() use the volatile read/write effects of VarHandle.getAndAdd, while get() has volatile-read semantics.

HotSpot must preserve those compiler-ordering constraints before x86-64 memory ordering, cache coherence, and locked RMW provide the hardware behavior.

Atomicity
  → Atomic Instruction
  → HotSpot / x86-64: LOCK XADD / LOCK CMPXCHG

Visibility
  → Cache Coherence
  → Typical x86-64 CPU: MESI-family (for example MESIF / MOESI)

Ordering
  → Compiler Ordering + Hardware Memory Ordering
  → HotSpot Memory Effects + x86-64 Locked RMW / Load rules
Enter fullscreen mode Exit fullscreen mode

LOCK XADD / LOCK CMPXCHG are typical x86-64 mappings, not a promise that every JVM version or architecture emits the same instruction sequence.

See OpenJDK AtomicInteger.java, Unsafe.java, and HotSpot atomicAccess_linux_x86.hpp.


2. How Does Go Implement sync/atomic?

2.1 Layers

Go API
Example: atomic.Int64.Add() / CompareAndSwap() / Load()
Role: declare single-variable Atomic operations
        │
        ▼
Go Implementation (sync/atomic)
Example: atomic.Int64 / AddInt64 / CompareAndSwapInt64
Role: express Add, CAS, Load, and related Atomic semantics
        │
        ▼
Go Runtime Atomic
Example: internal/runtime/atomic
Role: implement target-architecture Atomic primitives
        │
        ▼
x86-64 Hardware
Example: LOCK XADDQ / LOCK CMPXCHGQ / Load
Role: provide the hardware basis for Atomicity, Visibility, and Ordering
Enter fullscreen mode Exit fullscreen mode

2.2 A Complete atomic.Int64 Implementation Path

How atomic.Int64.Add travels through sync/atomic, internal/runtime/atomic, and x86-64 hardware to complete one atomic RMW.

Common Atomic operations can be expanded as:

Add, CompareAndSwap, and Load pass through sync/atomic and Runtime Atomic, then map to amd64 XADD, CMPXCHG, and Load.

2.3 Atomicity

On amd64, atomic addition can reach an XADD with the LOCK prefix:

LOCK
XADDQ AX, 0(BX)
Enter fullscreen mode Exit fullscreen mode

CAS can reach:

LOCK
CMPXCHGQ CX, 0(BX)
Enter fullscreen mode Exit fullscreen mode

So counter.Add(1) can use a hardware atomic-add operation directly rather than first being rewritten as a CAS loop.

For counter + ready, the compiler must also preserve the ordering required by the Go Memory Model. The amd64 implementation uses strongly ordered atomic instructions; other architectures can use different instruction and barrier combinations while preserving the same observable semantics.

2.4 Visibility and Ordering

The Go Memory Model defines synchronized-before for Atomic operations and requires all Atomic operations to behave as though they execute in some sequentially consistent order. The compiler and Runtime must preserve that observable behavior.

In the current non-race sync/atomic path, primitives enter internal/runtime/atomic; on amd64, Add and CAS then use locked RMW instructions.

Atomicity
  → Atomic Instruction
  → Go / amd64: LOCK XADDQ / LOCK CMPXCHGQ

Visibility
  → Cache Coherence
  → Typical x86-64 CPU: MESI-family (for example MESIF / MOESI)

Ordering
  → Compiler Ordering + Hardware Memory Ordering
  → Go Atomic semantics + amd64 Locked RMW / Load rules
Enter fullscreen mode Exit fullscreen mode

See sync/atomic, internal/runtime/atomic, and atomic_amd64.s.


3. How Does CPython Implement Internal Atomics?

This section discusses only the Atomics used inside the current CPython Runtime. It does not turn them into a public application-level Python API.

3.1 Layers

CPython Runtime
Example: Reference Count / Runtime State / Lock State
Role: create internal shared-state update requirements
        │
        ▼
CPython Atomic Abstraction
Example: _Py_atomic_add_* / _Py_atomic_compare_exchange_* / _Py_atomic_load_*
Role: express Atomic operations and Memory Order
        │
        ▼
C Compiler Atomic Builtin
Example: GCC / Clang __atomic_*
Role: lower Atomic semantics to the target architecture
        │
        ▼
x86-64 Hardware
Example: LOCK XADD / LOCK CMPXCHG / Load
Role: provide the hardware basis for Atomicity, Visibility, and Ordering
Enter fullscreen mode Exit fullscreen mode

3.2 A Complete CPython Atomic Implementation Path

How an internal CPython atomic add travels through _Py_atomic, GCC or Clang Atomic Builtins, and x86-64 hardware.

Common internal Atomic operations can be expanded as:

CPython Runtime expresses Add, Compare-Exchange, and Load/Store through _Py_atomic and Memory Order, then compiler builtins map them to x86-64.

3.3 Atomicity

On Linux x86-64 with GCC or Clang, an atomic addition can typically compile to LOCK XADD, while CAS can typically compile to LOCK CMPXCHG:

_Py_atomic_add_*
        ↓
__atomic_fetch_add
        ↓
LOCK XADD
Enter fullscreen mode Exit fullscreen mode
_Py_atomic_compare_exchange_*
        ↓
__atomic_compare_exchange_n
        ↓
LOCK CMPXCHG
Enter fullscreen mode Exit fullscreen mode

The exact instruction depends on compiler version, target architecture, operand width, and requested memory order, so these are representative paths rather than universal instruction listings.

3.4 Visibility and Ordering

CPython's _Py_atomic_* layer supports more than one memory order. Current implementations include default sequentially consistent operations as well as variants such as _relaxed, _acquire, and _release.

Visibility and Ordering therefore depend on the Memory Order selected at the call site:

CPython Runtime
  → choose _Py_atomic_* Memory Order
  → GCC / Clang __atomic_* order
  → x86-64 Load / Store / Locked RMW and architecture ordering
Enter fullscreen mode Exit fullscreen mode

Mapped back to the hardware model:

Atomicity
  → Atomic Instruction
  → x86-64: LOCK XADD / LOCK CMPXCHG

Visibility
  → Cache Coherence
  → Typical x86-64 CPU: MESI-family (for example MESIF / MOESI)

Ordering
  → Compiler Atomic Memory Order + Hardware Memory Ordering
  → __atomic_* order + x86-64 ordering rules
Enter fullscreen mode Exit fullscreen mode

This is an implementation contract of the CPython Runtime, not a Python-language Memory Model.

See pyatomic.h, pyatomic_gcc.h, and a real use site in Python/lock.c.


4. The Common Pattern Behind All Three Implementations

Java, Go, and CPython expose different APIs, but their Atomic implementations converge on the same three problems:

1. How can one update become indivisible?
2. If a conditional update fails, who decides whether to retry?
3. How do ordinary reads and writes around the atomic operation gain the required visibility and ordering?
Enter fullscreen mode Exit fullscreen mode

4.1 How Does a Single-variable Update Become Indivisible? — Atomic RMW

An ordinary read → modify → write contains several interleavable steps. Atomic Add, Exchange, CAS, and other RMW operations hand that shared-state update to a CPU Atomic Instruction:

Language / Runtime Atomic
        ↓
Compiler / Intrinsic
        ↓
Atomic RMW Instruction
        ↓
Cache Coherence
Enter fullscreen mode Exit fullscreen mode

On x86-64, typical examples are LOCK XADD and LOCK CMPXCHG.

4.2 What Happens When a Conditional Update Fails? — CAS / Retry

CAS performs one atomic "compare and conditionally write":

expected still matches
  → update succeeds

expected is stale
  → update fails
Enter fullscreen mode Exit fullscreen mode

CAS itself does not promise unlimited retries. compareAndSet() / CompareAndSwap() may simply return failure. updateAndGet(), some Runtime algorithms, or caller-written CAS loops are the layers that decide to reload, recompute, and retry.

CAS
= one atomic conditional update

CAS Loop
= CAS + algorithm-level retry after failure
Enter fullscreen mode Exit fullscreen mode

4.3 Why Can Atomics Also Establish Visibility and Ordering?

Indivisible RMW solves only Atomicity.

Java VarHandle Memory Effects, the Go Memory Model's Atomic rules, and the Memory Order chosen by CPython _Py_atomic_* also constrain:

Compiler Reordering
      +
CPU Memory Ordering
      +
Cache Coherence
Enter fullscreen mode Exit fullscreen mode

So the full path remains:

Language / Runtime defines Memory Semantics
        ↓
Compiler preserves ordering constraints
        ↓
CPU uses Atomic Instruction / Load / Store / Fence mechanisms
        ↓
Cache Coherence propagates shared state
        ↓
Program receives the corresponding Visibility and Ordering
Enter fullscreen mode Exit fullscreen mode

This is also where Atomics and the previous Mutex article converge at the bottom: both depend on Atomic Instructions, Cache Coherence, and Memory Ordering. Mutexes additionally need waiting and wakeup paths under serious contention; the Atomic operation itself has no Park / Wakeup flow.


5. Next: volatile

Atomic operations address the atomicity of RMW updates such as counter++. If no atomic update is needed and the goal is only to use ready to publish an already-written counter, Java provides volatile. The next article explains how it provides visibility and ordering, while also explaining why Go and Python do not have a corresponding volatile keyword.


This article was first published on ThinkerQAQ's personal blog and syndicated here by the author. The original article may be revised over time; please refer to the personal blog for the latest version.

Top comments (0)