Memory Ordering — Relaxed, Acquire/Release, SeqCst
Atomicity stops a torn word. Memory order controls which other accesses may move around that word. Relaxed, acquire/release, and seq_cst are the ladder, and a publish flag is the pattern to know.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
Question ladder
L1
What does relaxed guarantee?
Answer
The word itself is atomic. You may see a stale value. No happens-before edge is created for the surrounding data.
L2
What does a release store keep from moving?
Answer
Earlier reads and writes in that thread cannot move to after the release store.
L3
What does an acquire load keep from moving?
Answer
Later reads and writes in that thread cannot move to before the acquire load.
L4
When does the pair become happens-before?
Answer
When the acquire load reads the value the release store wrote, or a later value in the modification order.
L5
When is seq_cst required?
Answer
When the protocol needs one total order of several atomic operations that all threads agree on. Independent reads of independent writes are the classic litmus.
L6
What order does a mutex use?
Answer
Unlock is a release. Lock is an acquire. Writes inside the critical section happen-before the next locker.
L7
Why prefer an ordered atomic to a standalone fence?
Answer
The order sits on the location that matters. A thread fence is easier to put in the wrong place and harder to review.
Failure modes
Relaxed flag, non-atomic payload
The reader observes the flag and still reads a stale or partial payload. Atomicity of the flag was never the edge.
Seq_cst on every metrics increment
The counter is correct and slow on ARM. A relaxed fetch-add, or a shard, was the actual requirement.
Store-buffer litmus dismissed
Both threads store their flag and read the other as zero. The author calls it impossible because x86 usually hides it.
Wrong failure order on CAS
The success path releases and the failure path still reads payload data with only relaxed, so a retried attempt observes nothing.
Misconceptions
Relaxed atomics can tear.
They cannot tear. They also cannot publish neighboring writes. Those are different guarantees.
Acquire and release are a global order.
They synchronize the threads that use that one variable. Other atomics can still disagree. Seq_cst is the total order.
JavaScript Atomics expose memory_order enums.
Atomics.load and Atomics.store synchronize that indexed cell. They are not a C++ order menu. Do not pretend a plain array write is a release.
Interviewer traps
Reaching for seq_cst because the name sounds safest, on a single flag.
Single flag is release and acquire. Name the multi-variable protocol before you pay for seq_cst.
Explaining IRIW with a queue diagram.
Keep the litmus on two locations. The queue page is next.
Design scenario
Same prompt for every reader.
Requirements
Config readers that see the new epoch must see the blob. Metrics may be slightly stale and must not fence every request. The two-flag protocol must have one order all observers agree on.
Traffic / scale
Millions of relaxed counter updates per second. Config publishes a few times a minute. The two-flag protocol is rare and must be correct.
Latency
The counter cannot take a seq_cst fence per event on ARM. Config publish can pay for a release.
Consistency
No reader may observe the new config flag and the previous blob. Observers of the two flags must agree on their order.
Availability
A slow config publish must not stall the counter. A wrong order on the two flags is a correctness bug, not a latency trade.
Failure assumptions
- Production includes ARM.
- Developers only run the litmus on x86.
- Someone will mark the counter seq_cst 'to be safe.'
Constraints
- Do not seq_cst the per-request counter.
- Do not publish the blob with a relaxed flag.
Prompt
A process publishes a read-mostly config blob to many workers and increments a metrics counter on every request. A second protocol sets two ready flags from two threads, and a third thread must not observe a mix the total order would forbid.
API
Which call is relaxed fetch-add, which is release store, and which is seq_cst?
Data
What is the flag word for the blob, and what are the two flag words that need a total order?
Architecture
Where do you run an aarch64 stress or litmus, not only an x86 unit test?
How strong an order to pay for
Prefer
Release and acquire for one flag
The payload is written first. The flag is a release store. The reader acquire-loads the flag and only then touches the payload. Relaxed stays on counters that publish nothing.
- One edge, one variable, easy to point at in review.
- Cheaper than seq_cst on ARM, and enough for this shape.
- The mutex version of the same idea is unlock and lock.
Alternative
Seq_cst everywhere, or relaxed everywhere
Seq_cst hides a design that is really one flag, and it taxes the counter. Relaxed on the flag looks atomic and is not a publish.
- A relaxed flag can be visible before the payload.
- A seq_cst increment on a hot metric shows up as fences.
- x86 will not teach you which one you shipped.
Publish a payload with one flag
The order of the four steps is the protocol. Swapping the first pair is the bug.
- 1
Write the payload
The fields can be ordinary writes. They must happen before the release in this thread. - 2
Release-store the flag
Earlier writes cannot sink below this store. This is the publish. - 3
Acquire-load the flag
Later reads cannot rise above this load. Spin or branch until you see the published value. - 4
Read the payload
Only after the acquire succeeds. A relaxed load of the flag does not license this read.
Overview
Atomicity means the word is not torn. Ordering means other memory operations may or may not move across that word. The ladder:
| Order | What it allows | Use it for |
|---|---|---|
| Relaxed | Atomicity only. No synchronization edge | A statistic that may be stale and publishes nothing |
| Acquire (load) | Later accesses cannot move before this load | The reader side of a flag or pointer |
| Release (store) | Earlier accesses cannot move after this store | The writer side, after the payload |
| Acq_rel (RMW) | Both, on one read-modify-write | A CAS that both observes and publishes |
| Seq_cst | One total order of all seq_cst operations | Several atomics that must agree. The simple mental model |
The acquire/release pair creates happens-before only when the load sees the store's value, or a later value in that word's modification order. A load that still sees the old value has not acquired that publish.
Flow
- 1
Relaxed: atomicity only
- nextAcquire and release: one flag
- 2
Acquire and release: one flag
- nextSeqCst: one total order
- 3
SeqCst: one total order
- nextName the edge in the review
- 4
Name the edge in the review
Lesson map
Memory Ordering — Relaxed, Acquire/Release, SeqCst
Atomicity stops a torn word. Memory order controls which other accesses may move around that word. Relaxed, acquire/release, and seq_cst are the ladder, and a publish flag is the pattern to know.
Architecture. Architecture
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB a["Relaxed: atomicity only"] b["Acquire and release: one flag"] c["SeqCst: one total order"] d["Name the edge in the review"] a -->|Relaxed: atomicity only| b b -->|Acquire and release: one flag| c c -->|SeqCst: one total order| d
The canonical publish
Writer: store the payload, then flag.store(1, release). Reader: if flag.load(acquire), then use the payload. The release cannot float above the payload writes. The acquire cannot float below the payload reads. That is the whole pattern for a single flag.
The wrong pattern is a relaxed store of the flag, or a read of the payload before the acquire has seen the new value. Both can compile. Only one is a publish.
Flow
- 1
1. Write the payload
- next2. Release-store the flag
- 2
2. Release-store the flag
- next3. Acquire-load the flag
- 3
3. Acquire-load the flag
- next4. Read the payload
- 4
4. Read the payload
When seq_cst is actually the point
Seq_cst adds a single total order that every thread agrees on for the seq_cst operations. Acquire and release do not promise that. Two independent writes can be seen in different orders by different readers. That litmus is independent reads of independent writes (IRIW). If your protocol is "two flags, set by different threads, and a third thread must not see a split that the total order forbids," you want seq_cst or a redesign that uses one flag.
If the protocol is one flag that publishes one payload, acquire and release are enough and cheaper on ARM.
The store buffer
Thread 1 stores x = 1 and then loads y. Thread 2 stores y = 1 and then loads x. With plain non-atomics, or with relaxed atomics and no fence, both loads can return 0. Each core can forward its own store from the store buffer and not yet see the other core's store. Seq_cst, or the right fences, closes that window. This is the interview probe for "do you believe in reorderings?"
Flow
- 1
T1 stores x then loads y
- nextT2 stores y then loads x
- 2
T2 stores y then loads x
- nextBoth loads can read 0
- 3
Both loads can read 0
- nextSeqCst or a fence closes it
- 4
SeqCst or a fence closes it
What it costs
| Order | x86 TSO | ARM and POWER |
|---|---|---|
| Relaxed | Cheap | Cheap |
| Acquire / release | Often cheap, because TSO is already strong | Real barriers or acquire loads |
| Seq_cst | Heavier stores | Noticeably more expensive |
Write for the weakest machine you ship. A green x86 test does not promote a relaxed flag into a release.
Fences versus orders on the atomic
Prefer store(release) and load(acquire) on the location you mean. The next reader of the diff can see the edge. A standalone thread fence is for the awkward protocols (Dekker-style mutual exclusion is the textbook case) and is easier to place so that it orders the wrong accesses. If you cannot say which location the fence is for, you wanted an ordered atomic.
Relaxed counter, occasional stronger publish
High-rate metrics can fetch_add(relaxed) per event. A rare seq_cst, or a mutex around a snapshot, publishes a batch to an exporter. Do not put seq_cst on every increment because the exporter needs a consistent cut once a second. The same split applies to a config blob: relaxed is the wrong order on the flag, and seq_cst is wasted if one release would publish it.
Mutexes are this pattern
Unlock is a release. Lock is an acquire. That is why the writes in a critical section become visible to the next thread that locks the same mutex, without a separate flag. You do not get that edge from a relaxed atomic that you happened to touch nearby. Waiting on a condition variable is still the mutex cluster, not this page.
Sandbox
Python cannot show a CPU reorder. The model makes the bug explicit: a reader who looks at the flag before the payload write observes 0. The fenced version writes the payload first.
ExpectedThe release order returns 7. The early flag returns 0.
Press Run. Snippets must be self-contained — no network, files, or native modules.
ExpectedThe in-order publish sees 7. The early flag sees 0.
Press Run. Snippets must be self-contained — no network, files, or native modules.
JavaScript Atomics operations synchronize the indexed cell of a SharedArrayBuffer. They do not offer a C++ order enum. A normal property write is not a release store. If you are teaching this in the browser, keep the payload in the shared buffer, not in a plain object you hope will become visible.
Orders across languages
- C++ and Rust.
memory_order_relaxed,acquire,release,acq_rel,seq_cst. Rust spells themOrdering::Relaxedand the rest. - Java.
VarHandlehas plain, opaque, acquire/release, and volatile (seq_cst-like) modes. Avolatilefield is the strong end, not an RMW. - Go. The atomic API synchronizes. Read the memory model for the version you run before you assume a C++-style relaxed knob exists. Do not invent one.
- JavaScript.
Atomics.load,store,add, andcompareExchangeare the synchronization points for that index.
On CAS, set the success order and the failure order separately when the language asks. A failed compare that will only loop can be relaxed or acquire. A failed compare that then reads payload data needs the acquire.
Interview Q&A
Can a relaxed counter tear?
Answer
No. The word is still atomic. You can observe a stale value, and you get no happens-before for data around the counter. Tearing and ordering are different bugs.
Does mutex unlock imply release?
Answer
Yes. Unlock releases and lock acquires. Writes in the critical section happen-before the next successful lock of that mutex. That is the edge, not an extra flag.
What is the store-buffer litmus?
Answer
Two threads each store to their own variable and load the other. Both loads can see zero if those stores are still in store buffers. People who only run x86 call this impossible. It is the reason the order exists.
When is acquire/release not enough?
Answer
When two or more atomic variables need one total order that all threads observe the same way. A single publish flag is not that protocol. IRIW is.
compare_exchange and the failure order?
Answer
Success might need release or acq_rel because you published. Failure often needs only relaxed or acquire, depending on whether you touch other memory before you retry. Copying the success order onto the failure path is a common over-fence. Using relaxed on a failure path that reads a payload is a bug.
Go atomics and ordering?
Answer
Go's atomic operations synchronize with the memory model, historically in a seq_cst-shaped way for the atomic itself. Check the sync/atomic docs and the memory model for your Go version instead of mapping C++ enums one-for-one.
Why not a fence on either side of a plain store?
Answer
You can, and reviewers will mis-see it. An ordered atomic names the location. A fence orders a set of accesses that is easy to break when someone moves a line. Use the fence when the architecture note says the atomic form cannot express the protocol.
What do you write in the pull request?
Answer
The happens-before edge for every non-atomic field, the order on every atomic, and whether you ran a race detector. If the only test machine is x86, say so, and say what remains unproven on ARM.
Pitfalls
- Publishing with a relaxed flag.
- Putting seq_cst on a hot counter to avoid thinking about the flag.
- Reading the payload before the acquire load.
- Testing the litmus only on x86.
- Hand-rolling double-checked locking instead of
once,call_once, orsync.Once.
Take a config pointer and a request counter. Mark each access relaxed, release, acquire, or seq_cst. Then delete any seq_cst you cannot justify with a second atomic variable. If the pointer's order is relaxed, move it to release and say what the reader was allowed to see before.
Go Deeper
- Jeff Preshing, acquire and release semantics.
- C++ memory_order and Rust Ordering.
- Java VarHandle and the Go memory model.
- LLVM atomics for the fence the CPU actually emits.
- Herb Sutter's atomic weapons talk, linked from the hub, for the hardware picture.
Next: Lock-free structures.