Data Races & Happens-Before — UB, Tools & Mental Models
A data race is conflicting access with no happens-before edge. C++ and Rust call that undefined behavior. This page is the definition, the edges, and the tools that catch it.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
Question ladder
L1
Which two ingredients make a conflicting access?
Answer
Both threads touch the same location, and at least one of the accesses is a write.
L2
What extra ingredient makes it a data race?
Answer
There is no happens-before edge from one access to the other, and the location is not an atomic both sides use.
L3
Name three happens-before sources.
Answer
Thread start and join, unlock then a later lock of the same mutex, and a release store observed by an acquire load.
L4
Is happens-before transitive?
Answer
Yes. If A happens-before B and B happens-before C, then A happens-before C. That is how a whole critical section becomes visible.
L5
Does C++ volatile fix a data race?
Answer
No. C++ volatile is for memory-mapped I/O and for stopping some elision. Inter-thread synchronization is std::atomic or a mutex.
L6
What does ThreadSanitizer miss?
Answer
Logic races where every access is synchronized, and atomic protocols that are race-free but use the wrong memory order.
L7
How do C++, Java, and Go disagree?
Answer
C++ and Rust give undefined behavior. Java defines a memory model and still forbids racy code in practice. Go's model says programs should be race-free, and the race detector is the tool.
Failure modes
Flag seen, payload stale
The writer stores the flag before the payload is visible. The reader proceeds into old data. The edge was missing or relaxed.
Check outside the lock
The balance read is outside the mutex and the subtract is inside. There is no data race on the write, and the account still goes negative.
Sanitizer-clean and still wrong
Every access is atomic with a bad order, or every access is locked around the wrong protocol. TSan stays quiet.
32-bit torn 64-bit write
A plain 64-bit store on a 32-bit ABI is two writes. A concurrent reader observes a value nobody stored.
Misconceptions
If it did not crash, it was not a data race.
Undefined behavior includes 'it printed the right number today.' The compiler may delete a later check.
Java volatile is an atomic increment.
A volatile write happens-before a later volatile read of that variable. The increment is still a separate read and write.
The GIL is a memory model.
It is an implementation detail of one interpreter. Free-threaded Python and native extensions still need a real edge.
Interviewer traps
Using wall-clock order as happens-before.
An edge is unlock/lock, release/acquire, or thread lifetime. Earlier on a timeline is not enough.
Explaining Mesa versus Hoare while defining a data race.
Name the mutex unlock and lock edge, then point at the mutex cluster for waiting.
Design scenario
Same prompt for every reader.
Requirements
A collector that observes ready must observe the struct that was published with it. CI must fail on a data race in the tests that exist. A locked check-then-act bug must not be declared fixed just because TSan is quiet.
Traffic / scale
Thousands of short jobs per second in one process, with a handful of collector threads.
Latency
The publish path is one struct and one flag. Do not add a lock to the collector if a release and acquire flag is enough.
Consistency
No reader may observe a mix of old and new fields once ready is true.
Availability
A stuck collector must not be explained away as a data race if the bug is a missed wakeup. That is a different page.
Failure assumptions
- Tests run on x86 today.
- Production includes ARM.
- One code path still increments a shared int with +=.
Constraints
- Do not treat C++ volatile as the fix.
- Do not claim the race detector proves the business protocol.
Prompt
A worker writes a job result into a plain struct and sets a boolean ready. A collector polls ready and reads the struct. The service is correct on developer laptops and fails a soak test on ARM. CI does not run a race detector.
API
What does the writer store last, and with which order?
Data
Which fields stay non-atomic, and which word carries the edge?
Architecture
Which CI job runs TSan or -race, and which bug do you still test by hand?
Two bugs that share a nickname
Prefer
Name which bug you have
A data race is a missing edge on a location. A race condition is a bad protocol. The fixes are different, and the tools are different.
- Data race: atomic, lock, or stop sharing. Then run a race detector.
- Race condition: fix the check and the update so they are one critical section, or one transaction.
- You can have either bug without the other.
Alternative
Call every flake a race and add a sleep
A sleep changes the timing and leaves both bugs in the binary. The ARM deploy or the optimizer finds them later.
- Undefined behavior does not owe you a crash.
- A locked function can still check the balance outside the lock.
- A green x86 run does not order the store buffer.
Decide if the execution is a data race
Location, conflict, then the edge. The protocol comes second.
- 1
Find the location
One scalar, one pointer, or one field. Two different atomics are two locations. - 2
Confirm a conflict
At least one access writes. Two plain reads are not a data race. - 3
Look for an edge
Start, join, unlock then lock, or release then acquire that observes the store. If you cannot point at one, you have a data race. - 4
Then read the protocol
If every access is synchronized and the answer is still wrong, it is a race condition. The detector will not save you.
Overview
A portable definition, tight enough for an interview:
- Two threads both access location L.
- At least one access is a write.
- The accesses are not both atomic operations on the same atomic object.
- There is no happens-before from one access to the other.
If those hold in C++ or Rust, the program has undefined behavior. The compiler may assume the race does not exist and delete a later null check. Java's memory model is defined and still allows surprising values, so racy Java is not a design. Go's memory model tells you to be race-free, and the race detector is how you check.
A race condition is not that definition. Two threads can each hold the right lock and still withdraw the same dollar if the check happened outside the lock. That program can be data-race-free and still wrong.
Decisions
- 1
Two accesses, one write
- nextHappens-before edge?
- ?
Happens-before edge?
- YesSynchronized. Check the logic
- NoData race in C++ and Rust
- 3
Synchronized. Check the logic
- 4
Data race in C++ and Rust
Lesson map
Data Races & Happens-Before — UB, Tools & Mental Models
A data race is conflicting access with no happens-before edge. C++ and Rust call that undefined behavior. This page is the definition, the edges, and the tools that catch it.
Architecture. Architecture
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB a["Two accesses, one write"] b["Happens-before edge?"] c["Synchronized. Check"] d["Data race in C++ and Rust"] a -->|Two accesses, one write| b b -->|Yes| c b -->|No| d
Where happens-before comes from
- Thread lifetime. The parent happens-before the child starts. The child's last action happens-before the parent's join returns.
- Mutex. Unlock happens-before a later lock of the same mutex. That is the whole critical section, not only the lock word. The waiting rules live on mutexes.
- Atomics. A release store happens-before an acquire load that reads that value, or a later value in the modification order. Orders are the memory ordering page.
- Java volatile. A write happens-before a later read of the same volatile. That is not C++
volatile, and it is not an atomic increment. - Go. A channel send happens-before the matching receive. Mutexes and atomics follow the Go memory model.
The relation is transitive. If the child writes, then unlocks, and the parent later locks, the parent sees the child's writes. You do not need a separate edge per field inside that section.
Flow
- 1
Parent starts the child
- nextChild writes, then unlocks
- 2
Child writes, then unlocks
- nextParent locks the same mutex
- 3
Parent locks the same mutex
- nextParent sees the child writes
- 4
Parent sees the child writes
Store buffers
CPUs buffer stores and replay them later. Without a release and an acquire, another core can see flag == true and still see a stale payload. x86's total store order hides many of those mistakes and not all of them. ARM and POWER show them. Write the edge you can defend on the weakest machine you ship, not the laptop you type on.
Race condition with the lock in the wrong place
The unsafe withdraw reads the balance, then locks only the subtract. One hundred dollars and one hundred fifty withdrawals can drive the balance negative because every caller liked the same snapshot. The safe version checks and updates under one lock. This is a logic race. Making the integer atomic would not fix the check outside the update.
ExpectedThe unsafe snapshot goes negative. The safe loop stops at zero.
Press Run. Snippets must be self-contained — no network, files, or native modules.
ExpectedThe ordered publish returns 42. The reordered publish returns the stale 0.
Press Run. Snippets must be self-contained — no network, files, or native modules.
Tools
| Tool | Languages | Catches | Misses |
|---|---|---|---|
| ThreadSanitizer | C, C++, Rust | Many data races on the run you execute | Logic races. Some mis-ordered atomics |
Go -race | Go | Unsynchronized shared memory | A wrong channel protocol that is still race-free |
| jcstress, Lincheck | Java | Litmus outcomes under the Java memory model | Bugs you did not encode |
| Helgrind | C, C++ | Races and some lock-order problems | Often heavier than TSan |
| Stress plus ASan | Mixed | Crashes and some corruption | Silent lost updates |
Turn the race detector on in CI for the tests that actually share memory. A single local run is not a memory model. Model checking of a distributed protocol is a different cluster. This page stops at the process.
Language footnotes
- C++. A data race is undefined behavior. Use
std::atomic, a mutex, or do not share. - Rust. Safe code refuses data races.
unsafeplus atomics is the same discipline as C++. - Java. Prefer
java.util.concurrent.volatileis not an RMW.AtomicIntegeris. - Go. Be race-free. Channels for ownership.
sync/atomicwhen you mean a word. - Python. The GIL is not a specification. Free-threaded builds and extension modules need a real edge.
Immutability is the default, not the whole answer
| What you gain | What remains |
|---|---|
| No shared mutation, so no data race on that object | The cost of the copy, unless the structure shares nodes |
| Easy local reasoning | A race on which message wins |
| A good shape for a request handler | Metrics and caches still share a word and still need an atomic or a lock |
Interview Q&A
Is a torn 64-bit read on a 32-bit machine a data race?
Answer
If the location is non-atomic and another thread writes it, yes. Use atomic of a 64-bit integer, AtomicLong, or atomic.Int64. A torn value is one symptom. Undefined behavior is the C++ rule even when you do not observe the tear.
Does C++ volatile fix races?
Answer
No. It is for I/O registers and for keeping the compiler from deleting accesses to those registers. Inter-thread synchronization is std::atomic or a mutex. Java volatile is a different keyword with happens-before on that variable.
Can ThreadSanitizer false-positive?
Answer
On instrumented code, a report is usually a real race you did not believe. Suppressions need a written reason. A clean run only covers the interleavings you executed.
Is happens-before transitive?
Answer
Yes. That transitivity is why unlocking a mutex publishes every write in the critical section, not only the lock word.
Can a program have a race condition and no data race?
Answer
Yes. All accesses take the same lock, and the code still checks a balance before acquiring it. The detector stays quiet. The account does not.
Can a program have a data race and look correct?
Answer
Yes, especially on x86, especially under a GIL, especially before the compiler gets clever. "It printed 1 in the test" does not refute undefined behavior.
What does a Go send on a channel synchronize?
Answer
The send happens-before the corresponding receive. Values in the message are published by that edge. Memory you mutate on the side, not in the message, is still shared and can still race.
Why is a relaxed atomic flag not enough to publish a struct?
Answer
Relaxed gives atomicity of the flag word and no happens-before for the fields written before it. The reader can see the flag and a stale struct. You need a release store and an acquire load, or a lock. Orders are the next pages after CAS.
Pitfalls
- Treating "it worked on x86" as a happens-before proof.
- Using C++
volatilebecause the Java rule leaked into a C++ review. - Declaring victory when TSan is clean and the protocol is still check-then-act.
- Assuming two atomic fields are one atomic update.
- Hiding a shared write in a Python extension and blaming the GIL when it tears.
Take a writer that fills three fields and sets ready. Draw the four accesses and the one edge that publishes all three. Then move the ready store above the fields and say what a reader on ARM is allowed to see.
Go Deeper
- Jeff Preshing, the happens-before relation.
- The Go memory model and the race detector.
- ThreadSanitizer.
- Java Language Specification, chapter 17 and jcstress.
Next: CAS.