Atomics, barriers & memory ordering
Specify indivisible updates and the ordering needed to publish data safely.
“It worked on my CPU” is not a memory model.
Defined indivisible access.
Follow an instruction
Overlap improves throughput. Dependencies introduce bubbles unless the implementation can forward or do independent work.
What this model includes
Five-stage, single-issue teaching model. Dependent mode inserts two idle issue cycles per instruction; real forwarding and hazards vary.
What happens inside
Choose the guarantee
An atomic operation prevents a specified access from tearing and supports defined concurrency semantics. Relaxed atomics give atomicity without general publication ordering. Release/acquire can synchronize associated data when an acquire observes the relevant release under the language rules. Sequential consistency imposes a stronger ordering contract.
Respect all layers
Compiler reordering and hardware reordering are different concerns. A compiler barrier alone does not universally order hardware accesses. Device memory and DMA need platform-specific primitives. Lock-free does not mean wait-free, and a contended atomic can serialize a large number of threads.
What this means for your code
Low-level engineer
Reason with the exact language model and target architecture. Use established primitives instead of inventing lock-free protocols from intuition.
Software developer
A mutex is often the clearest correct choice. Reduce shared mutable state before weakening memory order for speed.
// C++ publication example; one producer and one consumer
std::atomic<bool> ready{false};
int payload = 0;
// producer
payload = 42;
ready.store(true, std::memory_order_release);
// consumer, after observing true
if (ready.load(std::memory_order_acquire)) use(payload);Read the actual specifications
These references supply the underlying contracts and implementation details. The diagrams here are simplified teaching models.