False sharing & cache-line contention
Independent variables can still share the same coherence unit.
Two threads can slow each other down while writing different variables.
Writes its own word.
Make a cache miss
A cache fetches a whole line. Locality reuses it; conflicting mappings can evict useful data even before total capacity is exhausted.
What this model includes
One read-only cache, 64-byte lines, LRU replacement, eight total lines, eight-byte elements, 32 accesses. No prefetching or multilevel effects.
What happens inside
See the unit of ownership
Coherence usually tracks cache lines, not individual source variables. If two threads repeatedly write different words in the same line, ownership can bounce between cores. This is false sharing. Actual sharing writes the same data; read-only sharing does not create the same invalidation pattern.
Change the layout deliberately
Separating frequently written per-thread data can reduce ownership traffic. Padding and alignment help when they match the line and allocation behavior, but enlarge memory use and may hurt cache locality. Aggregating local counters later can avoid a hot shared update.
What this means for your code
Low-level engineer
Confirm line contention with appropriate counters or tools. Allocation alignment and object stride both matter.
Software developer
Prefer local accumulators with a final merge for counters and reductions. Measure before padding every object.
Read the actual specifications
These references supply the underlying contracts and implementation details. The diagrams here are simplified teaching models.