Igor Ostrovsky demonstrates how microarchitectural cache behavior shapes program performance using C# microbenchmarks and visualizations. The centerpiece is a sweep that increments every Kth array element across varying array sizes and step sizes; runtimes plotted as a heatmap reveal dark bands where accesses cannot be held simultaneously in a 16‑way associative cache and a triangular region where total working set exceeds the 4MB L2 capacity. Specifics explain vertical lines from associativity conflicts (e.g., power‑of‑two steps like 256 and 512 cause many values to map to the same set), the role of a 64‑byte cache line (multiple small strides hit the same line cheaply), and why a 16‑way limit stops mattering under 4MB (16 × 262,144 = 4,194,304).
A later experiment exposes false sharing on multicore hardware: because coherence operates at cache‑line granularity, writes on nearby fields cause cross‑core invalidations and surprising timing differences. Simple increment patterns produce timings such as A++;B++;C++;D++ = 719 ms, A++;C++;E++;G++ = 448 ms, and A++;C++ = 518 ms, highlighting banking, instruction‑level parallelism, and other subtle effects. The practical takeaway is that cache size, line size, access pattern and coherence interactions dominate performance; careful measurement and informed layout often beat intuition when tuning real code.
Summary generated by AI from the linked article. hn.today is not affiliated with Hacker News or Y Combinator.