NVIDIA’s Olympus is a server core designed to drive single‑thread performance by maximizing work per cycle rather than by high clock speeds. It is a 10‑wide out‑of‑order design running around 3.3 GHz with very large OoO structures and an SMT2 implementation that lets it approach desktop-class throughput. Olympus shares a high‑level layout with Arm’s Cortex X925 but pushes larger reordering structures, uses ahead indexing in branch prediction, and implements a 16K entry L2 BTB with 4‑cycle latency. Branch prediction accuracy lands near the best Intel/AMD cores: tracking roughly 5-6K global history patterns with a few branches and up to ~48K with many branches, and it sustains two taken branches per cycle until instruction footprints exceed roughly 48 KB. Many predictor and BTB resources are partitioned per thread under SMT.
The frontend and backend emphasize high fetch and execution bandwidth for a single thread: a 64‑entry iTLB, 64 KB 4‑way I‑cache delivering 128 B/cycle, and high L2 code throughput. However, static partitioning of nearly all structures (schedulers, queues, execution ports, L2) halves per‑thread capacity under SMT and causes sharp fetch throughput drops. Backend features include very large register files (architectural allocation limited to ~606 registers), limited move‑elimination benefit, eight ALU pairs, six‑pipe FPU with a non‑scheduling prequeue, and four load pipelines. L1D nominal latency measures 2-4 cycles (value prediction suspected); store‑to‑load forwarding hits at 3-6 cycles for aligned cases, while other overlaps incur 12-13 cycle penalties and a notable anomaly around 32 B boundaries.
Summary generated by AI from the linked article. hn.today is not affiliated with Hacker News or Y Combinator.