This examines how page table structures and their memory cost become a real scalability problem when many address spaces map the same pages. The traditional defense of multi-level tree page tables is that adjacent entries pack well into cache lines and accelerate TLB fills, unlike hashed tables that scatter entries. That latency advantage hides the fact that each address space needs its own PTEs, so mappings scale with N. With 8-byte PTEs and 4 KiB pages, 512 identical mappings require 512 PTEs, which equals one full data page. Real-world cases illustrate the risk: a kernel metric was added after a user-space NIC caused large pgtables; a 2002 x86-32 PAE deployment exhausted low memory and required moving page tables; an Oracle SGA with ~1500 clients produced hundreds of gigabytes of PTEs; and a jemalloc/tcmalloc workload leaked page tables by using MADV_DONTNEED until a fix landed in 2025.
Databases and NUMA systems expose practical consequences and trade-offs. Postgres showed page tables exploding to tens of gigabytes and was fixed by using huge pages; ClickHouse measured per-connection page-table costs (tens of MiB each) reduced dramatically with huge pages. Huge pages, however, demand defragmented memory and can fail under load or cause CPU/pathology issues. NUMA adds latency when page tables live on remote nodes; approaches diverge between replicating whole page-table trees on each node (high memory and synchronization cost) and lazy per-PTE copying with node-specific replicas to reduce remote walks.
Summary generated by AI from the linked article. hn.today is not affiliated with Hacker News or Y Combinator.