Fernando Simões describes how enabling swap for a cgroup with a Go process produced surprising, long stop-the-world garbage collection pauses because GC metadata pages were swapped out. The test used a Go process that reads a blob (io.ReadAll) and proto.Unmarshal into a graph marked as scan, alongside a mostly-idle HTTP server in the same cgroup. Median GC pauses were ≈51 μs, but when metadata lived on slower storage the worst pause reached ~40 ms. A BPF trace showed 228 major page faults consuming ~39 ms of a 39.9 ms pause; stack traces point to runtime bookkeeping (spanSet.reset, finishsweep_m, gcStart, nextMarkBitArenaEpoch). The runtime stops the world at sweep- and mark-termination, and 312 such pauses occurred in 30 minutes under the test workload.
The root cause is that the runtime allocates and reuses metadata pages that the kernel evicts by age; when the GC touches them they cause costly swap-in work (PTE handling, do_swap_page, charging the cgroup, issuing bios and waiting for disk). Allocating a 511 KiB message usually costing 3-5 ms jumped to 105 ms on NVMe and 903 ms on a network volume, a heavy localized penalty. Green Tea GC in Go 1.26 did not materially change metadata access behavior. Full experiments, scripts and plots are published in an accompanying repository.
Summary generated by AI from the linked article. hn.today is not affiliated with Hacker News or Y Combinator.