A common programming shortcut - taking a large uniform integer and reducing it with modulo to pick among N choices - introduces subtle bias whenever the integer range isn’t an exact multiple of N. For example, mapping 10 uniformly distributed inputs to 3 outcomes yields counts of 4, 3, 3, so the first choice occurs 40% of the time instead of 33.3%. That happens because modulo doesn’t preserve the underlying distribution; non-linear operations on random variables do not commute with expectations (E[f(X)] ≠ f(E[X])). The correct primitive for bounds is a uniform random_between(l, h), but even that is only a partial solution to the common need: selecting discrete outcomes according to explicit probabilities.
Designing APIs around explicit distributions fixes the class of bugs and makes intent clearer. Provide a random_choice that accepts either floating probabilities or, better, relative integer weights (like Python’s choices) so users express ratios directly (e.g., 4,3,3). Implementing weighted selection is simple with random_between(0, sum-1) and range mapping. In practice, this approach improved testing coverage in a deterministic fuzzer: using weighted_choice for blob sizes (e.g., 88/4/4/4) lets tests exercise both fast and slow code paths, and the team removed ad hoc random_u64() usages in favor of weighted choices. Defaulting to explicit weighted selection prevents biased sampling and yields more maintainable, testable code.
Summary generated by AI from the linked article. hn.today is not affiliated with Hacker News or Y Combinator.