Dust introduces a zeroth-order training algorithm for transformers that perturbs activations instead of weights, creating a virtual population where each token in a sequence acts as an independent perturbation member. Gaussian noise is added to the outputs of each linear layer at every token; each token’s reward is the change in its token loss caused by that perturbation. Averaging reward-weighted perturbations gives an estimated error at layer outputs, and the outer product with the layer input yields a weight update. Attention modules receive credit via estimated errors at the attention outputs over current and future tokens, while a small set of residual mixing scalars are updated with ordinary weight-space ES. Practical measures to avoid interference and tune token-level credit assignment are part of the method.
Results show Dust is the first zeroth-order method competitive with backprop for transformer pretraining and, at large populations, can outperform backprop in several settings. Activation-space perturbation makes a single forward pass evaluate thousands of virtual members, yielding estimated efficiency gains of 10^3-10^4× over weight-space ES (extrapolated against EGGROLL from 1M tokens up). Larger models prove more population-efficient, with gradient estimates increasingly aligned with backprop as population size grows (tested up to 1B tokens). Dust does not yet aim to be universally compute-efficient enough to replace backprop, but demonstrates that search-based credit assignment can scale and sometimes exceed gradient-based training in compute-rich regimes.
Summary generated by AI from the linked article. hn.today is not affiliated with Hacker News or Y Combinator.