October 2026Correspondence to
[email protected]·Code·
TL;DR
We present the first zeroth-order method that is competitive with backprop at pretraining transformer language models. Dust perturbs activations (node perturbation) independently at every token, so each token is a virtual population member and one forward pass evaluates them all in parallel.
Dust approximates backprop closely at large population (i.e. substantially more compute) and in multiple settings even exceeds it. This hints that in a co...