News Score: Score the News, Sort the News, Rewrite the Headlines

Dust: Pretraining Transformers Without Backpropagation

October 2026Correspondence to [email protected]·Code· TL;DR We present the first zeroth-order method that is competitive with backprop at pretraining transformer language models. Dust perturbs activations (node perturbation) independently at every token, so each token is a virtual population member and one forward pass evaluates them all in parallel. Dust approximates backprop closely at large population (i.e. substantially more compute) and in multiple settings even exceeds it. This hints that in a co...

Read more at qlabs.sh

© News Score  score the news, sort the news, rewrite the headlines