🧠 Someone just pretrained transformers without backprop, and it's competitive.
Dust is a zeroth-order method. Per the authors, it perturbs activations at every token, so one forward pass evaluates a whole "population" in parallel. It's 10^3 to 10^4 times more efficient than weight-space evolution strategies, and sometimes beats backprop at large population sizes.
Bigger models got more population-efficient. A 243M model beat one 120x smaller.
Brutal compute bill though. Backprop isn't sweating yet.
Post #658
647