It might be beneficial while not being optimal on its own.
The obvious example is if it has different behaviour around local minima, it could be an altenate pathway out.
I have often wondered if doing training with radically different aproaches for the first few iterarions would avoid any method specific artifacts before the weights had time to denoise.
~85% accuracy on MNIST. Sigh.
How does it do on CIFAR-10, or even better, ImageNet?
Interesting research, not sure it's a backprop alternative.
===
EDIT: accuracy on MNIST is not ~90%. It's ~85%.
It might be beneficial while not being optimal on its own.
The obvious example is if it has different behaviour around local minima, it could be an altenate pathway out.
I have often wondered if doing training with radically different aproaches for the first few iterarions would avoid any method specific artifacts before the weights had time to denoise.
Their image classification benchmarks include both: https://pub.sakana.ai/pc-alm/assets/figures/benchmark_accura...
~74% on CIFAR-10. Still a far cry from backprop.
I didn't see ImageNet. TinyImageNet is something else.
New paper by Sakana.ai [1]
[1]: https://arxiv.org/abs/2605.31022