Even though this is way more expensive than backprop, could a hybrid approach where you fine tune an existing checkpoint that's been backpropped unlock further gains?
It would be cool to apply this to different stages and see if that affects the learning trajectory
The authors seem to be arguing that the full dimensional direction information is mostly wasted in large models and that a cobbled together random-subspace estimate is good enough.
Maybe taking this up one step to Hessian matrix estimation would improve the asymptotic speed over backprop which is already O(inference) per step, by opening the door for higher order methods.
Not necessarily, backprop is highly parallelizable since it is just a bunch of matrix mults.
Something like Dust skips the backward pass on backprop. But other techniques like
Neural Predictive Coding can be completely asynchronous, each "weight" can fire independent of those far away from it. Innocenti, et. al have shown that NPC gradients converge to backprop within a certain "regime".
The win with asynchronous techniques like NPC is that you do not need the extreme co-ordination that backprop requires and hence should be computationally much easier given the right device.
Although at this point the industry has so much money in the forward-backward pass system that I doubt a backprop successor would win unless someone makes NPC hardware feasible and can prove scaling up to billions of params
Even though this is way more expensive than backprop, could a hybrid approach where you fine tune an existing checkpoint that's been backpropped unlock further gains? It would be cool to apply this to different stages and see if that affects the learning trajectory
The authors seem to be arguing that the full dimensional direction information is mostly wasted in large models and that a cobbled together random-subspace estimate is good enough.
Maybe taking this up one step to Hessian matrix estimation would improve the asymptotic speed over backprop which is already O(inference) per step, by opening the door for higher order methods.
It sounds like this is less computationally efficient than backprop, but more easily parallelizable. Is that fair?
Not necessarily, backprop is highly parallelizable since it is just a bunch of matrix mults.
Something like Dust skips the backward pass on backprop. But other techniques like Neural Predictive Coding can be completely asynchronous, each "weight" can fire independent of those far away from it. Innocenti, et. al have shown that NPC gradients converge to backprop within a certain "regime".
The win with asynchronous techniques like NPC is that you do not need the extreme co-ordination that backprop requires and hence should be computationally much easier given the right device.
Although at this point the industry has so much money in the forward-backward pass system that I doubt a backprop successor would win unless someone makes NPC hardware feasible and can prove scaling up to billions of params