I understand the obsession with low bit quantization, but it is empirically quite evident that it is not possible to compress models to less than 4 bit per weight without severe loss of capabilities².
It may be nice as an experiment, but it is obviously a very inefficient route for model training: spending all the flops on a saturated model only to prune its capabilties.
²As to why, I have seen few explanations. But the empirical evidence is there.
Perf goes from 80% to 47% on Wikitext-2. Also no comparisons to FP4 solutions that are able to maintain or exceed perf on the same dataset 80% perf with a 4.25-4.5 big budget.
I think more meaningful thing here would be a hybrid solution that went down to sub-bit representations when the informational representation does not need it (for example later layers) that still maintains task performance
I don't think it's trying to be a useful implementation, but the significance they do provide is that they are able to improve on the relative loss at lower quants.
So, not something anyone would want to run currently, but an indicator that there is still more to squeeze out of lower precisions.
Trellis quantization is a far more approachable enhancement right now, but it doesn't cross the 1-bit barrier (and perhaps doesn't intend to).
It may be nice as an experiment, but it is obviously a very inefficient route for model training: spending all the flops on a saturated model only to prune its capabilties.
²As to why, I have seen few explanations. But the empirical evidence is there.
I think more meaningful thing here would be a hybrid solution that went down to sub-bit representations when the informational representation does not need it (for example later layers) that still maintains task performance
So, not something anyone would want to run currently, but an indicator that there is still more to squeeze out of lower precisions.
Trellis quantization is a far more approachable enhancement right now, but it doesn't cross the 1-bit barrier (and perhaps doesn't intend to).
https://huggingface.co/moondream/parakeet-redux