A Breakthrough in Neural Network Quantisation

A Breakthrough in Neural Network Quantisation

Neural Network Quantisation

Neural network quantisation is used to reduce the precision of the weights of a neural network. There are a few methods to do this, the simpliest and often satisfactory method is to perform symmetric or asymetric quantisation. This does yield generally not terrible results, and the inference cost is not significant.

A more modern and now more common method is to use GPTQ, AWQ, K-Quantisation or I-quantisation. Each of these methods are described in their papers. In my experince GPTQ is effective, which is expected given the iterative nature of the method, and K-quantisation can also be effective as well. The difference can often be small between these methods and the choice of which depends on too many factors, the best approach is to try different methods to find out which works best for the given task.

The breakthrough

I found a method of inreasing the effective bit-rate of the very low-bit neural networks, enabling major performance gains compared to bitnet and QAT. This is, as far as I am aware, a new stet-of-the-art in quantisation. I currently do not have the compute to show this, but on small scale experiments, this method performs far better any other method at 1-bit quantisation.

I will be seeking to continue this line of research, there are novel branches of research that come from this breakthrough, pushing the efficiency of neural networks, including LLMs, far beyond the current capabilities at extreme quantisation.