Large models have good performance, but they are often large and slow, making it difficult to deploy on mobile phones, edge devices, or services that require low latency. Is there a way for a small model to inherit the ability of a large model? Knowledge Distillation is a model compression technique that emerged for this purpose: it uses a pre trained large model (teacher) to guide a small model (student) to learn, allowing students to approach the teacher's level with a smaller scale.
Core idea: It's not just about learning "answers", but also about learning "ideas"
Traditional training only tells the model that 'this image is a cat', which belongs to a hard label. And the teacher model will output a complete probability distribution for each sample, such as "80% like cat, 15% like dog, 5% like fox", which is called a soft label. Compared to only providing the correct answer, soft labels also contain "dark knowledge" such as similarity between categories, which is more informative. The student model can learn better generalization ability than simply supervising by learning both the teacher's soft labels and the real labels of the data simultaneously.
Temperature parameter: soften the distribution
The output of the teacher model is often very "sharp" - the probability of the correct answer is close to 1, other classes are almost 0, and dark knowledge is submerged. Hinton et al. introduced a temperature parameter T and divided logits by T in softmax. The larger T, the smoother the distribution and the clearer the relative relationships between categories. Usually, a higher T (such as 3 to 5) is used during training, and then restored to 1 during inference.
How to design the loss function
The total loss of the student model is usually composed of two weighted parts: the difference between the student soft output and the teacher soft output (commonly known as KL divergence), and the cross entropy between the student output and the true label. The former is responsible for "imitating the teacher", while the latter is responsible for "not deviating". Only by combining the two can one learn hidden knowledge without deviating from the real goal.
What is the difference between quantification and pruning
There are roughly three paths to model compression: quantization (reducing weights from high precision to low precision), pruning (removing unimportant connections or structures), and distillation (using a novice to imitate a master). The three are not mutually exclusive and are often used in combination in practical deployment, such as distilling small models first, and then quantifying and further compressing them.
Typical Applications
In the field of natural language processing, DistillelBERT is a smaller and faster model obtained by distilling BERT, which significantly reduces the number of parameters and inference costs while retaining most of the performance. In addition, distillation is also used to compress the "integration" capability of multiple models into a single model, facilitating online deployment.
Limitations and Challenges
The capacity of student models is limited, and it is impossible for teachers to inherit 100% of their abilities;
The biases and errors of the teacher model itself will also be "taught" to students;
The distillation strategy (temperature, loss weight, alignment of intermediate layer features) requires repeated parameter tuning.
Summary
Knowledge distillation uses a "teacher led student" approach to condense the abilities of large models into small models, which is an important part of model compression and efficient deployment. It reminds us that the value of a model lies not only in its parameter size, but also in how to effectively transfer and reuse knowledge.
[Reference source]
Hinton, Vinyals, Dean, "Distilling the Knowledge in a Neural Network",arXiv:1503.02531,2015
Sanh et al., "DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter",arXiv:1910.01108,2019
Comprehensive compilation of industry information that has been publicly released