1、 Why is fine-tuning becoming increasingly expensive
Large models often have billions to billions of parameters. Traditional Full Fine tuning requires saving a weight equal to the original model for each task, as well as storing gradients and optimizer states in the same size of VRAM. Taking the 7B model as an example, full fine-tuning also requires tens of GB of video memory under mixed precision, which is a high threshold for most teams.
2、 The core idea of LoRA: low rank decomposition
LoRA (Low Rank Adaptation) was proposed by the Microsoft team in 2021. Its core assumption is that when the model adapts to a new task, the "change" in weights is low rank and does not need to be expressed using a full rank matrix.
So LoRA froze the original weight W and only added a bypass next to it: Δ W=B × A, where A and B are two small matrices that are "thin and long", and the rank r is much smaller than the original dimension (commonly r=4/8/16/64). During training, only update A and B, and the original W remains motionless. When reasoning, B × A can be directly merged and added back to W, so it will not increase any inference delay.
3、 Why does it save
-Reduce the number of trainable parameters to 0.01% to 1% of the original model;
-The demand for video memory has significantly decreased, and the 7B model can be fine tuned with a single 24GB graphics card;
-Each task only needs to save a few tens of MB of adapter, which can be plugged and switched;
-The base model remains unchanged, naturally avoiding catastrophic forgetting.
4、 Key hyperparameters
-Rank r: The larger the capacity, the stronger and easier it is to overfit, usually ranging from 8 to 64;
-Alpha: scaling factor, the actual update amount is roughly scaled by alpha/r;
-Target module: usually applied to the Q and V projections of the attention layer, and can also cover all linear layers;
-Dropout: Used to suppress overfitting on small datasets.
5、 Common variants
QLoRA performs LoRA on a 4-bit quantized base model, further lowering the threshold for video memory to a single consumer grade graphics card; DoRA, LoRA+, and others have made improvements in training stability and convergence speed.
6、 Suitable or not suitable
Suitable for: domain knowledge injection, writing style alignment, instruction following, and small sample task adaptation.
Not suitable for scenarios that require large-scale changes to the underlying capabilities of the model, or scenarios where both data and computing power are sufficient for direct training from scratch.
Conclusion
LoRA has transformed 'fine-tuning' from heavy assets to light assets: with less than 1% of parameters, it achieves a nearly full-scale fine-tuning effect. This is precisely why it has become the de facto standard in the open source community.
[Reference source]
-Edward J. Hu et al.: LoRA: Low Rank Adaptation of Large Language Models (arXiv: 2106.09685, 2021)
-Tim Dettmers et al.: QLoRA: Efficient Finetuning of Quantified LLMs (arXiv: 2305.14314, 2023)