Full Fine tuning is often not practical for a large model on one's computer to learn a specific task - it requires updating all parameters of the model, has high memory overhead, slow training, and requires storing a complete copy of the model for each task.
LoRA (Low Rank Adaptation) provides a more cost-effective approach: freezing the original model and only "hanging" a small number of trainable parameters next to it.
1、 The core idea of LoRA
The starting point of LoRA is that when the model adapts to new tasks, the change in parameters is actually "low rank", which means that this change does not need to be characterized by filling the entire parameter matrix.
The specific approach is to freeze certain weight matrices W in the model, and introduce two small matrices A and B to approximate the parameter increments by multiplying them. During training, only A and B are updated, and the original weights remain unchanged.
The benefits of doing so are very direct:
-The trainable parameters are significantly reduced, usually only accounting for a very small proportion of the original model;
-The cost of video memory and training has significantly decreased, and ordinary consumer grade graphics cards can also run;
-The training product is just a small adaptation file that is easy to save and distribute.
2、 Why is it so economical
The key lies in 'dimensionality reduction'. Assuming the original matrix is 4096 × 4096, training it directly requires millions of parameters; And LoRA decomposes it into two small matrices, 4096 × r and r × 4096. r usually takes very small values such as 8, 16, and 32, and the number of parameters drops to zero at once.
At the same time, because the original weights are frozen, there is no need to save the optimizer state for them during training, and the memory usage is significantly reduced.
3、 Typical usage process
1. Prepare data: Organize task data into "input-output" pairs, such as Q&A, instruction following, and format conversion.
2. Choose a base model: Select an open-source base model as the base.
3. Configure LoRA: Set hyperparameters such as rank r, target layer, learning rate, etc.
4. Training: Only update LoRA parameters and observe the loss curve.
5. Reasoning: Merge LoRA weights with the base model or dynamically load them at runtime.
Common open-source toolchains, such as Hugging Face's PEFT library, have encapsulated this process quite maturely, and training can begin with just a few lines of configuration.
4、 Suitable and unsuitable scenarios
LoRA excels in adapting at the level of "style" and "format": allowing models to learn specific tones, output structures, and the organization of professional terminology. It is also suitable for multitasking parallelism - the same base can be paired with multiple different LoRAs, switching on demand.
The boundary to note is that if the task is to inject a large amount of new factual knowledge into the model, LoRA is not the optimal solution, and this scenario is often more suitable for using Retrieval Augmentation (RAG) to place knowledge externally. In addition, a higher rank r is not necessarily better. An excessively high rank can increase costs and may also lead to overfitting.
5、 Relationship with RAG
LoRA and RAG often appear together. A common combination is to fine tune the model with LoRA to follow the response specifications required by the business, and then use RAG to provide it with real-time data that can be updated at any time. The former shapes' habits', while the latter provides' facts' with clear division of labor.
Summary
LoRA uses the simple idea of "low rank decomposition" to transform the fine-tuning of large models from the exclusive capability of big companies to something that even ordinary developers can handle. It does not change the knowledge core of the model, but it allows the model to be expressed in a more appropriate way, which is precisely why it has become popular in practice.
【 Reference source 】 Comprehensive compilation of industry information and mainstream technical documents that have been publicly released.