To add multiple abilities to the same large model, it is usually necessary to fine tune them separately and find ways to put the results together. In recent years, a practice called "model merging" has become increasingly popular: without any training, the weights of multiple fine-tuning models are directly "merged" into one, often retaining multiple abilities at the same time. This article explains its principles, common methods, and usage boundaries.
1、 Why merge models
Assuming you have a base model that has been fine tuned in code, mathematics, and writing. The traditional choice is either to use only one of them or to retrain with mixed data - the latter is costly and prone to losing sight of one aspect. Model merging provides a third way: overlaying the parameters of several fine tuned models according to a certain rule to obtain a new model that does not require gradients, raw data, or retraining.
2、 Common merging methods
-Weighted average: Directly average the weights of multiple models, or mix them with weighting coefficients. The implementation is the simplest, but when the update directions of different models are inconsistent, it is easy to "fight" with each other.
-Task Arithmetic: First, calculate the difference between each fine-tuning model and the base (i.e. the "increment" caused by fine-tuning), and then add or linearly combine these increments to adjust the strength of each ability.
-TIES Merging: specializes in handling parameter conflicts - cutting out small updates, unifying symbol directions, and then merging to reduce mutual cancellation.
-DARE: Randomly reset most of the fine-tuning increments to zero and rescale before merging, using sparser updates to reduce redundancy and conflicts.
-SLERP: Interpolation on a sphere, commonly used for smooth fusion between two models, is more stable than simple averaging.
3、 Why does it work
One intuitive explanation is that fine-tuning usually only involves small adjustments near the base parameters, and the adjustment directions for different tasks do not completely conflict. Therefore, these "increments" can be stacked in the parameter space. Of course, if two fine-tuning directions are sharply opposed, simply adding them together will indeed destroy the ability, which is exactly the problem that methods such as TIES and DARE need to solve.
4、 Precautions in practice
-It is best for the models involved in the merger to come from the same base (with consistent architecture and parameter naming), otherwise the parameters cannot correspond one-to-one.
-The merging result must be tested and cannot be assumed to necessarily improve; How to allocate weights usually requires trial.
-Merge cannot create capabilities that models do not originally possess out of thin air, it only integrates what has already been learned through fine-tuning.
-Due to its low threshold and almost zero cost, it has become a common source of model variants in the open source community.
5、 Summary
Model merging is a form of "low-cost reuse": integrating existing multiple fine-tuning results into a more universal model. It cannot replace high-quality fine-tuning, but in the open source ecosystem, it has become a practical tool for recombining the value of numerous community models.
【 Reference Source 】 Comprehensive compilation of industry information released publicly