Training a large model with billions or even billions of parameters, a single graphics card can no longer accommodate: weights cannot fit, gradients cannot fit, and optimizer states cannot fit. Breaking down training tasks onto hundreds or thousands of cards simultaneously is the problem that distributed training aims to solve. But 'parallelism' is not simply about cutting data into everything, it has at least three different splitting ideas.
1、 Data Parallelism (DP)
The easiest way to understand is to put a complete model on each card, divide a batch of data into several parts and distribute them to each card. Each card independently calculates the gradient, and then synchronizes the gradient sum through communication. Finally, everyone updates the parameters together.
The advantages are simple implementation and good scalability; The disadvantage is that each card needs to store a complete model and optimizer state, resulting in low memory efficiency. When the model itself is too large to fit on a single card, data parallelism becomes ineffective. In order to alleviate the situation, improvement schemes such as ZeRO/FSDP have emerged, which involve opening up the optimizer state, gradient, and parameters to each card.
2、 Tensor Parallelism (TP)
Split the matrix inside a single operator and calculate it on different cards. For example, a large matrix multiplication can be divided into columns or rows, with each card only accounting for a portion, and the results can be pieced together through communication.
Its advantage is that it can truly reduce the single card video memory, making it suitable for models with large single-layer memory; The disadvantage is frequent communication and extremely high bandwidth requirements between cards, usually only used on the NVLink high-speed interconnection within the same machine.
3、 Pipeline Parallelism (PP)
Layered segmentation: Place different layers of the model on different cards, and the data flows from front to back like a pipeline - the first group of cards calculates the first segment, and the intermediate results are handed over to the second group of cards to calculate the second segment.
The problem is that there will be "bubbles": the front one is stuck in the calculation, and the back one is stuck waiting. In order to reduce idling, in practice, a batch is divided into multiple micro batches to keep multiple production lines busy at the same time.
4、 Combination in Reality: 3D Parallel
Real large-scale model training rarely uses only one type of parallelism, but instead stacks three together: tensor parallelism within nodes (eating NVLink bandwidth), pipeline parallelism between nodes (low communication volume), and then data parallelism (increasing throughput). This is commonly referred to as' 3D parallelism '.
Summary in one sentence: Data parallel throughput expansion, tensor parallel memory saving, pipeline parallel cross node - understanding their trade-offs, also understands why large model training relies so heavily on clusters and high-speed interconnection.
[Reference source] Comprehensive compilation of industry information publicly released (including Megatron LM, ZeRO, and other public papers and technical documents).