Back to Home

Introduction to Mixed Precision Training: Why Training Large Models with FP16/BF16 is Fast and Memory Saving

September 23, 2026 at 08:01 AMSource: RunByAI0 comment(s)TechGuide

When training large models, memory and computing power are often the main bottlenecks. Mixed Precision Training is a widely adopted acceleration method that makes training faster and more memory efficient without sacrificing too much accuracy. This article explains its principles and common practices.

1、 Starting from Floating Point Numbers: FP32, FP16, and BF16

Computers use floating-point numbers to represent decimals, and the number of digits determines the range and accuracy of the representation. FP32 (single precision) uses 32-bit technology, with a wide range and high accuracy, and is the default choice for deep learning; FP16 (half precision) only uses 16 bits, which roughly halves the memory usage and bandwidth, and has faster computation. However, it can represent a smaller range and is prone to overflow; BF16 is also 16 bits, but it assigns more bits to the exponent and fewer bits to the tail, indicating a range comparable to FP32, with slightly lower accuracy and more stable performance in large model training.

2、 Why not use FP16 for everything?

Directly replacing the model and calculations with FP16 will encounter two problems: first, when the gradient is very small, it will overflow to 0; second, when the numerical range is small, it is easy to overflow to infinity. So the industry's approach is not to "replace everything", but to "mix and match".

3、 The core idea of mixed precision

Sovereign weights are still stored in FP32 to ensure sufficient accuracy during updates; Forward and backward calculations are completed using FP16/BF16, which speeds up intensive operations such as matrix multiplication; When updating the weights, add the gradient to the weight of FP32. This not only allows for the speed of low precision calculations, but also avoids the accumulation of accuracy loss.

4、 Loss Scaling

To alleviate the gradient underflow of FP16, loss scaling is commonly used: multiply the loss by a larger coefficient, and the gradient will be amplified accordingly to avoid becoming 0; Divide by the same coefficient before updating to restore. BF16 represents a large range and is usually insensitive to loss scaling, which is one of the reasons why it is more popular in large models.

5、 Several key points in practice

1. Hardware support: Newer GPUs and NPUs have specialized acceleration for FP16/BF16, resulting in significant benefits.

2. Stability: Some precision sensitive operators (such as normalization, softmax) frameworks will automatically fallback to FP32.

3. Framework support: PyTorch's torch.amp and TensorFlow's mixed precision API both encapsulate automatic type conversion and loss scaling, which can usually be enabled with just a few lines of code.

4. It can be used in conjunction with quantization, gradient checkpoint, and other techniques to further compress video memory.

6、 Summary

Mixed precision training is not simply about halving the precision, but rather allowing different numerical types to perform their respective duties: FP32 tube precision, FP16/BF16 tube speed. Understanding sovereignty, loss scaling, and automatic type conversion is the key to mastering them.

[Reference source] Comprehensive compilation of publicly released deep learning framework documents and technical materials.

deep learninglarge model
Discussion

Comments (0)

No comments yet. Be the first!

Leave a Comment