Back to Home

Introduction to Gradient Checkpointing: How to Swap Time for Video Memory When Training Large Models

September 25, 2026 at 01:32 PMSource: RunByAI0 comment(s)TechGuide

When training large models, video memory often becomes a bottleneck earlier than computing power. In addition to mixed precision and model parallelism, there is also a commonly underestimated path - gradient checking. Its core idea is only one sentence: trade time for video memory.

1、 What ate the video memory

To calculate gradients in backpropagation, the intermediate activation values of each layer in the forward process need to be used. The deeper the layers, the longer the sequence, and the larger the batch, the more noticeable the volume of these activation values stacked in the video memory. Many times, it is not the model parameters that burst the memory, but these intermediate results.

2、 Only save the skeleton, calculate when needed

The method of gradient checkpoint is to not save all intermediate activations in the forward direction, only leaving input in a few layers (checkpoints), and discarding the remaining intermediate results. Wait until backpropagation requires activation of a certain layer, then recalculate forward from the nearest checkpoint and temporarily fill it out. In other words, replacing 'save' with 'recalculate'.

3、 Costs and benefits

The cost is additional computation: the discarded activations have to go forward again in reverse, resulting in slower training; The denser the checkpoints are set and the more video memory is saved, the higher the cost of recalculation. The benefit is a significant decrease in video memory usage, making it possible to train larger models, longer contexts, or on smaller graphics cards. A common compromise in engineering is to set checkpoints every few layers to find an acceptable balance between video memory and time.

4、 When is it worth using

When you are forced to compress batch sizes to a very small size, or want to fit into longer sequences and larger models, gradient checkpoints are often the card of 'just adding a little more computing power can get you running'. It is mixed with precision ZeRO、 Parallel models are not conflicting, but often used in combination to compress large model training into limited video memory.

Summary

The wisdom of gradient checkpoint lies not in "saving", but in "switching": since computing power can be bought a little more, but video memory is often difficult to obtain, then a part of the video memory pressure should be replaced with affordable computing costs.

[Reference source] Comprehensive compilation of industry information publicly released.

deep learninglarge model
Discussion

Comments (0)

No comments yet. Be the first!

Leave a Comment