In supervised learning, models rely on a large number of "input label" samples to learn. But labeling data is expensive and slow, so researchers have turned their attention to more "self reliant" ways. Contrastive learning is one of the most representative approaches: instead of assigning specific categories to each image, the model is left to determine which samples are similar and which are different.
Its core idea is very simple - to bring closer and push farther. Taking a picture of a cat as an example, apply two different enhancements (cropping, color changing, flipping, etc.) to obtain two slightly different "views". These two views are still semantically the same cat, and contrastive learning requires the model to 'bring them closer' in the feature space; At the same time, treat other images (other cats, dogs, landscapes) as "negative samples" and ask the model to "push them away". The goal of training is to make similar things more similar and dissimilar things less similar.
Why can useful representations be learned in this way? To determine whether two views are "homologous", the model must ignore unimportant surface differences (such as brightness and angle) and capture truly stable semantic features. Over time, it learned not to memorize pixels, but to understand the content. This also explains a counterintuitive phenomenon: features learned through contrastive learning often perform well on downstream tasks, even if the labels for these tasks have never been seen before.
There are two key points in terms of technology. One is the design of 'data augmentation', which determines what the model considers' invariant '. The enhancement is too weak, the task is too easy, and I cannot learn anything; Excessive enhancement may disrupt semantics and make positive samples dissimilar to each other. The second is the quantity and quality of negative samples. Early methods required a large number of negative samples and high demands on video memory; Later on, solutions emerged that only used positive samples or employed techniques such as large batches and momentum encoders to reduce costs.
The value of contrastive learning is particularly prominent in scenarios where labels are scarce. Images, text, speech, and even multimodal data can be pre trained by constructing positive and negative sample pairs. It is also the common logic behind many "self supervised learning" methods, laying the foundation for later multimodal models.
Of course, it is not omnipotent. The construction of positive and negative samples carries strong prior assumptions, and improper design can introduce bias; The representations learned through contrastive learning may not necessarily be directly transferable to all tasks, and often require a small amount of annotation for fine-tuning. But it at least proves one thing: without giving an answer, the model can still learn a lot from the structure of the data itself.
Summary in one sentence: Comparative learning enables models to "distinguish similarities and differences", and distinguishing similarities and differences is the starting point of understanding.
[Reference source] Comprehensive compilation of industry information and publicly available materials from research institutions.