When it comes to AI generated images, many people think of diffusion models. But before the diffusion model became popular, another paradigm that profoundly influenced generative AI had already emerged - the Generative Adversarial Network (GAN). It was proposed by Ian Goodfellow and others in 2014, using a game of "left and right fighting" to teach machines to create realistic images out of thin air.
The cleverness of GAN lies in making two networks confront each other. One is a generator, responsible for "creating" from random noise and producing samples that look like real data; The other is the Discriminator, which is responsible for distinguishing whether the input is from real data or forged by the generator. The generator tries its best to deceive the discriminator, and the discriminator tries its best to detect forgery. Both sides become stronger in the confrontation.
It can be imagined as a competition between counterfeiters and appraisers: counterfeiters constantly improve their techniques, appraisers constantly update their vision, and ultimately counterfeiters create something that even appraisers find difficult to distinguish. The goal of training is to make the samples output by the generator as close to real data as possible in distribution.
With this approach, GAN has been widely used in image generation, image super-resolution, style transfer, face synthesis, and even used to synthesize training samples for tasks with insufficient data. It has been a representative technology for generative visual tasks for a long time.
However, the training of GAN is notoriously unstable. The generator and discriminator must maintain a delicate balance: if one is too strong, the other cannot learn, which can easily lead to "pattern collapse" - the generator repeatedly produces a few samples, resulting in a serious lack of diversity; In addition, it is highly sensitive to hyperparameters, and training often requires repeated tuning. This is also one of the reasons why diffusion models gradually became mainstream with more stable training and better diversity.
In summary, GAN opened the door to generative AI with the simple yet profound idea of "adversarial". Despite its unstable training, the game theory it left behind still inspires new model designs to this day.
[Reference source] Comprehensive compilation of industry information and publicly available materials from research institutions.