What would happen if you had a model first compress an image into a small segment of numbers, and then reconstruct the original image based solely on that segment of numbers? This is exactly the idea of an autoencoder: the encoder is responsible for compression, the decoder is responsible for restoration, and the low dimensional representation in the middle is called a "latent vector". But the hidden space learned by ordinary autoencoders is often chaotic, and decoding something from a random point is meaningless - so it is good at compression, but not very good at "generation".
Variational Autoencoder (VAE) takes a step forward on this basis: it no longer maps the input to a deterministic point, but to a probability distribution, typically described by mean and variance. During training, the model "samples" a latent vector from this distribution and hands it over to the decoder for reconstruction. In this way, the hidden space is constrained to be more regular and continuous, and decoding any point can yield a decent sample.
To enable backpropagation of the "sampling" step, VAE used a clever technique - reparameterization: the sampling is written as "mean+standard deviation x noise", where the noise is randomly taken from the standard normal distribution. This way, the gradient can smoothly pass through the random operation and be transmitted back to the encoder.
The loss function of VAE consists of two parts: one is the reconstruction error, which measures the difference between the decoding result and the original image; The other part is KL divergence, which is responsible for "pulling" the encoded distribution towards the standard normal distribution, keeping the hidden space smooth. The former makes the reconstruction clearer, while the latter makes the hidden space more organized. The two complement each other, forming the core trade-off in VAE training.
Its value lies in the fact that the hidden space is continuous and interpretable, which means that interpolation and arithmetic operations become meaningful - linear interpolation between two faces can result in a smooth transition of the middle face; Subtracting the 'calm face' from the 'smiling face' and adding the 'angry face' can even approximate the 'angry smile'. This structured generation capability is the reason why VAE is still widely used today.
Of course, VAE also has its shortcomings. Reconstructed images often tend to be blurry, as common reconstruction losses (such as mean square error) tend to output an "average" and details are easily lost. Improvement directions include introducing perceptual loss, adversarial loss (combined with GAN), or combining with diffusion models. Nowadays, diffusion models are more popular on the main stage of image generation, but the hidden space idea of VAE is still an indispensable part of many systems - for example, latent diffusion models (such as Stable Diffusion) rely on a VAE encoder to compress images into the hidden space before generation.
In summary, GAN relies on adversarial realism, diffusion models rely on gradual denoising, and VAE relies on probabilistic compression - it contributes a regular, continuous, and computable hidden space to the generative model.
【 Reference Source 】 Comprehensive compilation of research literature and publicly available technical materials on variational autoencoders.