Why reduce dimensions
Real data often has dozens, hundreds, or even thousands of features: a small image has thousands of pixels, and a text vector may have hundreds of dimensions. As the dimensionality increases, troubles arise - computational and storage costs rise, visualization becomes difficult to start with, and samples in high-dimensional spaces become extremely sparse, leading to the gradual "failure" of distance measurement. This is known as the "curse of dimensionality". Dimensionality Reduction is the process of finding ways to use fewer dimensions while preserving as much useful information as possible from the original data.
Two approaches to dimensionality reduction
One is feature selection: selecting the most useful columns from the original features without changing their meaning, such as retaining only the features with the highest variance or most relevant to the target. The second is feature extraction (projection): recombining original features into new, fewer dimensions, which are usually functions of the original features and often no longer have intuitive meanings. PCA and t-SNE both belong to the latter.
PCA: Find the direction with the highest amount of information
The idea of Principal Component Analysis (PCA) is to project data onto a low dimensional subspace and maximize the variance after projection - the larger the variance, the more information is retained. After data centralization, it obtains a set of mutually orthogonal principal components through the eigenvectors of the covariance matrix (or singular value decomposition SVD): the first principal component has the largest variance, followed by the second principal component, and so on. Taking the first k components completes the dimensionality reduction.
The advantages of PCA are clear mathematical foundation, reversibility (able to approximate reduction), and fast computation; The disadvantage is that it is a linear method that can only capture linear structures and is sensitive to feature scales, usually requiring standardization first. It is also commonly used for denoising and compression.
T-SNE and UMAP: In order to 'see'
T-SNE (t-distributed random neighborhood embedding) is a nonlinear dimensionality reduction method that focuses on visualization. It brings together "similar" points in high dimensions in low dimensions (usually reduced to two dimensions) and pushes away "dissimilar" points. It excels at spreading complex high-dimensional clusters on a plane and is a commonly used tool for neural network feature visualization. But it calculates slowly, and the distance between clusters in low dimensional graphs is not reliable, which cannot be used to determine whether cluster A is really farther than cluster B. UMAP is a later method that is faster and generally better at maintaining global structure, and has been widely used in recent years.
How to choose
The purpose is to compress, denoise, and accelerate downstream models: prioritize PCA (linear, controllable, reversible). The purpose is to draw high-dimensional data into a graph for people to see: using t-SNE or UMAP. The data scale is large and needs to balance local and global structures: UMAP is often used directly, or PCA is used to reduce it to tens of dimensions before running t-SNE (this is a common two-step approach).
Common Misconceptions
1. Consider t-SNE graph as the "truth" - it mainly reflects local proximity relationships, and the size and spacing of clusters are unstable, which may change with a random seed graph.
2. Forget standardization - PCA is extremely sensitive to dimensions, and non standardization can allow features with larger values to dominate the results.
3. Thinking that dimensionality reduction will definitely improve accuracy - it is more used for compression, denoising, and visualization; The lower the dimension, the better. Cutting off useful information will actually make the effect worse.
one-sentence summary
Dimensionality reduction is the process of "slimming down" and "translating" high-dimensional data: PCA uses linear projection to preserve maximum variance, while t-SNE and UMAP provide non-linear visualization services. Understanding their applicable boundaries is often more important than memorizing formulas.
[Reference source] Comprehensive compilation of self published textbooks and industry materials, including official documents from scikit learn on PCA and t-SNE, as well as published papers and reviews on PCA, t-SNE, UMAP, and other methods.