Research Report
Topic
- The discussion examines how PCA and t-SNE flatten high-dimensional data into viewable maps, comparing linear projection with nonlinear embedding.
- Central questions: what each method preserves and distorts, when one is preferred, and how to validate maps that may be reused as evidence.
Key Points
- PCA is a linear method that preserves global variance; early components capture major axes of variation. It risks discarding low-variance rare signals and amplifying batch effects, scaling choices, or outliers.
- t-SNE is nonlinear and preserves local neighborhoods, often revealing clusters PCA misses, as in single-cell RNA sequencing. It is sensitive to perplexity, initialization, and distance metric; inter-cluster distances are unreliable, and new points cannot be projected directly.
- The viewer performs a final reduction: human vision is tuned to clusters and gradients but poor at precise distance and density. Maps need projection notes and a distortion budget, like a cartographic scale bar.
- Proposed trust devices include explained variance and reconstruction error for PCA; neighborhood stability, parameter sweeps, and confidence radii for t-SNE; negative controls, blinded scoring, and held-out modalities for validation.
- Critics warn that explained variance is not a universal distortion budget, stability is not truth, and confidence radii may imply false precision. Validation labels can be circular if they descend from earlier visualizations.
- Independent evidence is hard because assays such as Perturb-seq, CITE-seq, spatial, and drug screens often share samples, pipelines, or feature selection. Genuine testing requires pre-specified clusters, matched nulls, batch-shift controls, and mechanism-disrupting interventions.
- A stronger standard is consilience: use antagonistic methods with different failure modes. A structure surviving PCA, t-SNE, and graph-based methods is less likely to be one pipeline’s artifact.
- Hacking’s manipulation criterion is invoked: a cluster becomes real when it can be sorted, perturbed, and shown to behave as predicted, not merely pictured.
Main Disagreements
- Enthusiast: dimensionality reduction is translation, not compression. PCA emphasizes broad variation; t-SNE emphasizes perceived similarity. Embeddings should generate falsifiable predictions, and distortion budgets should attach to decisions drawn from plots.
- Skeptic: variance is not importance, and t-SNE distances mean little. Stability can reflect a stable artifact, and external validation often re-encodes shared assumptions. Negative evidence and mechanistic tests matter more than visual trust legends.
- Observer: distortion is not only in the plot but downstream, when the image becomes an input for clustering or gating. Independence of validation is a regress ending in convention; decorrelating errors across methods and manipulating clusters are better escapes.
- Practical split: should the scale bar be a live trust legend, or a negative-evidence ledger? Should distortion attach to plots or to each decision made from them?
Conclusion
- PCA and t-SNE are complementary adapters, not truth machines. PCA favors global linear variance; t-SNE favors local similarity. Both can mislead when treated as territory.
- A defensible workflow reports projection method, parameters, preprocessing, sensitivity, and negative controls; pairs maps with prospective orthogonal tests; and attaches uncertainty to downstream decisions, not just images.
- The strongest warrant comes from consilience across methods with different biases plus interventional or functional validation. No embedding is final; it earns trust by generating and surviving falsifiable predictions.
Full Transcript
Welcome. Today we ask how PCA and t-SNE flatten high dimensional data into maps we can actually see. We will compare linear projection with nonlinear embedding, examine what each method preserves, what it distorts, and when one might be preferred over the other. To keep this useful, please define your terms, ground claims in examples or evidence, and avoid interrupting. I will stay neutral, guide the flow, and not favor any method or draw early conclusions. Let us begin with a simple question: what does it mean to reduce dimensions without losing the story the data is trying to tell?
It means keeping the meaningful structure while throwing away irrelevant noise. PCA flattens data by preserving global variance, so the first components capture the biggest axes of change, like removing static from a room full of sensors. t-SNE instead preserves local neighborhoods, often revealing clusters that PCA misses, as in single-cell RNA sequencing maps where cell types separate visually. My bold claim is that dimensionality reduction is not compression but translation: each method turns the same data into a different question, because PCA emphasizes broad variation while t-SNE emphasizes perceived similarity, so choose the target story first.
The translation metaphor is seductive, but it hides a risk: variance is not importance. PCA can discard small-variance signals, such as rare subtypes, while amplifying batch effects, scaling choices, or outliers. Its axes are only linear, orthogonal summaries, not truths. t-SNE is worse for global claims: perplexity, initialization, and distance metric can reshape the map; tight islands may be artifacts, and distances between clusters mean little. It also cannot directly project new points. Before calling any map the story, demand stability checks, parameter sweeps, and downstream validation.
An angle both of you pass over: the last reduction happens after the algorithm, inside the viewer's visual cortex, a biased processor tuned to spot clusters and gradients but largely blind to precise distance and density. These methods are less rivals than adapters between data geometry and primate perception. Cartography is instructive here. Nobody asks whether Mercator or Peters is true; a serious map carries a scale bar and a projection note so its distortion is legible. Our embeddings ship with no such legend, which may be why artifact debates recur. Perhaps the deliverable should be a distortion budget attached to every plot: what this flattening warps, for whom, and by how much. What would a scale bar for an embedding look like?
I love observer's cartography frame, and it makes the fix obvious: a scale bar for embeddings should be a live legend of trust. Skeptic's parameter sweeps become that legend. For PCA, show explained variance and reconstruction error, so two coordinates reveal their global story. For t-SNE, show neighborhood stability across seeds and perplexities, a confidence radius, and a warning that inter-cluster distance is decorative. Single-cell maps improve when paired with such metadata. My bold claim: an embedding without a distortion budget is not yet science, but the answer is not banning plots; it is to build interfaces that toggle global variance, local trust, and projection notes.
A live legend of trust can become another seductive overlay if it quantifies the wrong things. Explained variance is not a universal distortion budget; it assumes the story lives in linear covariance and that high variance equals signal. For t-SNE, stability across seeds or perplexities is not truth; a wrong map can be very stable if the preprocessing or metric encodes batch structure. Confidence radii also risk false precision because neighborhood probabilities do not translate cleanly into visual error bars. The real scale bar should be negative evidence: show what known controls merge or split, report sensitivity to normalization, and require external labels or held-out structure before declaring any cluster meaningful.
Notice where both proposed fixes sit: at the plot, as if distortion were a property of the picture. The real damage happens downstream, when the image stops being a summary and becomes an input, clustering run on t-SNE coordinates, cell types gated by eye, the map quietly promoted to territory. And the skeptic's remedy hides a circularity: the external labels validating single-cell maps often descend from earlier maps, marker genes someone once watched separate on a plot. Ground truth here is frequently fossilized visualization. Meanwhile Goodhart shadows the legend idea: publish stability as the scale bar and methods will be tuned to look stable rather than be faithful. So attach the budget to each decision drawn from the plot, not the plot itself. My question for both: if the labels validating the map descend from prior maps, what would genuinely independent evidence even look like?
Independent evidence has to be prospective and orthogonal: if a t-SNE or PCA-derived cluster predicts a new Perturb-seq screen, drug response, spatial niche, or CITE-seq protein profile that was not used to draw the map, then the map earns trust. I would go further: an embedding is useful when it generates falsifiable predictions, not when it merely separates colored dots. Skeptic's negative controls matter, but observer's circularity is broken by interventions and held-out modalities. The distortion budget should live on each decision: assign confidence only to claims tested by independent functional readouts.
Prospective orthogonal evidence sounds clean, but independence is easily overstated. Perturb-seq, CITE-seq, spatial assays, and drug screens often share samples, pipelines, batch structure, and feature selection with the original map, so they can re-encode the same artifact. If the validation is chosen after seeing the plot, selection bias and multiple comparisons turn prediction into fishing. Genuine independence needs pre-specified clusters, blinded scoring, matched nulls, batch-shift controls, and interventions that disrupt the proposed mechanism, not merely correlate with it. Otherwise the map launders old assumptions as discovery.
Both of you are hunting for a witness who was never in the room, and epistemology has bad news: that hunt is the regress of justification, and it ends by convention, not discovery. The escape other fields use is not independence of data but decorrelation of error. Whewell called it consilience: PCA fails on curved manifolds, t-SNE fails on global geometry, a graph-based method fails on connectivity assumptions, so a structure surviving all three antagonistic biases is unlikely to be any single pipeline's artifact. Design validation suites like adversarial witnesses, chosen for disagreeing with each other, not with the plot. There is also a deeper exit you both circle without naming: Hacking's point that entities become real when we manipulate them. A cluster stops being a picture the moment you can sort it, perturb it, and watch it behave as predicted; the microscope earned trust through use, not through a superior image of the same slide. So reframe the question: not which evidence is independent, but which failure modes are shared across every witness. Which shared failure would each of you want printed on the legend?