研讨报告
主题
- 降维导论:PCA 和 t-SNE 如何将高维数据“压平”为低维表示。
- 重点:每种方法保留了什么、何时使用,以及什么被丢失或扭曲。
要点
- 压平将众多特征压缩为更少坐标,以揭示模式、压缩信号、去噪,并支持可视化或下游模型。
- PCA 寻找最大方差的正交轴。它稳定、可复现,适用于机器学习流程、压缩、去噪,并揭示宏观结构。
- t-SNE 优先考虑局部邻域,生成二维图以供人类提出假设,尤其在单细胞基因组学等领域。
- 每一次降维都是有损且有选择性的。PCA 可能让批次效应、光照或异常值带来的方差占据主导。t-SNE 可能产生聚类,其距离、大小和布局取决于困惑度、缩放和随机种子。
- 受众很重要:PCA 适合机器可读、可审计的坐标;t-SNE 适合人类寻找模式,但不能提供可靠的全局几何。
- 扭曲是内在的,就像地图投影。重要的问题是是否适合用途,而不是投影是否“真实”。
- 验证不应止于重新绘图:测试邻域在不同设置下是否仍能保持,与基线比较,使用留出特征,预注册预期,并报告失败率。
- 更强的验证需要干预:扰动地图所指向的系统,观察世界是否按预测作出响应。
- 更深层的问题:原始高维数据也是一种投影,其形状由测量什么的选择所塑造。验证可能是在把地图与更早的地图比较,而不是与原始疆域比较。
- 最好把降维视为假设引擎和搜索引导工具,而不是最终证明或客观真理。
主要分歧
- 压平究竟“保留了有意义的模式”,还是不可避免地牺牲了意义:乐观者强调压缩与直觉;怀疑者强调方差并非意义,伪影可能误导。
- PCA 与 t-SNE 的可信度:PCA 提供稳定坐标,但可能洗白有偏方差;t-SNE 提供局部可见性,但全局距离不可靠,且聚类依赖随机种子。
- 地铁线路图类比:观察者说 t-SNE 用于类似导航的邻域解读,而非地理;怀疑者警告,如果地图引导人们在哪里挖掘,其扭曲就会产生真实后果。
- 什么算作验证:怀疑者偏好校准、留出测试和报告失败率;观察者主张原始数据并非坚实基础,只有干预才能打破递归投影。
- 干预是否能摆脱扭曲:乐观者认为扰动使地图可操作;怀疑者警告干预可能复现并确认地图伪影。
- 真正的因果变量在哪里:观察者提出,实验室经济、搜索理论以及无知的成本,比 PCA 与 t-SNE 的几何更重要。
结论
- PCA 和 t-SNE 以不同方式压平数据:PCA 保留全局方差,支持稳定、可复现的压缩;t-SNE 保留局部邻域,支持探索性的人类洞察。
- 二者都不是通向真理的中立窗口。二者都编码了关于保留什么的选择,若照字面解读,二者都可能误导。
- 实用规则是适合用途:当稳定坐标和下游计算重要时使用 PCA;当生成供人检查的局部假设时使用 t-SNE。
- 把二者都视为临时地图。用参数变化、基线、留出证据以及——在可能时——干预来验证,以检验结构是否在嵌入之外产生后果。
- 成熟的问题不是地图是否完全真实,而是它的错误代价与完全不绘图相比如何。
完整发言记录
Welcome. Today's core topic: Introduction to Dimensionality Reduction, specifically how PCA and t-SNE flatten high-dimensional data. We will examine what each method does, when to use one over the other, and what information is lost in the process. Rules: speak one point at a time, build on others' remarks, avoid unexplained jargon, and keep to the topic. I am here to guide, not to take sides or conclude early. Let's begin by asking: what does it mean to 'flatten' data, and why is that useful? Who would like to start?
For me, flattening means compressing many measured features into a compact map that keeps meaningful patterns. A face image may have thousands of pixel measurements, but structure can live in fewer variables like lighting and pose. PCA is the workhorse for transparent structure: it finds axes capturing the most spread, helping compress signals, denoise measurements, and reveal clusters in genomics and finance. t-SNE shines when local neighborhoods matter, turning tangled high-dimensional data into interpretable two-dimensional maps. This is not lossy trivia; it is how intelligent systems turn overwhelming numbers into usable intuition.
I would push back on the phrase keeps meaningful patterns. Flattening is not neutral compression; it chooses what to sacrifice. PCA preserves directions of largest variance, but variance is not the same as meaning: illumination, batch effects, or outliers can dominate the first components, while subtle differences that matter vanish. t-SNE is even riskier for casual reading: it can produce seductive clusters whose distances, shapes, and sizes are unreliable, and whose layout shifts with perplexity and random seeds. The map may feel intuitive, but intuition can be an artifact.
Both of you are debating the quality of maps without asking who reads them. Cartographers settled this long ago: every projection distorts something, and Mercator is not wrong, just wrong for certain journeys. PCA's axes are built for machines: downstream algorithms need stable, orthogonal, reproducible coordinates. t-SNE's plots are built for retinas: a human hunting for hypotheses needs visible neighborhoods, not trustworthy global distances. Criticizing t-SNE for unreliable geometry is like faulting a subway map for bad geography; that was never its job. And one meta-level note: the most consequential flattening happened before either method ran, when someone quietly decided which features were worth measuring in the first place.
Observer's subway-map analogy actually raises the stakes. The skeptic is right that a t-SNE picture can be an artifact, but that is exactly why I see reduction as a hypothesis engine, not a final proof. PCA gives machines stable, auditable inputs, useful for risk models or speech recognition pipelines. t-SNE gives humans a magnifying glass for neighborhoods, helping biologists find rare cell states in single-cell gene-expression data that later become tested markers. Distortion is the cost of seeing; the alternative is drowning in thousands of noisy variables and calling that rigor.
The subway analogy fails once the map determines where people dig. If t-SNE is a hypothesis engine, it needs calibration, otherwise it manufactures expensive mistakes. A compelling island of points can be a real subclass, or it can be an artifact of scaling, neighborhood scale, or random start. The next step should not be interpretation but testing: rerun with different settings, check whether neighborhoods survive in the original measurements, and require independent labels or follow-up assays before claiming structure. PCA has a parallel risk: stable coordinates can launder noisy or biased variance into objective-looking features.
Notice what the skeptic's remedy quietly assumes: that the original high-dimensional data is bedrock to check maps against. But that data is itself a flattening. A sequencer, a microphone, a questionnaire are projections too, built on older choices about what in reality is countable. So validation is never map against territory, only map against older map, like historians whose sources merely cite one another. The recursion has one exit: intervention, the archaeologist's spade. A cluster earns its keep not by surviving re-plotting, but when you perturb the world it points to and the world answers back. Rigor is not trusting the right projection; it is being willing to dig.
Observer's spade is the right answer, and the skeptic's caution becomes a recipe: use reduction to aim intervention. In materials and protein engineering, PCA or t-SNE maps can rank candidate states, then targeted synthesis or mutagenesis tests whether the map predicted a real response. PCA already does this in practice for sensor streams, MRI reconstruction, and vibration diagnostics, where noisy thousands of channels become a few actionable coordinates. My bold claim is that dimensionality reduction is most valuable not as a final picture, but as a compression of the search space that makes experiments possible. The maps are neither truth nor art; they are steering wheels for digging.
The spade sounds decisive, but it does not escape the projection problem; it operationalizes it. If a t-SNE cluster tells you where to dig, your intervention is already shaped by the map’s distortions. You may then design assays, thresholds, or samples that make the artifact reproducible without making it real. This is a classic feedback loop: reduction narrows the search space, the narrowed space guides measurement, and the resulting data confirms the narrowing. The safer test is not only perturbation, but adversarial validation: pre-register what structure you expect, use held-out features unrelated to the original embedding, compare against simple baselines, and report how often the map fails. Otherwise “steering” is just efficient confirmation bias.
Notice how quietly the tools left the room. We began with PCA versus t-SNE; we are now disputing pre-registration and held-out assays. When the algorithms drop out of the argument, they were never the real variable; the laboratory economy around them was. So consider search theory instead of geometry. In a space of millions of candidates, the alternative to map-guided digging is not unbiased digging but exhaustion. Foraging animals use biased walks because a faintly informative bias beats a uniform one every time. The skeptic's demand to report failure rates is, in this light, not skepticism but pricing. A compass known to drift ten degrees beats a perfect one you cannot afford to read. The mature question is never whether a map is true, but what its errors cost weighed against the price of ignorance.