Manifolds and Tangent Spaces: What ‘The Data Lies on a Manifold’ Actually Claims
Published:
Locally flat, globally not
Take the circle $S^1$. Near any point it looks like an interval; globally it is compact and closes up, so no continuous injective map to $\mathbb{R}$ captures it. That gap between local and global is the entire content of the definition.
A topological $d$-manifold is a Hausdorff, second-countable space in which every point has a neighbourhood homeomorphic to an open subset of $\mathbb{R}^d$. Such a homeomorphism $\varphi: U \to \mathbb{R}^d$ is a chart, a local coordinate system; a collection of charts covering the space is an atlas; and the manifold is smooth when every transition map $\varphi_j\circ\varphi_i^{-1}$ is $C^\infty$.
Why not use one chart? For most interesting spaces you cannot: $S^2$ is compact and any nonempty open subset of $\mathbb{R}^2$ is not, so a single homeomorphism onto one is impossible. Two stereographic charts, one omitting each pole, suffice. Longitude/latitude is a chart too, with the familiar defect that it degenerates at the poles and tears at the date line.
Keep three examples in mind: $S^2$; the torus $T^2 = S^1\times S^1$, also 2-dimensional but not homeomorphic to the sphere since it has a hole; and the Swiss roll, a flat rectangle rolled up inside $\mathbb{R}^3$, intrinsically flat despite looking dramatic.
The tangent space
Collect the velocities $\gamma’(0)$ of all smooth curves with $\gamma(0) = p$. The result is a $d$-dimensional real vector space, the tangent space $T_pM$, spanned in a chart by the coordinate directions $\partial/\partial x^1,\dots,\partial/\partial x^d$.
Two things make it useful. It is a genuine vector space, so all the machinery of Euclidean space applies locally even though the manifold itself has no notion of addition. And it is where derivatives live: the differential of a smooth $f: M\to N$ at $p$ is a linear map \(df_p: T_pM\to T_{f(p)}N\), so gradients and Jacobians on manifolds are maps between tangent spaces.
The manifold hypothesis
The claim is empirical, not mathematical: high-dimensional data concentrates near a low-dimensional manifold in the ambient space. A $224\times 224$ RGB image lives in $\mathbb{R}^{150528}$, yet photographs vary along far fewer directions — pose, illumination, identity, background. Three strands of evidence support it, none a proof:
- Intrinsic-dimension estimates. Nearest-neighbour estimators such as Levina and Bickel’s maximum-likelihood method return dimensions in the tens for standard image datasets; Pope et al. (2021) report roughly $40$ for ImageNet against $150{,}528$ ambient, and find that lower estimated intrinsic dimension goes with easier learning — fewer samples for the same accuracy.
- Testability. Fefferman, Mitter and Narayanan (2016) turned the hypothesis into a decidable statistical question: given samples, test whether they lie near a manifold of bounded dimension, volume and reach. It can fail, so it is a hypothesis rather than a slogan.
- What works in practice. Methods that assume a manifold — Isomap, LLE, their descendants — recover structure PCA cannot.
Where the hypothesis is weaker than it sounds
Data lies near a manifold, not on one — noise thickens the support, which is why “reach” appears in the formal statements. The support may also be a union of manifolds of different dimensions, making a single latent dimension the wrong model. And topology matters: an autoencoder with latent space $\mathbb{R}^k$ and continuous encoder and decoder is precisely one chart. If the data manifold is a circle of rotations, no such pair can be a homeomorphism, so the model must either tear the circle or fold it — the failure that hyperspherical and other structured-latent VAEs were built to avoid.
Recap
- A \(d\)-manifold is locally homeomorphic to \(\mathbb{R}^d\); charts are local coordinates, an atlas covers the space, smoothness constrains the transition maps.
- One chart is often impossible: \(S^2\) is compact and open subsets of \(\mathbb{R}^2\) are not; two stereographic charts suffice.
- \(T_pM\) is a \(d\)-dimensional space of curve velocities — the local linearisation, distinct at each point, not a subset of \(M\).
- The manifold hypothesis is empirical: intrinsic dimension of image data measures in the tens against \(10^5\)-plus ambient dimensions, and lower intrinsic dimension correlates with easier learning.
- Caveats: data is near, not on; the support may be a union of manifolds; a single \(\mathbb{R}^k\) latent space is one chart, so nontrivial topology breaks it.
References
- Lee, J. M. Introduction to Smooth Manifolds, 2nd ed. Springer, 2012.
- Levina, E., & Bickel, P. J. Maximum Likelihood Estimation of Intrinsic Dimension. NeurIPS 2004.
- Fefferman, C., Mitter, S., & Narayanan, H. Testing the Manifold Hypothesis. Journal of the American Mathematical Society, 29(4), 2016.
- Pope, P., Zhu, C., Abdelkader, A., Goldblum, M., & Goldstein, T. The Intrinsic Dimension of Images and Its Impact on Learning. ICLR 2021.
- Tenenbaum, J. B., de Silva, V., & Langford, J. C. A Global Geometric Framework for Nonlinear Dimensionality Reduction. Science, 290(5500), 2000.
- Davidson, T. R., Falorsi, L., De Cao, N., Kipf, T., & Tomczak, J. M. Hyperspherical Variational Auto-Encoders. UAI 2018.
