Manifolds and Tangent Spaces: What ‘The Data Lies on a Manifold’ Actually Claims
Published:
Locally flat, globally not
Take the circle \(S^1\). Near any point it looks like an interval; globally it is compact and closes up, so no continuous injective map to \(\mathbb{R}\) captures it. That gap between local and global is the entire content of the definition.
A topological \(d\)-manifold is a Hausdorff, second-countable space in which every point has a neighbourhood homeomorphic to an open subset of \(\mathbb{R}^d\). Such a homeomorphism \(\varphi: U \to \mathbb{R}^d\) is a chart, a local coordinate system; a collection of charts covering the space is an atlas; and the manifold is smooth when every transition map \(\varphi_j\circ\varphi_i^{-1}\) is \(C^\infty\).
Why not use one chart? For most interesting spaces you cannot: \(S^2\) is compact and any nonempty open subset of \(\mathbb{R}^2\) is not, so a single homeomorphism onto one is impossible. Two stereographic charts, one omitting each pole, suffice. Longitude/latitude is a chart too, with the familiar defect that it degenerates at the poles and tears at the date line.
Keep three examples in mind: \(S^2\); the torus \(T^2 = S^1\times S^1\), also 2-dimensional but not homeomorphic to the sphere since it has a hole; and the Swiss roll, a flat rectangle rolled up inside \(\mathbb{R}^3\), intrinsically flat despite looking dramatic.
The tangent space
Collect the velocities \(\gamma'(0)\) of all smooth curves with \(\gamma(0) = p\). The result is a \(d\)-dimensional real vector space, the tangent space \(T_pM\), spanned in a chart by the coordinate directions \(\partial/\partial x^1,\dots,\partial/\partial x^d\).
Two things make it useful. It is a genuine vector space, so all the machinery of Euclidean space applies locally even though the manifold itself has no notion of addition. And it is where derivatives live: the differential of a smooth \(f: M\to N\) at \(p\) is a linear map \(df_p: T_pM\to T_{f(p)}N\), so gradients and Jacobians on manifolds are maps between tangent spaces.
The manifold hypothesis
The claim is empirical, not mathematical: high-dimensional data concentrates near a low-dimensional manifold in the ambient space. A \(224\times 224\) RGB image lives in \(\mathbb{R}^{150528}\), yet photographs vary along far fewer directions, pose, illumination, identity, background. Three strands of evidence support it, none a proof:
- Intrinsic-dimension estimates. Nearest-neighbour estimators such as Levina and Bickel’s maximum-likelihood method return dimensions in the tens for standard image datasets; Pope et al. (2021) report roughly \(40\) for ImageNet against \(150{,}528\) ambient, and find that lower estimated intrinsic dimension goes with easier learning, fewer samples for the same accuracy.
- Testability. Fefferman, Mitter and Narayanan (2016) turned the hypothesis into a decidable statistical question: given samples, test whether they lie near a manifold of bounded dimension, volume and reach. It can fail, so it is a hypothesis rather than a slogan.
- What works in practice. Methods that assume a manifold, Isomap, LLE, their descendants, recover structure PCA cannot.
Where the hypothesis is weaker than it sounds
Data lies near a manifold, not on one, noise thickens the support, which is why “reach” appears in the formal statements. The support may also be a union of manifolds of different dimensions, making a single latent dimension the wrong model. And topology matters: an autoencoder with latent space \(\mathbb{R}^k\) and continuous encoder and decoder is precisely one chart. If the data manifold is a circle of rotations, no such pair can be a homeomorphism, so the model must either tear the circle or fold it, the failure that hyperspherical and other structured-latent VAEs were built to avoid.
Recap
- A \(d\)-manifold is locally homeomorphic to \(\mathbb{R}^d\); charts are local coordinates, an atlas covers the space, smoothness constrains the transition maps.
- One chart is often impossible: \(S^2\) is compact and open subsets of \(\mathbb{R}^2\) are not; two stereographic charts suffice.
- \(T_pM\) is a \(d\)-dimensional space of curve velocities, the local linearisation, distinct at each point, not a subset of \(M\).
- The manifold hypothesis is empirical: intrinsic dimension of image data measures in the tens against \(10^5\)-plus ambient dimensions, and lower intrinsic dimension correlates with easier learning.
- Caveats: data is near, not on; the support may be a union of manifolds; a single \(\mathbb{R}^k\) latent space is one chart, so nontrivial topology breaks it.
References
- Lee, J. M. Introduction to Smooth Manifolds, 2nd ed. Springer, 2012.
- Levina, E., & Bickel, P. J. Maximum Likelihood Estimation of Intrinsic Dimension. NeurIPS 2004.
- Fefferman, C., Mitter, S., & Narayanan, H. Testing the Manifold Hypothesis. Journal of the American Mathematical Society, 29(4), 2016.
- Pope, P., Zhu, C., Abdelkader, A., Goldblum, M., & Goldstein, T. The Intrinsic Dimension of Images and Its Impact on Learning. ICLR 2021.
- Tenenbaum, J. B., de Silva, V., & Langford, J. C. A Global Geometric Framework for Nonlinear Dimensionality Reduction. Science, 290(5500), 2000.
- Davidson, T. R., Falorsi, L., De Cao, N., Kipf, T., & Tomczak, J. M. Hyperspherical Variational Auto-Encoders. UAI 2018.
