The curse of dimensionality is a fundamental mathematical obstacle in organizing and analyzing data with thousands of variables—that is, data that is represented in high dimensions.
Computationally, it means that a brute force exhaustive search is not possible, since the sheer size of a search space explodes exponentially as more variables or dimensions are added.
“This is a fact, it is not something that will change,” said Eric Vanden-Eijnden, a mathematician at New York University’s Courant Institute. Dr. Vanden-Eijnden gave the analogy of a simple grid search: “In one dimension, you have two choices, forward or backward. In two dimensions, you have four choices—forward, backward, left, right. But in \(1000\) dimensions you have \(2^{1000}\) choices, which is an astronomically large number, much more than all the atoms in the known universe.”
Generative models, however, can successfully operate in thousands of dimensions, because the data they are trained on do not fill that vast space uniformly: they concentrate on lower-dimensional structures. That concentration is the blessing hidden inside the curse. Flows and diffusions—the models used for generating images, molecules, and scientific data—exploit it, and so do large language models.
But there is a challenge. Nobody yet knows how to explain this mathematically. Why these models locate the right low-dimensional structure, how much data they need, and when they fail to, are all open questions.
(In recognition of support from the Fernholz Foundation, the program honors John Tukey, the Princeton and Bell Labs mathematician and statistician who pioneered computer-based data analysis and introduced the terms “bit” and “software.”)
Over two weeks of lectures and hands-on generative-model coding, students explored how mathematics—including stochastic calculus and optimal transport—is directly relevant to cutting-edge AI; how it reveals hidden structure in high-dimensional data, thereby making learning in high dimensions feasible.
dimensionality.png929.36 KB Figure 1. Generative models based on flows and diffusions learn to turn noise into data: in this case flowers. Sampling different noises allows the generation of different flowers, thereby effectively learning the flower distribution from the data. Adapted from Michael Albergo, Nicholas M. Boffi, and Eric Vanden-Eijnden, “Stochastic interpolants: A unifying framework for flows and diffusions,” Journal of Machine Learning Research 26.209 (2025): 1-80.
“They were excited to see that what they learn in their first- and second-year graduate courses is directly relevant to cutting-edge AI,” said co-organizer and mathematician Jianfeng Lu at Duke University.
Dr. Vanden-Eijnden explained: “The data are samples from an unknown distribution, which we want to learn in a constructive way, so that we can generate more samples from it. LLMs themselves are generative models in exactly this sense. They learn the distribution of token sequences and sample from it, one token at a time. For images, molecules, and scientific data, the generation is instead performed by flows and diffusions, which transport a simple distribution onto the unknown data distribution. And analyzing the associated transports is precisely where the mathematics—stochastic calculus, optimal transport—becomes indispensable”
For many attendees, the program underscored that foundational research remains essential to AI, even in an era dominated by industry and tech giants. “The mathematical concepts at the heart of generative models originated in academic research,” Dr. Vanden-Eijnden noted. “And academia continues to play a crucial role in understanding the limits and capabilities of these methods and in bringing them into a broader range of scientific domains.”
One participant, Nadia Khoury, a Ph.D. student in mathematics at Penn State University, said that the summer school affirmed her research path. “It showed me that the mathematical ideas I have been studying during my PhD can connect very naturally to some of the most active questions in AI today,” Ms. Khoury said.
She added: “At a time when AI is developing so quickly, I think it is important not only to use these models, but also to understand them more deeply. The experience made me more excited about bringing tools from probability and PDEs into the study of generative AI.”
As Dr. Vanden-Eijnden noted: “These early career researchers will not just be black-box users of AI tools, they are the future stewards of the mathematical foundations and frontiers of AI.”
For Caleb Adeyemi Adeleye, a doctoral candidate at the University of Delaware in applied mathematics, the program prompted him to investigate how generative models might advance his research on the modeling of hollow-fiber membrane systems for carbon capture. “Generative models could be used as physics-informed models to learn maps from simple latent distributions to physically meaningful solution fields,” he said.
Another attendee, John Turnage, a graduate student in mathematics at the University of Utah, studies how to build data-driven surrogate models for scientific problems. “In operator learning, I start by building a deterministic surrogate: given an input function, it predicts an output function,” he said. “But the surrogate inevitably misses something, and that residual error is itself some function.”
Mr. Turnage provided a synopsis:
“The idea would be to combine a surrogate with a generative correction. First, you’d learn the main operator. Then you’d learn a distribution over the part the surrogate missed, so the model can produce not just one corrected solution, but an ensemble of corrected solutions. There are serious mathematical difficulties here, but if we can quantify when the models are stable, how much data they need, and how their errors propagate then they could become a more flexible tool for scientific uncertainty quantification.”
And indeed, understanding the curse of dimensionality and structure in data is essential to trustworthy and reliable AI, especially in scientific applications.