Videos

Sam Eisenstat - Concepts, information, and objectivity

September 2, 2026
Abstract
Humans are able to learn language from each other, and we observe, both behaviourally and via interpretability methods, that different neural net models learn representations that correspond, at least to some degree, to each other and to human concepts. So far, these have been empirical observations, but we'd like to have a theoretically grounded understanding of what it is about the world and our representations that allows such correspondences. In this work, we suggest a model using latent variables. We introduce a set of latent variables, where the different latent variables can be distinguished by which observed variables they contribute to. Under some natural information-theoretic hypotheses, we prove an approximate uniqueness theorem, showing that two such latent variable models admit an approximate isomorphism.