Videos

Josh Batson - Interpretability: From Art Towards Science

August 31, 2026
Abstract
If deep neural networks at initialization are like physical systems -- model-able in terms of field theory and random matrices -- then trained neural networks are like biological systems -- full of idiosyncrasy reflecting the world they evolved in, path dependent accidents of evolution, and robust structures enabling complex, general behavior. Our team at Anthropic has found a plenitude of these "biological" or even "psychological" style phenomena, and mapped some of them to specific vectors, subspaces, and matrices. But today, we lack a strong theory of what about the data+model+optimization gives rise to these structures. I will present a number of phenomena along with concrete open problems amenable to mathematical and toy-model analysis, suggesting a path from artisanal analysis of model mechanism towards a science of emergence.