Videos

Jacob Hilton - Miniature mathematical mysteries: uninterpreted models with fewer than 1,500 parameters

August 31, 2026
Abstract
What would it mean to completely understand the internal computation of a frontier model? We approach this question by starting with models containing only dozens to hundreds of parameters. These models are trained on simple algorithmic tasks, such as identifying the position of the second-largest number in a sequence, but they can be remarkably challenging to fully interpret. We argue for a practical intermediate target: producing "deductive" (i.e., sample-free) estimates of the model’s loss that are at least as accurate as sampling-based estimates given the same computational budget. This target can be surprisingly difficult to achieve, and offers a concrete way for theoretical research to clarify and extend the limits of ambitious mechanistic interpretability.