Paul Riechers - The shape of beliefs and abstraction in neural networks
September 4, 2026
Abstract
We find that next-token prediction induces predictable and interpretable low-dimensional fractal structures associated with beliefs and abstraction. We start simply with carefully crafted synthetic training data meant to test and falsify various alternatives. We progressively refine our ansatz for the shape of beliefs and abstraction, establishing the following five points: (i) Models learn to effectively perform Bayesian updates over the latent states of a world model during inference. (ii) This world model goes beyond the classical computational paradigm. (iii) Neural networks have an inductive bias to factor their world into approximately conditionally independent parts, and represent these parts in orthogonal subspaces. (iv) Multiple ergodic components induce sparsity. (v) Shared parts among distinct components induce abstraction. We leverage these insights from the toy-model setting to discover and surgically steer beliefs and abstractions in LLMs.