Mahdi Soltanolkotabi - Interpreting Generative Models for Better Steering: From Visual Generation to Verifiable Reasoning
September 3, 2026
Abstract
Despite their impressive capabilities, generative models continue to struggle with basic forms of visual reasoning, including reliably generating a specified number of objects. We study the internal dynamics of diffusion models and how visual concepts and scene structure emerge over the denoising trajectory. These insights lead to an early-time steering method that intervenes when global semantic content is first formed. We then study reinforcement learning with verifiable rewards and show how asymmetrically weighting positive and negative feedback can improve visual reasoning. Our results demonstrate how interpretability can provide a principled foundation for steering visual generation and enhancing visual and verifiable reasoning.