Neural Mechanics: Symmetry and Broken Conservation Laws in Deep Learning Dynamics
Daniel Kunin, J. Sagastuy-Brena, Surya Ganguli, D. L. K. Yamins, Hidenori Tanaka
Mechanistic Swarm Interpretability: What Is the Agent Condition? · Part 1
This post opens our series on Mechanistic Swarm Interpretability: understanding how the personas of individual agents and their interactions shape collective personas and behavior.
Blog Author: Hidenori Tanaka · Sep 2026
From steam engines to transistors, the history of industrial revolutions is a history of understanding and harnessing emergence. We build a ladder of conceptual frameworks, from neurons to personas to swarms, to understand and align emerging superintelligence.
What laws govern how neural networks learn?
We study how symmetry and symmetry breaking shape learning dynamics, and how concepts and algorithms emerge and compete during training.
Daniel Kunin, J. Sagastuy-Brena, Surya Ganguli, D. L. K. Yamins, Hidenori Tanaka
Hidenori Tanaka, Daniel Kunin
Yongyi Yang, Core Francisco Park, Ekdeep Singh Lubana, Maya Okawa, Wei Hu, Hidenori Tanaka
Maya Okawa, Ekdeep Singh Lubana, R. P. Dick, Hidenori Tanaka
Core Francisco Park, Ekdeep Singh Lubana, Itamar Pres, Hidenori Tanaka
What structures underlie model behavior, and how can we steer it?
We seek low-dimensional structures underlying model behavior and study how they change under fine-tuning, in-context prompting, and activation steering.
Ekdeep Singh Lubana, Eric Bigelow, R. P. Dick, David Krueger, Hidenori Tanaka
Core Francisco Park, Andrew Lee, Ekdeep Singh Lubana, Yongyi Yang, Maya Okawa, Kento Nishi, Martin Wattenberg, Hidenori Tanaka
Core Francisco Park, Maya Okawa, Andrew Lee, Hidenori Tanaka, Ekdeep Singh Lubana
Equal advising: Hidenori Tanaka and Ekdeep Singh Lubana.
Eric J. Bigelow, Ekdeep Singh Lubana, Robert P. Dick, Hidenori Tanaka, Tomer D. Ullman
Equal contribution: Hidenori Tanaka and Tomer D. Ullman.
Kento Nishi, Rahul Ramesh, Maya Okawa, Mikail Khona, Hidenori Tanaka, Ekdeep Singh Lubana
Equal contribution: Hidenori Tanaka and Ekdeep Singh Lubana.
Ultimately, the success of Physics of Intelligence should be measured by its ability to forecast the future. We aim to make concrete, qualitative predictions about emerging AI safety risks, drawing on a scientific foundation of intelligence to make the unimaginable intelligible.
Jun 2026
Interactive
An interactive exhibit on body, alarm, attention, and action. Developed under the supervision of Prof. Mai Uchida and shown at a Boston Museum of Science event.
Mechanistic Swarm Interpretability
How do individual agents and their interactions shape collective beliefs and behavior?
We trace how messages change individual states, how those changes propagate, and which interventions alter collective behavior.
When Is Collective Intelligence a Lottery? Multi-Agent Scaling Laws for Memetic Drift in LLMs
Hidenori Tanaka
Hidenori has previously worked on the evolutionary dynamics of self-replicating artificial materials and the safety challenges of containing CRISPR gene drives.
H. Tanaka, H. A. Stone, D. R. Nelson
PNAS · 2017
H. Tanaka, Z. Zeravcic, M. P. Brenner
Physical Review Letters · 2016 · Editors’ Suggestion