PIRAMID: Building Scientific Foundations for Mechanistic Interpretability

PIRAMID uses the tools of statistical physics to create a scientific, theory-driven approach to mechanistic interpretability, aiming for faithful and scalable AI transparency.

MiHiR SEN
MiHiR SEN
·4 min read
PIRAMID is a research division using statistical physics to build a scientific foundation for mechanistic interpretability. By dividing its efforts across theory, applications, and data validation, it aims to create faithful, scalable tools for understanding and intervening in AI systems.

Mechanistic interpretability aims to reverse-engineer the internal workings of neural networks. It's a crucial field for AI safety, as understanding how models think is the first step to ensuring they align with human values. But a persistent challenge is the gap between theory and practice. We need more than persuasive, ad-hoc explanations of model behavior. We need tools grounded in a scientific understanding of data, learning, and representations. That's the driving principle behind PIRAMID, an internal research division using the tools and techniques of statistical physics to build scientific foundations for ambitious mechanistic interpretability.

A Three-Pronged Research Loop

PIRAMID's central premise is that scalable alignment requires interpretability tools that develop alongside a scientific understanding of structure. To reflect this, they divide their attention across three synergistic research teams, forming a loop that mirrors physics' division of labor: theory, empirics, and phenomenology.

The Advancements in Learning Theory team aims to produce an end-to-end toy-to-real case study where a theoretically motivated statistic predicts and explains generalization within the next few months. Their core hypothesis is that the microscopic state of a trained network is too irregular to reason about directly, while coarse loss or benchmark performance measures obscure the structure we care about. Instead, they consider aggregate statistical measures of neural network distributions as the correct level of abstraction: coarse enough to theoretically describe regular structure, fine enough to uncover mechanistic detail that can guide interpretability work. They spend significant time thinking about how to build theories that are "structurally portable."

The Interpretability Applications team's goal is to develop a suite of tools that recover mechanistic structure which reflects the causal and hierarchical nature of learning and computation. These tools are physics-informed in the sense of being grounded in theoretical hypotheses about how networks learn structure across scales. They will validate progress on pragmatic downstream tasks like data attribution, backdoor detection, and sandbagging, as well as intermediate measures of faithfulness grounded in a physics-informed understanding. Tool failures, in turn, reveal where theories or data models are incomplete.

Finally, the Data Models and Validation Methods team aims to construct analytically tractable, physics-inspired data models that capture key aspects of natural data, such as hierarchy, sparsity, criticality, and power-law statistics. Because the latent structure is known by construction in these synthetic datasets, researchers can test whether theory predicts, or tools recover, the structure the model actually learns, rather than a plausible but unfaithful explanation. Within the next year, they aim to turn these datasets into a public benchmark with stronger faithfulness guarantees than current evaluation practice.

Why Statistical Physics?

Physicists often identify the variables that matter most at a particular "level of description," quantities that govern large-scale behavior like phase transitions and other emergent properties. This results in treating model components as "mesoscopic," in the sense of an intermediate description. The presence of structurally relevant randomness, for which statistical physics is the canonical framework, is a common thread across PIRAMID's research groups.

Statistical physics gives us the language to formalize and the tools to track the mechanistic role of this randomness. Instead of accounting for every microscopic detail, it allows us to focus on the components and interactions that matter most. This aligns with PIRAMID's core hypothesis that relevant structure does not live at one fixed level of description. It may appear as local features, global directions, hierarchical latent variables, circuits, training phases, basins, or weight-space perturbations. Some details matter at one scale and wash out at another, and some mechanisms are distributed across many units rather than localized in a single component.

The Operative Target: Faithfulness

While full reverse engineering of an AI system is neither necessary nor sufficient for safety, PIRAMID sees it as a proxy for the kind of faithful mechanistic transparency that would make scalable alignment more feasible. A faithful explanation must track the mechanism the model actually learns and uses, not just correlate with behavior. Future systems may differ from today's models in terms of architecture, continual learning, or memory. A method tied only to today's empirical probes may fail when the model generalizes in a new way or moves into a regime where our probes no longer behave as expected. However, one grounded in broader principles governing learning and computation has a better chance of transferring.

The physics-informed approach aims to guide us toward principled definitions of faithfulness and a structural understanding of what a network learns, grounded in data, training dynamics, and representational geometry. PIRAMID is actively building this bridge, connecting statistical physics with AI interpretability to create a field that can rise to the challenge of building safe and reliable AI systems. Their goal is to make AI systems sufficiently transparent to support high-confidence, well-founded interventions.