What is the purpose of interpretability?

2026-07-13

I gave a five-minute lightning talk at the Mechanistic Interpretability Workshop at ICML this year in Seoul. Many thanks to the organizers, who encouraged us to speak about our "high-level vision" and "what the field should be doing differently". My talk ended up being indeed quite high-level and personal. I've reproduced a slightly expanded version of it below.

Title slide: What is the purpose of interpretability?

Today I'll share some very high-level thoughts on my hopes for the field of interpretability—what we could become, or fail to become, in the coming years. Many of these thoughts are what come up for me when reflecting on the call for "pragmatic interpretability" which was of course the topic of so much discussion at this workshop at NeurIPS in December.1

Slide: 'Interpretability' is a very deep science

I'll start with something I believe, which is that the task of understanding neural networks is a very deep science. I believe that if we froze AI progress, we could study today's models for decades, possibly centuries, and that that project would become as culturally significant as any other field of basic science.

Slide: Two kinds of understanding

Why is there a deep science here? Well, I think there are two kinds of understanding one could want of neural networks, which together would constitute a "theory" of deep learning.

The first kind of understanding is to describe what algorithms our networks have learned. This is the conventional task of interpretability. We might even want to ambitiously2 describe, in tremendous detail, these algorithms!

But the second kind of understanding is to have good explanations for why a network has learned what it has learned.3 Part of this story will be about optimization and learning dynamics,4 but I think many of us also expect it to be a story about the structure of data, and how some essential aspects of the structure of data determine the internal structure that networks must learn in order to optimally predict that data. This is of course the perspective of many on the Simplex team.5

Slide: The structure of 'data' relates to the structure of the world and of our minds

But then things become really interesting, especially for language models trained on text, because language is very close to thought.6 The data generating process for language is the process of human minds thinking. And not even a single person thinking, but rather all the humans across history who have contributed to our corpus of text. And so this question of understanding why networks learn what they learn, in some manner, is entangled with facts about the structure of language, and of thought, and maybe even facts about cultural evolution.

For instance, one class of theories for the origin of neural scaling laws is that they come from some power-law distribution over how often different concepts are present in the data.7 Now, the occurrence frequency of a concept in data might seem like a mundane fact, but for a concept to be present in a text corpus means it was present in some person's mind when they produced that text. So this is ultimately a claim about something like the relative abundance at which different memes8 have been replicated across minds, and how often they contribute to thought in those minds.9 This may not ultimately be the right way of thinking about the structure of the data that gives rise to neural scaling laws, but whatever the right story is, I suspect it will teach us something new about the world.10

And even if there isn't a close correspondence between the internal computations that language models learn and the computations that underlie our own thought, that would present its own opportunities for understanding something about the diversity of possible minds in the world.

Slide: Charles Darwin and his notebook sketch of the tree of life

I expect there are Darwin-sized contributions to be made in putting these pieces together, in working out a more universal cognitive science. But Darwin spent years observing the diversity of life and reflecting on those observations. Imagine telling him to "time-box his curiosity" to two weeks.11 It would make me very sad if someone who might be in a position to make a contribution of this nature were dissuaded from doing that work on the grounds it is not pragmatic.

Slide: Dr. Louise Banks at the whiteboard in Arrival (2016), pointing at 'What is your purpose on Earth'

I really love the movie Arrival12,13, which is about scientists trying to communicate with aliens who have landed on Earth. There's a scene where Dr. Banks here is criticized for moving too slowly—for starting with such basic words instead of jumping to ask the question those around her pragmatically care about: "What is your purpose on Earth?" She explains that the aliens would first need to understand what a question is, and the idea of a purpose or a goal, and that we would need to understand enough of their language to understand their answer.

It strikes me that the safety question that some of us are interested in answering for AI systems today is similar: "is this system aligned", "what are its values", "what is its purpose on Earth?" But correctly asking and answering this question seems wrapped up in many subtle questions about the cognition of AI systems and the nature of their beliefs and goals. And I worry that if we don't address these more basic parts of the question first, instead jumping too far ahead too quickly, the field will continue to flail around. Trying to take shortcuts in solving a very deep scientific problem is a recipe for chaos and stagnation and cuts against our pragmatic goals.

Slide: PCA projections of transformer embeddings trained on modular addition, juxtaposed with heptapod logograms from Arrival
PCA projections of transformer embeddings across many training runs on modular addition versus the glyphs of the heptapod language in Arrival. To be clear, this comparison is obviously completely scientifically meaningless—it is just meant to be evocative. Note that I manually (randomly) varied the relative sizes of the circular representations when making this image.14

In the story of Arrival, the process of learning the alien language ultimately had implications far beyond the pragmatic goal that motivated that effort. The language transforms how the learner sees the world. My hope is that interpretability will do something similar, if only we'd let it.

This talk/post was influenced by conversations with many folks, including David Bau, Atticus Geiger, and Adam Shai and the rest of the Simplex team. I especially want to thank David for his "In defense of curiosity" talk at the previous Mechanistic Interpretability Workshop at NeurIPS, and Adam for his "Don't give up on ambitious interpretability" post.

Changes: July 24, 2026—removed a sentence and added a reference to Saphra's "Interpretability creationism".

Notes and References

  1. Nanda et al. (2025). A pragmatic vision for interpretability. Alignment Forum. [link] ↩︎
  2. Gao, L. (2025). An ambitious vision for interpretability. Alignment Forum. [link]
    Shai, A. (2025). Don't give up on ambitious interpretability. Blog post. [link] ↩︎
  3. My friend Jamie Simon joked that this could be called something like "super ambitious interpretability" or "even more ambitious interpretability". ↩︎
  4. Simon, Kunin, et al. (2026). There will be a scientific theory of deep learning. arXiv preprint arXiv:2604.21691. [link]
    See also the developmental interpretability agenda:
    Hoogland, J., Gietelink Oldenziel, A., Murfet, D., & van Wingerden, S. (2023). Towards developmental interpretability. AI Alignment Forum. [link]
    and of course:
    Saphra, N. (2022). Interpretability creationism. Blog post. [link] ↩︎
  5. See e.g. Shai, A. S., Marzen, S. E., Teixeira, L., Gietelink Oldenziel, A., & Riechers, P. M. (2024). Transformers represent belief state geometry in their residual stream. arXiv preprint arXiv:2405.15943. [link] ↩︎
  6. Though cf. Fedorenko, E., Piantadosi, S. T., & Gibson, E. A. F. (2024). Language is primarily a tool for communication rather than thought. Nature, 630, 575–586. [link] ↩︎
  7. See my post On neural scaling and the quanta hypothesis, and the extended related work discussion there. ↩︎
  8. Dawkins, R. (1976). The selfish gene. Oxford University Press. ↩︎
  9. See also A short note on interpretability and minds. ↩︎
  10. Though for a more sober perspective, see:
    Barkeshli, M., Alfarano, A., & Gromov, A. (2026). On the origin of neural scaling laws: from random graphs to natural language. arXiv preprint arXiv:2601.10684. [link] ↩︎
  11. This is a somewhat unfair response to the "pragmatic interpretability" post, which in the details is a pretty reasonable document and only argues that more folks on the margin should orient their work towards concrete pragmatic goals, not everyone. ↩︎
  12. Villeneuve, D. (2016). Arrival. Based on Ted Chiang's Story of Your Life (1998). ↩︎
  13. I loved David Bau's metaphor of Venetian glassmakers in his In defense of curiosity. I'm using Arrival here to make a similar point. ↩︎
  14. These transformers were trained with either the setup used in
    Liu, Z., Kitouni, O., Nolte, N., Michaud, E. J., Tegmark, M., & Williams, M. (2022). Towards understanding grokking: An effective theory of representation learning. NeurIPS 2022. [link]
    OR
    Nanda, N., Chan, L., Lieberum, T., Smith, J., & Steinhardt, J. (2023). Progress measures for grokking via mechanistic interpretability. ICLR 2023. [link]
    I forget which, I made this artsy plot at least a few years ago. ↩︎
@misc{michaud2026purpose,
  author = {Michaud, Eric J.},
  title = {What is the purpose of interpretability?},
  year = {2026},
  howpublished = {\url{https://ericjmichaud.com/interp-purpose/}},
}