What is the purpose of interpretability?
2026-07-13
I gave a five-minute lightning talk at the Mechanistic Interpretability Workshop at ICML this year in Seoul. Many thanks to the organizers, who encouraged us to speak about our "high-level vision" and "what the field should be doing differently". My talk ended up being indeed quite high-level and personal. I've reproduced a slightly expanded version of it below.
This talk/post was influenced by conversations with many folks, including David Bau, Atticus Geiger, and Adam Shai and the rest of the Simplex team. I especially want to thank David for his "In defense of curiosity" talk at the previous Mechanistic Interpretability Workshop at NeurIPS, and Adam for his "Don't give up on ambitious interpretability" post.
Changes: July 24, 2026—removed a sentence and added a reference to Saphra's "Interpretability creationism".
Notes and References
- Nanda et al. (2025). A pragmatic vision for interpretability. Alignment Forum. [link] ↩︎
-
Gao, L. (2025). An ambitious vision for interpretability. Alignment Forum. [link]
Shai, A. (2025). Don't give up on ambitious interpretability. Blog post. [link] ↩︎ - My friend Jamie Simon joked that this could be called something like "super ambitious interpretability" or "even more ambitious interpretability". ↩︎
-
Simon, Kunin, et al. (2026). There will be a scientific theory of deep learning. arXiv preprint arXiv:2604.21691. [link]
See also the developmental interpretability agenda:
Hoogland, J., Gietelink Oldenziel, A., Murfet, D., & van Wingerden, S. (2023). Towards developmental interpretability. AI Alignment Forum. [link]
and of course:
Saphra, N. (2022). Interpretability creationism. Blog post. [link] ↩︎ - See e.g. Shai, A. S., Marzen, S. E., Teixeira, L., Gietelink Oldenziel, A., & Riechers, P. M. (2024). Transformers represent belief state geometry in their residual stream. arXiv preprint arXiv:2405.15943. [link] ↩︎
- Though cf. Fedorenko, E., Piantadosi, S. T., & Gibson, E. A. F. (2024). Language is primarily a tool for communication rather than thought. Nature, 630, 575–586. [link] ↩︎
- See my post On neural scaling and the quanta hypothesis, and the extended related work discussion there. ↩︎
- Dawkins, R. (1976). The selfish gene. Oxford University Press. ↩︎
- See also A short note on interpretability and minds. ↩︎
-
Though for a more sober perspective, see:
Barkeshli, M., Alfarano, A., & Gromov, A. (2026). On the origin of neural scaling laws: from random graphs to natural language. arXiv preprint arXiv:2601.10684. [link] ↩︎ - This is a somewhat unfair response to the "pragmatic interpretability" post, which in the details is a pretty reasonable document and only argues that more folks on the margin should orient their work towards concrete pragmatic goals, not everyone. ↩︎
- Villeneuve, D. (2016). Arrival. Based on Ted Chiang's Story of Your Life (1998). ↩︎
- I loved David Bau's metaphor of Venetian glassmakers in his In defense of curiosity. I'm using Arrival here to make a similar point. ↩︎
-
These transformers were trained with either the setup used in
Liu, Z., Kitouni, O., Nolte, N., Michaud, E. J., Tegmark, M., & Williams, M. (2022). Towards understanding grokking: An effective theory of representation learning. NeurIPS 2022. [link]
OR
Nanda, N., Chan, L., Lieberum, T., Smith, J., & Steinhardt, J. (2023). Progress measures for grokking via mechanistic interpretability. ICLR 2023. [link]
I forget which, I made this artsy plot at least a few years ago. ↩︎
@misc{michaud2026purpose,
author = {Michaud, Eric J.},
title = {What is the purpose of interpretability?},
year = {2026},
howpublished = {\url{https://ericjmichaud.com/interp-purpose/}},
}