Word
mechanistic interpretability
Definition
n. the research field that tries to understand AI models by taking apart their internal circuits, like reverse-engineering computer code.
n. a subfield of machine learning research dedicated to reverse-engineering the internal parameters and computational circuits of neural networks into human-understandable algorithms.
Examples
“Researchers use mechanistic interpretability to see exactly how a transformer solves a math puzzle.”
“Safety laboratories rely on mechanistic interpretability to detect whether an autonomous system is harboring deceptive intentions.”
“By treating neural weights as compiled binary programs, mechanistic interpretability seeks to uncover discrete computational circuits responsible for in-context learning.”
Examples
simple
“Researchers use mechanistic interpretability to see exactly how a transformer solves a math puzzle.”
contextual
“Safety laboratories rely on mechanistic interpretability to detect whether an autonomous system is harboring deceptive intentions.”
complex
“By treating neural weights as compiled binary programs, mechanistic interpretability seeks to uncover discrete computational circuits responsible for in-context learning.”
Real-World Examples
“Anthropic is a leader in this effort to bring to light models' internal deliberations, called mechanistic interpretability, a deceptively boring designation for a critical task.” “Using a method known as mechanistic interpretability, they trained a smaller model to recognize telltale activations across the agents' weights.” Etymology
Coined in the late 2010s by machine learning researchers including Chris Olah, combining mechanistic (from Greek mekhane, 'machine') with interpretability.
Etymology adapted from Wiktionary, available under CC BY-SA 4.0.