Embodied AI Glossary中文

Interpretability

可解释性(机制可解释性)Advanced

Studying how a neural network's internals produce its output; mechanistic interpretability breaks that computation into human-understandable pieces.

A deep network has hundreds of millions of parameters, and the computation from input to output is unreadable to a person. Interpretability research tries to answer 'why did the model produce this output,' with methods ranging from attention heatmaps and saliency maps to mechanistic interpretability: reverse-engineering the network, like a program, to find which concepts (features) it represents internally and how those features wire together into circuits that carry out a computation. One difficulty is that a single neuron often participates in representing several concepts at once; in May 2024, Anthropic used dictionary learning (a sparse decomposition method) to extract millions of features from Claude 3 Sonnet, including one that activates for the text and images of 'Golden Gate Bridge' across many languages. For robots, this matters for debugging and safety: in 2025 a Berkeley team found directions inside π0-FAST and OpenVLA corresponding to concepts like 'fast/slow' and 'high/low,' and adjusting these activations at inference time changes the robot's actions in real time, with no fine-tuning or reward signal needed.

ExampleThe Berkeley team ran π0-FAST on a UR5 arm carrying a toy: amplifying the internal activation associated with 'fast' made the arm move faster, and amplifying the one associated with 'low' lowered the peak height of the carrying trajectory, all without ever changing the model's weights.

Also called
Mechanistic Interpretability, Explainable AI
Related
Vision-Language-Action Model · Large Language Model · Neural Network · Embodied Safety · Uncertainty Estimation
Sources
Mapping the Mind of a Large Language Model(Anthropic, 2024-05)
Mechanistic interpretability for steering vision-language-action models (arXiv:2509.00328)
As of
2025-08

See it in the full glossary →