Mechanistic Interpretability

Parent: LLM Models and APIs · researched 2026-06-03T00:22:13.503Z· 16 sources · 12 concepts · skill mechanistic-interpretability

Reverse-engineering the internal computation of neural networks (chiefly transformer LLMs) into human-understandable mechanisms — the features a model represents and the circuits that combine them. A

Mechanistic Interpretability — SAEs, Circuits, Steering

1. The goal and the two paradigms

2. Superposition and polysemanticity

3. Sparse autoencoders (SAEs) and dictionary learning

4. The SAE critique wave (you MUST carry this)

5. Transcoders and crosscoders

6. Circuit discovery and causal analysis

7. Attribution graphs and circuit tracing

8. Lenses and probing

9. Activation steering and representation engineering

10. Auto-interpretability and evaluation

11. Tooling stack

12. Interpretability for alignment and safety

Sources

Children

Frontier under this node: Activation steering & representation engineering (RepE, ActAdd), Attribution graphs & circuit tracing (Biology of an LLM), Auto-interpretability & evaluation (SAEBench, Neuronpedia), Circuit discovery (induction heads, IOI, activation/path patching, ACDC, EAP), Interpretability for alignment & safety (auditing, deception probes, the interpretability illusion), Logit lens / tuned lens / linear probes / concept erasure, Sparse autoencoders & dictionary learning (Gated/TopK/JumpReLU, Gemma Scope), Sparse feature circuits, Superposition & polysemanticity (linear representation hypothesis), The SAE critique wave (random-Transformer baselines, non-canonical, deprioritization), Tooling (TransformerLens, SAELens, nnsight/NDIF, Neuronpedia, Gemma Scope), Transcoders & cross-layer transcoders/crosscoders

← the whole tree · 3D view· how to read this page