Mechanistic Interpretability in Frontier LLMs: Reverse-Engineering Circuits and Sparse Autoencoders
Can we look inside the black box of transformer models? Mechanistic interpretability is using Sparse Autoencoders (SAEs) to decompose billions of polysemantic weights into human-interpretable circuits.









