Mechanistic Interpretability: Reverse-Engineering the Algorithms Inside Neural Networks
Take a simple completion: “Alice gave Bob the book because wanted it.” Suppose the model assigns a high target logit $y$ to the corr...
AI Interpretabilitymechanistic interpretabilitytransformer circuits