-
Tracing Query-Conditioned Attention Heads in In-Context Learning
Research notes on identifying and interpreting query-conditioned attention heads in transformer language models.
-
Zero-Shot Inference and Relation Vectors
Applying the Function Vectors methodology to zero-shot inference and investigating why the relation vector fails on incorrectly answered examples.
-
Reproduction of Geometry of Truth
Notes on reproducing "Geometry of Truth" with causal intervention and visualization.
-
Reproduction of Refusal Direction in LLMs
Notes on reproducing "Refusal in Language Models Is Mediated by a Single Direction" with causal intervention on the residual stream.