Reproduction of Geometry of Truth
TL;DR
I reproduced the main experiments from “Geometry of Truth in Language Models” (arXiv:2310.06824), including localization, PCA visualization, generalization, and NIE-based intervention analysis. In my reproduction, Llama-2-13b shows the same overall pattern as the paper: a clearer linear truth structure in PCA, better cross-dataset generalization when probes are trained on two datasets, and substantially stronger intervention effects. In contrast, Llama-2-7b and smaller models do not show the same clean structure; instead, PCA suggests a more discrete clustered structure, and the weak NIE values suggest that the truth-related representation is not yet causally organized in the same way. The implemented repository is available here.
Motivation
In the reproduction of refusal direction in LLMs, the concept of intervention was restricted to refusal behavior. If the linear structure truly exists, it should be applicable to other concepts as well.
Core Idea
The paper investigates the linear structure of the truth feature by applying PCA to the activations extracted from localized points.
Experimental Setup
The experiments are conducted on various models.
The dataset is imported directly from the codebase of the paper.
1. Localizing the activation of truth feature
To localize causally implicated hidden states, the paper runs patching experiments on cached residual stream activations. Starting from a prompt p_F whose final statement is false, it constructs a corresponding prompt p_T by making the same statement true. Residual stream activations are cached at every token position i and layer l for p_T.
Then, for each (i, l), the model is run on p_F while swapping in the cached residual stream activation from p_T (and letting this affect downstream computation). For each intervention, it measures causal influence by the difference in log probability between the tokens “TRUE” and “FALSE,” and collects the most causally influential activation spots into groups.
2. Computing PCA of truth feature
Guided by the localized “truth” hidden states, the paper visualizes representation geometry with PCA. For each dataset, it takes the most downstream residual stream activation from the truth group (for example, LLaMA-2-13B uses the layer-15 residual stream over the end-of-sentence punctuation token, see the figure 1 in the paper).
On curated true/false datasets which have little variation in non-truth factors, the top principal components reveal clear linear structure: true and false statements separate in the projection.
3. Probing and generalization experiments
For the generalization experiments, the paper finds truth directions and evaluates them on other datasets with different topics or surface forms. To find the directions, in addition to logistic regression (LR), the paper introduces mass-mean (MM) probing, which uses a difference-in-means direction together with a correction term intended to mitigate interference from features that are not orthogonal to the truth feature. See the section 5.1 in the paper for more details.
The main goal is to test whether the direction captures a more general notion of truth rather than features tied to one template. The paper also studies whether training on a dataset together with its “opposite” dataset, such as cities + neg_cities or larger_than + smaller_than, improves transfer to other datasets.
4. Measuring NIE for intervention
To test whether the learned probe directions are causally implicated in model outputs, the paper performs intervention experiments on the localized truth-related hidden states. For false-to-true and true-to-false settings, it adds or subtracts the probe direction at the selected token positions and layers, then measures how much the model’s output moves toward labeling the statement as TRUE or FALSE.
The paper summarizes this causal effect with the normalized indirect effect (NIE). Intuitively, an NIE near 0 means that the intervention had little effect on the model’s prediction, while an NIE near 1 means that the intervention shifted the model’s output as strongly as moving from a genuinely false statement to a genuinely true one, or vice versa.
Using this metric, the paper argues that MM identifies directions that are more causally implicated in model outputs than LR, even when its classification accuracy is slightly lower than LR.
Implementation
This implementation relies on the TransformerLens and nnsight libraries for convenient activation caching and interventions. The experiments are conducted on various models. The list of models is shown below:
| Family | Model |
|---|---|
| Meta Llama | meta-llama/Llama-2-7b-hf |
| Meta Llama | meta-llama/Llama-2-13b-hf |
| Meta Llama | meta-llama/Llama-3.1-8B |
| Meta Llama | meta-llama/Llama-3.2-1B |
| Meta Llama | meta-llama/Llama-3.2-3B |
| Pythia | EleutherAI/pythia-160m |
| Pythia | EleutherAI/pythia-410m |
| Pythia | EleutherAI/pythia-1b |
| Pythia | EleutherAI/pythia-1.4b |
| Pythia | EleutherAI/pythia-2.8b |
| Pythia | EleutherAI/pythia-6.9b |
| Qwen | Qwen/Qwen1.5-1.8B |
| Gemma | google/gemma-2b-it |
Results
Paper Results
The paper's result for Llama-2-13b.
The paper's result for Llama-2-70b.
Reproduced Results
The reproduced results are shown below:
The truth feature is localized at layer 12 of Llama-2-13b, as in the paper.
The PCA of the truth feature shows a consistent linear structure, as in the paper.
Extended Results
The experiment is extended to the Llama-2-7b model to test whether the linear structure exists in the smaller model.
The truth feature is localized at layer 12 of Llama-2-7b, but the linear structure does not form.
Pythia-160m's PCA shows a similar structure to Llama-2-7b.
The results for other models, more layers are available here.
Generalization
The generalization experiments were evaluated at layer 10 for Llama-2-7b and at layer 15 for Llama-2-13b. The reproduced figures are shown side by side with the corresponding figures from the paper.
Llama-2-7b
Generalization results reported in the paper.
Reproduced results at layer 10.
Llama-2-13b
Generalization results reported in the paper.
Reproduced results at layer 15.
The exact numbers are not perfectly reproduced. However, when the probe is trained on two datasets, probing performance improves in both models, so this reproduction supports the paper’s main claim about generalization.
NIE
NIE was measured on Llama-2-7b by intervening on tokens from layers 5 to 10 and on Llama-2-13b by intervening on tokens from layers 8 to 14. Across both models, MM generally produced stronger intervention results than LR, and the difference is especially clear on Llama-2-13b.
| Train set | Probe | Llama-2-13b false-to-true | Llama-2-13b true-to-false | Llama-2-7b false-to-true | Llama-2-7b true-to-false |
|---|---|---|---|---|---|
cities | LR | 0.138 | 0.100 | -0.064 | -0.005 |
cities | MM | 0.663 | 0.770 | 0.014 | 0.032 |
cities_combined | LR | 0.205 | 0.353 | -0.014 | 0.002 |
cities_combined | MM | 0.697 | 0.811 | 0.017 | -0.015 |
larger_than | LR | 0.197 | 0.169 | -0.081 | -0.012 |
larger_than | MM | 0.491 | 0.600 | -0.071 | 0.003 |
larger_than_combined | LR | 0.070 | 0.070 | -0.007 | 0.000 |
larger_than_combined | MM | 0.214 | 0.332 | -0.007 | 0.000 |
Note that intervention does not meaningfully work on the 7B model: the NIE values stay close to zero and are much smaller than those of the 13B model. This is consistent with the earlier observation that the linear truth structure does not clearly form in the smaller model.
Interpretation
What the result suggests
The reproduced results suggest that the linear structure of the truth feature clearly exists in the larger models; however, it does not form in the smaller models.
The intervention results make this interpretation more concrete. In PCA, Llama-2-13b shows a qualitatively clear separation between true and false statements, while Llama-2-7b and Pythia-160m do not show the same clean linear organization. This qualitative gap is reflected in the causal results: the 13B model shows large NIE values, whereas the 7B model stays near zero across datasets.
The generalization results point in the same direction. Even though the exact numbers are not perfectly reproduced, probes trained on two datasets generalize better than probes trained on a single dataset, which suggests that when using MM probing, interference by features that are not orthogonal to the truth feature is mitigated.
Enjoy Reading This Article?
Here are some more articles you might like to read next: