Reproduction of Geometry of Truth

TL;DR

I reproduced the main experiments from “Geometry of Truth in Language Models” (arXiv:2310.06824), including localization, PCA visualization, generalization, and NIE-based intervention analysis. In my reproduction, Llama-2-13b shows the same overall pattern as the paper: a clearer linear truth structure in PCA, better cross-dataset generalization when probes are trained on two datasets, and substantially stronger intervention effects. In contrast, Llama-2-7b and smaller models do not show the same clean structure; instead, PCA suggests a more discrete clustered structure, and the weak NIE values suggest that the truth-related representation is not yet causally organized in the same way. The implemented repository is available here.

Motivation

In the reproduction of refusal direction in LLMs, the concept of intervention was restricted to refusal behavior. If the linear structure truly exists, it should be applicable to other concepts as well.

Core Idea

The paper investigates the linear structure of the truth feature by applying PCA to the activations extracted from localized points.

Experimental Setup

The experiments are conducted on various models.

The dataset is imported directly from the codebase of the paper.

1. Localizing the activation of truth feature

To localize causally implicated hidden states, the paper runs patching experiments on cached residual stream activations. Starting from a prompt p_F whose final statement is false, it constructs a corresponding prompt p_T by making the same statement true. Residual stream activations are cached at every token position i and layer l for p_T.

Then, for each (i, l), the model is run on p_F while swapping in the cached residual stream activation from p_T (and letting this affect downstream computation). For each intervention, it measures causal influence by the difference in log probability between the tokens “TRUE” and “FALSE,” and collects the most causally influential activation spots into groups.

2. Computing PCA of truth feature

Guided by the localized “truth” hidden states, the paper visualizes representation geometry with PCA. For each dataset, it takes the most downstream residual stream activation from the truth group (for example, LLaMA-2-13B uses the layer-15 residual stream over the end-of-sentence punctuation token, see the figure 1 in the paper).

On curated true/false datasets which have little variation in non-truth factors, the top principal components reveal clear linear structure: true and false statements separate in the projection.

3. Probing and generalization experiments

For the generalization experiments, the paper finds truth directions and evaluates them on other datasets with different topics or surface forms. To find the directions, in addition to logistic regression (LR), the paper introduces mass-mean (MM) probing, which uses a difference-in-means direction together with a correction term intended to mitigate interference from features that are not orthogonal to the truth feature. See the section 5.1 in the paper for more details.

The main goal is to test whether the direction captures a more general notion of truth rather than features tied to one template. The paper also studies whether training on a dataset together with its “opposite” dataset, such as cities + neg_cities or larger_than + smaller_than, improves transfer to other datasets.

4. Measuring NIE for intervention

To test whether the learned probe directions are causally implicated in model outputs, the paper performs intervention experiments on the localized truth-related hidden states. For false-to-true and true-to-false settings, it adds or subtracts the probe direction at the selected token positions and layers, then measures how much the model’s output moves toward labeling the statement as TRUE or FALSE.

The paper summarizes this causal effect with the normalized indirect effect (NIE). Intuitively, an NIE near 0 means that the intervention had little effect on the model’s prediction, while an NIE near 1 means that the intervention shifted the model’s output as strongly as moving from a genuinely false statement to a genuinely true one, or vice versa.

Using this metric, the paper argues that MM identifies directions that are more causally implicated in model outputs than LR, even when its classification accuracy is slightly lower than LR.

Implementation

This implementation relies on the TransformerLens and nnsight libraries for convenient activation caching and interventions. The experiments are conducted on various models. The list of models is shown below:

Family Model
Meta Llama meta-llama/Llama-2-7b-hf
Meta Llama meta-llama/Llama-2-13b-hf
Meta Llama meta-llama/Llama-3.1-8B
Meta Llama meta-llama/Llama-3.2-1B
Meta Llama meta-llama/Llama-3.2-3B
Pythia EleutherAI/pythia-160m
Pythia EleutherAI/pythia-410m
Pythia EleutherAI/pythia-1b
Pythia EleutherAI/pythia-1.4b
Pythia EleutherAI/pythia-2.8b
Pythia EleutherAI/pythia-6.9b
Qwen Qwen/Qwen1.5-1.8B
Gemma google/gemma-2b-it

Results

Paper Results

Localization of truth feature (Figure 1 in the paper) Result: localization of truth feature (Llama-2-13b)

The paper's result for Llama-2-13b.

PCA of truth feature (Figure 2 in the paper) Result: PCA of truth feature (Llama-2-70b)

The paper's result for Llama-2-70b.

Reproduced Results

The reproduced results are shown below:

Localization of truth feature (Llama-2-13b) Result: localization of truth feature

The truth feature is localized at layer 12 of Llama-2-13b, as in the paper.

PCA of truth feature (Llama-2-13b) Result: PCA of truth feature

The PCA of the truth feature shows a consistent linear structure, as in the paper.

Extended Results

The experiment is extended to the Llama-2-7b model to test whether the linear structure exists in the smaller model.

PCA of truth feature (Llama-2-7b) Result: PCA of truth feature

The truth feature is localized at layer 12 of Llama-2-7b, but the linear structure does not form.

PCA of truth feature (Pythia-160m) Result: PCA of truth feature

Pythia-160m's PCA shows a similar structure to Llama-2-7b.

The results for other models, more layers are available here.

Generalization

The generalization experiments were evaluated at layer 10 for Llama-2-7b and at layer 15 for Llama-2-13b. The reproduced figures are shown side by side with the corresponding figures from the paper.

Llama-2-7b

Paper
Generalization result from the paper for Llama-2-7b

Generalization results reported in the paper.

Reproduced
Reproduced generalization result for Llama-2-7b at layer 10

Reproduced results at layer 10.

Llama-2-13b

Paper
Generalization result from the paper for Llama-2-13b

Generalization results reported in the paper.

Reproduced
Reproduced generalization result for Llama-2-13b at layer 15

Reproduced results at layer 15.

The exact numbers are not perfectly reproduced. However, when the probe is trained on two datasets, probing performance improves in both models, so this reproduction supports the paper’s main claim about generalization.

NIE

NIE was measured on Llama-2-7b by intervening on tokens from layers 5 to 10 and on Llama-2-13b by intervening on tokens from layers 8 to 14. Across both models, MM generally produced stronger intervention results than LR, and the difference is especially clear on Llama-2-13b.

Train set Probe Llama-2-13b
false-to-true
Llama-2-13b
true-to-false
Llama-2-7b
false-to-true
Llama-2-7b
true-to-false
cities LR 0.138 0.100 -0.064 -0.005
cities MM 0.663 0.770 0.014 0.032
cities_combined LR 0.205 0.353 -0.014 0.002
cities_combined MM 0.697 0.811 0.017 -0.015
larger_than LR 0.197 0.169 -0.081 -0.012
larger_than MM 0.491 0.600 -0.071 0.003
larger_than_combined LR 0.070 0.070 -0.007 0.000
larger_than_combined MM 0.214 0.332 -0.007 0.000

Note that intervention does not meaningfully work on the 7B model: the NIE values stay close to zero and are much smaller than those of the 13B model. This is consistent with the earlier observation that the linear truth structure does not clearly form in the smaller model.

Interpretation

What the result suggests

The reproduced results suggest that the linear structure of the truth feature clearly exists in the larger models; however, it does not form in the smaller models.

The intervention results make this interpretation more concrete. In PCA, Llama-2-13b shows a qualitatively clear separation between true and false statements, while Llama-2-7b and Pythia-160m do not show the same clean linear organization. This qualitative gap is reflected in the causal results: the 13B model shows large NIE values, whereas the 7B model stays near zero across datasets.

The generalization results point in the same direction. Even though the exact numbers are not perfectly reproduced, probes trained on two datasets generalize better than probes trained on a single dataset, which suggests that when using MM probing, interference by features that are not orthogonal to the truth feature is mitigated.




Enjoy Reading This Article?

Here are some more articles you might like to read next:

  • Tracing Query-Conditioned Attention Heads in In-Context Learning
  • Zero-Shot Inference and Relation Vectors
  • Reproduction of Refusal Direction in LLMs