<?xml version="1.0" encoding="utf-8"?><feed xmlns="http://www.w3.org/2005/Atom" ><generator uri="https://jekyllrb.com/" version="3.10.0">Jekyll</generator><link href="https://jasteinberg.github.io/feed.xml" rel="self" type="application/atom+xml" /><link href="https://jasteinberg.github.io/" rel="alternate" type="text/html" /><updated>2026-09-18T02:11:44+00:00</updated><id>https://jasteinberg.github.io/feed.xml</id><title type="html">Julia Steinberg</title><subtitle>Physics PhD working on alignment and interpretability.</subtitle><entry><title type="html">Truth Directions: Signal-to-Noise and Geometry of Recoverability</title><link href="https://jasteinberg.github.io/blog/2026/truth-directions-snr/" rel="alternate" type="text/html" title="Truth Directions: Signal-to-Noise and Geometry of Recoverability" /><published>2026-09-10T00:00:00+00:00</published><updated>2026-09-10T00:00:00+00:00</updated><id>https://jasteinberg.github.io/blog/2026/truth-directions-snr</id><content type="html" xml:base="https://jasteinberg.github.io/blog/2026/truth-directions-snr/"><![CDATA[<h2 id="introduction">Introduction</h2>

<p>The linear representation hypothesis states that a model encodes high-level
concepts as directions in activation space. This reduces reading a concept to a dot
product with a single vector, and steering along it to adding a multiple of that vector to the
residual stream. <a href="https://arxiv.org/abs/2310.06824">Marks &amp; Tegmark (2023)</a> showed that in sufficiently large models, a linear direction corresponding to an abstract notion of truth emerges that applies across diverse sets of inputs. They estimated this direction with a difference-in-means (“mass-mean”) direction fit on true/false statements and showed that it separates held-out statements, transfers across datasets, and is causally implicated under intervention.</p>

<p>Previously, <a href="https://arxiv.org/abs/2212.03827">Burns et al. (2022)</a> had proposed finding a
linear truth direction <em>without</em> labels, by demanding logical consistency, a method they call
Contrast-Consistent Search (CCS). However,
<a href="https://www.alignmentforum.org/posts/bWxNPMy5MhPnQTzKz/what-discovering-latent-knowledge-did-and-did-not-find-4">Roger (2023)</a> showed empirically that the CCS estimator insufficiently constrained.  Untrained, randomly initialized probes already reach about \(75\%\) accuracy on the “easy” dataset  once the CCS convention of inverting the sign of a below-chance probe is applied. More than twenty mutually orthogonal “truth probes” reach accuracies similar to the one returned by the CCS estimator implying that CSS has not found the optimal linear probe. Additionally, due to the inversion convention in choosing the sign, comparing to a random baseline is misleading because the accuracy of this baseline will always be either greater than or equal to one half. Furthermore,
<a href="https://arxiv.org/abs/2312.10029">Farquhar et al. (2023)</a> then showed that because
arbitrary binary features are optimal under that consistency loss, nothing in it
selects for knowledge, and in practice unsupervised probes recover whatever feature is
<em>most prominent</em> in the representation implying that in the unsupervised case a linear direction can appear to decode truth while actually corresponding to a salient direction that happens to correlate with the label on the particular dataset. Thus the common thread across these observations implies that the decoding of truth inherits a signal-versus-noise problem, which I will show is also relevant for the supervised case. Consequently, on a benchmark where chance-level structure is this strong, the ability of a truth probe to separate the classes is a much weaker result than it first appears.</p>

<div class="tldr gray">

  <p>A mass-mean “truth direction” is an estimator for the decodable truth direction which separates true and false statements. However, in the low signal-to-noise regime, it can also end up returning the most salient direction in activation sapce, i.e. the direction of largest within-class variance. On a dataset where the class gap is small relative to the within class spread along the axis connecting the centroids, the finite-sample mean difference is dominated by noise along that axis, aligning the estimator along the salient axis regardless of whether it has large overlap with the recoverable truth direction. On <code class="language-plaintext highlighter-rouge">counterfact</code>, the decodable truth direction does not lie along the salient axis, so the estimator returns a nuisance direction. A truth direction and a salient axis look alike on a benchmark and require a signal-to-noise reading to tell them apart. This post takes up the question of recoverability: <em>when</em> does a model contain a linear truth direction a probe can actually recover, which rises above a random-direction null and carries signal beyond the single most salient axis?</p>

  <p>Systems neuroscience has spent decades asking what a downstream reader can recover from a population of noisy neural units. In that spirit I treat probing for a truth direction as a readout problem: I take the probe as a linear readout, quantify its separation with a detection-theoretic \(d'\), and benchmark that \(d'\) against an explicit random-direction null across the Pythia scale ladder. This reading yields three results:</p>

  <p><strong>(1) Apparent separation is trivial.</strong> Two effects inflate a probe’s score before any truth content enters. Fitting below Cover’s capacity (\(N \ll 2d\) throughout) buys separation from shuffled labels alone, and a random direction inherits a share of the real class gap, so the chance level rises with the very separation being measured. The null, not the raw score, is the bar.</p>

  <p><strong>(2) The estimator returns a nuisance direction when the signal is weak.</strong> It returns the dominant activation axis \(\hat v_1\) rather than a truth direction, and steering along what it returns then moves behavior with the wrong sign. Recoverability comes down to whether the class gap grows with depth until it dominates the spread along that axis. On <code class="language-plaintext highlighter-rouge">cities</code> it does, on <code class="language-plaintext highlighter-rouge">counterfact</code> it never does. That pattern, the alignment to \(\hat v_1\) and the decoding alike, replicates on OLMo-2-1B across a different architecture and corpus. The steering measurement is on Pythia alone.</p>

  <p><strong>(3) The failure is in the estimator, not the model:</strong> the offending direction is identified in advance from the within-class spectrum, not from the steering outcome. Removing it, without leaving the linear class, raises the held-out AUROC above the random decoding null and corrects the sign of the steering behavior, at the same layer.</p>

</div>

<h2 id="notation">Notation</h2>

<p>I start by fixing a model and a layer \(L\). Each statement is passed through the model once and
summarized by the residual-stream activation at its final token,
\(x \in \mathbb{R}^{d}\), where \(d\) is the model’s hidden width. Each statement carries
a truth label \(y \in \{0, 1\}\) and a dataset consists of \(N\) such pairs \((x_i, y_i)\).</p>

<p>Throughout, a <strong>hat</strong> marks a quantity estimated from a finite sample, and its absence
marks the population quantity it estimates, with expectations taken over the
data-generating distribution of activations at the probed layer. A table of every symbol
used in the post is collected in the notation appendix. The distinction is crucial:
almost every trap in this post is the gap between an estimate and its population target,
\(\hat\theta\) and \(\theta\), or the eigenvalues of \(\hat C\) and of \(\Sigma\), which
agree only in the large-sample limit. Estimates are computed on a training split of
\(N_{\text{train}} = N_0 + N_1\) activations, \(N_0\) of them with \(y = 0\) and \(N_1\)
with \(y = 1\). Held-out data is used only to
evaluate a fitted direction, never to determine it. I define the population class means
and their estimators as</p>

\[\mu_0 = \mathbb{E}[x \mid y=0], \qquad
\hat\mu_0 = \frac{1}{N_0}\sum_{i:\,y_i=0} x_i .\]

\[\mu_1 = \mathbb{E}[x \mid y=1], \qquad
\hat\mu_1 = \frac{1}{N_1}\sum_{i:\,y_i=1} x_i .\]

<p>I define the <strong>class-mean gap</strong> as
\(\delta = \mu_1 - \mu_0\), with \(\hat\delta = \hat\mu_1 - \hat\mu_0\), and the
<strong>within-class covariance</strong> as the class-weighted average of the two
class-conditional covariances,</p>

\[\Sigma \;=\; \pi_0 \operatorname{Cov}[x \mid y=0] \;+\; \pi_1 \operatorname{Cov}[x \mid y=1],
\qquad \pi_0 = \Pr[y = 0], \quad \pi_1 = \Pr[y = 1].\]

<p>Its estimator centers each class on <em>its own</em> mean, so that the between-class shift
does not leak into the noise estimate:</p>

\[\hat C \;=\; \frac{1}{N_{\text{train}}}
\left[
\sum_{i:\, y_i = 0} (x_i - \hat\mu_0)(x_i - \hat\mu_0)^{\top}
\;+\;
\sum_{i:\, y_i = 1} (x_i - \hat\mu_1)(x_i - \hat\mu_1)^{\top}
\right]
\;=\; \hat\pi_0 \hat C_0 + \hat\pi_1 \hat C_1 ,\]

\[\hat\pi_0 = \frac{N_0}{N_{\text{train}}},
\qquad \hat\pi_1 = \frac{N_1}{N_{\text{train}}},\]

<p>with \(\hat C_0\) and \(\hat C_1\) the sample covariance of each class (dividing by
\(N_0\) and \(N_1\)). Balanced classes
are the special case \(\pi_0 = \pi_1 = \tfrac{1}{2}\), where \(\Sigma\) reduces to the
plain average \(\tfrac{1}{2}\big(\operatorname{Cov}[x\mid y{=}0] +
\operatorname{Cov}[x\mid y{=}1]\big)\).</p>

<p>For real activations, these population quantities have no closed form: they are the
\(N_{\text{train}}\to\infty\) limits under the data-generating distribution and are never
obtained directly. Their influence is detected operationally. A training-fit direction
is scored on held-out data against the random-direction null, and a shortfall, or a
wrong-signed causal effect, is the signature of the estimate having locked onto a
finite-sample artifact rather than the population target. The one exception is the
whitening schematic below, whose activations are drawn from a <em>known</em> Gaussian, so there
\(\Sigma\) and the optimal direction \(\Sigma^{-1}\delta\) are known by construction rather
than estimated.</p>

<h2 id="definitions">Definitions</h2>

<p><strong>Probe.</strong> A probe is defined as a unit vector \(u \in \mathbb{R}^{d}\), \(\lVert u \rVert = 1\), that reads the scalar \(z = u^{\top} x\).</p>

<p><strong>Steering.</strong> Steering is the inference-time counterpart of reading. Rather than
taking a direction’s inner product with the activation, one adds a multiple of it to the residual
stream at a chosen layer during the forward pass, \(x \mapsto x + h\,c\,w\), and measures
how the model’s output moves. The weights are untouched, only the hidden state is
displaced.</p>

<p>The two operations differ in kind and not only in notation. Reading projects the space
onto a line and asks where an activation already sits along it. Steering moves the
activation itself, to a different point of the residual stream. What motivates the move
is the expectation that the destination is semantically different from the origin, so
that displacing along a truth direction should carry the representation to where the
model treats the statement as true. That expectation is an assumption rather than a
definition, and the steering null below exists to test it. Steering along \(\hat v_1\), a direction fit without any reference to the labels,
tests it most directly.</p>

<p>Generic steering directions are written \(w\) and generic readout directions \(u\). The
specific directions steered along below carry their own names.
The unit \(c\) and coefficient \(h\) are fixed in <em>The scale of an intervention</em> below.</p>

<p><strong>Estimator.</strong> An estimator is the rule that turns a labelled sample into a direction.
Two appear throughout, the mass-mean rule and the Fisher rule defined next, and they
differ in the rule alone, not in the data they see. A direction carries a hat because it
is the output of such a rule on a finite sample, so a claim that a direction is <em>wrong</em>
is a claim about a rule and the sample it saw, never about the activations themselves.</p>

<p><strong>Mass-mean (difference-in-means) direction.</strong> The mass-mean is defined as</p>

\[\hat\theta \;=\; \frac{\hat\delta}{\lVert \hat\delta \rVert}.\]

<p>i.e., the unit vector from the false-class centroid to the true-class centroid. It fits two
averages and thresholds \(z\) at the midpoint
\(\hat b = \tfrac{1}{2}(\hat\mu_1 + \hat\mu_0)^{\top}\hat\theta\).</p>

<p><strong>Fisher (whitened) direction.</strong> The Fisher direction is defined as</p>

\[\hat\theta_{\mathrm{F}} \;\propto\; \hat\Sigma^{-1} \hat\delta,\]

<p>and is normalized to unit length. The multiplication of \(\hat\delta\) by \(\hat\Sigma^{-1}\) downweights components of the mean shift that lie along
high-variance nuisance directions. Here \(\hat\Sigma\) is the Ledoit–Wolf shrinkage estimate
\(\hat\Sigma_{\text{LW}} = (1-\rho)\,\hat C + \rho\,\frac{\operatorname{tr}\hat C}{d} I\),
with \(\hat C\) the within-class estimate above and \(\rho \in [0,1]\) the closed-form
Ledoit–Wolf intensity. Shrinkage is needed because the raw \(\hat C\) is singular
whenever \(N_{\text{train}} &lt; d\), the relevant regime here. The Ledoit–Wolf intensity
is chosen to minimize the error of \(\hat\Sigma\) itself, not to maximize \(d'\) of the
whitened direction, which makes every whitened number in this post a lower bound as described in the
methods appendix.</p>

<p><strong>Separation, \(d'\).</strong> For a probe direction \(u\), one can project both class-conditional clouds
onto \(u\): each collapses to one dimension, with means \(u^{\top}\mu_0\), \(u^{\top}\mu_1\)
and variances \(u^{\top}\operatorname{Cov}[x\mid y{=}0]\,u\),
\(u^{\top}\operatorname{Cov}[x\mid y{=}1]\,u\). The separation \(d'(u)\) is the class-mean
gap projected along \(u\), in units of the projected within-class spread:</p>

\[d'^2(u) \;=\; \frac{(u^{\top}\delta)^2}{u^{\top}\Sigma u},
\qquad d'(u) \;=\; +\sqrt{d'^2(u)} .\]

<p>The numerator is the projected class-mean gap, the denominator the projected
within-class variance, which for balanced classes is the pooled variance
\(\tfrac{1}{2}\big(u^{\top}\operatorname{Cov}[x\mid y{=}1]\,u +
u^{\top}\operatorname{Cov}[x\mid y{=}0]\,u\big)\), since \(\Sigma\) is the class-weighted
average of the two. Writing it this way makes the object a Rayleigh quotient in \(u\)
from the start, which is the form every later result takes.</p>

<p>This is the separation measured in units of its own noise, the detection-theoretic
sensitivity of an ideal observer discriminating two Gaussians. Defining the square
first makes the <strong>sign-blindness</strong> explicit rather than asserted: \(d'\) depends on the
mean difference only through \((u^{\top}\delta)^2\), so it cannot distinguish a direction
from its negation.</p>

<p>\(d'\) is a property of a <em>direction</em>, so every reported value has to name the estimator
that produced the direction. I subscript throughout:</p>

\[d'_{\mathrm{mm}} = d'(\hat\theta), \qquad
d'_{\mathrm F} = d'(\hat\theta_{\mathrm F}), \qquad
d'_{\perp} = d'(\hat\theta_{\perp}),\]

<p>for the mass-mean, whitened and projection-out directions defined below. The optimal
\(d'\) over directions is the Mahalanobis separation. The Rayleigh quotient in \(u\) is
maximized at \(u \propto \Sigma^{-1}\delta\), which is the Fisher direction defined
above, and its maximum value is</p>

\[d'_{\mathrm M} \;=\; \max_{u} d'(u) \;=\; \sqrt{\delta^{\top}\Sigma^{-1}\delta} .\]

<p>So the whitened direction is not a heuristic correction to the mass-mean one. It is the
solution of the optimization that \(d'\) poses, and the mass-mean direction coincides
with it only when \(\Sigma\) is a multiple of the identity. Its sample estimate is</p>

\[\hat d'_{\mathrm M} \;=\; \sqrt{\hat\delta^{\top}\hat\Sigma^{-1}\hat\delta} ,\]

<p>with \(\hat\Sigma\) the shrunk estimate. It upper-bounds the three sample directions
when all are scored in sample, but it estimates the population maximum rather than
attaining it, and it is a fitted in-sample quantity at \(N \ll 2d\). The distinction
matters numerically, and in a way that is easy to get wrong. At <code class="language-plaintext highlighter-rouge">counterfact</code> layer
\(28\), \(\hat d'_{\mathrm M}\) \(= 1.03\) on the full set, while the fitted whitened direction achieves
\(d'_{\mathrm F} = 0.38\) held-out. It is tempting to read the first as what the
geometry is capable of and the gap as what the estimator loses, but that reading is
incorrect. Recomputing
\(\hat d'_{\mathrm M}\) with the labels shuffled, under the identical recipe, gives \(0.79 \pm 0.01\).
Three quarters of the \(1.03\) is present with no signal in the labels at all, and the
same holds at layer \(24\) (\(0.58\) against \(0.48\) shuffled) and layer \(32\)
(\(4.96\) against \(3.98\)). \(\hat d'_{\mathrm M}\) inflates for exactly the reason developed in <em>How
small is too small?</em> below, and it inflates more as the shrinkage intensity falls. On
<code class="language-plaintext highlighter-rouge">cities</code>, by contrast, the same check at layer \(24\) gives \(11.97\) against \(3.56\)
shuffled, so the comparison does discriminate a real ceiling from an inflated one. In
this post \(\hat d'_{\mathrm M}\) is therefore an in-sample bound, never a ceiling on what a held-out
probe should reach. Unless marked otherwise, every \(d'\) below is held-out.</p>

<p><strong>AUROC (area under the receiver-operating-characteristic curve).</strong> The AUROC is the probability
that a randomly chosen true statement projects above a
randomly chosen false one, \(\mathrm{AUROC}(u) = \Pr[\,z_1 &gt; z_0\,]\) for
\(z_1 \sim p(z\mid y{=}1)\), \(z_0 \sim p(z\mid y{=}0)\) independent. Unlike \(d'\), it is
sign-aware: \(\mathrm{AUROC}(-u) = 1 - \mathrm{AUROC}(u)\). The name is a radar-era
inheritance and the quantity is described further in the appendix <em>AUROC and its
relation to \(d'\)</em>.</p>

<p>\(\mathrm{AUROC}\) and \(d'\) are linked when the class-conditionals are Gaussian and the classes balanced, so that \(\Sigma\) is the plain average of the two class covariances. Then
\(z_1 - z_0 \sim \mathcal{N}\big(u^{\top}\delta,\; 2\,u^{\top}\Sigma u\big)\), so</p>

\[\mathrm{AUROC} \;=\; \Pr[z_1 - z_0 &gt; 0]
\;=\; \Phi\!\left(\frac{u^{\top}\delta}{\sqrt{2\,u^{\top}\Sigma u}}\right)
\;=\; \Phi\!\left(\frac{d'}{\sqrt{2}}\right),\]

<p>with \(\Phi\) the standard normal CDF. So \(d' = 1\) corresponds to
\(\mathrm{AUROC} = \Phi(0.71) \approx 0.76\), and \(d' = 3\) to \(\approx 0.98\). Equal class
variances are not required: for balanced classes the variance of \(z_1 - z_0\) is
\(2\,u^{\top}\Sigma u\) whether or not the two class covariances agree. Away from the
Gaussian case the identity is only a guide, which is why both are reported. On the data
here it holds to within \(3\%\) under shuffled labels even where the projections carry
excess kurtosis above \(1\), as <em>How small is too small?</em> reports.</p>

<h2 id="the-geometry">The geometry</h2>

<p><strong>The mass-mean direction.</strong> Geometrically, obtaining \(\hat\theta\) is the most naive thing one
could do: project onto the line joining the two class centroids. There is no fitting
beyond two averages, so there is little room to launder a spurious feature in through
an optimizer. Therefore, whatever separation it finds
is a property of the data, not of an optimizer’s freedom to search.</p>

<p><strong>Cover’s theorem: the counting baseline.</strong> Before asking how <em>well</em> a direction
separates the classes, the first question is how surprising it is that a separating
direction exists at all. Cover’s function-counting theorem (1965) makes the
accounting exact: for \(N\) points in general position in \(\mathbb{R}^{d}\), the number
of the \(2^{N}\) possible binary labelings that a hyperplane through the origin can
separate is</p>

\[C(N,d) = 2\sum_{k=0}^{d-1}\binom{N-1}{k}.\]

<p>General position means every subset of at most \(d\) points is linearly independent.
Activations satisfy this generically, and the count is exact under it. The <em>fraction</em>
of labelings that are linearly separable is
\(f(N,d) = C(N,d)/2^{N}\) (a probe with a bias term is an affine hyperplane, the
homogeneous case one dimension up, and at \(d\sim10^{3}\) the distinction is
immaterial). That fraction stays near \(1\) for \(N \lesssim d\) and falls through
\(\tfrac{1}{2}\) at the separating capacity \(N = 2d\). Probing sits at or below that
capacity, since \(d_{\text{model}}\) runs from a few hundred to a few thousand across the
ladder, while a probing set is hundreds to low thousands. This means that a large fraction of
labelings are linearly separable: the true/false one, but equally a random relabeling
of it. In this regime the linear separability of the truth dichotomy is close to
information-free. It reflects the ambient dimension more than anything the model has
learned. The question is therefore never <em>whether</em> a separating direction exists, but
<em>how far above the null</em> the particular mass-mean direction lands.</p>

<p>The distinction between these three baselines, Cover, the null, and the mass-mean estimator, is crucial. Cover counts labelings, not directions: it says how many of the \(2^{N}\) dichotomies admit <em>some</em> separating hyperplane, which upper-bounds what any label-informed fit could achieve however it searches. The null below asks what a <em>fixed</em>, <em>label-agnostic</em> direction achieves. The mass-mean estimator sits between them, label-informed but fit only through two class means, and so under no obligation to recover a separating hyperplane even when Cover guarantees one exists. The experiment measures where between those two baselines the estimator lands.</p>

<p>Cover also sets the scale. \(N/2d\) is the natural unit for the empirical question, and \(N \lesssim 2d\) is the regime in which the dimensional slack is available to any fit. The shuffled-label control asks how much of that slack the mass-mean rule converts into apparent signal when the labels carry none.</p>

<p><strong>The random-direction null.</strong> Roger’s and Farquhar’s results show that a <em>nonzero</em>
\(d'\) is not, by itself, evidence of anything. A random unit vector \(u\) is not
orthogonal to the class gap. Its projection inherits a share of the real mean shift, of
order \(\lVert\delta\rVert/\sqrt{d}\), so it produces some separation with no fitting at
all, and that separation grows with the true separation rather than sitting at a fixed
background. The meaningful quantity is therefore not \(d'\) but \(d'\) relative to its null. Draw
many random unit directions \(u \sim \mathrm{Unif}(S^{d-1})\), form the distribution of
their \(d'\) (equivalently AUROC), and take its 95th percentile,</p>

\[p_{95} \;=\; Q_{0.95}\big[\,\mathrm{AUROC}(u)\,\big] .\]

<p>A truth direction is <em>recoverable</em> only insofar as its own score clears that quantile,
\(\mathrm{AUROC}(\hat\theta) &gt; p_{95}\). That is the observation in Roger’s post
turned into an instrument, a step
<a href="https://arxiv.org/abs/2312.01037">Mallen &amp; Belrose (2023)</a> took first, resolving each
random probe’s sign on its source distribution and scoring transfer against quantiles
of \(10^7\) random-probe AUROCs. The per-layer null, the \(d'\) scale, and the margin
criterion are what is added here. The report that random probes reach ~75% on the “easy”
datasets becomes a chance level measured for each model, layer, and dataset, and
every claim below is stated relative to it.</p>

<p>For a direction \(u\) evaluated on held-out data against the null of its own layer and
dataset, the margin \(m(u)\) is given by</p>

\[m(u) \;=\; \mathrm{AUROC}(u) \;-\; p_{95},\]

<p>with \(p_{95}\) as defined above, computed on the same held-out points. A direction
counts as recovered when \(m &gt; 0\), a fixed 5% false-positive rate against random
directions. Layer selection maximizes \(m\) over depth,
\(L^{*} = \arg\max_{L} m(u_L)\), each layer scored against its own null, not the raw
AUROC. The same construction on the \(d'\) scale gives the equivalent criterion: a
random direction’s sign is arbitrary and \(\mathrm{AUROC}(-u) = 1 - \mathrm{AUROC}(u)\),
so each null draw’s AUROC is <em>folded</em> to \(\max\{a, 1-a\} \ge \tfrac{1}{2}\), and the
folded AUROC is monotone in the sign-blind \(d'\) through the Gaussian link above.</p>

<p><strong>Whitening: the Fisher direction.</strong> The mass-mean probe ignores the <em>shape</em> of the
within-class noise. If the class-conditional covariance \(\Sigma\) is anisotropic,
with large nuisance variance along some directions, the optimal linear discriminant is
not \(\mu_1-\mu_0\) but the Fisher/LDA direction</p>

\[\theta_{\mathrm{F}} \;\propto\; \Sigma^{-1}(\mu_1 - \mu_0).\]

<p>This correction is not introduced here. It is the second half of mass-mean probing
itself. Marks &amp; Tegmark set \(\theta_{\mathrm{mm}} = \mu_1 - \mu_0\) and then, for IID
evaluation, read with \(\sigma(\theta_{\mathrm{mm}}^{\top}\Sigma^{-1}x)\), noting that
this coincides with linear discriminant analysis. Their Appendix E gives the same
construction in terms of Mahalanobis whitening, and their \(\Sigma\) is the within-class
covariance defined exactly as above. What is at issue is <em>where</em> the correction is
applied. They are explicit that \(\Sigma^{-1}\) is there to tilt the decision boundary,
while \(\theta_{\mathrm{mm}}\) remains the candidate feature direction, one which may be
non-orthogonal to that boundary. Their intervention experiments steer along \(\theta_{\mathrm{mm}}\). In their framework, the whitened vector is a readout and the raw
mean-difference is the feature. The steering results below report a regime in which
that assignment is inverted.</p>

<p>Whitening should raise \(d'\) when the signal is partly buried
under structured noise, and <em>hurt</em> when \(n \ll d\) and \(\hat\Sigma\) is ill-conditioned, i.e.
where the inverse amplifies estimation error. I estimate \(\Sigma\) with
Ledoit–Wolf shrinkage for exactly that regime, and report plain and whitened side
by side, since the gap between them is itself diagnostic of how anisotropic the truth
geometry is.</p>

<p>The readout framing puts this in familiar company: whitened comparisons of
representational geometry have become standard in work on artificial networks
(<a href="https://arxiv.org/abs/2007.02789">Diedrichsen et al., 2021</a>), and <a href="https://arxiv.org/abs/2602.20273">Ying et al.
(2026)</a> measure a Mahalanobis cosine between truth
probes for exactly the reason above: these representations are anisotropic enough
that a raw inner product misleads. A follow-up by <a href="https://arxiv.org/abs/2606.19603">Ying, Hase &amp; Kriegeskorte
(2026)</a> proves the link between that cosine and
readout quality in closed form. For balanced classes with Gaussian projections, a probe’s
held-out AUROC is \(\Phi(s/\sqrt{2})\) in its signal-to-noise ratio \(s\), and its
Mahalanobis cosine to the Fisher direction is a softsign in the same \(s\). Their \(s\)
is the \(d'(u)\) of this post computed with the pooled within-class covariance, and their
Fisher distance is \(d'_{\mathrm M}\). One of the conditions under which they show the law
fails, a difference-of-means probe far from the Fisher direction, is the regime studied
below.</p>

<p><img src="/assets/figures/truth_whitening_schematic.png" alt="Whitening in two dimensions, on synthetic data where $$\Sigma$$ is known by
construction rather than estimated. Both panels carry the same class-mean gap. Only the
within-class noise differs. Left, isotropic noise: the mass-mean direction
$$\hat\theta \propto \hat\delta$$ (black) and the Fisher direction $$\hat\theta_{\mathrm F}
\propto \hat\Sigma^{-1}\hat\delta$$ (gold) coincide, and whitening buys nothing,
$$d'_{\mathrm{mm}} = d'_{\mathrm F} = 2.06$$. Right, the same gap with the noise stretched
along $$x_1$$ and compressed along $$x_2$$ (variances $$9$$ and $$0.35$$ against $$1$$ and
$$1$$): the mean gap now has a large component along the high-variance axis, so
$$\hat\theta$$ tilts into it and reads $$d'_{\mathrm{mm}} = 1.02$$, while $$\hat\Sigma^{-1}$$
divides each component by its variance and rotates $$\hat\theta_{\mathrm F}$$ onto the
low-variance axis that carries the class separation, $$d'_{\mathrm F} = 2.55$$. That the
Fisher value exceeds the isotropic $$2.06$$ is due to the compression along $$x_2$$, not
the stretch. The recovery is geometric: no extra data and no
richer function class, only a change in which direction of the same plane is
read." /></p>

<p><strong>Superposition: the salience knob.</strong> Since features share the residual stream, the
directions of <em>largest variance</em> need not be the directions of <em>interest</em>. The
dominant principal components may encode whatever is most prominent on this dataset
(sentence length, topic, token identity) rather than truth. Here a salient direction means a leading principal component of the activations, i.e. one of the directions of largest variance. This is Farquhar’s
“most prominent feature” stated geometrically, suggesting a direct test: fit
the mass-mean direction after projecting out the top-\(k\) principal components of the
activations, taken on the total covariance and so including any between-class shift,
and analyze \(d'\) as a function of \(k\). If truth is merely contained in the
salient subspace, the signal collapses as soon as the leading components are
removed. However, if it occupies its own low-variance subspace, \(d'\) survives. This
superposition probe, \(d'\) against the number of top directions removed, is the
salience-versus-truth confound turned into a measurement. The measurement itself is
reported in <em>The rogue dimension</em> below, once the null is in place.</p>

<h2 id="drawing-a-random-direction">Drawing a random direction</h2>

<p>I build both nulls in this post from <em>random unit vectors</em>. To sample \(u\) uniformly from the unit sphere
\(S^{d-1}\), I draw a standard Gaussian and normalize:</p>

\[\xi \sim \mathcal{N}(0, I_d), \qquad u = \frac{\xi}{\lVert \xi \rVert}.\]

<p>This is uniform because the isotropic Gaussian density depends on \(\xi\) only through
\(\lVert \xi \rVert\), so it is invariant under every rotation. Normalizing leaves a
distribution on the sphere with the same invariance, and the uniform distribution is
the only one with that property. (Normalizing each coordinate independently, or
sampling coordinates uniformly in \([-1,1]\) and normalizing, does <em>not</em> give a uniform
direction, since it concentrates near the cube’s diagonals.)</p>

<p>Two facts about Gaussian random vectors in high dimension enter here. The overlap of such a vector \(u\) with any <em>fixed</em> unit vector \(v\) is small, \(u^{\top} v \sim \mathcal{N}(0, 1/d)\) to a
good approximation, so typically \(\lvert u^{\top} v\rvert \approx 1/\sqrt{d}\). Likewise, the
overlap with the <em>coordinate</em> axes is \(O(1/\sqrt{d})\). This implies that a random direction is nearly orthogonal to any direction fixed in advance, and by a union bound to every
direction on a list at once, so long as the list is not exponentially long in \(d\).</p>

<p><strong>The decoding null.</strong> To construct the decoding null, I draw \(u\), compute
\(\mathrm{AUROC}(u^{\top}x)\) on held-out data, and fold as above. Repeating this \(n\)
times gives the null distribution whose 95th percentile is the \(p_{95}\) of the margin
criterion.</p>

<p>A random direction does not achieve \(\mathrm{AUROC} = \tfrac{1}{2}\) merely up to
sampling noise. Its projection inherits a share of the real mean shift:
\(u^{\top}\delta \sim \mathcal{N}(0, \lVert\delta\rVert^2/d)\), so</p>

\[d'(u) \;\approx\; \frac{\lvert u^{\top}\delta\rvert}{\sqrt{u^{\top}\Sigma u}} .\]

<p>When \(\Sigma\) is close to isotropic, \(u^{\top}\Sigma u \approx \operatorname{tr}\Sigma / d\)
and the factors of \(d\) cancel, leaving \(d'(u) \approx \lvert z \rvert \cdot
\lVert\delta\rVert / \sqrt{\operatorname{tr}\Sigma}\) with \(z \sim \mathcal{N}(0,1)\). The
null is therefore <em>large exactly when the true separation is large</em>. It is not a
fixed background, and it must be recomputed for every layer and dataset.</p>

<p><img src="/assets/figures/truth_null_distribution.png" alt="The decoding null on `cities`, pythia-2.8b layer $$28$$, drawn rather than asserted:
the histogram of $$d'$$ achieved by $$400$$ random unit directions, with its 95th
percentile marked as the bar a probe must clear, and the mass-mean direction's $$d'$$
set against it. Both are computed on the full set here, so $$\hat\theta$$ reads the
in-sample $$3.14$$ rather than the held-out $$3.02$$ the sweep reports for the same
layer. The figure is about the ratio, not the number. Recoverability is the margin above
the percentile, not the raw $$d'$$." class="fig-single" /></p>

<p><strong>The steering null.</strong> A steering direction \(w\) is added to the activation rather than
projected onto it. I draw \(w\) the same way, add it to the residual stream, and measure
how the model’s behavior moves: this measures what an arbitrary direction <em>causes</em>,
rather than what it <em>reads</em>. A direction can clear the decoding null without clearing
the steering null or vice versa. Throughout this post an effect, decoding or steering,
is called <em>significant</em> when it clears the 95th percentile of its null, and the word is
used in no other sense.</p>

<h2 id="the-scale-of-an-intervention">The scale of an intervention</h2>

<p><strong>The behavioral score.</strong> What steering moves is the model’s preference between two
completions of the same prompt. Each statement is split at its final entity into a
prompt and a pair of one-entity completions. On <code class="language-plaintext highlighter-rouge">cities</code> the prompt is “The city of
Krasnodar is in” with true completion “ Russia” against a paired false country. On
<code class="language-plaintext highlighter-rouge">counterfact</code>, the shipped relation template, e.g. “.NET Framework is created by” with
“ Microsoft” against “ Google”. The score is</p>

\[\ell(x) \;=\; \log P(\text{true completion} \mid \text{prompt})
\;-\; \log P(\text{false completion} \mid \text{prompt}),\]

<p>summed over the completion’s tokens. Both completions are scored against the <em>same</em>
prompt, so the prompt’s own likelihood cancels and only the preference between the two
entities remains.</p>

<p><img src="/assets/figures/truth_behavioral_score.png" alt="Left: the behavioral score. A statement is split into a prompt and a contrastive pair of completions, each scored as a log-probability against the same prompt; $$\ell$$ is their difference. Right: the steered passes at $$\pm h$$ define $$\Delta(\pm h)$$, whose odd part $$A$$ is the evidence of a signed direction and whose even part $$S$$ collects generic disruption." /></p>

<p>Steering displaces the residual stream, \(x \mapsto x + h\, c\, w\). The unit \(c\) sets
the scale, and should be chosen so that the effects of interest appear at \(h \sim 1\).
Much smaller displacements cannot move a statement across the decision boundary, and
much larger ones push the activation far outside the range the model ever sees, where
the response is large, idiosyncratic, and mostly independent of which direction was
pushed. The steering null widens sharply in that regime, which is a useful diagnostic
in its own right. If a random direction of the same length produces as much behavioral
change as the chosen direction, \(h\) is too large.</p>

<p>Taking \(c = \sigma_w = \operatorname{std}(w^{\top} x)\), the spread of activations along
\(w\), seems natural but is a trap: \(\sigma_w\) depends on the direction, so different
conditions of an experiment receive pushes of different magnitude. It also depends on the
layer, so a fixed \(h\) means different things at different depths.</p>

<p>The right unit is the <strong>class-mean gap</strong> \(c = \lVert\hat\delta\rVert\), which is a property of the
dataset and layer rather than of the direction. Then \(h = 1\) displaces an activation
by exactly the distance between the two class means, so that along \(\hat\theta\) it
carries the false-class centroid onto the true-class centroid. Every displacement is
norm-matched
and each layer is comparable.</p>

<p><strong>Decomposition.</strong> The
shift in \(\ell\) at coefficient \(h\), averaged over the evaluation set, is
\(\Delta(h) = \big\langle\, \ell(x + h\, c\, w) - \ell(x) \,\big\rangle\). I express the symmetric and antisymmetric parts of the shift as:</p>

\[A \;=\; \tfrac{1}{2}\big[\Delta(+h) - \Delta(-h)\big], \qquad
S \;=\; \tfrac{1}{2}\big[\Delta(+h) + \Delta(-h)\big].\]

<p>Only the antisymmetric part \(A\) is evidence of a <em>direction</em>: pushing along \(\hat\theta\)
should make the model favor the true completion and pushing against it the false one,
so a genuine truth direction responds as an odd function of \(h\). A generic perturbation degrades
the computation whichever way it points, contributing to \(S\) and leaving \(A\) near zero.</p>

<p><strong>Seeds and draws.</strong> I define a <em>seed</em> as a resampling of the class-stratified train/test split. Averaging over seeds gives the standard error on the given direction’s
effect, that is, the reproducibility of the measurement. I define a <em>draw</em> as a resampling of a fresh random direction (\(u\) for a decoding null, \(w\) for a steering null). Since no
fitting is involved, the drawn direction does not depend on the split, and the draws
simply provide the null distribution. A tight seed error bar says nothing about whether the effect sits inside the null.</p>

<h2 id="setup">Setup</h2>

<p><strong>Models.</strong> I use the Pythia suite (Biderman et al., 2023) at 70m, 410m, 1.4b, and 2.8b
parameters. They share architecture, data and training order and differ only in
scale, which is what makes it the right ladder for an emergence question. Residual
widths are \(d = 512, 1024, 2048, 2560\).</p>

<p><strong>Data.</strong> I use the twelve curated true/false sets from
<a href="https://github.com/saprmarks/geometry_of_truth">Marks &amp; Tegmark</a>. Nine form a
<strong>main tier</strong>, <code class="language-plaintext highlighter-rouge">cities</code>, <code class="language-plaintext highlighter-rouge">neg_cities</code>, <code class="language-plaintext highlighter-rouge">larger_than</code>, <code class="language-plaintext highlighter-rouge">smaller_than</code>,
<code class="language-plaintext highlighter-rouge">cities_cities_conj</code>, <code class="language-plaintext highlighter-rouge">cities_cities_disj</code>, <code class="language-plaintext highlighter-rouge">common_claim_true_false</code>,
<code class="language-plaintext highlighter-rouge">companies_true_false</code> and <code class="language-plaintext highlighter-rouge">counterfact_true_false</code>, subsampled class-balanced to
\(N = 1198\), the size of the smallest member rounded to a class-balanced count, so that \(d'\), transfer, and the ratio
\(N/2d\) are directly comparable across datasets. Two translation sets
(<code class="language-plaintext highlighter-rouge">sp_en_trans</code>, <code class="language-plaintext highlighter-rouge">neg_sp_en_trans</code>) have only \(N = 354\) and are reported separately. The <code class="language-plaintext highlighter-rouge">likely</code> set is included as a
<strong>distractor</strong>: it is constructed so that textual plausibility decorrelates from
truth, so a direction fitted on it isolates the plausibility axis on its own. These
datasets, the papers that introduced them, and their known failure modes are catalogued
in a companion <a href="https://jasteinberg.github.io/reviews/truth-probes-map/">map of the truth-probing
literature</a>.</p>

<p><strong>Activations.</strong> I pass each statement through the model once, and the residual
stream is read at the final real token (right-padded, located from the attention
mask) at every layer. Layer \(0\) is the embedding, before any transformer block. On
templated statements the last-token embedding is nearly constant within a dataset, so
its \(d'\) is degenerate and it is excluded from layer selection.</p>

<p><strong>Estimation, held out.</strong> I fit every quantity on a class-stratified training half
and evaluate on the held-out half: the mass-mean direction, the within-class
covariance used for whitening, and the PCA basis used by the superposition probe.
The estimator is not prone to overfitting in the usual sense, since it has
almost no capacity, fitting only two class means. However, with \(d\) comparable to \(N\)
the estimated mean difference absorbs \(O(\sqrt{d/N})\) of noise, and scoring it
in-sample biases \(d'\) upward through exactly the fluctuations that defined the direction. In practice the correction is
large. On a synthetic control with isotropic noise, a planted separation of \(d' = 1\),
and the post’s own \(d = 2560\) and \(N = 1198\), the in-sample estimate returns \(4.2\)
and the held-out estimate \(0.24\). Both follow from the derivation in <em>How small is too
small?</em>: the noise in \(\hat\theta\) is aligned with the sample fluctuations that
produced it, so in-sample \(d'^{2} \simeq 1 + 4d/N_{\text{train}}\), while held-out
\(d' \simeq \cos(\hat\theta, \theta) \simeq (1 + 4d/N_{\text{train}})^{-1/2}\). The
simulation matches the in-sample prediction to within \(2\%\) and the overlap
\(\cos(\hat\theta, \theta)\) to within \(4\%\) for \(d/N_{\text{train}}\) from \(0.2\) to \(14\).
The held-out \(d'\) itself scatters more, \(-11\%\) to \(+16\%\) about its prediction,
because it is estimated on a finite held-out half: its standard deviation across
repetitions is about \(0.08\), so a twenty-repetition mean is fixed only to about
\(\pm 0.02\), and every deviation sits within two of those.</p>

<p>Two consequences follow. First, the held-out \(d'\) is
<em>attenuated</em>: the fitted \(\hat\theta\) is misaligned with the true direction by
\(O(\sqrt{d/N_{\text{train}}})\), so the reported \(d'\) underestimates the intrinsic
separation on average, and the attenuation is largest for the widest model. Second, \(d'\) is
sign-blind while AUROC is not, so \(\hat\theta\) is oriented on the training half and
that sign is held fixed on the test half. Layer selection then maximizes held-out
AUROC above the layer’s own null, which cannot reward an anti-predictive direction.
Selecting a layer by an SNR is itself the field’s practice. <a href="https://arxiv.org/abs/2407.12831">Bürger et
al.</a> take the layer maximizing the ratio of
between- to within-class variance, and later work adopts the recipe unchanged
(<a href="https://aclanthology.org/2025.findings-acl.38.pdf">Bao et al.</a> select the layer
for every model in their study by that ratio, crediting Bürger et al. and
<a href="https://www.anthropic.com/research/probes-catch-sleeper-agents">MacDiarmid et al.</a>
for the technique. <a href="https://arxiv.org/abs/2604.03754">Poulis et al.</a> App. B.1 write
out the form Bürger and Bao are using, a ratio of squared mean differences to
variances averaged across dimensions). That is \(d'^{2}\) averaged over coordinates,
with no direction and no null defined. The margin used here is the same ratio along one
direction, read against what random directions give.</p>

<p><strong>The null.</strong> At every layer I draw \(200\) random unit directions for the decoding
null, evaluate them on the <em>same held-out points</em>, and fold as above. The null is
recomputed per layer rather than once globally: where the class gap grows with depth
relative to the total spread, a random direction inherits more of it, and the null
rises with the signal it is meant to calibrate.</p>

<p><strong>Whitening.</strong> \(\hat\Sigma_{\text{LW}}\) as defined above. At \(N_{\text{train}}
\approx 600\) against \(d = 2560\) the shrinkage is what makes the inverse usable.</p>

<h2 id="where-the-direction-is-recoverable">Where the direction is recoverable</h2>

<p>With the null, \(d'\), and the class-mean-gap unit in place, the reproduction question
becomes concrete: at what scale, at what depth, and across which datasets is a truth direction
actually recoverable?</p>

<p>Across the Pythia ladder it emerges with scale. On <code class="language-plaintext highlighter-rouge">cities</code> the best-layer mass-mean
AUROC climbs from \(0.53\) at pythia-70m, inside the random-direction null (\(p_{95} = 0.54\)) and so
not recoverable, to \(0.85\) at 410m, \(0.88\) at 1.4b, and \(0.97\) at 2.8b, with \(d'\)
rising from \(0.07\) to \(2.93\). At 70m the whitened direction does not clear the null at any
layer either, the best value being \(0.504\) on <code class="language-plaintext highlighter-rouge">cities</code> and \(0.520\) on <code class="language-plaintext highlighter-rouge">counterfact</code>.
What fails there is the whole linear class, not one estimator within it. Whether a
70m model represents truth in some form this class cannot express is not tested here.</p>

<p><img src="/assets/figures/truth_emergence.png" alt="Emergence of the truth direction across the Pythia scale ladder, `cities`: best-layer held-out AUROC of the mass-mean direction (red) and the whitened direction (blue) against the random-direction null (gray band, up to its 95th percentile). At 70m both directions sit inside the null. From 410m on, both clear it, and the mass-mean margin widens with $$d_{\text{model}}$$. Layers selected by the largest margin above each layer's own null." class="fig-single" /></p>

<p>Within a single model, the truth direction is a property of depth. Sweeping the layers of pythia-2.8b on
<code class="language-plaintext highlighter-rouge">cities</code>, the mass-mean AUROC sits inside the null through the early layers, lifts clear
of it around the middle of the network, and plateaus high across the late layers,
peaking at \(0.975\) at layer \(28\). The null itself widens with depth, its 95th
percentile climbing from \(0.54\) early to \(0.75\) deep, so recoverability is
again the <em>margin</em> above the null, not the raw number: a deep-layer AUROC in the low
\(0.7\)s can still sit inside it. Layer selection maximizes the margin, not the AUROC,
and the two peak one layer apart: the AUROC at layer \(28\) (\(0.975\) against a null of
\(0.744\)), the margin at layer \(29\) (\(0.973\) against \(0.721\)). The scale ladder
above reports layer \(29\), since the margin is what carries the claim.</p>

<p><img src="/assets/figures/truth_layer_sweep.png" alt="Layer sweep, pythia-2.8b on `cities`: held-out AUROC of the mass-mean direction (red) against the random-direction null (gray band, to its 95th percentile) as a function of depth. The curve sits inside the null through the early layers and lifts clear around the middle of the network. The null widens with depth, so the selected layer (star, $$L=29$$) maximizes the margin above each layer's own null rather than the raw AUROC. Layer 0 is the embedding, drawn hollow and excluded from selection." class="fig-single" /></p>

<p><strong>The direction transfers, but only within a polarity and a family.</strong> Fitting the
mass-mean direction on one main-tier set and evaluating it on another reproduces the
within-dataset signal on the diagonal, \(0.93\) to \(0.98\) for the single-frame sets and
down to \(0.57\) for <code class="language-plaintext highlighter-rouge">counterfact</code>, but the off-diagonal is <em>organized</em>, not merely weak.
Three structures are visible in the full nine-by-nine matrix. First, between a
statement type and its logical negation the transfer is anti-predictive: <code class="language-plaintext highlighter-rouge">cities</code>
scores AUROC \(0.08\) on <code class="language-plaintext highlighter-rouge">neg_cities</code> and <code class="language-plaintext highlighter-rouge">larger_than</code> scores \(0.07\) on
<code class="language-plaintext highlighter-rouge">smaller_than</code>, both far below chance, so the same vector reads truth backwards once
polarity flips. The <code class="language-plaintext highlighter-rouge">neg_cities</code> direction is anti-predictive not on <code class="language-plaintext highlighter-rouge">cities</code> alone but
on the compound and claim-like sets as well, \(0.02\) to \(0.33\) across five of them.
Second, the sets fall into two families that do not speak to each other. The numeric
comparisons transfer to nothing outside their own pair, \(0.32\) to \(0.64\) in both
directions, while <code class="language-plaintext highlighter-rouge">cities</code>, the two compound sets, <code class="language-plaintext highlighter-rouge">common_claim</code>, <code class="language-plaintext highlighter-rouge">companies</code> and
<code class="language-plaintext highlighter-rouge">counterfact</code> form a block that transfers within itself at \(0.52\) to \(0.94\). Third,
that block is asymmetric, and the asymmetry belongs to the receiver rather than to the
direction. The <code class="language-plaintext highlighter-rouge">cities</code> column is uniformly high, \(0.84\) to \(0.94\) from every other
member, because a transferred score is \(\Phi(d'/\sqrt2)\) with the <em>target’s</em> gap and
noise inside \(d'\), and <code class="language-plaintext highlighter-rouge">cities</code> has the largest gap in the benchmark. A direction that
reads its own set at \(0.57\) still reads <code class="language-plaintext highlighter-rouge">cities</code> at \(0.87\). Transfer AUROC scores the
receiver’s signal-to-noise as much as the direction’s alignment, so the matrix has to
be read by rows and columns together. The mass-mean direction is genuinely there, but a
single vector approximates a structure carrying at least a separate polarity axis and
a family axis. The below-chance cells are what the decomposition of
<a href="https://arxiv.org/abs/2407.12831">Bürger et al. (2024)</a> predicts: a direction fit on
affirmative statements alone is a mixture \(t_A = \alpha\, t_G + \beta\, t_P\) of a
general truth direction and a polarity-sensitive one that reads truth forwards on
affirmatives and backwards on negations, so under negation the \(t_P\) term
anti-correlates with the label, and an AUROC of \(0.08\) rather than \(\tfrac{1}{2}\)
says that term dominates the mixture at this layer. The comparison is direct, not
analogical: their <code class="language-plaintext highlighter-rouge">cities</code>/<code class="language-plaintext highlighter-rouge">neg_cities</code> pair is the Marks &amp; Tegmark pair used here.
<a href="https://arxiv.org/abs/2604.03754">Poulis et al. (2026)</a> measure the same signature
in Llama, affirmative-trained probes anti-predictive on negations
(AUROC \(\approx 0\)) at the depths where the polarity direction carries most of the
truth-related variance.</p>

<p><img src="/assets/figures/truth_transfer.png" alt="Cross-dataset transfer of the mass-mean direction across the nine main-tier sets, pythia-2.8b. Each direction is fit on its source's training half at the source's selected layer and scored on each target's held-out half at the target's selected layer, so the diagonal is held-out too. Rows are the fitting set, columns the evaluation set. Three structures: negation pairs are *anti*-predictive (`cities`$$\leftrightarrow$$`neg_cities` $$0.02$$ and $$0.08$$, `larger_than`$$\leftrightarrow$$`smaller_than` $$0.07$$ both ways, and the `neg_cities` row is cold against the whole `cities` family); the numeric pair is an island, near chance with everything outside itself; and the six remaining sets form a warm block whose `cities` column is uniformly high because `cities` carries the largest class gap, so even a weakly aligned direction reads it well. The color map diverges about chance because the below-chance cells are a finding, not noise." class="fig-single" /></p>

<p>The sweep covers all twelve sets, not only the nine of the transfer matrix. Every
templated set with a single frame clears its null comfortably (plain AUROC
\(0.93\)–\(0.98\) at its selected layer). The two free-form sets sit at the bottom
of the table (\(d'_{\mathrm{mm}} = 0.64\) and \(0.20\)). <code class="language-plaintext highlighter-rouge">companies_true_false</code>
shows the largest plain/whitened gap in the benchmark, a second instance of the
rogue-dimension pattern, decoded rather than steered. The full table and
per-dataset commentary are in the appendix <em>Results on the full twelve-dataset benchmark</em>.</p>

<p><strong>The plausibility axis has the opposite depth profile.</strong> The <code class="language-plaintext highlighter-rouge">likely</code> set carries no
truth labels at all. Its two classes are the most likely and hundredth-most-likely
final token, so a direction fitted on it reads textual probability alone.
This axis is easy to find. On pythia-2.8b the mass-mean direction reaches held-out
\(\mathrm{AUROC} = 0.890\) (\(d'_{\mathrm{mm}} = 1.75\)) at layer \(12\), against a null 95th percentile of
\(0.593\), a wider margin than <code class="language-plaintext highlighter-rouge">counterfact</code> achieves at any depth. What separates it
from the truth sets is the depth at which it is recoverable. Plausibility is already
readable in the embedding (\(0.700\) at layer \(0\)), peaks at layers \(11\) and \(12\) (the AUROC at \(11\), the margin at \(12\)), and then
decays through the second half of the network to \(0.657\) by layer \(28\). <code class="language-plaintext highlighter-rouge">cities</code>
runs the other way: inside the null through the early layers, clear of it by
mid-network, peaking at \(0.973\) at layer \(29\). In one model, the direction that
reads probability and the direction that reads truth are strongest at opposite ends of
the depth axis, which is a reason to doubt that the deep-layer truth direction is
plausibility in disguise.</p>

<p><img src="/assets/figures/truth_plausibility_depth.png" alt="Depth profiles of the plausibility axis and the truth axis, plotted as the margin $$m = \mathrm{AUROC} - p_{95}$$ against fractional depth $$L/L_{\max}$$. The margin rather than the raw score, because the null widens with depth and AUROC is therefore not comparable across layers. Left, pythia-2.8b: `likely` peaks at layer $$12$$ ($$m = 0.297$$) and `cities` at layer $$29$$ ($$m = 0.252$$), at opposite ends of the depth axis, and the two curves cross near $$L/L_{\max} = 0.7$$. Right, pythia-410m (dotted) and pythia-1.4b (solid): the ordering reverses — `cities` peaks at layers $$11$$ and $$7$$ of $$24$$, before `likely` at layers $$15$$ and $$13$$ — and the late-depth truth structure appears only as a secondary rise over the final layers, still climbing at layer $$24$$, which is the last block at both widths, so whether it would peak is not resolved by this sweep. Stars mark the selected layer, hollow circles the embedding, which is excluded from selection." /></p>

<p>One limit on that reading, and one further test, should be stated. The limit is that
the profiles are cleanest at 2.8b. At 410m and
1.4b the plausibility peak stays mid-network (\(L = 15\) and \(L = 13\) of \(24\)) but
<code class="language-plaintext highlighter-rouge">cities</code> peaks earlier rather than later (\(L = 11\) and \(L = 7\)), so the ordering is
reversed, and the late-depth structure survives only as a secondary rise over the final
layers (\(m = 0.153\) at 410m layer \(24\), \(0.144\) at 1.4b layer \(24\), the last block
at both widths, where the margin is still rising). The test is the one Marks &amp;
Tegmark designed the set for, the cross-dataset one: take a
direction fitted on truth statements and evaluate it on <code class="language-plaintext highlighter-rouge">likely</code>. A direction that had
been reading plausibility all along should separate it. However, none of them do. Across all nine main-tier sets, a direction fitted at that set’s own
selected layer and evaluated on held-out <code class="language-plaintext highlighter-rouge">likely</code> gives AUROC between \(0.44\) and
\(0.60\), and not one clears the random-direction null on <code class="language-plaintext highlighter-rouge">likely</code> at the corresponding
layer (\(p_{95} \approx 0.58\)–\(0.60\)). The largest is <code class="language-plaintext highlighter-rouge">larger_than</code> at \(0.598\) against a
null of \(0.600\). <code class="language-plaintext highlighter-rouge">cities</code> gives \(0.559\) and <code class="language-plaintext highlighter-rouge">counterfact</code> \(0.480\). The reverse direction
agrees: the <code class="language-plaintext highlighter-rouge">likely</code>-fitted direction scores \(0.43\)–\(0.54\) on the nine truth sets. So
the plausibility axis and the truth directions are separately readable and mutually
uninformative. This is what the distractor was constructed to detect, and it comes
out clean.</p>

<h2 id="how-small-is-too-small">How small is too small?</h2>

<p>Cover’s theorem marks \(2d\) as the scale below which separability stops being
informative. The shuffled-label control measures what that costs in practice: fit the
mass-mean direction to a <em>random permutation</em> of the labels and record how much
separation comes back.</p>

<p>The control is the <em>control task</em> of
<a href="https://aclanthology.org/D19-1275/">Hewitt &amp; Liang (2019)</a>, and the gap between
real-task and control-task performance is what they call <em>selectivity</em>. Their diagnosis
and their prescription concern probe <strong>capacity</strong>: a probe expressive enough to memorize
the control task is too expressive to trust, so a smaller one should be used. That
prescription has no purchase here. The mass-mean probe fits two class means and nothing
else, and on <code class="language-plaintext highlighter-rouge">counterfact</code> at \(N = 100\) it still separates shuffled labels with AUROC
\(0.80\) on pythia-2.8b (\(0.795 \pm 0.042\) over sixteen permutations), which is what
the true labels reach in sample at the same \(N\). There is no capacity to shrink. The
freedom is in the ambient dimension.</p>

<p>This control is the one place in the post where a direction is scored on the points it
was fitted to. Everywhere else a direction is fitted on a training split and scored
held-out. Here that would defeat the purpose, since on held-out data a shuffled-label
direction scores \(\tfrac{1}{2}\) by construction, and the quantity of interest is
exactly how much separation the estimator manufactures from the noise it was fitted
to. Scoring in sample is what exposes it. Pooling the four models, the excess AUROC a
mass-mean direction extracts from pure noise scales as</p>

\[\mathrm{AUROC}_{\text{shuffled}} - \tfrac{1}{2}
\;\approx\; 0.045 \left(\frac{N}{2d}\right)^{-0.49}.\]

<p>The exponent follows from the estimator alone. \(\hat\theta\) is a difference of two
sample means, each concentrating at the \(\sqrt{N}\) rate, so the noise it carries has
magnitude \(O(\sqrt{d/N})\) whatever the activations look like, and the excess must
fall as \(N^{-1/2}\) at fixed \(d\). Measured across four models and \(N\) from \(100\)
to \(32{,}000\), it comes back \(-0.49\). The phenomenon is <em>dimensional slack</em>, the
room that \(d \gg N\) leaves for a direction built from a sample’s own fluctuations to
separate that sample, since the fluctuations that define the direction are the ones
being scored. It is not capacity.</p>

<p>Inverting the law gives the condition for noise alone to contribute less than
\(\varepsilon\) of excess AUROC,</p>

\[N \;&gt;\; 2d\left(\frac{a}{\varepsilon}\right)^{1/\lvert b\rvert}
\;=\; 2d\left(\frac{0.045}{\varepsilon}\right)^{2.0} ,\]

<p>which at the widths in this post and the standard 7B width gives the following. The
table uses the unrounded fit, \(a = 0.0446\) and \(b = -0.487\), and the rounded values
reproduce each entry to within about \(2\%\).</p>

<table>
  <thead>
    <tr>
      <th>\(d_{\text{model}}\)</th>
      <th>\(N\) for \(\varepsilon = 0.05\)</th>
      <th>\(\varepsilon = 0.02\)</th>
      <th>\(\varepsilon = 0.01\)</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td>\(512\)</td>
      <td>\(809\)</td>
      <td>\(5{,}305\)</td>
      <td>\(22{,}024\)</td>
    </tr>
    <tr>
      <td>\(2{,}048\)</td>
      <td>\(3{,}233\)</td>
      <td>\(21{,}220\)</td>
      <td>\(88{,}093\)</td>
    </tr>
    <tr>
      <td>\(2{,}560\)</td>
      <td>\(4{,}041\)</td>
      <td>\(26{,}525\)</td>
      <td>\(110{,}117\)</td>
    </tr>
    <tr>
      <td>\(4{,}096\)</td>
      <td>\(6{,}465\)</td>
      <td>\(42{,}440\)</td>
      <td>\(176{,}186\)</td>
    </tr>
    <tr>
      <td>\(8{,}192\)</td>
      <td>\(12{,}930\)</td>
      <td>\(84{,}879\)</td>
      <td>\(352{,}372\)</td>
    </tr>
  </tbody>
</table>

<p>The curated true/false datasets in this literature hold \(N \approx 1{,}200\)
statements. For a \(d = 4096\) model, the width of the 7B models these probes are
usually run on, noise alone buys in-sample AUROC \(0.55\) until \(N \approx 6{,}500\)
and \(0.52\) until \(N \approx 42{,}000\). Every dataset in the benchmark is between five
and thirty times too small for an in-sample AUROC in the seventies to mean what it
appears to mean. Held-out scoring removes the inflation but not its source. The fitted
direction is the same object either way, and at these sizes it is mostly noise: the
synthetic control in <em>Setup</em> puts its overlap with the true direction at \(0.24\) for
the post’s own \(d/N_{\text{train}}\). Held-out evaluation reports that attenuated
direction’s honest score, which is why held-out \(d'\) errs low where in-sample \(d'\)
errs high.</p>

<p>The newest state-of-the-art lie detector sits deeper still in this regime. The topic
datasets of Bürger et al. run from \(N = 164\) (<code class="language-plaintext highlighter-rouge">animal_class</code>) to \(1{,}496\) (<code class="language-plaintext highlighter-rouge">cities</code>)
at \(d = 4096\), so \(N/2d\) falls between \(0.02\) and \(0.18\). Their
leave-one-topic-out protocol is held-out, which is the right defense, and the table says
how much it defends against: at these sizes the slack is available in full to anything
fit or diagnosed in sample.</p>

<p>The control also behaves correctly at the other end of the ladder. On pythia-70m the
true and shuffled curves lie on top of each other at every \(N\), which is what a model
with nothing recoverable in it should give: on this dataset neither the plain nor the
whitened direction clears the random-direction null at any layer, so there is no signal
for the true labels to add.</p>

<p>The prefactor is where the dataset enters. What it quantifies is the number of
directions the noise effectively occupies. A spectrum concentrated on a few axes leaves
a random labeling less room to find a separator than a flat one does, so the ambient
width \(d\) is the right count only when the spectrum is flat. In general the count is
the participation ratio of the within-class covariance, and the amplitude should
collapse across datasets once \(N\) is measured in units of it.</p>

<p>A second sweep tests the exponent and the prefactor across datasets. The control is
scored in sample, so no half has to be held back and the grid can run to the full set.
Repeating it on pythia-2.8b at each dataset’s own best layer, over a seven-point
geometric grid from \(N = 100\) to that dataset’s total with sixteen label permutations
per point, and fitting
\(\mathrm{AUROC}_{\text{shuffled}} - \tfrac{1}{2} \approx a\,(N/2d)^{b}\) on each, gives</p>

<table>
  <thead>
    <tr>
      <th>dataset</th>
      <th>\(N_{\max}\)</th>
      <th>\(a\)</th>
      <th>\(b\)</th>
      <th>\(R^{2}\)</th>
      <th>\(\mathrm{PR}\)</th>
      <th>\(\hat\lambda_1/\operatorname{tr}\hat C\)</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td><code class="language-plaintext highlighter-rouge">counterfact_true_false</code></td>
      <td>\(3000\)</td>
      <td>\(0.045\)</td>
      <td>\(-0.490 \pm 0.019\)</td>
      <td>\(0.993\)</td>
      <td>\(30.3\)</td>
      <td>\(0.11\)</td>
    </tr>
    <tr>
      <td><code class="language-plaintext highlighter-rouge">cities</code></td>
      <td>\(1496\)</td>
      <td>\(0.053\)</td>
      <td>\(-0.446 \pm 0.009\)</td>
      <td>\(0.998\)</td>
      <td>\(34.6\)</td>
      <td>\(0.11\)</td>
    </tr>
    <tr>
      <td><code class="language-plaintext highlighter-rouge">larger_than</code></td>
      <td>\(1980\)</td>
      <td>\(0.027\)</td>
      <td>\(-0.484 \pm 0.014\)</td>
      <td>\(0.996\)</td>
      <td>\(8.9\)</td>
      <td>\(0.28\)</td>
    </tr>
    <tr>
      <td><code class="language-plaintext highlighter-rouge">sp_en_trans</code></td>
      <td>\(354\)</td>
      <td>\(0.064\)</td>
      <td>\(-0.413 \pm 0.044\)</td>
      <td>\(0.947\)</td>
      <td>\(16.8\)</td>
      <td>\(0.23\)</td>
    </tr>
  </tbody>
</table>

<p>The exponents look scattered, and taken at face value <code class="language-plaintext highlighter-rouge">cities</code> sits six standard errors
from \(-\tfrac12\). They are scattered not by dataset but by how far up in \(N\) each
grid reaches, because the power law is asymptotic and its small-\(N\) end is shallower
than \(-\tfrac12\). Refitting <code class="language-plaintext highlighter-rouge">counterfact</code> over truncated windows makes this explicit:</p>

\[b = -0.385 \pm 0.029 \;\; (N \le 400), \qquad
-0.445 \pm 0.019 \;\; (N \le 1500), \qquad
-0.490 \pm 0.019 \;\; (\text{all } N).\]

<p>Each of the other three then matches <code class="language-plaintext highlighter-rouge">counterfact</code> restricted to its own reach.
<code class="language-plaintext highlighter-rouge">cities</code> stops at \(1496\) and gives \(-0.446\) against \(-0.445\), <code class="language-plaintext highlighter-rouge">larger_than</code>
reaches \(1980\) and gives \(-0.484\) against \(-0.465\), and <code class="language-plaintext highlighter-rouge">sp_en_trans</code> stops at
\(354\) and gives \(-0.413\) against \(-0.385\), within its own error. With the fitting
window matched the four agree, so the dataset-independence of the exponent is measured
rather than only argued from the estimator, and the best estimate of the asymptotic
value is the one with the longest lever arm, \(-0.490 \pm 0.019\) over \(1.5\) decades.
That is the predicted \(N^{-1/2}\) to within half a standard error, and both it and its
amplitude \(0.045\) reproduce the pooled four-model fit above.</p>

<p>For the amplitude, \(a\) is the wrong statistic to compare across datasets, since each
\(a\) is defined at its own fitted \(b\). It is also not a free parameter. Under shuffled
labels there is no signal, so within a class \(x \sim \mathcal{N}(0, \Sigma)\), and the
mass-mean direction is a signed sum of the samples,</p>

\[\hat\theta \;=\; \hat\mu_{+} - \hat\mu_{-} \;=\; \frac{2}{N}\sum_i \epsilon_i x_i ,
\qquad \epsilon_i = \pm 1 ,\]

<p>so \(\hat\theta \sim \mathcal{N}\!\left(0, \tfrac{4}{N}\Sigma\right)\) and</p>

\[\mathbb{E}\lVert\hat\theta\rVert^{2} = \frac{4}{N}\operatorname{tr}\Sigma ,
\qquad
\mathbb{E}\,\hat\theta^{\top}\Sigma\,\hat\theta = \frac{4}{N}\operatorname{tr}\Sigma^{2} .\]

<p>The in-sample separation it produces is the projected class gap, which is its own
squared length, over the projected within-class spread,</p>

\[d'_{\text{in}} \;=\; \frac{\lVert\hat\theta\rVert^{2}}{\sqrt{\hat\theta^{\top}\Sigma\hat\theta}}
\;\simeq\; \frac{2}{\sqrt{N}}\,\frac{\operatorname{tr}\Sigma}{\sqrt{\operatorname{tr}\Sigma^{2}}}
\;=\; 2\sqrt{\frac{\mathrm{PR}}{N}} ,\]

<p>which is where the participation ratio enters:
\(\mathrm{PR} = (\operatorname{tr}\Sigma)^{2}/\operatorname{tr}\Sigma^{2}\) is the count
of directions the noise occupies, and it equals \(d\) only for a flat spectrum. With
\(\mathrm{AUROC} = \Phi(d'/\sqrt{2})\) and \(\Phi(x) \simeq \tfrac12 + x/\sqrt{2\pi}\) at
small argument,</p>

\[\mathrm{AUROC}_{\text{shuffled}} - \tfrac{1}{2}
\;\simeq\; \frac{d'_{\text{in}}}{2\sqrt{\pi}}
\;=\; \frac{1}{\sqrt{\pi}}\sqrt{\frac{\mathrm{PR}}{N}} ,\]

<p>so the collapse constant is predicted rather than fitted:</p>

\[C \;\equiv\; \Bigl(\mathrm{AUROC}_{\text{shuffled}} - \tfrac{1}{2}\Bigr)\sqrt{\frac{N}{\mathrm{PR}}}
\;=\; \frac{1}{\sqrt{\pi}} \;=\; 0.564 .\]

<p>Measured, \(C\) is flat to about \(\pm 10\%\) within a dataset over a thirtyfold range in
\(N\), and runs from \(0.51\) to \(0.65\) across <code class="language-plaintext highlighter-rouge">counterfact</code>, <code class="language-plaintext highlighter-rouge">cities</code> and
<code class="language-plaintext highlighter-rouge">larger_than</code>, against a factor \(2.0\) in \(a\) and \(3.9\) in \(\mathrm{PR}\). The
dependence on the spectrum runs the way the count predicts and against the intuitive
direction: <code class="language-plaintext highlighter-rouge">larger_than</code> is the most concentrated set, \(\mathrm{PR} = 8.9\) with
\(28\%\) of the within-class variance on one axis, and it carries the smallest amplitude
of the three.</p>

<p><img src="/assets/figures/truth_shuffled_collapse.png" alt="Left: the shuffled-label law. In-sample excess AUROC of the mass-mean direction under shuffled labels on `counterfact`, four Pythia widths, against $$N/2d$$, with the pooled fit $$0.045\,(N/2d)^{-0.49}$$ and Cover's capacity marked. The three pythia-70m points with negative excess are not drawn. The 70m points sit below the pooled line throughout, which is the across-model form of the question the right panel settles across datasets: whether the ambient width is the right denominator. Right: the effective-dimension collapse on pythia-2.8b, each dataset at its own best layer, against $$N/\mathrm{PR}$$. The line is $$\pi^{-1/2}\sqrt{\mathrm{PR}/N}$$ with no free parameter. `counterfact`, `cities` and `larger_than` fall on it across a thirtyfold range in $$N$$ and a fourfold range in $$\mathrm{PR}$$; `sp_en_trans` sits above it by $$36$$–$$60\%$$. Error bars are the standard error over sixteen permutations." /></p>

<p><code class="language-plaintext highlighter-rouge">sp_en_trans</code> does not join the collapse. It sits at \(C = 0.77\) to \(0.90\), well above
\(1/\sqrt{\pi}\), where the other three bracket it. The derivation factors into two
independent steps, and measuring them separately localizes the miss. The distributional
step holds: \(\mathrm{AUROC}\) tracks \(\Phi(d'_{\text{in}}/\sqrt{2})\) to within \(3\%\)
on all four datasets, this one included, despite its projections carrying the largest
excess kurtosis of the four (\(+0.8\) to \(+1.2\) against \(-0.7\) to \(+0.2\) elsewhere).
The discrepancy is in the second-moment step: the measured \(d'_{\text{in}}\) runs
\(1.47\) to \(1.62\) times \(2\sqrt{\mathrm{PR}/N}\) here, against \(1.03\) to \(1.17\)
on the other three.</p>

<p>Two explanations were tested which both fail. Near-duplicate items would inflate an in-sample fit by making the effective sample smaller than \(N\). However, <code class="language-plaintext highlighter-rouge">sp_en_trans</code> has \(344\) distinct Spanish
words across its \(354\) statements, while <code class="language-plaintext highlighter-rouge">cities</code> has exactly two statements per
subject and collapses normally. Under-measurement of
\(\mathrm{PR}\) is the other candidate, since \(\mathrm{PR}\) is itself a sample quantity that rises 
without saturating below \(N \approx 1000\), and <code class="language-plaintext highlighter-rouge">sp_en_trans</code> can only be read at
\(N = 354\). But restricting each of the other three to a pool of \(354\) statements and re-reading \(\mathrm{PR}\) from that pool leaves them between \(1.08\) and \(1.15\), nowhere near \(1.5\).
The anomaly is a property of <code class="language-plaintext highlighter-rouge">sp_en_trans</code> rather than of its size. It sits in the second-moment step rather than in the shape of the projection, and it is not explained here.</p>

<p>The sweep varies the dataset at fixed model and identifies the quantity that
sets the prefactor at a given width. It does not test whether \(\mathrm{PR}\) should
replace \(2d\) across the scale ladder, where the two co-vary, and the pythia-70m points
below the pooled line in the figure are that question in its across-model form. The
inversion table therefore remains a guide at the ambient width, read upward for a
dataset whose within-class noise is spread out and downward for one whose noise is
concentrated.</p>

<h2 id="a-steering-effect-with-the-wrong-sign">A steering effect with the wrong sign</h2>

<p>The decoding analysis says when a truth direction is <em>readable</em>. Steering asks the
separate question of whether it is <em>causal</em>, that is, whether displacing the residual stream
along \(\hat\theta\) moves the model’s behavior toward the true completion. The two
need not agree. On <code class="language-plaintext highlighter-rouge">counterfact_true_false</code> at the deep layers of pythia-2.8b the
mass-mean direction is neither readable nor correctly causal. For the readout, the held-out AUROC at layer \(28\) is \(0.502\) against a null of
\(0.560\), and it stays inside the null at every layer except for the last. As an intervention
it produces a significant effect with the wrong sign.</p>

<p>As defined above, steering is reported as the antisymmetric response \(A\) in
class-gap units. A genuine truth direction produces \(A &gt; 0\), and the random-direction
steering null fixes the bar \(A\) must clear. Across ten seeds at layer \(28\), at a
displacement of one class gap, the mass-mean direction returns</p>

\[A(1) &lt; 0 \;\text{ in } 8/10 \text{ seeds, each at } p \le 0.003,
\qquad \operatorname{med}_{\text{seeds}} A(1) = -0.062,
\qquad \sigma_{\text{null}} = 0.012,\]

<p>where \(p\) is the fraction of the \(400\)-draw random-direction null of \(A(1)\) at
least as extreme in the same direction, evaluated seed by seed. The seed median sits
five null standard deviations below zero, and the two positive seeds sit inside the
null. The linear steering susceptibility \(\chi\), the through-origin slope of \(A\)
against \(h\) over \(h \le 4\), is \(-0.023\) at this layer. It is a diluted summary of
the same effect. The response is not linear in \(h\), and by \(h = 4\), four class gaps
out, \(A\) has already fallen back toward zero, which is the regime <em>The scale of an
intervention</em> warned about. Layer \(24\) gives the same sign less cleanly: \(7/10\)
seeds negative with a median \(A(1) = -0.037\) against a null standard deviation of
\(0.015\), and one seed strongly positive. The effect is not null. It is
<em>significantly wrong-signed</em>. Displacing an activation toward the
true-class centroid, along the very vector the estimator returns for truth, makes the model
measurably <strong>less</strong> likely to produce the true completion.</p>

<p>Because \(\hat\theta\) reads at chance here, that significance carries the whole claim.
A direction sitting inside its decoding null is expected to do nothing under
intervention, and a small negative point estimate on its own would not be
distinguishable from noise. What makes this a result rather than a null is that the
effect clears a \(400\)-draw random-direction null in the wrong direction, reproducibly
across seeds and at two layers.</p>

<p>The slope of behavioral response against steering coefficient is not a new summary.
It is the <em>steerability</em> of <a href="https://arxiv.org/abs/2407.12404">Tan et al.
(2024)</a>, who fit exactly this line through a
logit-difference propensity curve. Their intervention displaces the last token
position only, where the harness here displaces every position, so magnitudes are not
directly comparable across the two. What is added here is the split into even and odd
parts, so that degradation and signed effect are separated, and the random-direction
null that says which susceptibilities are distinguishable from chance.</p>

<p>Read naively, this is a failure of the linear picture. Displacing along the direction
fit to read truth does not raise the probability of the true completion. It lowers it.
That reading would put
<code class="language-plaintext highlighter-rouge">counterfact</code> in the column of cases where the truth direction is an artifact of the
readout, epiphenomenal to the behavior. It does not even earn that description. At this
depth \(\hat\theta\) sits inside its decoding null, so there is no separation for the
steering result to be epiphenomenal to. However, I use the rest of this post to argue
that the naive reading is wrong: the wrong sign is a property of the <em>estimator</em>, not of
the activation geometry. What that means precisely is that the same activations, at the
same layer, contain a direction along which displacement moves the model toward the true
completion. But the mass-mean rule does not return that direction. It returns \(\hat v_1\),
the axis of largest within-class variance, with which \(\hat\theta\) shares a cosine of
\(0.994\) at this depth. The weight of that claim rests on the correction rather than on
a mechanism for the wrong sign: replacing the mass-mean rule by the Fisher rule, inside
the same linear class, recovers the correct sign, while what makes the plain direction negative at layer \(28\) turns out to lie
beyond first order and is not settled here. Change the rule, either by downweighting that axis or by
deleting it outright, and the correct sign comes back out of the same data. The claim is
about which part of the geometry the estimator points at, not about how much truth the
layer encodes. How well the corrected direction reads is a separate question, quantified
in the next section by its held-out AUROC against the decoding null.</p>

<p><img src="/assets/figures/truth_steering_signflip.png" alt="Whitening reverses the sign of causal steering. Left: antisymmetric
steering response $$A(h)$$ on `counterfact_true_false` / pythia-2.8b at $$L=28$$. The plain mass-mean
direction (red) drives $$A$$ significantly negative, so steering toward the true centroid
suppresses the true completion, while the whitened direction (blue), differing only
by the covariance correction, drives it positive. Gray band: 5th–95th percentiles
of the random-direction null, so a point outside it in its own direction is a
one-sided clearance at 5%. Right: steering susceptibility $$\chi$$, the through-origin slope of $$A$$ against $$h$$,
across depth, median over ten seeds with inter-quartile bars, the median rather than
the mean because the plain direction's seeds are bimodal, eight clustered near $$-0.024$$ and
two positive, so a mean lands on a value no seed exhibits. Plain is wrong-signed at
$$L=24, 28$$ ($$2$$–$$3/10$$ seeds positive), whitened is correct-signed throughout ($$\ge 9/10$$), with
$$L=20$$ the crossover." /></p>

<h2 id="the-rogue-dimension">The rogue dimension</h2>

<p><strong>Definition.</strong> I call \(\hat v_1\) a <strong>rogue dimension</strong> when it dominates the
<em>within-class</em> covariance, that is, \(\hat\lambda_1/\operatorname{tr}\hat C \approx 1\) or equivalently \(\mathrm{PR} \approx 1\) and \(\hat\lambda_1/\hat\lambda_2 \gg 1\). This
condition is a statement about the noise geometry alone and is basis-free and independent
of any probe. The estimator only becomes involved through the separate question of
whether \(\hat\theta\) has aligned with it.</p>

<p>This is deliberately not the same object as the <strong>massive activations</strong> of Sun et al.,
which are individual coordinates whose magnitude far exceeds the median, nor the
<strong>rogue dimensions</strong> of Timkey &amp; van Schijndel, which are coordinates that dominate
cosine similarity. Those are properties of the mean and of a basis. This is a property
of the covariance and of no particular basis. The distinction is important here. At
<code class="language-plaintext highlighter-rouge">counterfact</code> layer \(8\) the two coincide: \(\hat v_1\) carries \(86\%\) of its mass on a
single coordinate, and that coordinate’s mean activation is \(1446\times\) the median
across coordinates, a massive activation by their criterion. By layer \(28\) the
eigenvector has delocalized, \(37\%\) on its largest coordinate and \(90\%\) across five,
so the rogue dimension is no longer any one neuron. <code class="language-plaintext highlighter-rouge">cities</code> at layer \(28\) still <em>has</em>
the massive activation, the same coordinate at \(563\times\) the median, and has no rogue
dimension at all. Its leading eigenvector holds half a percent of its mass on any
coordinate. A massive activation is nearly constant across statements,
so it inflates the mean without inflating the within-class covariance. It produces a
rogue dimension only when the remaining variance is small enough for it to dominate.</p>

<p>The diagnosis is in the spectrum of the within-class covariance. I diagonalize
\(\hat C = \sum_i \hat\lambda_i \hat v_i \hat v_i^{\top}\) with sample eigenvalues
\(\hat\lambda_1 \ge \hat\lambda_2 \ge \dots\) and eigenvectors \(\hat v_i\),
and measure where the mass-mean direction sits relative to its leading eigenvector.
On <code class="language-plaintext highlighter-rouge">counterfact</code> at pythia-2.8b, layer \(28\):</p>

\[\frac{\hat\lambda_1}{\operatorname{tr}\hat C} = 0.980, \qquad
\frac{\hat\lambda_1}{\hat\lambda_2} = 535, \qquad
\lvert\cos(\hat\theta, \hat v_1)\rvert = 0.994 .\]

<p>A single eigendirection carries \(98\%\) of the within-class
variance. It is \(535\) times larger than the next, and the mass-mean direction is
almost perfectly aligned with it. Therefore, the estimator has not returned a truth direction, it has returned \(\hat v_1\), the dominant axis of the within-class noise. On this
dataset there is essentially no class-gap signal for \(\hat\delta\) to lock onto, so the
finite-sample \(\hat\delta\) is dominated by its projection onto the highest-variance
axis. This axis is the salient direction of the superposition test above. When the class
gap is negligible, the leading direction of the total activation covariance and the
within-class \(\hat v_1\) coincide, so the two diagnostics see the same axis. They
separate only once the gap grows. The participation ratio
\(\mathrm{PR} = (\sum_i\hat\lambda_i)^2/\sum_i\hat\lambda_i^2\) puts the collapse on a
dimension-free scale: \(\mathrm{PR} \to 1\) when one eigenvalue dominates the spectrum
and \(\mathrm{PR} \to d\) when the spectrum is flat, so it counts the directions the
noise effectively occupies.</p>

<p>Read across depth, those observables say something sharper than the layer-\(28\)
snapshot. What decides recoverability is <em>not</em> whether a dominant axis exists. At layer
\(8\), <code class="language-plaintext highlighter-rouge">cities</code> is collapsed too, with
\(\hat\lambda_1/\operatorname{tr}\hat C = 0.723\), \(\lvert\cos(\hat\theta,\hat
v_1)\rvert = 0.914\) and \(d'_{\mathrm{mm}} = 0.11\). That is <code class="language-plaintext highlighter-rouge">counterfact</code>’s condition,
not a milder version of it. The two sets differ in what becomes of that condition at
later layers. By layer \(28\) it no longer holds for <code class="language-plaintext highlighter-rouge">cities</code>. The participation ratio
climbs \(1.9 \to 29.7\), the leading axis falls to \(11\%\) of the variance, the alignment
drops to \(0.159\), and \(d'_{\mathrm{mm}}\) reaches \(3.14\). <code class="language-plaintext highlighter-rouge">counterfact</code> never moves. Its participation
ratio stays at \(1.0\)–\(1.04\) at every layer, its alignment stays near \(1\), and
\(d'_{\mathrm{mm}} = 0.09\) throughout. These observables are computed on the full set
rather than the held-out half, since they describe the geometry rather than a probe’s
performance, and the held-out \(d'_{\mathrm{mm}}\) at <code class="language-plaintext highlighter-rouge">counterfact</code> layer \(28\) is
\(0.080\) against the \(0.089\) quoted here. The rogue dimension is the default
condition, not the pathology. Nor is it a property of <code class="language-plaintext highlighter-rouge">counterfact</code> alone:
<code class="language-plaintext highlighter-rouge">companies_true_false</code> and <code class="language-plaintext highlighter-rouge">common_claim_true_false</code> carry the same signature at layers
\(24\) and \(28\), \(\mathrm{PR}\) within \(0.05\) of \(1\) and alignment above \(0.8\),
and escape it by layer \(31\) where their class gap has grown, as the appendix <em>Results
on the full twelve-dataset benchmark</em> records.
Recoverability is whether the class gap ever grows large enough to pull \(\hat\delta\)
off that axis and clear the null.</p>

<p>Read against the decoding null, the two datasets separate cleanly, and the correction
makes the truth direction recoverable on <code class="language-plaintext highlighter-rouge">counterfact</code>. On <code class="language-plaintext highlighter-rouge">cities</code> the
mass-mean direction leaves the null band at layer \(14\) and stays above every
subsequent layer, climbing above \(0.95\) from layer \(25\) and peaking at \(0.975\) at
layer \(28\). The whitened direction clears the null earlier, at layer \(7\).
On <code class="language-plaintext highlighter-rouge">counterfact</code> the mass-mean direction stays
inside the null band at every layer but the last, where it reaches \(0.566\) against a null of
\(0.556\), while the whitened direction leaves the band at layer \(26\) and climbs to
\(0.716\) against \(0.552\) at layer \(31\). The layer-\(32\) value is the maximum of a
\(5\%\) test taken over \(33\) layers, so it is a selected extreme rather than a
recovery. The rank-one direction \(\hat\theta_\perp\), measured on a six-layer grid in the rogue-dimension
sweep, tracks the whitened direction: on <code class="language-plaintext highlighter-rouge">counterfact</code> it sits at chance through layer \(24\) and
reaches \(0.574\) at layer \(28\), and on <code class="language-plaintext highlighter-rouge">cities</code> it is already at \(0.921\) by layer
\(12\). Deleting the axis and downweighting it do the same work.</p>

<p>There is a sharper way to explain why the plain probe cannot leave the band on
<code class="language-plaintext highlighter-rouge">counterfact</code>. When \(\delta\) lies along \(\hat v_1\) and \(\hat v_1\) carries nearly all
of \(\hat C\), a random direction \(u\) reads the class gap and the noise through the
same coordinate, \(u^{\top}\delta \approx \lVert\delta\rVert u_1\) and
\(u^{\top}\hat C u \approx \hat\lambda_1 u_1^{2}\), so</p>

\[d'(u) \;\approx\; \frac{\lVert\delta\rVert\,\lvert u_1\rvert}{\sqrt{\hat\lambda_1}\,\lvert u_1\rvert}
\;=\; \frac{\lVert\delta\rVert}{\sqrt{\hat\lambda_1}} \;=\; d'_{\mathrm{mm}} ,\]

<p>and \(u_1\) cancels. For every random direction the projected gap and the projected
noise are carried by the same component, so their ratio is the one \(\hat\theta\) itself
gives, and the null collapses onto the mass-mean value. The sweep shows it
to the third decimal. On <code class="language-plaintext highlighter-rouge">counterfact</code> the held-out \(d'_{\mathrm{mm}}\) is \(0.082\),
\(0.081\), \(0.081\), \(0.080\) at layers \(8\), \(20\), \(24\), \(28\), and the median of
the random-direction null at the same layers is \(0.083\), \(0.081\), \(0.079\), \(0.078\).
The plain probe is not merely inside its null. Its \(d'\) is indistinguishable from
that of a random direction, and the spread above that median, \(p_{95}\) from \(0.094\) to \(0.139\), is what the two percent of variance off the axis contributes. That is what locked to the axis means
operationally, and it is why more data would not lift \(\hat\theta\) clear of the band
while the alignment holds: the null and the estimate move together.</p>

<p><img src="/assets/figures/truth_depth_arms.png" alt="Held-out AUROC against depth for both estimators on pythia-2.8b, with the per-layer random-direction null shaded from $$1/2$$ to its 95th percentile. A curve inside the band is not recoverable, and height above the band is the margin $$m$$ that layer selection maximizes. Left, `counterfact`: the mass-mean direction sits inside the band until layer $$32$$, while the whitened direction leaves it at layer $$26$$ and reaches $$0.716$$ at layer $$31$$. Right, `cities`: the mass-mean direction leaves the band for good at layer $$14$$ and the whitened one at layer $$7$$, saturating near $$0.98$$ and $$0.99$$. Stars are the rank-one corrected direction $$\hat\theta_\perp$$, which projects $$\hat v_1$$ out of the estimator and is measured on a six-layer grid rather than at every layer. Dashed lines mark the first layer from which each direction stays clear of its null. Note the different vertical scale of the achievement: the same correction that lifts `counterfact` from chance to $$0.72$$ is barely needed where the class gap is strong." /></p>

<p><img src="/assets/figures/truth_clusters.png" alt="Within-class geometry on `cities` (top) and `counterfact` (bottom), pythia-2.8b layer 28. Left: activations in the plane of the top two within-class principal axes $$x\cdot\hat v_1$$, $$x\cdot\hat v_2$$, colored by truth label — on `counterfact` a single axis carries nearly all the variance. Middle: the mass-mean projection $$x\cdot\hat\theta$$, with its separation $$d'$$ and its alignment $$\lvert\cos(\hat\theta,\hat v_1)\rvert$$, near 1 on `counterfact` where the estimator has collapsed onto the leading variance axis. Right: the within-class eigenvalue spectrum, $$\hat\lambda_1/\hat\lambda_2$$. The eleven points standing clear of the bulk on `counterfact` — the scattered group near $$x\cdot\hat v_1 \approx -200$$ in the bottom-left panel, away from the dense blob at $$\approx -1100$$, and the small bar near $$x\cdot\hat\theta \approx 200$$ in the bottom-middle one — are the statements that do not receive the massive activation. They are what the leading eigendirection is measuring, and what sets the axis range of both panels." /></p>

<p>\(\hat v_1\) is also <em>well</em> estimated, which is what makes the collapse of \(\hat\theta\)
onto it a feature of \(\Sigma\) rather than an accident of the sample. The reason is not
the spectral gap by itself. With \(d = 2560\) and \(N = 1198\) the sample covariance is
rank-deficient, so no deterministic perturbation bound applies without first controlling
\(\lVert\hat C - \Sigma\rVert\), and in this regime a sample eigenvector is in
general an attenuated estimate of its population axis. What controls the attenuation is
the spike strength relative to the aspect ratio. Writing \(\gamma = d/N = 2.14\) and
\(\ell\) for the ratio of \(\hat\lambda_1\) to the bulk scale, the relative bias of
\(\hat\lambda_1\) and the squared overlap deficit of \(\hat v_1\) are both of order
\(\gamma/\ell\). Taking \(\hat\lambda_2\) as the bulk gives \(\ell = 535\) and
\(\gamma/\ell = 4\times10^{-3}\), and taking the mean of the nonzero bulk eigenvalues gives
\(\ell = 5.8\times10^{4}\). Either way the spike sits far above the threshold \(1 + \sqrt\gamma = 2.46\), in this
ratio convention, below which the leading eigenvector would carry no information about
its population axis at all.</p>

<p>Those asymptotics assume rows with bounded fourth moments. The next section shows that
this spike is generated by eleven statements out of \(1198\) carrying a single massive
coordinate, which violates that assumption, so the spike-strength argument is indicative
rather than tight here. That \(\hat v_1\) is a population axis and not a finite-sample
artifact rests instead on two direct checks. Identifying the eleven statements
coordinate-first, by the size of that one coordinate and without forming the
covariance, returns the same eleven at every layer, and a two-point variance formula
built from those eleven predicts \(97\)–\(99\%\) of \(\hat\lambda_1\). A direction that
can be derived by two disjoint routes is not a finite-sample accident. \(\hat\theta\), by
contrast, is a poor estimate of \(\theta\), precisely because the class-gap signal is
weak.</p>

<p>The obvious objection is that this is a fact about Pythia. However, the same observables on
<a href="https://huggingface.co/allenai/OLMo-2-0425-1B">OLMo-2-1B</a>, a different architecture and
training corpus at a third of the parameters, reproduce the pattern. On <code class="language-plaintext highlighter-rouge">counterfact</code> the
mass-mean direction stays aligned with the leading axis at every depth except the final
layer, and the plain probe
never clears its null while whitening does, and on <code class="language-plaintext highlighter-rouge">cities</code> the alignment falls, the
participation ratio climbs, and the plain probe clears the null as it does in Pythia. The
numbers are in the appendix <em>Replication on OLMo-2-1B</em>.</p>

<p>The superposition probe set up in <em>The geometry</em> reads the same object from the other
side, and on Pythia it separates the datasets the same way. On <code class="language-plaintext highlighter-rouge">cities</code> the separation
decays steadily once the leading components go, \(d'_{\mathrm{mm}} = 2.93 \to 0.25\), and
<code class="language-plaintext highlighter-rouge">larger_than</code> decays likewise. On <code class="language-plaintext highlighter-rouge">neg_cities</code> it is untouched until the top two are
removed (\(3.03\) at \(k = 2\)) and only then collapses, so its truth direction sits below
the most salient axes rather than in them. <code class="language-plaintext highlighter-rouge">counterfact_true_false</code> does neither. Its
\(d'_{\mathrm{mm}}\) <em>rises</em>, \(0.20 \to 0.44\), so stripping the leading directions makes
the truth signal <strong>better</strong>. <code class="language-plaintext highlighter-rouge">companies_true_false</code> exhibits the same effect more strongly, \(d'_{\mathrm{mm}} = 0.17 \to 1.09\)
at \(k = 8\) before falling back to \(0.55\) at \(k = 64\). When removing the leading
direction <em>improves</em> recovery, that direction is the rogue dimension seen through the
superposition probe rather than the spectrum. Datasets without a rogue dimension lose
signal when the same components are stripped.</p>

<p><img src="/assets/figures/truth_superposition.png" alt="The salience knob: $$d'$$ of the mass-mean direction after projecting out the top-$$k$$ principal components, at each dataset's best layer on pythia-2.8b. On `cities`, `neg_cities` and `larger_than` the separation decays as the leading directions are removed — the truth signal is partly contained in the salient subspace. On `counterfact` it rises instead, from $$0.20$$ to $$0.44$$, so the leading directions are not carrying the truth signal but obscuring it." class="fig-single" /></p>

<h3 id="where-the-rogue-dimension-comes-from">Where the rogue dimension comes from</h3>

<p>The definition in the previous section separates the rogue dimension, a property of
the within-class covariance, from the massive activation, a property of one coordinate. On <code class="language-plaintext highlighter-rouge">counterfact</code> they coincide at layer \(8\), where
\(\hat v_1\) carries \(86\%\) of its mass on the massive coordinate, and differ by layer
\(28\), where the eigenvector has delocalized across five coordinates. On <code class="language-plaintext highlighter-rouge">cities</code> at
layer \(28\) the massive activation is present and the rogue dimension is absent. However, the
relation is closer than coincidence. On <code class="language-plaintext highlighter-rouge">counterfact</code> the leading axis of the
covariance is <em>generated</em> by the statements on which the massive activation drops.</p>

<p>A coordinate that is exactly constant across statements contributes nothing to
\(\hat C\), however large it is. A coordinate that takes a large value \(a\) on a
fraction \(1-p\) of statements and a much smaller value \(b\) on the remaining \(p\)
contributes the variance of a two-point distribution,</p>

\[\operatorname{Var}[x_j] \;=\; p(1-p)\,(a-b)^{2},\]

<p>which for small \(p\) and large \(a-b\) can be enormous. At <code class="language-plaintext highlighter-rouge">counterfact</code> layer \(8\),
four coordinates qualify as massive activations under the magnitude criterion of
Sun et al. On each, \(1187\) of the
\(1198\) statements sit at a near-constant value and eleven sit far below it. Coordinate \(1793\) reads \(1076.3 \pm 2.7\) on the majority and
\(38.4 \pm 2.1\) on the eleven. With \(p = 11/1198\) the formula predicts \(9800\) against a measured
coordinate variance of \(9788\). Summed over the four coordinates it gives \(11{,}271\)
against \(\hat\lambda_1 = 11{,}413\), and \(\hat v_1\) carries \(98.6\%\) of its mass in
their span. At layer \(28\) the same holds with eight such coordinates: \(7249\) against
\(\hat\lambda_1 = 7500\), with \(96.6\%\) of \(\hat v_1\) in their span. In this case \(\hat v_1\) carries no truth content. It is, to within a few
percent, the <em>indicator of which statements failed to receive the massive activation</em>.
The mass-mean estimator’s \(0.994\) alignment with it is an alignment with an eleven-out-of-1198
membership function.</p>

<p>The eleven statements are identifiable coordinate-first, by a rule that never touches the
covariance, the class means, or the labels, and it returns the same eleven at
every layer. Refitting with them excluded collapses the geometry at layer \(28\), taking
\(\hat\lambda_1/\hat\lambda_2\) from \(535\) to \(1.1\) and
\(\lvert\cos(\hat\theta,\hat v_1)\rvert\) from \(0.994\) to \(0.063\), and it leaves
held-out \(d'_{\mathrm F}\) unchanged at \(0.385\). At that layer whitening leaves
nothing further for deletion to recover. Nothing rescues the plain estimator, whose
\(d'_{\mathrm{mm}}\) stays inside its null under every removal regime tested. How the
eleven are identified, what removing them does to \(d'\) at each layer, and how that
comparison depends on the Ledoit–Wolf intensity are given in the appendix
<em>Identification and removal of the eleven outlier statements</em>.</p>

<p>The massive coordinates themselves are not dataset-specific. The same coordinates
qualify on <code class="language-plaintext highlighter-rouge">cities</code> and <code class="language-plaintext highlighter-rouge">counterfact</code> alike. What differs is whether any statement
drops these coordinates. On <code class="language-plaintext highlighter-rouge">cities</code> none do, so the coordinates stay constant, contribute nothing to
\(\hat C\), and leave no rogue dimension behind. That is the earlier observation that
<code class="language-plaintext highlighter-rouge">cities</code> carries the same massive activation without the same pathology, now with a
mechanism rather than a coincidence. The incidence is specific to the dataset but the
mechanism is general. The per-layer incidence on both models, including pythia-1.4b, is
in the appendix <em>Identification and removal of the eleven outlier statements</em>.</p>

<p>While \(\hat\theta\) and \(\hat v_1\) are aligned geometrically, it is unclear whether this is causal. The steering sweep includes a condition that displaces along \(\hat v_1\) itself, the
leading eigenvector of the within-class covariance. Because \(\hat v_1\) is fit after both class
means are removed, it is defined purely by within-class scatter and carries no
information about which statements are true. At layer \(28\) the antisymmetric response
to \(\hat v_1\) is \(-0.051\) against \(\hat\theta\)’s \(-0.047\). At layer \(20\) the two agree to
within \(10^{-4}\) (\(-0.0390\) against \(-0.0391\)). At every depth measured they track each
other to within a few thousandths. They decode alike too, at held-out
\(\mathrm{AUROC} = 0.511\) for \(\hat v_1\) against \(0.514\) for \(\hat\theta\) in the
six-layer rogue-dimension sweep, both inside the null. The \(0.994\) alignment is a
geometric statement. The steering agreement is its behavioral form: a direction fit
without any reference to the labels moves the model as much as \(\hat\theta\) does and in
the same direction. That establishes that truth content is not what carries the effect.
It does not add independent evidence about the size of the effect, since under
\(\chi(w) = c\,(w\cdot g)\), derived below, two directions with cosine \(0.994\) must
produce nearly equal susceptibilities.</p>

<p>This reframes the anomaly. Steering along \(\hat\theta \approx \hat v_1\) is not steering
along truth. It is displacing the activation along the dominant axis of the within-class
noise, which perturbs the forward pass in a way that happens to suppress the true completion.
A nuisance direction carries no information about the label, so nothing requires its
behavioral effect to come out positive. Nothing requires it to be reproducible either,
but here it is: the same negative sign appears across ten seeds, at two <code class="language-plaintext highlighter-rouge">counterfact</code>
layers, and at mid-depth on <code class="language-plaintext highlighter-rouge">cities</code>. Whatever produces it is systematic rather than
arbitrary, and the label-blindness of \(\hat v_1\) says only that truth is not what
produces it.</p>

<p>If that account is right, removing the contribution of \(\hat v_1\) should recover a
correctly signed causal effect, by correcting the estimator’s alignment with the rogue
dimension rather than by enriching the function class.</p>

<h3 id="what-steering-measures">What steering measures</h3>

<p>Before testing that prediction, it is worth asking what the steering number reports. The
susceptibility is a linear functional of the direction pushed. Writing
\(g = \big\langle \nabla_x \ell(x)\big\rangle\) for the mean gradient of the behavioral
score over the evaluation set, the odd part of the response gives</p>

\[\chi(w) \;=\; \left.\frac{\mathrm{d}A}{\mathrm{d}h}\right|_{h\to 0}
\;=\; c\,\big(w \cdot g\big)\]

<p>to leading order, the even terms having been removed by the antisymmetrization. So
steering in the small-\(h\) window reports the overlap of \(w\) with the single vector
\(g\) along which the model’s true-versus-false margin moves. It does not test whether
\(w\) is a truth direction.</p>

<p>Every susceptibility quoted above is therefore a
projection of one vector. This is why the \(\hat v_1\) and \(\hat\theta\) measurements agreeing
to within \(10^{-4}\) at layer \(20\) is a restatement of their \(0.994\) alignment rather
than independent confirmation.</p>

<p>Decomposing in the plane of the next section,
\(\hat\theta = \cos\varphi\,\hat v_1 + \sin\varphi\,\hat e_2\), linearity gives</p>

\[\chi(\hat\theta) \;=\; \cos\varphi\;\chi(\hat v_1) \;+\; \sin\varphi\;\chi(\hat e_2),\]

<p>and at \(\cos\varphi = 0.994\) the first term dominates unless \(\chi(\hat e_2)\) is two
orders of magnitude larger. One consequence holds before any measurement: a
direction that merely disrupts the computation contributes to \(S\), not to \(A\), since
\(A\) is odd by construction. A wrong-signed \(A\) therefore requires a <em>signed</em> channel
onto the margin, \(\hat v_1 \cdot g &lt; 0\), which the estimator would inherit in
proportion to \(\cos\varphi\). Whether that channel exists is a question the gradient
answers directly.</p>

<p>This also makes the assumption behind the correction explicit. Projecting out
\(\hat v_1\) recovers the sign <strong>only if \(g\) has little weight along \(\hat v_1\)</strong>, i.e.
the behaviorally causal direction and the salient axis are close to orthogonal in
activation space. When they are not, \(\chi(\hat v_1)\) mixes a truth term with a nuisance
term inseparably, so deleting the axis would delete real causal signal along with the
artifact.</p>

<p><strong>Measured.</strong> One backward pass per example gives \(\nabla_x \ell\), and \(g\) follows by
averaging over the same \(250\) contrastive pairs the steering measurements use. Because the
gradient is computed per pair, every overlap below carries a bootstrap interval over
those pairs rather than resting on the point estimate. I quote \(95\%\) intervals from
\(10{,}000\) bootstrap resamples under a fixed seed, and
the scale to keep in mind is that a random direction gives \(\lvert\cos\rvert \approx
1/\sqrt{d} = 0.020\) at \(d = 2560\).</p>

<p>At layer \(28\), \(\cos(g,\hat v_1) = -0.011\) with \(95\%\) confidence interval
\([-0.038, +0.020]\). The overlap is consistent with zero. That has two consequences. Projecting \(\hat v_1\) out
removes essentially none of the behavioral channel, which is what licenses the
correction. A channel consistent with zero cannot produce the wrong sign at first
order, so \(A(\hat v_1) = -0.051\) at this layer is not a linear effect. It is beyond first order in the displacement.</p>

<p>At layer \(20\) the overlap is resolved: \(\cos(g,\hat v_1) = -0.049\), confidence
interval \([-0.061, -0.030]\), and \(\cos(g,\hat\theta)\) is the same through the \(0.994\)
alignment. The linear prediction \(c\,(\hat\theta\cdot g) = -0.078\) then lands
within about \(20\%\) of the measured \(A(0.5)/0.5 = -0.064\). The through-origin susceptibility
hides this, because the response reverses sign with displacement, from \(-0.032\) at
\(h = 0.5\) (\(7/10\) seeds negative, \(2.5\) null standard deviations) through zero near
\(h = 1\) to \(+0.016\) at \(h = 2\) (\(9/10\) seeds positive), so the slope fit over
\(h \le 4\) comes out null. The two layers therefore differ in kind. At layer \(20\) the
wrong sign is a first-order effect confined to the linear window. At layer \(28\) the
same prediction gives \(-0.002\) against a measured \(-0.039\) at \(h = 0.5\), and the
wrong sign persists to \(h = 2\) before its magnitude falls at \(h = 4\).</p>

<p>The overlap with \(g\) also separates the two estimators. On <code class="language-plaintext highlighter-rouge">counterfact</code> the whitened
direction has a positive overlap at every layer from \(20\) on, \(+0.040\)
\([+0.017, +0.056]\) at \(L20\), \(+0.065\) \([+0.034, +0.079]\) at \(L24\) and \(+0.064\)
\([+0.028, +0.080]\) at \(L28\), while the plain direction is consistent with zero at
\(L24\) and \(L28\). On <code class="language-plaintext highlighter-rouge">cities</code> it is the plain direction that has the positive overlap
at layer \(28\), \(+0.050\) \([+0.032, +0.062]\), while the whitened direction does not. In
both datasets the direction that steers correctly is the direction whose overlap with
\(g\) is positive, and the overlap is measured without displacing an activation. The
sign and not the magnitude is what carries this: at <code class="language-plaintext highlighter-rouge">counterfact</code> layer \(20\) both
directions have resolved overlaps, the whitened one positive and the plain one negative.</p>

<p>The rank-one correction gives the sharpest test of this, because
\(\hat\theta_\perp\) is exactly \(\hat e_2\), so its susceptibility is predicted by
\(\cos(g, \hat e_2)\) alone, measured before anything is steered. On <code class="language-plaintext highlighter-rouge">counterfact</code> the
prediction has the right sign at all six layers of the rogue-dimension sweep, and where
the overlap is resolved the steering is significant and agrees with it. At layer \(8\)
the overlap is negative, \(-0.048\) \([-0.062, -0.025]\), the prediction is
\(c\,(\hat e_2\cdot g) = -0.12\), and \(\hat\theta_\perp\) steers significantly
wrong-signed, \(A(1) = -0.104\) over three seeds (\(p = 0.02\) against a \(100\)-draw
null). At layers \(24\) and \(28\) the overlap is positive, \(+0.059\) and \(+0.073\),
and \(\hat\theta_\perp\) steers correctly, \(+0.032\) and \(+0.043\) against a predicted
\(+0.054\) at both. At layers \(12\) through \(20\) the overlap’s interval spans zero and
the steering stays inside its null. Removing \(\hat v_1\) corrects the sign only where
what remains of the class gap overlaps \(g\) positively.</p>

<p><code class="language-plaintext highlighter-rouge">cities</code> also reproduces the mid-depth dip from the other side. Its plain overlap turns
significantly <em>negative</em> at layer \(20\), \(-0.035\) \([-0.041, -0.027]\), which is the same
window where the steering sweep finds \(\chi = -0.0062\), an independent confirmation of
an anomaly that the steering measurement alone left uncertain.</p>

<p>There is a geometric reading of the two directions that makes the inversion of Marks &amp;
Tegmark’s assignment in <em>The geometry</em> less paradoxical than it first appears. A
readout is a covector and a steering direction is a vector. To first order a
displacement couples to one covector, the score gradient \(g\), so a direction steers
correctly when it overlaps \(g\). Marks &amp; Tegmark’s assignment is to steer along \(\delta\)
and read with \(\Sigma^{-1}\delta\). This is right whenever \(\delta\) overlaps \(g\), which
is the case on <code class="language-plaintext highlighter-rouge">cities</code>. On <code class="language-plaintext highlighter-rouge">counterfact</code>, \(\hat\delta\) lies along a nuisance axis with
no overlap with \(g\). \(\Sigma^{-1}\) removes that axis and restores the overlap.
This does not show that the model’s own readout is closer to \(\Sigma^{-1}\delta\) than
to \(\delta\) in general. On <code class="language-plaintext highlighter-rouge">cities</code> the plain direction carries the overlap but the
whitened one does not.</p>

<p>Two cautions on reading these magnitudes, the curvature at the large-\(h\) end and how
much of \(g\) the plane captures, are given in the methods appendix.</p>

<h2 id="the-correction-is-a-rotation-in-a-plane">The correction is a rotation in a plane</h2>

<p>Two corrections to the estimator, whitening and projecting out \(\hat v_1\), make the
same prediction, and neither adds capacity. Whitening replaces
\(\hat\theta \propto \hat\delta\) with the Fisher direction
\(\hat\theta_{\mathrm F} \propto \hat\Sigma^{-1}\hat\delta\), which downweights the
mean-shift component along high-variance axes, \(\hat v_1\) chief among them. The blunter
control simply projects \(\hat v_1\) out, \(\hat\theta_\perp \propto (I - \hat v_1 \hat v_1^{\top})\hat\delta\),
removing only the rogue dimension.</p>

<p>The two look like corrections of different rank, since the projection is exactly rank
one while \(\hat\Sigma^{-1}\) is not. However, in this regime they act inside the same plane. Letting
\(P = \mathrm{span}\{\hat v_1, \hat e_2\}\) with
\(\hat e_2 = \hat\delta_\perp / \lVert \hat\delta_\perp \rVert\) defined as the unit vector along the
part of the class gap orthogonal to the rogue axis, \(\delta_1\) and \(\delta_\perp\) can be written as
\(\delta_1 = \hat v_1^{\top}\hat\delta\) and \(\delta_\perp = \lVert\hat\delta_\perp\rVert\) respectively. \(\hat\Sigma\) can be restricted to \(P\) as \(\mathrm{diag}(\lambda_1, \bar\lambda)\). For a unit direction in the plane \(u(\varphi) = \cos\varphi\,\hat v_1 + \sin\varphi\,\hat e_2\),
the separation it achieves is a Rayleigh quotient in the single angle \(\varphi\):</p>

\[d'^{2}(\varphi) \;=\;
\frac{(\delta_1\cos\varphi + \delta_\perp\sin\varphi)^{2}}
{\lambda_1\cos^{2}\varphi + \bar\lambda\sin^{2}\varphi}.\]

<p>The mass-mean estimator sits at \(\tan\varphi_{\mathrm{mm}} = \delta_\perp/\delta_1\) and the
Fisher direction has plane components \((\delta_1/\lambda_1,\ \delta_\perp/\bar\lambda)\),
so the two are related by a single lever arm \(\kappa \equiv \lambda_1/\bar\lambda\):</p>

\[\tan\varphi_{\mathrm F} \;=\; \kappa \tan\varphi_{\mathrm{mm}} .\]

<p>This shows that whitening is simply a rotation within \(P\). The measured
\(\lvert\cos(\hat\theta,\hat v_1)\rvert = 0.994\) puts \(\varphi_{\mathrm{mm}} = 6.3^{\circ}\).
With \(\kappa = \hat\lambda_1/\hat\lambda_2 = 535\) this gives
\(\varphi_{\mathrm F} = 89.0^{\circ}\). These are full-sample values of \(\hat C\). The whitening
actually applied uses the shrunk \(\hat\Sigma\) of the training half, where
\(\kappa = 463\) and \(\varphi_{\mathrm{mm}} = 6.8^{\circ}\), and it returns the same
\(89.0^{\circ}\) and the same predicted gain below. The whitened direction is orthogonal to \(\hat v_1\)
to within a degree, which is why \(\hat\theta_{\mathrm F}\) and \(\hat\theta_\perp\) produce the same sign flip:
in this regime they are the same vector, and the full-rank correction reduces to the
rank-one one.</p>

<p><img src="/assets/figures/truth_plane_rotation.png" alt="Left: the plane $$\mathrm{span}{\hat v_1, \hat e_2}$$ at the `counterfact` numbers,
with the within-class noise ellipse ($$\hat\lambda_1/\hat\lambda_2 = 535$$, so a
$$23{:}1$$ axis ratio). The mass-mean direction (black) lies $$6.3^{\circ}$$ off the rogue
axis, inside the long axis of the noise. $$\hat\Sigma^{-1}$$ swings the Fisher direction
(gold) to $$89.0^{\circ}$$ which is effectively orthogonal to it. Right: the Rayleigh quotient
$$d'(\varphi)$$ normalized by its maximum, in the collapsed regime ($$\kappa = 535$$,
$$r = 0.11$$) and the resolved regime ($$\kappa = 1.42$$, $$r = 6.20$$). Circles mark the
mass-mean angle and squares the whitened angle. Where the spectrum has one dominant eigenmode the
curve is a cliff and the mass-mean estimator sits at its foot; where the class gap has
grown the curve is broad and both directions already sit near the
top." /></p>

<p>The quotient also predicts the size of the gain. Writing
\(r = \tan\varphi_{\mathrm{mm}}\),</p>

\[\frac{d'^{2}_{\mathrm F}}{d'^{2}_{\mathrm{mm}}}
\;=\; \frac{(\kappa^{-1} + r^{2})(\kappa + r^{2})}{(1+r^{2})^{2}}
\;\approx\; \frac{1 + \kappa r^{2}}{(1+r^{2})^{2}},\]

<p>which at \(\kappa = 535\) gives \(2.7\times\). Measured on <code class="language-plaintext highlighter-rouge">counterfact</code> at layer \(28\),
\(d'_{\mathrm{mm}} = 0.080 \to d'_{\mathrm F} = 0.384\), a factor of \(4.8\). The formula
predicts the right order and underestimates the measured gain. Inverting it for an
effective \(\kappa\) is worked through in the methods appendix, under <em>The
\(\kappa_{\mathrm{eff}}\) inversion</em>.</p>

<p>Two conditions make the reduction valid, and together they are the criterion for this
failure mode. The spectrum must have one dominant eigenmode, \(\mathrm{PR} \approx 1\)
with \(\lambda_1/\lambda_2 \gg 1\), or there is no plane to reduce to. Additionally,
\(\varphi_{\mathrm{mm}}\) must be small, or the estimator is not at the foot of the cliff
and there is nothing to correct. Both are read off the activations and the fitted direction, and together they diagnose
that the estimator has collapsed onto \(\hat v_1\). They do not by themselves say which
way either direction will steer. That is set by the overlaps with the score gradient,
\(\hat v_1\cdot g\) for the mass-mean direction and \(\hat e_2\cdot g\) for the
correction, which are also measured without displacing an activation. At <code class="language-plaintext highlighter-rouge">counterfact</code>
layers \(8\) through \(16\) the spectrum meets the criterion more strongly than at layer
\(28\), yet the corrections do not restore the sign there, because \(\hat e_2\cdot g\) is
negative or unresolved. The diagnosis from the spectrum and the sign from the gradient
are both available before any steering is run. The diagnosis has been tested on one
dataset (<code class="language-plaintext highlighter-rouge">counterfact</code>) that satisfies the criterion, and one (<code class="language-plaintext highlighter-rouge">cities</code>) that does not,
and held in both. Two further main-tier sets,
<code class="language-plaintext highlighter-rouge">companies_true_false</code> and <code class="language-plaintext highlighter-rouge">common_claim_true_false</code>, satisfy the criterion at layers
\(24\) and \(28\) but cannot be tested causally with this harness. Their statements
have no relation structure from which to build a pair of completions, so the
behavioral score of <em>The scale of an intervention</em> cannot be formed. The alternative
would be to score the model’s own judgment of whether a statement is true, but
unsteered that judgment separates true from false at AUROC \(0.57\) on <code class="language-plaintext highlighter-rouge">common_claim</code>,
against \(0.78\) on <code class="language-plaintext highlighter-rouge">cities</code>, so there is almost no behavior there for a displacement to
move. So they meet the criterion, but the prediction it makes for them cannot be
checked. <code class="language-plaintext highlighter-rouge">cities</code> at layer \(28\) violates both conditions (\(\mathrm{PR} = 29.7\),
\(\lvert\cos\rvert = 0.159\)). There the formula predicts a gain of \(1.00\times\) at
\(\kappa = 1.42\) and \(r = 6.2\), since both directions already sit on the broad top of
the quotient, and the measured \(1.22\times\) is a small gain from the rest of the
spectrum, outside the plane. The formula stops applying where the regime ends, and the
regime control below draws the same boundary.</p>

<p>Both corrections flip the sign. Steering along the whitened direction at layer \(28\)
gives</p>

\[A(1) &gt; 0 \;\text{ in } 10/10 \text{ seeds}, \qquad
\operatorname{med}_{\text{seeds}} A(1) = +0.025 = 2.1\,\sigma_{\text{null}}, \qquad
p = 0.010,\]

<p>with \(8/10\) seeds individually clearing the null at \(5\%\). Layer \(24\) gives the
same picture: \(10/10\) positive, \(7/10\) clearing, \(p = 0.020\). Here \(\chi\)
coincides with \(A(1)\) at \(+0.025\), because the whitened response is linear in \(h\)
over the whole sweep, \(0.013\), \(0.025\), \(0.050\), \(0.099\) at \(h = 0.5, 1, 2, 4\),
where the plain direction’s response fell off at \(h = 4\). The corrected effect is
smaller than the wrong-signed one it replaces, two null standard deviations against
five, but it is monotone where that one was not. The rank-one correction
\(\hat\theta_\perp\), measured independently in the rogue-dimension sweep, agrees. At
layer \(28\) its three seeds give \(A(1) = +0.045\), \(+0.036\) and \(+0.048\), each at
\(p \le 0.0025\) against the same \(400\)-draw null, so on its own it carries \(A\) from
significantly negative to significantly positive, and it raises held-out decoding AUROC
at layer \(28\) from \(0.51\), the mass-mean value inside the null, to \(0.57\). Layer
\(20\) is the crossover. Whitened steering there is correctly signed in \(9/10\) seeds
but does not clear the null, \(p = 0.12\), which corresponds to a transition region rather
than a clean effect.</p>

<h2 id="the-regime-control">The regime control</h2>

<p>Marks &amp; Tegmark steer along the feature direction and report that difference-in-means
directions are the most causally implicated of the probes they compare. At the deep
<code class="language-plaintext highlighter-rouge">counterfact</code> layers that ordering reverses. The mass-mean direction, which their
framework nominates as the feature, steers the model away from the true completion,
and the whitened direction, which it demotes to a decision boundary, steers the model
toward it. The rogue-dimension diagnosis explains why. There \(\hat\theta\) has collapsed
onto \(\hat v_1\) and is not tracking a feature at all, so the \(\Sigma^{-1}\) meant to
sharpen a readout is instead doing the work of recovering the direction.</p>

<p>The inversion is a property of a regime, not a refutation of Marks &amp; Tegmark, which is demonstrated by <code class="language-plaintext highlighter-rouge">cities</code> as a the control. There the two directions behave as mass-mean
probing intends. At pythia-2.8b layer \(28\), the same model and depth at which
<code class="language-plaintext highlighter-rouge">counterfact</code> inverts, the plain direction steers correctly in \(10/10\) seeds
(\(\chi = +0.030\)) while the whitened direction is weak and inconsistent
(\(\chi = -0.003\), \(4/10\)). On pythia-1.4b the contrast is starker still, with
\(\chi = +0.46\) plain against \(-0.013\) whitened at layer \(12\). Read across the whole
depth sweep, the two datasets mirror each other. On <code class="language-plaintext highlighter-rouge">cities</code>, plain steering clears
its null at three layers, \(12\) and \(28\) clearly and \(24\) weakly (\(\chi = +0.009\),
\(10/10\) seeds positive, \(p = 0.020\)), and is correctly signed at all three. On
<code class="language-plaintext highlighter-rouge">counterfact</code>, plain steering clears its null at two layers, \(24\) and \(28\), and is
wrong-signed at both. At every other layer of either dataset it sits inside the null.
At layer \(28\) the two datasets share the model, the depth and the estimator, and
their significant effects point in opposite directions. When no single eigenmode
dominates the within-class covariance, the mean difference <em>is</em> the causal feature and
the inverse covariance only adds estimation noise, which is what Marks &amp; Tegmark report, on datasets of exactly this kind. The inversion is confined to the regime where \(\hat\delta\) has collapsed, and the rogue dimension decides which regime a dataset is in.</p>

<p>Two points in the <code class="language-plaintext highlighter-rouge">cities</code> panel deserve naming, since they are visible and read at
first glance like counterexamples. At layers \(16\) and \(20\) the plain direction becomes
mildly negative with \(\chi = -0.0031\) and \(-0.0062\). Both estimates are consistent across
seeds (\(9/10\) and \(8/10\) negative), so the sign is not noise. However, both sit an order of magnitude below the \(+0.0296\) the same direction produces
at layer \(28\), and neither clears the steering null. The regime
claim is that the <em>significant</em> effects on <code class="language-plaintext highlighter-rouge">cities</code> are correctly signed, not that
every layer’s point estimate is positive.</p>

<p><img src="/assets/figures/truth_regime_control.png" alt="The causal direction is set by the presence or absence of a rogue dimension. Steering susceptibility $$\chi$$ on pythia-2.8b, plain (red) against whitened (blue), median over seeds with inter-quartile bars. Left, `cities`: the plain difference-in-means direction carries the causal effect and whitening degrades it which is the intended behavior of mass-mean probing. The mild negative excursions of the plain direction at layers $$16$$ and $$20$$ do not clear the steering null and are an order of magnitude below its layer-$$28$$ effect. Right, `counterfact_true_false`: at depth the assignment inverts, the plain direction steers significantly wrong-signed while the whitened direction steers correctly. This is the same model, estimator, and protocol and only the within-class geometry differs." /></p>

<h2 id="discussion">Discussion</h2>

<p>The results above establish that, when the within-class spectrum meets a specific
criterion, the mass-mean estimator returns a nuisance direction rather than a truth
direction, and that a correction inside the linear class restores the sign of the
steering effect along the corrected direction. Two things
should be noted before comparing to other reports. The first is the magnitude of the
corrected effect, and the second is what its success does and does not imply about the
function class.</p>

<p>The corrected effect is modest. At one class gap the plain direction’s median response
is \(-0.062\), five null standard deviations, and the whitened direction’s is \(+0.025\),
two. The correction flips the sign of the effect but returns less than half its size.
The decoding gain is likewise real but small in absolute terms. Whitening carries
held-out AUROC at layer \(28\) from \(0.502\), inside the null, to \(0.624\) against a
null of \(0.560\), and the whitened direction stays clear of its null from layer \(26\)
to the end of the network, peaking at \(0.716\) against \(0.552\) at layer \(31\). Set
beside <code class="language-plaintext highlighter-rouge">cities</code>, where the same direction reads \(0.99\), this is a weak readout. The overall
claim is a corrected sign and a confirmed mechanism, not a recovered truth direction of
practical use.</p>

<p>A wrong-signed steering effect invites a natural reading that the linear function
class is too weak: the direction one can fit is not expressive enough to move behavior,
and a richer, perhaps nonlinear, intervention is required. The rogue-dimension account
makes a different claim about this <em>particular</em> failure. The class is adequate, but the
mass-mean estimator points at \(\hat v_1\), a massive-activation direction, instead of at
truth. The two readings make different predictions, and the data separate them. A
rank-one correction within the same linear class, with no richer classifier and no
nonlinear probe, converts a wrong-signed effect that clears its null into a
correctly signed one that clears its null. For the <code class="language-plaintext highlighter-rouge">counterfact</code> failure at layers \(24\) and \(28\) the function class was never the
bottleneck. The estimator’s alignment with a massive-activation direction was.</p>

<h3 id="relation-to-other-reports">Relation to other reports</h3>

<p>Four recent reports also examine why a fitted direction fails to steer, and the
rogue-dimension account can be compared against each. Braun et al. and Ying et al.
locate the failure in the direction. For Braun et al. steering is unreliable when the
per-example activation differences do not point the same way, so that their mean is
not representative of any of them. For Ying et al. steering is wrong-signed when the
direction is fitted across many kinds of truth and mixes truth with sycophancy. Torop et al.
find a direction that discriminates well and steers backwards. Liu finds a direction
that decodes but does not steer at all. This post shares a predictor with the first
two and differs on the diagnosis, and against the last two it supplies the remaining
case, a direction that neither decodes nor steers correctly.</p>

<p><a href="https://arxiv.org/abs/2505.22637">Braun et al. (2025)</a> study contrastive activation
addition across thirty-six behaviors and find two predictors of whether steering
works: the mean pairwise cosine similarity among the per-example activation
differences, which measures whether those differences share a direction, and the
separability of positive from negative activations along the difference-of-means
line. Their separability index is the
discriminability \(d'\) of the projection onto that line, which is the
\(d'_{\mathrm{mm}}\) of this post. While these two reports share a predictor, they differ in
what they find at low \(d'_{\mathrm{mm}}\). Every dataset in Braun et al.’s study steers with a
net positive effect, with low separability showing up as per-example scatter around
that mean. In those cases they conclude that the behavior is not represented by a
coherent linear direction.
<code class="language-plaintext highlighter-rouge">counterfact</code> at layers \(24\) and \(28\), with \(d'_{\mathrm{mm}} = 0.08\), sits at the
unsteerable end of their scale but steers with a negative mean that clears its null. This implies
that the behavior is linearly represented, since the whitened direction steers it correctly at
the same layer. What has failed is the estimator. The superposition result above is the
decoding-side form of the same failure, a high-variance direction interfering with the
target signal, whose removal improves recovery. The shallow <code class="language-plaintext highlighter-rouge">counterfact</code> layers, \(L \le 16\), are closer to the regime Braun et al.
describe, but not the same one. Whitening leaves the effect inside the null there, and
projecting out \(\hat v_1\) does not restore the sign: at layer \(8\) it steers
significantly wrong-signed. That is not an absence of linear signal. It is the sign of
\(\hat e_2\cdot g\), which at that depth is negative, as <em>What steering measures</em> shows. The per-sample form of the anomaly is documented by
<a href="https://arxiv.org/abs/2407.12404">Tan et al. (2024)</a>, who find that for several
concepts close to half the inputs steer in the direction opposite to the one intended.
The <code class="language-plaintext highlighter-rouge">counterfact</code> effect at layers \(24\) and \(28\) is that variance surfacing as a
wrong-signed dataset-level mean that clears its null, with a proposed mechanism.</p>

<p><a href="https://arxiv.org/abs/2602.20273">Ying et al. (2026)</a> compare two kinds of truth
direction: domain-specific ones, each fitted within a single truth type, and a
<em>domain-general</em> one recovered by concept erasure across many truth types. Steering
along a domain-specific direction improves truthfulness on held-out factual questions.
Steering along the domain-general one consistently degrades it. Their account is one
of mixture: a direction trained to span many domains conflates factual variance with
sycophancy-related variance, so intervening along it moves several things at once.
That mechanism is not available here, since the estimator is fitted on the
<code class="language-plaintext highlighter-rouge">counterfact</code> domain alone and there is no second domain to mix in. Within that one
domain the class-mean difference is small compared with the spread along the leading
within-class eigenmode, and the estimator returns that eigenmode rather than the truth
direction. The two results therefore agree that a direction fitted to truth can steer
against it but disagree on the mechanism and on the remedy. Theirs is to fit within a
single truth type. The remedy here is to keep the fitted direction and project out one
eigendirection before reading it.</p>

<p><a href="https://arxiv.org/abs/2608.02957">Torop, Masoomi &amp; Dy (2026)</a> find <em>inverted steering
vectors</em> in the attention-head outputs of Gemma 3 12B, Qwen 2.5 14B and Olmo 3 7B:
mean-difference directions that discriminate the concept well, \(\mathrm{AUC} \ge 0.85\)
by their candidate cutoff, and whose positive steering consistently suppresses it
across inputs. They separate this from Tan’s per-input anti-steerable examples and
from the low-discriminability failures Braun et al. describe. Their inversion comes
with a direction that decodes well while the <code class="language-plaintext highlighter-rouge">counterfact</code> inversion comes with a nuisance direction
that fails to decode, since \(\hat\theta\) sits inside its decoding null. The
mechanism found here, alignment with a label-blind variance axis, is unavailable when
the direction is discriminative. The two are the same failure of sign at opposite ends
of decodability, and the diagnostics differ accordingly. They diagnose it through a downstream inner-product response, while this post reads the spectrum of the within-class covariance. Whether the
rogue-dimension criterion says anything about their cases is open.</p>

<p><a href="https://arxiv.org/abs/2605.05715">Liu (2026)</a> finds an overthinking failure in medical
QA that is linearly decodable at \(71.6\%\) balanced accuracy. Liu subsequently finds that five
families of fixed residual-stream linear steering, across twenty-nine configurations,
all move the behavior by nothing measurable. This is the decode-without-cause shape
arrived at from the other side, and Liu attributes it to representational
entanglement. Two things separate it from the case here. The effect described by Liu is null
rather than wrong-signed, so there is no sign to restore. The entanglement is
inferred from the failure, whereas the nuisance direction responsible here is named in advance
from the within-class spectrum and then removed. A null steering result cannot
distinguish an entangled direction from one that is absent, while a significantly wrong-signed result with an identified direction can.</p>

<h3 id="summary">Summary</h3>

<p>A mass-mean truth direction can either be a truth direction or the estimator’s
projection onto the most salient axis of the activations. On a benchmark the two
are told apart only by a signal-to-noise reading. Three results follow. In-sample
separation is inflated by dimensional slack that scales as \(N^{-1/2}\) with a prefactor
set by the effective dimension of the noise, meaning that for the datasets in this literature the
null and not the raw score is the bar. When the class gap is weak the mass-mean
estimator collapses onto the leading within-class eigenvector. On <code class="language-plaintext highlighter-rouge">counterfact</code>
that direction fails to decode and steers with a significant wrong sign, five null
standard deviations deep at one class gap. The failure is in the estimator: the same activations used to obtain the mass-mean
direction contain a direction that steers correctly, and it can be reached without
leaving the linear class, either by the Fisher direction or by projecting out the
leading eigenvector.</p>

<p>The results have limits. The criterion that diagnoses this failure, a within-class spectrum with one dominant
eigenmode and the mass-mean direction aligned to it, has one confirmed positive and one
confirmed negative. It identifies the collapse but not the sign of a correction, which
is set by the overlap of the corrected direction with the score gradient. The
steering measurements are on Pythia alone, since OLMo replicates the geometry and the
decoding but was not steered. The wrong sign at layer \(28\) is beyond first order in
the displacement and its mechanism is not settled here. The corrected effect is modest. The
whitened direction’s response at one class gap is two null standard deviations, where
the wrong-signed response of the mass-mean direction was five, so the correction
restores a sign rather than supplying a large causal lever. On one of four datasets the
shuffled-label amplitude departs from the parameter-free prediction by \(36\)–\(60\%\) for
reasons that were localized but not explained.</p>

<p>This post has treated steering as a diagnostic. The steering results themselves are analyzed in a second, forthcoming post, Steering Vectors and the Limits of Linear Response. There the object of interest is the vector \(g\) itself. Starting from the identity \(\chi(w) = c\,(w\cdot g)\), I ask how much of \(g\) lies along the high-variance directions of the within-class covariance, what bound it puts on any linear steering direction, how much of that bound the best direction actually reaches, and how much of \(g\) published steering vectors capture.</p>

<h2 id="appendix-auroc-and-its-relation-to-d">Appendix: AUROC and its relation to \(d'\)</h2>

<p>The main text defines \(\mathrm{AUROC}(u) = \Pr[z_1 &gt; z_0]\) as the probability that a
random true statement outscores a random false one. This is the quantity the
random-direction null resamples. The name refers to a different construction, the
<em>area under the receiver-operating-characteristic curve</em>. This appendix builds that
curve, shows its area equals the scoring probability, and records the Gaussian special
case behind \(\mathrm{AUROC} = \Phi(d'/\sqrt{2})\). The same Gaussian link, together
with a closed form for the Mahalanobis cosine as a function of \(d'\), is derived by
<a href="https://arxiv.org/abs/2606.19603">Ying, Hase &amp; Kriegeskorte (2026)</a>, whose
signal-to-noise ratio is \(d'(u)\) under the pooled within-class covariance.</p>

<p><img src="/assets/figures/auroc_two_definitions.png" alt="AUROC two ways. Left: the class-conditional score densities $$p_0, p_1$$ with a threshold $$t$$ cutting each into its false-positive and true-positive tail; the mean gap in noise units is $$d'$$. Right: sweeping $$t$$ traces the ROC curve, whose shaded area equals $$\Pr[z_1 &gt; z_0]$$; the dot is the ROC point for the threshold at left, and the dashed diagonal is chance." /></p>

<p>The <em>receiver operating characteristic</em> is an inheritance from WWII
radar and quantified how well a receiver’s operator could separate signal (an
aircraft echo) from noise as the detection threshold was varied. The same curve was
adopted wholesale by signal detection theory and then by psychophysics and
statistics, which is why a truth-probe analysis and a 1940s radar set share an
acronym.</p>

<p>Fixing a direction \(u\) and reading the scalar \(z = u^{\top}x\), a classifier
is obtained by thresholding, \(\hat y = \mathbb{1}[z &gt; t]\), and sweeping the threshold
\(t\) from \(+\infty\) to \(-\infty\) traces out two functions of \(t\):</p>

\[\mathrm{TPR}(t) = \Pr[z &gt; t \mid y=1], \qquad
\mathrm{FPR}(t) = \Pr[z &gt; t \mid y=0].\]

<p>Plotting the point \((\mathrm{FPR}(t), \mathrm{TPR}(t))\) as \(t\) sweeps traces out a curve
from \((0,0)\) to \((1,1)\). A direction that separates the classes lifts the curve above
the diagonal: at every false-positive rate the true-positive rate is higher, so the
curve passes near the top-left corner. A useless direction gives the diagonal
\(\mathrm{TPR} = \mathrm{FPR}\) itself. The AUROC is the area
under this curve. It is \(1\) for perfect separation,
\(\tfrac{1}{2}\) for the diagonal, and \(\mathrm{AUROC}(-u) = 1 - \mathrm{AUROC}(u)\) since changing the sign of \(u\) swaps the conditional distributions.</p>

<p>The geometric area and the
probabilistic \(\Pr[z_1 &gt; z_0]\) are the same number, which the main text
leans on. Writing the class-conditional densities as \(p_1(z) = p(z\mid y=1)\) and
\(p_0(z) = p(z\mid y=0)\), with CDFs \(F_1, F_0\), one can parametrize the curve by \(t\): the
height is \(\mathrm{TPR}(t) = 1 - F_1(t)\) and the horizontal coordinate is
\(\mathrm{FPR}(t) = 1 - F_0(t)\), so \(\mathrm{d}(\mathrm{FPR}) = -p_0(t)\,\mathrm{d}t\).
The area, integrated as \(\mathrm{FPR}\) runs \(0 \to 1\) (i.e. \(t\) runs \(+\infty \to
-\infty\)), is</p>

\[\mathrm{AUROC}
= \int_0^1 \mathrm{TPR}\;\mathrm{d}(\mathrm{FPR})
= \int_{-\infty}^{\infty} \big(1 - F_1(t)\big)\, p_0(t)\,\mathrm{d}t .\]

<p>Reading \(1 - F_1(t) = \Pr[z_1 &gt; t]\) for an independent draw \(z_1 \sim p_1\), the integral
averages this over \(t \sim p_0\):</p>

\[\mathrm{AUROC}
= \mathbb{E}_{z_0 \sim p_0}\!\big[\Pr[z_1 &gt; z_0]\big]
= \Pr[z_1 &gt; z_0],\]

<p>the probability that a random positive outscores a random negative — the Wilcoxon–
Mann–Whitney identity. Ties contribute a boundary term of measure zero for
continuous \(z\), and half a count apiece if one insists on discrete scores. This is
also why AUROC is <em>calibration-free</em>: it depends only on the ordering of scores, not
their scale or location, so any strictly monotone reparametrization of \(z\) leaves it
unchanged — in particular it needs no choice of threshold, unlike accuracy.</p>

<p>When the class-conditionals of \(z\) are Gaussian and the classes
balanced, \(z_y \sim \mathcal{N}(u^{\top}\mu_y,\; u^{\top}\Sigma u)\), the difference \(z_1 - z_0 \sim
\mathcal{N}(u^{\top}\delta,\; 2\,u^{\top}\Sigma u)\), so \(\Pr[z_1 &gt; z_0] = \Phi\big(u^{\top}\delta/\sqrt{2\,u^{\top}\Sigma u}\big) =
\Phi(d'/\sqrt2)\), recovering the identity quoted in the definitions. Away from that
case the map between \(d'\) and AUROC is only approximate, which is why the main text
reports both rather than deriving one from the other.</p>

<h2 id="appendix-results-on-the-full-twelve-dataset-benchmark">Appendix: results on the full twelve-dataset benchmark</h2>

<p>The main text carries <code class="language-plaintext highlighter-rouge">cities</code> and <code class="language-plaintext highlighter-rouge">counterfact</code> through the steering and
rogue-dimension analysis and the nine main-tier sets through the transfer matrix. The
sweep covers all twelve. At pythia-2.8b, each at its own selected layer:</p>

<table>
  <thead>
    <tr>
      <th>dataset</th>
      <th>\(L\)</th>
      <th>plain AUROC</th>
      <th>\(d'_{\mathrm{mm}}\)</th>
      <th>whitened AUROC</th>
      <th>null \(p_{95}\)</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td><code class="language-plaintext highlighter-rouge">neg_cities</code></td>
      <td>25</td>
      <td>0.978</td>
      <td>3.03</td>
      <td>0.981</td>
      <td>0.683</td>
    </tr>
    <tr>
      <td><code class="language-plaintext highlighter-rouge">sp_en_trans</code></td>
      <td>27</td>
      <td>0.980</td>
      <td>2.68</td>
      <td>0.974</td>
      <td>0.673</td>
    </tr>
    <tr>
      <td><code class="language-plaintext highlighter-rouge">neg_sp_en_trans</code></td>
      <td>28</td>
      <td>0.974</td>
      <td>2.83</td>
      <td>0.994</td>
      <td>0.677</td>
    </tr>
    <tr>
      <td><code class="language-plaintext highlighter-rouge">cities</code></td>
      <td>29</td>
      <td>0.973</td>
      <td>2.93</td>
      <td>0.983</td>
      <td>0.721</td>
    </tr>
    <tr>
      <td><code class="language-plaintext highlighter-rouge">smaller_than</code></td>
      <td>31</td>
      <td>0.974</td>
      <td>2.79</td>
      <td>1.000</td>
      <td>0.786</td>
    </tr>
    <tr>
      <td><code class="language-plaintext highlighter-rouge">larger_than</code></td>
      <td>28</td>
      <td>0.929</td>
      <td>2.05</td>
      <td>1.000</td>
      <td>0.735</td>
    </tr>
    <tr>
      <td><code class="language-plaintext highlighter-rouge">cities_cities_disj</code></td>
      <td>29</td>
      <td>0.836</td>
      <td>1.42</td>
      <td>0.836</td>
      <td>0.586</td>
    </tr>
    <tr>
      <td><code class="language-plaintext highlighter-rouge">cities_cities_conj</code></td>
      <td>26</td>
      <td>0.810</td>
      <td>1.30</td>
      <td>0.930</td>
      <td>0.597</td>
    </tr>
    <tr>
      <td><code class="language-plaintext highlighter-rouge">common_claim_true_false</code></td>
      <td>31</td>
      <td>0.757</td>
      <td>0.64</td>
      <td>0.722</td>
      <td>0.611</td>
    </tr>
    <tr>
      <td><code class="language-plaintext highlighter-rouge">companies_true_false</code></td>
      <td>31</td>
      <td>0.699</td>
      <td>0.17</td>
      <td>0.855</td>
      <td>0.585</td>
    </tr>
    <tr>
      <td><code class="language-plaintext highlighter-rouge">counterfact_true_false</code></td>
      <td>32</td>
      <td>0.566</td>
      <td>0.20</td>
      <td>0.701</td>
      <td>0.556</td>
    </tr>
    <tr>
      <td><code class="language-plaintext highlighter-rouge">likely</code></td>
      <td>12</td>
      <td>0.890</td>
      <td>1.75</td>
      <td>0.929</td>
      <td>0.593</td>
    </tr>
  </tbody>
</table>

<p>The two translation sets sit at \(N = 354\) rather than \(1198\) and clear their null by
as wide a margin as the single-frame sets. The ordering is informative. Every templated set with
a single frame clears its null comfortably. The two free-form sets are the bottom of
the table: <code class="language-plaintext highlighter-rouge">common_claim_true_false</code> at \(d'_{\mathrm{mm}} = 0.64\) and <code class="language-plaintext highlighter-rouge">counterfact_true_false</code> at
\(0.20\). <code class="language-plaintext highlighter-rouge">companies_true_false</code> is the interesting entry — a plain \(d'_{\mathrm{mm}}\) of \(0.17\), at
the floor with <code class="language-plaintext highlighter-rouge">counterfact</code>, but a whitened AUROC of \(0.855\), the largest gap between
the two estimators anywhere in the table. Together with its rising superposition curve, that
makes it a second instance of the rogue-dimension pattern, decoded rather than steered:
the dataset ships only statements and labels, with no contrastive completions from
which to build a behavioral score, so it carries no causal measurement. The spectrum confirms
it. At layers \(24\) and \(28\), <code class="language-plaintext highlighter-rouge">companies_true_false</code> has \(\mathrm{PR} = 1.01\) and
\(1.02\), \(\hat\lambda_1/\hat\lambda_2\) near \(2000\) and \(1250\), and
\(\lvert\cos(\hat\theta,\hat v_1)\rvert = 0.98\) and \(0.93\), with \(d'_{\mathrm{mm}}\) of
\(0.02\) and \(0.03\). That is <code class="language-plaintext highlighter-rouge">counterfact</code>’s signature at the same depths. By layer
\(31\), where selection lands, it has escaped: \(\mathrm{PR} = 3.2\), the alignment down
to \(0.18\), \(d'_{\mathrm{mm}}\) up to \(0.51\). <code class="language-plaintext highlighter-rouge">common_claim_true_false</code> is a third
instance, locked at layers \(24\) and \(28\) (\(\mathrm{PR} = 1.02\) and \(1.05\),
alignment \(0.95\) and \(0.83\)) and escaping by layer \(31\) (\(\mathrm{PR} = 12.1\),
alignment \(0.21\)). <code class="language-plaintext highlighter-rouge">cities_cities_conj</code>, by contrast, never locks (\(\mathrm{PR}\) from
\(20\) to \(42\) across the same layers, alignment below \(0.11\)). So within Pythia the
rogue dimension is a property of at least three of the nine main-tier sets at
mid-depth, and what separates the other two from <code class="language-plaintext highlighter-rouge">counterfact</code> is that their class gap
grows enough to leave the axis before the network ends, where <code class="language-plaintext highlighter-rouge">counterfact</code>’s does not.</p>

<p>One caution about the <code class="language-plaintext highlighter-rouge">companies_true_false</code> row is that \(d'\) of \(0.17\) maps to \(\mathrm{AUROC} \approx 0.55\)
under the Gaussian identity, not the \(0.699\) measured — the largest
departure in the table, and a sign that the projected class-conditionals on this
dataset are far from that idealization.</p>

<h2 id="appendix-identification-and-removal-of-the-eleven-outlier-statements">Appendix: identification and removal of the eleven outlier statements</h2>

<p>The eleven statements are identifiable without reference to \(\hat v_1\). Selecting outliers by
their projection onto \(\hat v_1\) and then recomputing \(\hat v_1\) without them would be
circular since removing the top of a distribution shortens it. Instead, the set is
defined coordinate-first, without referring to the covariance, the class means,
or the labels. A coordinate is massive if \(\lvert\operatorname{med}_i x_{ij}\rvert\)
exceeds both an absolute threshold and a large multiple of the median activation, and a
statement is flagged if it deviates from the modal value on any massive coordinate. This criterion returns exactly the eleven statements of the projection-based set, at
every layer measured. Layer-invariance makes it a property of the
input rather than of any particular representation.</p>

<p>Refitting on a training half with the eleven excluded and scoring on
the full held-out half leaves the mass-mean direction where
it was and moves the whitened one substantially at early and middle depth:</p>

<table>
  <thead>
    <tr>
      <th>layer</th>
      <th>\(d'_{\mathrm F}\), all</th>
      <th>\(d'_{\mathrm F}\), eleven removed</th>
      <th>null \(p_{95}\)</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td>\(8\)</td>
      <td>\(0.010\)</td>
      <td>\(0.129\)</td>
      <td>\(0.094\)</td>
    </tr>
    <tr>
      <td>\(12\)</td>
      <td>\(0.015\)</td>
      <td>\(0.162\)</td>
      <td>\(0.092\)</td>
    </tr>
    <tr>
      <td>\(16\)</td>
      <td>\(0.031\)</td>
      <td>\(0.163\)</td>
      <td>\(0.095\)</td>
    </tr>
    <tr>
      <td>\(20\)</td>
      <td>\(0.005\)</td>
      <td>\(0.385\)</td>
      <td>\(0.097\)</td>
    </tr>
    <tr>
      <td>\(24\)</td>
      <td>\(0.063\)</td>
      <td>\(0.136\)</td>
      <td>\(0.132\)</td>
    </tr>
    <tr>
      <td>\(28\)</td>
      <td>\(0.384\)</td>
      <td>\(0.385\)</td>
      <td>\(0.128\)</td>
    </tr>
  </tbody>
</table>

<p>\(d'_{\mathrm{mm}}\) stays inside its null under every regime, so nothing here rescues the
plain estimator. The layer-\(28\) row is the informative one. Removing the eleven moves
the geometry enormously there — \(\hat\lambda_1/\hat\lambda_2\) falls from \(535\) to
\(1.1\) and \(\lvert\cos(\hat\theta,\hat v_1)\rvert\) from \(0.994\) to \(0.063\) while \(d'_{\mathrm F}\) does not move at all. At that depth whitening has the same
effect on \(d'_{\mathrm F}\) as deleting the contaminated statements, but from layers
\(8\) through \(20\) whitening alone leaves \(d'_{\mathrm F}\) inside its null, while
explicit removal lifts it above the null by factors of \(1.4\)–\(4\). Layer \(24\) sits between the two
regimes.</p>

<p>The difference between the two regimes is controlled by the shrinkage intensity. Sweeping
\(\rho\) on the uncleaned data reproduces most of the effect of deleting the eleven:
at layer \(16\), \(d'_{\mathrm F}\) runs from \(0.031\) at the Ledoit–Wolf
\(\rho = 0.137\) to \(0.530\) at \(\rho = 10^{-4}\). What limits \(\hat\Sigma^{-1}\)
at these depths is therefore not the contamination itself but the regularization. At
the Ledoit–Wolf intensity the estimate is pulled far enough toward the isotropic
target that the inverse barely downweights the leading directions. The methods appendix states this as a general point about \(\rho\) — in the note that
the whitened numbers are a lower bound — rather than as one about these eleven
statements.</p>

<p>The massive coordinates are not dataset-specific. At pythia-2.8b the same
four coordinates qualify at layers \(8\)–\(16\), the same seven at layer \(20\) and the
same eight at layers \(24\)–\(28\), on <code class="language-plaintext highlighter-rouge">cities</code> and <code class="language-plaintext highlighter-rouge">counterfact</code> alike. What differs is
whether any statement drops them: on <code class="language-plaintext highlighter-rouge">cities</code>, none do, at any layer in either model.
The coordinates therefore stay constant, contribute nothing to \(\hat C\), and leave
no rogue dimension behind. On pythia-1.4b no coordinate qualifies through layer \(15\). One does at
layers \(18\) and \(21\), where three <code class="language-plaintext highlighter-rouge">counterfact</code> statements drop it, none of them among the eleven and two of
the three shared between the layers, and <code class="language-plaintext highlighter-rouge">cities</code> again has none. There \(\hat\lambda_1/\hat\lambda_2\) moves only from
\(1.5\) to \(1.4\). So the incidence is specific and the mechanism is general.
Wherever a near-constant massive activation is dropped by a small minority of inputs,
the within-class covariance acquires a rogue dimension whose scale is set by
\(p(1-p)(a-b)^2\), and where it is dropped by none, it acquires nothing. I do not know
what distinguishes the eleven. They span both labels, all end in a period, and their
token lengths sit inside the bulk of the distribution.</p>

<h2 id="appendix-replication-on-olmo-2-1b">Appendix: replication on OLMo-2-1B</h2>

<p>To check that the rogue dimension is not a fact about Pythia, I ran the observables of
<em>The rogue dimension</em> on <a href="https://huggingface.co/allenai/OLMo-2-0425-1B">OLMo-2-1B</a>, which has a
different architecture, a different training corpus and a third of the parameters. <code class="language-plaintext highlighter-rouge">counterfact</code>
stays locked to the leading axis at every depth except the final layer,
\(\lvert\cos(\hat\theta,\hat v_1)\rvert\) between \(0.64\) and \(0.94\) from layer \(2\) to layer
\(15\) (it falls to \(0.04\) at layer \(16\)), with \(d'_{\mathrm{mm}} \approx 0.1\). Its plain probe never clears its own null (held-out
\(\mathrm{AUROC} = 0.529\) at its best layer, against a null 95th percentile of \(0.547\)),
while whitening lifts it to \(0.647\). <code class="language-plaintext highlighter-rouge">cities</code> clears the null, as it does in Pythia: the alignment
falls to \(\lvert\cos\rvert \approx 0.3\), the participation ratio climbs \(1.4 \to 17.7\), and
the plain probe reaches \(0.903\). The superposition probe agrees from the other side.
Stripping the leading components <em>raises</em> \(d'_{\mathrm{mm}}\) on <code class="language-plaintext highlighter-rouge">counterfact</code>
(\(0.07 \to 0.31\)) and destroys it on <code class="language-plaintext highlighter-rouge">cities</code> (\(1.85 \to 0.54\)).</p>

<h2 id="appendix-methods-in-detail">Appendix: methods in detail</h2>

<p>Everything below describes what the code does, not what the method ideally would
do. Where the two differ the difference is stated. Script names refer to the
public repository,
<a href="https://github.com/jasteinberg/rl-alignment-repo">jasteinberg/rl-alignment-repo</a>,
which holds the scripts below, the shared library they import, and the committed
artifacts each number is read from.</p>

<p><strong>Models and activations.</strong> Pythia 70m, 410m, 1.4b and 2.8b, and OLMo-2-1B for the
cross-family check, run in <code class="language-plaintext highlighter-rouge">float16</code> on Apple silicon (MPS) through HuggingFace
<code class="language-plaintext highlighter-rouge">transformers</code>. The activation \(x\) for a statement is the residual stream at the
<strong>final token</strong>, taken from every layer in one pass via <code class="language-plaintext highlighter-rouge">output_hidden_states</code>.
Layer \(L\) means the output of block \(L\), and layer \(0\) is the embedding.</p>

<p><strong>Environment.</strong> Python 3.11.15, <code class="language-plaintext highlighter-rouge">torch</code> 2.10.0 with the MPS backend,
<code class="language-plaintext highlighter-rouge">transformers</code> 4.57.6, <code class="language-plaintext highlighter-rouge">numpy</code> 1.26.4, <code class="language-plaintext highlighter-rouge">scipy</code> 1.17.1, <code class="language-plaintext highlighter-rouge">scikit-learn</code> 1.8.0, on
macOS 26.5.2 (arm64). The backend matters more than usual: MPS kernels have changed
numerical behavior between <code class="language-plaintext highlighter-rouge">torch</code> releases, so exact reproduction of third-decimal
figures should pin these versions.</p>

<p><strong>Data.</strong> The nine main-tier sets are capped at \(1199\) and class-balanced, giving
\(N = 1198\), the two translation sets sit at \(N = 354\) (their natural size), and
<code class="language-plaintext highlighter-rouge">likely</code> is carried as a distractor. Caps are
applied before splitting, so every dataset in a comparison contributes the same \(N\).</p>

<p><strong>Transfer.</strong> A direction is fit on the source’s training half at the source’s
selected layer and scored on the target’s held-out half at the target’s selected
layer, so the two ends of a transfer sit at different depths whenever the selected
layers differ. The diagonal is scored the same way and is held-out.</p>

<p><strong>Split and estimators.</strong> One class-stratified \(50/50\) split at seed \(0\)
(<code class="language-plaintext highlighter-rouge">split_indices</code>), fit on the training half and scored on the held-out half unless a
quantity is explicitly marked in-sample. The three directions are</p>

\[\hat\theta \propto \hat\delta, \qquad
\hat\theta_{\mathrm F} \propto \hat\Sigma^{-1}\hat\delta, \qquad
\hat\theta_\perp \propto (I - \hat v_1\hat v_1^{\top})\hat\delta ,\]

<p>each normalized, and each sign-oriented so that its training-half AUROC is at least
\(\tfrac12\). This is necessary because \(d'\) is sign-blind and AUROC is not.
\(\hat\Sigma\) is the within-class covariance with each class centered on its own mean,
under Ledoit–Wolf shrinkage toward \((\operatorname{tr}\hat C/d)\,I\). The raw \(\hat C\)
is singular whenever \(N_{\text{train}} &lt; d\), which is most of this study.</p>

<p><strong>Nulls.</strong> The decoding null draws \(200\) random unit directions per layer
(Gaussian, then normalized) and scores them <strong>on the same held-out points</strong> as the
fitted direction, so the comparison is not confounded by the split. (The
null-distribution figure above draws its own \(400\) in-sample directions on the full
set, as an illustration of the distribution’s shape. The \(200\)-draw held-out null is
the one every quoted number is scored against.) The steering null
displaces along random unit directions exactly as the probe directions are displaced. Its draw
count is \(400\) at <code class="language-plaintext highlighter-rouge">counterfact</code> layers \(20\), \(24\), and \(28\) and <code class="language-plaintext highlighter-rouge">cities</code> layer
\(28\), and \(100\) at every other cell, including <code class="language-plaintext highlighter-rouge">cities</code> layers \(12\) and \(24\). The decoding null is reported as the \(95\)th percentile of the folded AUROC.
Steering clearances are <em>signed</em>: every steering \(p\) in this post is a one-sided
rank of the direction’s mean \(A(1)\) among the null draws in that direction, and
“clears the null” means \(p &lt; 0.05\) — the same test the figure bands denote as the
signed 5th–95th percentiles. Every claim of signal in this post is a claim about
that margin, never about the raw score.</p>

<p><strong>Layer selection.</strong> The reported layer maximizes the margin \(m\) of <em>The geometry</em> —
held-out plain AUROC minus the layer’s own null \(p_{95}\), not raw AUROC — so a layer
with a large null cannot win by inflation. Layer \(0\) is excluded: on templated statements the final-token embedding
is nearly constant within a dataset, giving a degenerate \(d'\). This margin criterion
follows the variance-ratio selection of Bürger et al. and MacDiarmid et al. as adopted
by Bao et al., with the null margin replacing the raw ratio.</p>

<p><strong>Superposition probe.</strong> PCA of the activations, project out the top \(k\) components
for \(k \in \{0,1,2,4,8,16,32,64\}\), refit the mass-mean direction in the residual
subspace, and report \(d'_{\mathrm{mm}}\) against \(k\). The PCA basis is fit on the
training half only.</p>

<p><strong>Massive coordinates and droppers.</strong> Two conventions for “massive” appear in this
post. In the rogue-dimension section a
coordinate’s magnitude is quoted as \(\lvert\text{mean}_i\, x_{ij}\rvert\) divided by the
median of that quantity across coordinates — the \(1446\times\) and \(563\times\) figures.
The dropper test instead calls coordinate \(j\) massive when
\(\lvert\operatorname{med}_i x_{ij}\rvert\) exceeds both \(100\) in absolute terms and
\(100\times\) the median activation magnitude \(\operatorname{med}_{ij}\lvert x_{ij}\rvert\),
and flags statement \(i\) as a dropper when
\(\lvert x_{ij} - \operatorname{med}_i x_{ij}\rvert &gt; \tfrac12\lvert\operatorname{med}_i x_{ij}\rvert\)
for some massive \(j\). The two ratios agree to about one percent on this data
(\(1446\) against \(1456\) at layer \(8\) and \(560\) against \(564\) at layer \(28\)). The
relative factor is \(100\times\) rather than the \(1000\times\) of Sun et al. because the
dominant coordinate runs \(1456\)–\(1693\times\) the median at layers \(8\)–\(20\) but
only \(564\)–\(587\times\) at layers \(24\)–\(28\), so the literal constant stops firing at
exactly the depths under discussion. Every value between the two gives the same flagged
set.</p>

<p><strong>Shuffled-label control.</strong> The labels are permuted and the mass-mean direction refit.
This control is scored <strong>in-sample by design</strong>: a held-out shuffled direction scores
\(\tfrac12\) by construction, which would hide exactly the inflation the control exists
to expose. The pooled fit \(0.045\,(N/2d)^{-0.49}\) is run on <code class="language-plaintext highlighter-rouge">counterfact</code> across all
four models with one permutation per \((\text{model}, N)\) point, thirty-six points in
all, of which the three pythia-70m points whose in-sample excess came out negative are
dropped because the fit is in log space. The per-dataset fits are a separate sweep on pythia-2.8b alone, at each
dataset’s own best layer, over a seven-point geometric grid from \(N = 100\) to that
dataset’s own total, with sixteen independent label permutations per \(N\). Because the
control is in-sample there is no held-out half to reserve, so the grid runs to the
complete set rather than stopping at a shared cap. The within-class spectrum is recorded
both on the full set and on the first subsample at each \(N\), so the participation ratio
can be checked for drift with sample size.</p>

<p><strong>Steering.</strong> The intervention adds \(h\,c\,w\) to the block output at layer \(L\) at
<strong>every token position</strong>, for \(h \in \{0.5, 1, 2, 4\}\) and both signs, with
\(c = \lVert\hat\delta\rVert\) estimated on the training half, so \(h = 1\) displaces
an activation by the distance between class means. The behavioral score is
\(\ell = \log P(\text{true completion}) - \log P(\text{false completion})\), summed over
completion tokens given the prompt. Pairs are drawn as \(250\) per seed from a fixed pool
of \(400\). Ten seeds per cell, except the rogue-dimension sweep of \(\hat v_1\) and
\(\hat\theta_\perp\), which has three.
\(A\) and \(S\) are the odd and even parts of the response in \(h\), and
\(\chi\) is the through-origin least-squares slope of \(A\) against \(h\) over the four
magnitudes, so \(\chi\) is weighted toward the large-\(h\) end and is not a pure
\(h\to0\) derivative. Rank \(p\) values are computed on \(A\) at \(h = 1\), each seed
against the null distribution of \(A(1)\) over random directions, and not on \(\chi\).
Ranking the seed-median \(\chi\) against the null of \(\chi\) gives \(p = 0.018\) (\(0.060\) for the seed mean) for the plain direction at
<code class="language-plaintext highlighter-rouge">counterfact</code> layer \(28\), because the slope averages in the turned-back response at
\(h = 4\).</p>

<p><strong>Score gradient.</strong> \(g = \langle \nabla_x \ell \rangle\) is computed by replacing the
block-\(L\) output with a leaf tensor that requires grad, so backpropagation runs only
through layers above \(L\), and by <strong>summing the gradient over token positions</strong> — the
quantity that matches an intervention applied at every position. A loss scale of
\(1024\) prevents <code class="language-plaintext highlighter-rouge">float16</code> underflow in the backward pass. The score \(\ell\), the pairs,
the pool and the seed selection are taken from the steering harness unchanged, so the
two measurements refer to the same object. Per-pair gradients are retained, and every
overlap is reported with a \(95\%\) bootstrap interval over pairs, \(10{,}000\) resamples.</p>

<p><strong>Scripts.</strong></p>

<table>
  <thead>
    <tr>
      <th>claim</th>
      <th>script</th>
      <th>artifact</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td>scale ladder, layer sweeps, transfer, superposition, Cover</td>
      <td><code class="language-plaintext highlighter-rouge">snr_sweep.py</code></td>
      <td><code class="language-plaintext highlighter-rouge">snr_sweep.json</code></td>
    </tr>
    <tr>
      <td>per-dataset shuffled-label fits and their spectra</td>
      <td><code class="language-plaintext highlighter-rouge">cover_by_dataset.py</code></td>
      <td><code class="language-plaintext highlighter-rouge">cover_by_dataset.json</code></td>
    </tr>
    <tr>
      <td>the shuffled-label law and collapse figure</td>
      <td><code class="language-plaintext highlighter-rouge">fig_shuffled_collapse.py</code></td>
      <td><code class="language-plaintext highlighter-rouge">truth_shuffled_collapse.png</code></td>
    </tr>
    <tr>
      <td>\(d'_{\text{in}}\) against \(2\sqrt{\mathrm{PR}/N}\), and AUROC against \(\Phi(d'/\sqrt2)\)</td>
      <td><code class="language-plaintext highlighter-rouge">cover_gaussian_check.py</code></td>
      <td><code class="language-plaintext highlighter-rouge">cover_gaussian_check.json</code></td>
    </tr>
    <tr>
      <td>pool-size control on the collapse constant</td>
      <td><code class="language-plaintext highlighter-rouge">cover_pool_control.py</code></td>
      <td><code class="language-plaintext highlighter-rouge">cover_pool_control.json</code></td>
    </tr>
    <tr>
      <td>spectra, PR, alignment, \(d'_{\mathrm M}\)</td>
      <td><code class="language-plaintext highlighter-rouge">geometry_observables.py</code></td>
      <td><code class="language-plaintext highlighter-rouge">geometry_observables.json</code></td>
    </tr>
    <tr>
      <td>the same observables on <code class="language-plaintext highlighter-rouge">companies</code>, <code class="language-plaintext highlighter-rouge">common_claim</code>, <code class="language-plaintext highlighter-rouge">conj</code></td>
      <td><code class="language-plaintext highlighter-rouge">geometry_extra_datasets.py</code></td>
      <td><code class="language-plaintext highlighter-rouge">geometry_extra_datasets.json</code></td>
    </tr>
    <tr>
      <td>shuffled-label floor of the in-sample \(\hat d'_{\mathrm M}\)</td>
      <td><code class="language-plaintext highlighter-rouge">check_dm_insample.py</code></td>
      <td><code class="language-plaintext highlighter-rouge">check_dm_insample.json</code></td>
    </tr>
    <tr>
      <td>synthetic in-sample versus held-out \(d'\) at planted \(d' = 1\)</td>
      <td><code class="language-plaintext highlighter-rouge">check_insample_attenuation.py</code></td>
      <td><code class="language-plaintext highlighter-rouge">check_insample_attenuation.json</code></td>
    </tr>
    <tr>
      <td>rogue-dimension steering and decoding of \(\hat v_1\), \(\hat\theta_\perp\)</td>
      <td><code class="language-plaintext highlighter-rouge">check_rogue_dimension.py</code></td>
      <td><code class="language-plaintext highlighter-rouge">rogue_dimension.json</code></td>
    </tr>
    <tr>
      <td>steering sweeps and \(\chi\)</td>
      <td><code class="language-plaintext highlighter-rouge">steer_confirm2.py</code>, <code class="language-plaintext highlighter-rouge">chi_whitening_analysis.py</code></td>
      <td><code class="language-plaintext highlighter-rouge">steer_ckpt/</code></td>
    </tr>
    <tr>
      <td>distractor transfer</td>
      <td><code class="language-plaintext highlighter-rouge">transfer_to_likely.py</code></td>
      <td><code class="language-plaintext highlighter-rouge">transfer_likely.json</code></td>
    </tr>
    <tr>
      <td>unsteered judgment readout on <code class="language-plaintext highlighter-rouge">cities</code>, <code class="language-plaintext highlighter-rouge">common_claim</code>, <code class="language-plaintext highlighter-rouge">counterfact</code></td>
      <td><code class="language-plaintext highlighter-rouge">check_model_verdicts.py</code></td>
      <td><code class="language-plaintext highlighter-rouge">check_model_verdicts.json</code></td>
    </tr>
    <tr>
      <td>score gradient and overlaps</td>
      <td><code class="language-plaintext highlighter-rouge">score_gradient.py</code></td>
      <td><code class="language-plaintext highlighter-rouge">score_gradient_*.json</code></td>
    </tr>
    <tr>
      <td>OLMo cross-family geometry</td>
      <td><code class="language-plaintext highlighter-rouge">geometry_olmo.py</code></td>
      <td><code class="language-plaintext highlighter-rouge">geometry_olmo.json</code></td>
    </tr>
    <tr>
      <td>OLMo decoding, nulls, superposition</td>
      <td><code class="language-plaintext highlighter-rouge">snr_sweep.py</code> run on OLMo-2-1B</td>
      <td><code class="language-plaintext highlighter-rouge">olmo_goNogo.json</code></td>
    </tr>
    <tr>
      <td>massive coordinates, droppers, cleaned held-out \(d'\)</td>
      <td><code class="language-plaintext highlighter-rouge">outlier_check.py</code></td>
      <td><code class="language-plaintext highlighter-rouge">outlier_check.json</code></td>
    </tr>
    <tr>
      <td>shrinkage intensity and its decomposition</td>
      <td><code class="language-plaintext highlighter-rouge">extract_shrinkage.py</code>, <code class="language-plaintext highlighter-rouge">shrinkage_decomposition.py</code></td>
      <td><code class="language-plaintext highlighter-rouge">shrinkage_intensity.json</code>, <code class="language-plaintext highlighter-rouge">shrinkage_decomposition.json</code></td>
    </tr>
    <tr>
      <td>\(\rho\) sweep and cross-validated \(\rho\)</td>
      <td><code class="language-plaintext highlighter-rouge">shrinkage_sweep.py</code>, <code class="language-plaintext highlighter-rouge">shrinkage_cv.py</code></td>
      <td><code class="language-plaintext highlighter-rouge">shrinkage_sweep.json</code>, <code class="language-plaintext highlighter-rouge">shrinkage_cv.json</code></td>
    </tr>
    <tr>
      <td>\(400\)- and \(100\)-draw steering nulls</td>
      <td><code class="language-plaintext highlighter-rouge">extend_null.py</code></td>
      <td><code class="language-plaintext highlighter-rouge">steer_ckpt/*__NULL.json</code></td>
    </tr>
    <tr>
      <td>fixed-seed bootstrap intervals on the overlaps</td>
      <td><code class="language-plaintext highlighter-rouge">regen_gradient_ci.py</code></td>
      <td><code class="language-plaintext highlighter-rouge">score_gradient_*.json</code></td>
    </tr>
    <tr>
      <td>the twelve datasets</td>
      <td><code class="language-plaintext highlighter-rouge">fetch_geometry_of_truth.py</code></td>
      <td> </td>
    </tr>
    <tr>
      <td>figures</td>
      <td><code class="language-plaintext highlighter-rouge">fig_*.py</code>, <code class="language-plaintext highlighter-rouge">make_cluster_figures.py</code></td>
      <td><code class="language-plaintext highlighter-rouge">truth_*.png</code></td>
    </tr>
  </tbody>
</table>

<p><strong>Reading the gradient decomposition.</strong> The identity \(\chi(w) = c\,(w \cdot g)\)
comes with two caveats. First, it is a small-\(h\) statement,
while the measured \(\chi\) is fit through the origin over
\(h \in \{0.5, 1, 2, 4\}\). The two agree to two percent on <code class="language-plaintext highlighter-rouge">cities</code> at layer
\(28\) — \(c\,(\hat\theta \cdot g) = +0.029\) against a measured \(\chi = +0.030\) —
and disagree by an order of magnitude on <code class="language-plaintext highlighter-rouge">counterfact</code>, \(-0.002\) against \(-0.023\).
Where the class signal is tiny, the fitted susceptibility picks up curvature at the
large-\(h\) end, so the decomposition should be read for signs and orderings
rather than magnitudes.</p>

<p>Second, the rotation plane of the main text, \(\mathrm{span}\{\hat v_1, \hat e_2\}\),
contains the estimators but not the causal direction. At the four cells where a rank test is quoted, only
\(5\)–\(7\%\) of \(g\)’s norm falls in the plane (the range over all twelve
layer–dataset cells is \(0.5\)–\(10\%\)), against the \(\sqrt{2/d} = 0.028\) a random
2-plane would capture. That is a factor of two above chance, not zero and not most of
it. The plane describes what the estimator does, not the geometry of the causal
channel.</p>

<p><strong>The \(\kappa_{\mathrm{eff}}\) inversion.</strong> The gain formula of <em>The correction is a
rotation in a plane</em> predicts \(2.7\times\) at \(\kappa = 535\) against a measured
\(4.8\times\) on <code class="language-plaintext highlighter-rouge">counterfact</code> at layer \(28\). Inverting for \(\kappa\) gives
\(\kappa_{\mathrm{eff}} \approx 1.9\times10^{3}\), well above
\(\hat\lambda_1/\hat\lambda_2\). That is consistent with the class gap carrying most of
its orthogonal weight below the second eigendirection, which the two-dimensional
reduction cannot resolve. That inversion is
estimator-dependent. Read against the in-sample \(\hat d'_{\mathrm M}\) instead,
\(d'_{\mathrm{mm}} = 0.089 \to \hat d'_{\mathrm M} = 1.03\) (both on the full set) is a factor of \(11.6\) and would imply
\(\kappa_{\mathrm{eff}} \approx 1.1\times10^{4}\), but \(\hat d'_{\mathrm M}\) is not a clean reading. Under
shuffled labels the identical recipe returns \(\hat d'_{\mathrm M} = 0.79\), so most of the
\(11.6\) is finite-sample inflation of the inverse rather than signal, and the held-out
inversion is the one to trust. The qualitative claim holds either way: \(\kappa_{\mathrm{eff}}\) sits far above
\(\hat\lambda_1/\hat\lambda_2\).</p>

<p><strong>Known limitations.</strong> Every steering cell carries ten seeds except the three-seed
rogue-dimension sweep. The steering null
carries \(400\) draws at the four cells listed under <em>Nulls</em> and \(100\) elsewhere, so at a
\(100\)-draw cell the rank \(p\) resolves only to \(0.01\) and a clearance near
\(p = 0.05\) rests on five draws. Marginal clearances there should be read with
that resolution in mind. Inference
is <code class="language-plaintext highlighter-rouge">float16</code> throughout, which is adequate for the effect sizes here but not for the
third decimal. The geometry observables are computed on the full set rather than the
held-out half, since they describe the data rather than a probe’s performance. And a
single split at seed \(0\) underlies the decoding numbers. The steering and gradient
results carry seed variation, the decoding ones do not.</p>

<p><strong>The whitened numbers are a lower bound.</strong> Every \(d'_{\mathrm F}\) in this post uses
Ledoit–Wolf’s shrinkage intensity, which minimizes
\(\mathbb{E}\lVert\hat\Sigma - \Sigma\rVert_F^2\). That is not the objective the post
reports: what matters here is \(d'\) of the direction \(\hat\Sigma^{-1}\hat\delta\), which
depends on the <em>inverse</em> and on one particular direction in it. The two diverge, and not
subtly. Sweeping \(\rho\) on <code class="language-plaintext highlighter-rouge">counterfact</code> at pythia-2.8b, \(d'_{\mathrm F}\) is larger at
some smaller \(\rho\) than at \(\rho_{\mathrm{LW}} \approx 0.137\) at every layer measured.
Choosing \(\rho\) by five-fold cross-validation inside the training half and scoring once
on the held-out half gives \(0.530\) against \(0.031\) at layer \(16\) and \(0.441\)
against \(0.384\) at layer \(28\), with <code class="language-plaintext highlighter-rouge">cities</code> improving more modestly
(\(3.68 \to 4.39\) at layer \(28\)). I have kept Ledoit–Wolf throughout rather than
switching estimators mid-study, so the whitened direction here understates what whitening can
do. This does not weaken any of the comparisons the post draws, since they all ask whether
whitening beats the plain estimator and clears the null. However, it does mean the cross-validated
selection is not itself reliable at every depth: at layers \(8\) and \(12\) it lands on
the opposite end of the grid and returns values inside the null, because five-fold on a
training half of about \(600\) rows scores \(d'\) on roughly \(120\) points, which cannot
resolve the ridge. Fixing that properly — via repeated cross-validation, or an objective
smoother than \(d'\) — is left open.</p>

<h2 id="appendix-notation">Appendix: notation</h2>

<p>A <strong>hat</strong> marks a quantity estimated from a finite sample and its absence marks the
population quantity it estimates.</p>

<p><em>Data and model.</em></p>

<table>
  <thead>
    <tr>
      <th>symbol</th>
      <th>meaning</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td>\(L\)</td>
      <td>layer index</td>
    </tr>
    <tr>
      <td>\(d\)</td>
      <td>residual-stream width</td>
    </tr>
    <tr>
      <td>\(x \in \mathbb{R}^{d}\)</td>
      <td>activation at the final token of a statement</td>
    </tr>
    <tr>
      <td>\(y \in \{0,1\}\)</td>
      <td>truth label</td>
    </tr>
    <tr>
      <td>\(N\), \(N_{\text{train}}\), \(N_0\), \(N_1\)</td>
      <td>dataset size, training split, per-class counts</td>
    </tr>
    <tr>
      <td>\(\pi_0\), \(\pi_1\), \(\hat\pi_0\), \(\hat\pi_1\)</td>
      <td>class priors</td>
    </tr>
  </tbody>
</table>

<p><em>Population and sample geometry.</em></p>

<table>
  <thead>
    <tr>
      <th>symbol</th>
      <th>meaning</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td>\(\mu_0\), \(\mu_1\), \(\hat\mu_0\), \(\hat\mu_1\)</td>
      <td>class-conditional means</td>
    </tr>
    <tr>
      <td>\(\delta = \mu_1 - \mu_0\)</td>
      <td>class-mean gap</td>
    </tr>
    <tr>
      <td>\(\Sigma\)</td>
      <td>within-class covariance (class-weighted)</td>
    </tr>
    <tr>
      <td>\(\hat C\)</td>
      <td>sample within-class covariance, each class centered on its own mean</td>
    </tr>
    <tr>
      <td>\(\hat\Sigma\)</td>
      <td>Ledoit–Wolf shrinkage estimate of \(\Sigma\), with intensity \(\rho\)</td>
    </tr>
    <tr>
      <td>\(\hat\lambda_i\), \(\hat v_i\)</td>
      <td>eigenvalues and eigenvectors of \(\hat C\), ordered \(\hat\lambda_1 \ge \hat\lambda_2 \ge \cdots\); \(\hat\Sigma\) shares the eigenvectors, with eigenvalues \((1-\rho)\hat\lambda_i + \rho\operatorname{tr}\hat C/d\)</td>
    </tr>
    <tr>
      <td>\(\mathrm{PR}\)</td>
      <td>participation ratio, \((\sum_i\hat\lambda_i)^2/\sum_i\hat\lambda_i^2\)</td>
    </tr>
    <tr>
      <td>\(k\)</td>
      <td>number of leading principal components projected out (superposition probe)</td>
    </tr>
  </tbody>
</table>

<p><em>Directions.</em> All are unit vectors.</p>

<table>
  <thead>
    <tr>
      <th>symbol</th>
      <th>meaning</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td>\(u\)</td>
      <td>a direction that is <strong>read</strong>, giving the scalar \(z = u^{\top}x\)</td>
    </tr>
    <tr>
      <td>\(w\)</td>
      <td>a direction that is <strong>steered along</strong>, added to the activation</td>
    </tr>
    <tr>
      <td>\(\hat\theta\)</td>
      <td>mass-mean (difference-in-means) direction, \(\hat\delta/\lVert\hat\delta\rVert\)</td>
    </tr>
    <tr>
      <td>\(\hat\theta_{\mathrm F}\)</td>
      <td>Fisher direction, \(\propto \hat\Sigma^{-1}\hat\delta\)</td>
    </tr>
    <tr>
      <td>\(\hat\theta_\perp\)</td>
      <td>projection-out control, \(\propto (I - \hat v_1\hat v_1^{\top})\hat\delta\)</td>
    </tr>
    <tr>
      <td>\(\hat e_2\)</td>
      <td>unit vector along \(\hat\delta_\perp\), the class gap orthogonal to \(\hat v_1\)</td>
    </tr>
    <tr>
      <td>\(\varphi\), \(\kappa\), \(r\)</td>
      <td>angle from \(\hat v_1\) in the plane, lever arm \(\lambda_1/\bar\lambda\), \(r = \tan\varphi_{\mathrm{mm}}\)</td>
    </tr>
  </tbody>
</table>

<p><em>Decoding.</em></p>

<table>
  <thead>
    <tr>
      <th>symbol</th>
      <th>meaning</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td>\(z = u^{\top}x\)</td>
      <td>projected score</td>
    </tr>
    <tr>
      <td>\(u^{\top}\mu_0\), \(u^{\top}\mu_1\), \(u^{\top}\Sigma u\)</td>
      <td>projected class means, projected within-class variance</td>
    </tr>
    <tr>
      <td>\(d'\)</td>
      <td>separation in units of its own noise, sign-blind, a property of a direction</td>
    </tr>
    <tr>
      <td>\(d'_{\mathrm{mm}}\), \(d'_{\mathrm F}\), \(d'_{\perp}\)</td>
      <td>\(d'\) along \(\hat\theta\), \(\hat\theta_{\mathrm F}\), \(\hat\theta_\perp\), held-out unless marked</td>
    </tr>
    <tr>
      <td>\(d'_{\mathrm M}\)</td>
      <td>\(\sqrt{\delta^{\top}\Sigma^{-1}\delta} = \max_u d'(u)\), the Mahalanobis separation</td>
    </tr>
    <tr>
      <td>\(\hat d'_{\mathrm M}\)</td>
      <td>\(\sqrt{\hat\delta^{\top}\hat\Sigma^{-1}\hat\delta}\), its in-sample estimate with shrunk \(\hat\Sigma\)</td>
    </tr>
    <tr>
      <td>\(\mathrm{AUROC}\)</td>
      <td>\(\Pr[z_1 &gt; z_0]\), sign-aware, \(\Phi(d'/\sqrt{2})\) for Gaussian classes, balanced</td>
    </tr>
    <tr>
      <td>\(C(N,d)\), \(f(N,d)\)</td>
      <td>Cover’s separable-dichotomy count and its fraction \(C/2^{N}\)</td>
    </tr>
  </tbody>
</table>

<p><em>Steering.</em></p>

<table>
  <thead>
    <tr>
      <th>symbol</th>
      <th>meaning</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td>\(h\), \(c\)</td>
      <td>steering coefficient and its unit, \(c = \lVert\hat\delta\rVert\)</td>
    </tr>
    <tr>
      <td>\(\ell(x)\)</td>
      <td>behavioral score, \(\log P(\text{true}\mid x) - \log P(\text{false}\mid x)\)</td>
    </tr>
    <tr>
      <td>\(\Delta(h)\)</td>
      <td>mean shift in \(\ell\) under \(x \mapsto x + h\,c\,w\)</td>
    </tr>
    <tr>
      <td>\(A\), \(S\)</td>
      <td>antisymmetric and symmetric parts of \(\Delta\)</td>
    </tr>
    <tr>
      <td>\(g\)</td>
      <td>mean gradient of the behavioral score, \(\langle\nabla_x \ell\rangle\)</td>
    </tr>
    <tr>
      <td>\(\chi\)</td>
      <td>steering susceptibility, \(\mathrm{d}A/\mathrm{d}h\) at \(h \to 0\), equals \(c\,(w\cdot g)\); measured as the through-origin slope of \(A\) against \(h\)</td>
    </tr>
    <tr>
      <td>seed / draw</td>
      <td>resampling of the train/test split / of a random direction</td>
    </tr>
  </tbody>
</table>

<h2 id="references">References</h2>

<p>A fuller, annotated version of this bibliography — organized as a reader’s map of how these results tension against each other — is at <a href="/reviews/truth-probes-map/">the geometry of truth probes</a>. The grouped list below gives locators only.
<strong>The geometry — separability, capacity, and readout</strong></p>

<ul>
  <li><strong>Cover, <em>Geometrical and Statistical Properties of Systems of Linear Inequalities with Applications in Pattern Recognition</em></strong> — IEEE Trans. Electronic Computers <strong>EC-14</strong>(3):326–334 (1965) <a href="http://hebb.mit.edu/courses/9.641/2002/readings/Cover65.pdf">PDF</a>.</li>
  <li><strong>Diedrichsen, Berlot, Mur, Schütt, Shahbazi &amp; Kriegeskorte, <em>Comparing Representational Geometries Using Whitened Unbiased-Distance-Matrix Similarity</em></strong> — <a href="https://arxiv.org/abs/2007.02789">arXiv:2007.02789</a>, <em>Neurons, Behavior, Data Analysis, and Theory</em> (2021).</li>
</ul>

<p><strong>Truth / honesty directions — reproduction targets</strong></p>

<ul>
  <li><strong>Marks &amp; Tegmark, <em>The Geometry of Truth: Emergent Linear Structure in LLM Representations of True/False Datasets</em></strong> — <a href="https://arxiv.org/abs/2310.06824">arXiv:2310.06824</a>, COLM 2024.</li>
  <li><strong>Bürger, Hamprecht &amp; Nadler, <em>Truth is Universal: Robust Detection of Lies in LLMs</em></strong> — <a href="https://arxiv.org/abs/2407.12831">arXiv:2407.12831</a>, NeurIPS 2024.</li>
  <li><strong>Burns, Ye, Klein &amp; Steinhardt, <em>Discovering Latent Knowledge in Language Models Without Supervision</em> (CCS)</strong> — <a href="https://arxiv.org/abs/2212.03827">arXiv:2212.03827</a>, ICLR 2023.</li>
</ul>

<p><strong>The critiques — identifiability / which direction did you actually find (the SNR angle)</strong></p>

<ul>
  <li><strong>Farquhar, Varma, Kenton, Gasteiger, Mikulik &amp; Shah (DeepMind), <em>Challenges with Unsupervised LLM Knowledge Discovery</em></strong> — <a href="https://arxiv.org/abs/2312.10029">arXiv:2312.10029</a> (2023).</li>
  <li><strong>Roger, <em>What Discovering Latent Knowledge Did and Did Not Find</em></strong> — <a href="https://www.alignmentforum.org/posts/bWxNPMy5MhPnQTzKz/what-discovering-latent-knowledge-did-and-did-not-find-4">AlignmentForum, 2023</a>.</li>
  <li><strong>Mallen &amp; Belrose, <em>Eliciting Latent Knowledge from Quirky Language Models</em></strong> — <a href="https://arxiv.org/abs/2312.01037">arXiv:2312.01037</a> (2023).</li>
  <li><strong>Bao et al., <em>Probing the Geometry of Truth: Consistency and Generalization</em></strong> — <a href="https://aclanthology.org/2025.findings-acl.38.pdf">ACL Findings 2025</a>.</li>
  <li><strong>Ying, Ravfogel, Kriegeskorte &amp; Hase, <em>The Truthfulness Spectrum Hypothesis</em></strong> — <a href="https://arxiv.org/abs/2602.20273">arXiv:2602.20273</a> (2026).</li>
  <li><strong>Ying, Hase &amp; Kriegeskorte, <em>Comparing Linear Probes with Mahalanobis Cosine Similarity</em></strong> — <a href="https://arxiv.org/abs/2606.19603">arXiv:2606.19603</a> (2026).</li>
  <li><strong>Poulis, Crovella &amp; Terzi, <em>Testing the Limits of Truth Directions in LLMs</em></strong> — <a href="https://arxiv.org/abs/2604.03754">arXiv:2604.03754</a> (2026).</li>
</ul>

<p><strong>Controls and methodology</strong></p>

<ul>
  <li><strong>Hewitt &amp; Liang, <em>Designing and Interpreting Probes with Control Tasks</em></strong> — <a href="https://arxiv.org/abs/1909.03368">arXiv:1909.03368</a>, EMNLP 2019.</li>
  <li><strong>MacDiarmid et al. (Anthropic), <em>Simple Probes Can Catch Sleeper Agents</em></strong> — <a href="https://www.anthropic.com/research/probes-catch-sleeper-agents">Anthropic Alignment Note, 2024</a>.</li>
</ul>

<p><strong>Representation geometry — the rogue-dimension lineage</strong></p>

<ul>
  <li><strong>Sun, Chen, Kolter &amp; Liu, <em>Massive Activations in Large Language Models</em></strong> — <a href="https://arxiv.org/abs/2402.17762">arXiv:2402.17762</a>, COLM 2024.</li>
  <li><strong>Timkey &amp; van Schijndel, <em>All Bark and No Bite: Rogue Dimensions in Transformer Language Models Obscure Representational Quality</em></strong> — <a href="https://arxiv.org/abs/2109.04404">arXiv:2109.04404</a>, EMNLP 2021 (pp. 4527–4546).</li>
</ul>

<p><strong>Models</strong></p>

<ul>
  <li><strong>Biderman et al., <em>Pythia: A Suite for Analyzing Large Language Models Across Training and Scaling</em></strong> — <a href="https://arxiv.org/abs/2304.01373">arXiv:2304.01373</a>, ICML 2023.</li>
</ul>

<p><strong>Steering directions</strong></p>

<ul>
  <li><strong>Tan, Chanin, Lynch, Paige, Kanoulas, Garriga-Alonso &amp; Kirk, <em>Analysing the Generalisation and Reliability of Steering Vectors</em></strong> — <a href="https://arxiv.org/abs/2407.12404">arXiv:2407.12404</a>, NeurIPS 2024.</li>
  <li><strong>Braun, Eickhoff, Krueger, Bahrainian &amp; Krasheninnikov, <em>Understanding (Un)Reliability of Steering Vectors in Language Models</em></strong> — <a href="https://arxiv.org/abs/2505.22637">arXiv:2505.22637</a>, ICLR 2025 Workshop on Foundation Models in the Wild.</li>
  <li><strong>Torop, Masoomi &amp; Dy, <em>Inverted Detection and Control in Steering Vectors</em></strong> — <a href="https://arxiv.org/abs/2608.02957">arXiv:2608.02957</a> (2026).</li>
  <li><strong>Liu, <em>Decodable but Not Corrected by Fixed Residual-Stream Linear Steering: Evidence from Medical LLM Failure Regimes</em></strong> — <a href="https://arxiv.org/abs/2605.05715">arXiv:2605.05715</a> (2026).</li>
</ul>

<hr />

<p><em><small>Prose edited with the assistance of Claude (Anthropic), which also checked the numbers against the analysis artifacts and ran supplementary experiments. I ran the main experiments and vouch for every claim here.</small></em></p>]]></content><author><name></name></author><summary type="html"><![CDATA[Treating the recovery of a linear truth direction as a signal-to-noise problem: when is such a direction recoverable at all from a language model's activations, and what does the estimator return when it is not?]]></summary></entry><entry><title type="html">SARSA vs. Q-learning on the cliff revisited</title><link href="https://jasteinberg.github.io/blog/2026/sarsa-vs-qlearning/" rel="alternate" type="text/html" title="SARSA vs. Q-learning on the cliff revisited" /><published>2026-07-08T00:00:00+00:00</published><updated>2026-07-08T00:00:00+00:00</updated><id>https://jasteinberg.github.io/blog/2026/sarsa-vs-qlearning</id><content type="html" xml:base="https://jasteinberg.github.io/blog/2026/sarsa-vs-qlearning/"><![CDATA[<h2 id="introduction">Introduction</h2>

<p>Comparing the SARSA and Q-learning algorithms used to train an agent to walk on a cliff environment is a rite of passage when learning RL. The textbook
result shows that the learned SARSA policy results in the agent taking the safe path far from the edge and the Q-learning policy results in the agent taking the optimal-but-risky path on the edge of the cliff. These algorithms differ only in which next-state action-value enters the bootstrap target — both still <em>behave</em> the same $\varepsilon$-greedy way. The usual demonstration runs SARSA and Q-learning on <em>separate</em> sets of trajectories, which mixes algorithmic and initialization effects together. To isolate them, the experiment has to control for initialization and exploration at once.</p>

<div class="tldr gray">
  <p>In this post I compare the SARSA and Q-learning algorithms by properly controlling for initialization and exploration.
I find the difference resolves into two numbers that should be considered separately: SARSA earns the higher behavior return by exploring at a safe distance from the cliff, while Q-learning reaches the better greedy policy ($-12$ vs $-16$) because its target ignores the cost of exploration. The real payoff is a confound-free comparison of the two update rules — not variance reduction, which washes out of the run total here.</p>
</div>

<p><em>Code and executed notebooks:
<a href="https://github.com/jasteinberg/rl-alignment-repo">rl-alignment-repo</a>.</em></p>

<p><strong>Related work.</strong> The cliff-walking task is Example 6.6 of
<a href="http://incompleteideas.net/book/the-book-2nd.html">Sutton &amp; Barto</a>; what follows
is a methodological re-examination of a standard example. Pairing runs with
<em>common random numbers</em> is a classical variance-reduction technique from
simulation; in reinforcement learning it appears as the fixed-seed value estimate
of PEGASUS (<a href="https://arxiv.org/abs/1301.3878">Ng &amp; Jordan, 2000</a>) and is used explicitly for variance reduction in
<a href="https://arxiv.org/abs/1502.05477">TRPO</a> (Schulman et al., 2015). That RL
comparisons are highly sensitive to the choice of seed, and that reporting
interval estimates over many runs rather than single trajectories is the honest
alternative, is argued forcefully by
<a href="https://arxiv.org/abs/1709.06560">Henderson et al. (2018)</a> and
<a href="https://arxiv.org/abs/2108.13264">Agarwal et al. (2021)</a>.</p>

<h2 id="the-cliff">The cliff</h2>
<p>The cliff-walking gridworld (Sutton &amp; Barto, Example 6.6) is a $4 \times 12$ grid, where the agent starts at a state $S$ at the bottom-left and attempts to reach a goal state $G$ at the
bottom-right. The entire bottom row between $S$ and $G$ is a cliff, which the agent must learn to avoid. The agent can take four deterministic actions (up/down/left/right) where every step that does not go over the cliff costs $-1$. Stepping into the cliff costs $-100$ and teleports the agent back to $S$. Episodes are
undiscounted ($\gamma = 1$) and end when the agent reaches $G$, where the terminating
transition carries reward $0$. So a path of $L$ moves returns $-(L-1)$ — one smaller in
magnitude than the Sutton &amp; Barto bookkeeping, where the final step into $G$ also costs
$-1$. I use this convention throughout, so the optimal edge path scores $-12$, not the
textbook $-13$.</p>

<p>The shortest route runs one row above the cliff — up once, right eleven
times, down once — for a return of $-12$. The catch is that this optimal path
<em>hugs</em> the cliff: one wrong step down and the agent pays $-100$. Staying safe means
climbing further from the edge and paying for the extra length — the tension the
SARSA and Q-learning algorithms resolve differently.</p>

<h2 id="one-td-error-two-backups">One TD error, two backups</h2>

<p>Both Q-learning and SARSA are tabular TD(0) control algorithms. They differ only in the selection of the
bootstrap target. Writing the temporal-difference error $\delta_t$ gives the following expressions for $\delta_t^{\,\mathrm{SARSA}}$ and $\delta_t^{\,\mathrm{Q}}$:</p>

\[\delta_t^{\,\mathrm{SARSA}} = r_{t+1} + \gamma\, Q(s_{t+1}, a_{t+1}) - Q(s_t, a_t)\]

\[\delta_t^{\,\mathrm{Q}} = r_{t+1} + \gamma\, \max_{a} Q(s_{t+1}, a) - Q(s_t, a_t)\]

<p>with the common update</p>

\[Q(s_t,a_t) \leftarrow Q(s_t,a_t) + \alpha\, \delta_t\]

<p>While SARSA bootstraps from $a_{t+1}$, the action the behavior policy will <em>actually</em>
take next, Q-learning bootstraps from the greedy value, pretending the agent will act greedily next regardless of what it actually does. This means that SARSA’s target carries the cost of exploration while Q-learning’s does not.</p>

<p>This results in SARSA evaluating and improving
the policy it actually follows — the $\varepsilon$-soft policy — and reaching the <em>optimal $\varepsilon$-soft policy</em> as a fixed point. By contrast, Q-learning reaches a fixed point corresponding to the optimal <em>greedy</em> policy, independent of how it behaves. The cliff is simply a
place where those two fixed points visibly disagree, because the optimal
$\varepsilon$-soft policy keeps a safety margin away from the $-100$ region
that the greedy optimum drops.</p>

<h2 id="asymptotic-comparison">Asymptotic comparison</h2>

<p>With a fixed $\varepsilon &gt; 0$, Q-learning’s greedy policy converges to the
cliff-edge optimum ($-12$), while SARSA’s settles on the maximally cautious
route along the top row ($-16$): at $\varepsilon = 0.1$ the per-step cost of
being cliff-adjacent, $\approx \tfrac{\varepsilon}{4}\times 100 = 2.5$, is large
enough that its optimal $\varepsilon$-soft policy backs all the way off the edge. With the
constant $\alpha = 0.5$ used here, neither set of $Q$-values literally
converges — they reach a stationary distribution of $O(\alpha)$ width around
the fixed point — but the greedy <em>policy</em> stabilizes well before the values
stop fluctuating, which is why I read the asymptotic numbers off greedy
rollouts rather than the running values. However, the quantity usually
studied is the <em>online</em> return — the sum of rewards collected while behaving
$\varepsilon$-greedily. There, SARSA wins: it rarely falls, while Q-learning walks the edge and occasionally falls off, incurring a cost of $-100$.</p>

<p>The subtlety is worth stating explicitly: “SARSA is safer” is a claim about <em>learning
under sustained exploration</em>, not about the asymptotic policy. If $\varepsilon$ is annealed to $0$, both SARSA and Q-learning converge to the same greedy optimum and the gap between their expected rewards closes. But unless behavior return and greedy-evaluation return (defined below) are kept separate, the two results get conflated.</p>

<h2 id="shared-seeds-do-not-imply-shared-exploration">Shared seeds do not imply shared exploration</h2>

<p>Cliff-walking has deterministic transitions, so the randomness lives entirely in
three places: the $Q$-initialization, the $\varepsilon$-greedy exploration
draws, and tie-breaking among equal-value actions. Two common flaws in analysis are:</p>

<p><strong>Under-control.</strong> Seed one global RNG, run SARSA, then run Q-learning. They
consume random numbers at different rates and in different states, so after the
first step where their policies differ, their streams are effectively unrelated.
So using the same seed guarantees nothing about whether the two agents faced
the same exploration.</p>

<p><strong>Not enough trials.</strong> A single seed is one sample, and the cliff is
high-variance — a couple of unlucky falls swing the online return by tens of
points. One run can make either algorithm look better by chance.</p>

<h2 id="the-fair-experiment">The fair experiment</h2>

<p>Two steps reconcile these issues. First, <strong>isolate randomness into named streams</strong> so
each source is reproducible and independently ablatable — one generator for
initialization, one for exploration, one for tie-breaking, spawned from a single
<code class="language-plaintext highlighter-rouge">SeedSequence</code>. Second, <strong>pair the two algorithms with common random numbers
(CRN)</strong> so for each seed, SARSA and Q-learning receive the <em>same</em> init and exploration
streams. They begin in identical states and share exploration draws up to the first policy
divergence.</p>

<p>This is the classic variance-reduction trick. The estimand is the gap
$\Delta = R_{\mathrm{S}} - R_{\mathrm{Q}}$, and</p>

\[\mathrm{Var}(\Delta) = \mathrm{Var}(R_{\mathrm{S}}) + \mathrm{Var}(R_{\mathrm{Q}}) - 2\,\mathrm{Cov}(R_{\mathrm{S}}, R_{\mathrm{Q}}).\]

<p>In principle, sharing randomness makes $\mathrm{Cov}(R_{\mathrm{S}}, R_{\mathrm{Q}})$
positive and shrinks the variance of the <em>difference</em> without biasing it — free
variance reduction for the exact quantity of interest. How much you actually
get depends on how long the shared randomness survives, and on the cliff that is
the interesting part.</p>

<p>In practice the cancellation is limited. CRN synchronizes the streams only until
the policies diverge, and on the cliff they diverge early and stay diverged —
after that the two agents visit different states and consume the shared draws in
different contexts. Measured directly, paired over master seeds, the per-episode correlation is
modest ($\rho \approx 0.25$) and falls to essentially zero in the run-total online
return: a run sums hundreds of episodes, the shared randomness couples the two
agents only in the short pre-divergence window of each, and so each run-total
averages over many near-independent trajectories. The variance reduction on the
quantity I actually plot is therefore small — expected, and not the reason to pair. The reason is the part that
does <em>not</em> depend on the cancellation surviving: pairing removes the
differing-exploration confound <em>by construction</em>, so the gap $\Delta$ is a clean
comparison of the two update rules at matched exploration rather than an artifact
of one agent happening to explore into the cliff more often. With enough seeds the
gap converges to the same value either way — pairing just makes each comparison
honest, which is what a controlled experiment is for.</p>

<p>In this experiment I use the same $\alpha$, $\varepsilon$, number of episodes, same initialization of $Q$, and the same environment for both algorithms. The <em>only</em> difference between the Q-learning and SARSA runs comes from the greedy versus the on-policy selection of $a_{t+1}$ in the Q update. I then average over many seeds (the cliff requires $M \gtrsim 100$ for
clean bands), and report a 95% confidence band rather than a single curve. Finally, I evaluate the learned greedy policies <em>separately</em> with greedy rollouts at $\varepsilon = 0$ so the asymptotic path is reported distinctly from the
behavior return.</p>

<h2 id="the-core-of-the-experiment">The core of the experiment</h2>

<p>The whole comparison rests on one detail: at every step both algorithms draw the
next action $a_{t+1}$ from the <em>same</em> exploration stream, so the two runs stay
coupled (common random numbers) and differ only in which value enters the
bootstrap target — the on-policy $a_{t+1}$ for SARSA, the greedy $\max_a$ for
Q-learning.</p>

<div class="language-python highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="c1"># a2 is drawn every step in BOTH algorithms, from a shared exploration
# stream, so the paired runs stay aligned (common random numbers)
</span><span class="n">a2</span> <span class="o">=</span> <span class="n">egreedy</span><span class="p">(</span><span class="n">Q</span><span class="p">,</span> <span class="n">s2</span><span class="p">,</span> <span class="n">eps</span><span class="p">,</span> <span class="n">rng_explore</span><span class="p">,</span> <span class="n">rng_tie</span><span class="p">)</span>
<span class="k">if</span> <span class="n">algo</span> <span class="o">==</span> <span class="s">"sarsa"</span><span class="p">:</span>
    <span class="n">target</span> <span class="o">=</span> <span class="n">r</span> <span class="o">+</span> <span class="n">gamma</span> <span class="o">*</span> <span class="n">Q</span><span class="p">[</span><span class="n">s2</span><span class="p">[</span><span class="mi">0</span><span class="p">],</span> <span class="n">s2</span><span class="p">[</span><span class="mi">1</span><span class="p">],</span> <span class="n">a2</span><span class="p">]</span> <span class="o">*</span> <span class="p">(</span><span class="ow">not</span> <span class="n">done</span><span class="p">)</span>    <span class="c1"># on-policy a_{t+1}
</span><span class="k">else</span><span class="p">:</span>  <span class="c1"># q-learning
</span>    <span class="n">target</span> <span class="o">=</span> <span class="n">r</span> <span class="o">+</span> <span class="n">gamma</span> <span class="o">*</span> <span class="n">Q</span><span class="p">[</span><span class="n">s2</span><span class="p">[</span><span class="mi">0</span><span class="p">],</span> <span class="n">s2</span><span class="p">[</span><span class="mi">1</span><span class="p">]].</span><span class="nb">max</span><span class="p">()</span> <span class="o">*</span> <span class="p">(</span><span class="ow">not</span> <span class="n">done</span><span class="p">)</span>  <span class="c1"># greedy value
</span><span class="n">Q</span><span class="p">[</span><span class="n">s</span><span class="p">[</span><span class="mi">0</span><span class="p">],</span> <span class="n">s</span><span class="p">[</span><span class="mi">1</span><span class="p">],</span> <span class="n">a</span><span class="p">]</span> <span class="o">+=</span> <span class="n">alpha</span> <span class="o">*</span> <span class="p">(</span><span class="n">target</span> <span class="o">-</span> <span class="n">Q</span><span class="p">[</span><span class="n">s</span><span class="p">[</span><span class="mi">0</span><span class="p">],</span> <span class="n">s</span><span class="p">[</span><span class="mi">1</span><span class="p">],</span> <span class="n">a</span><span class="p">])</span>
</code></pre></div></div>

<p>The full implementation — the named RNG streams spawned from a single
<code class="language-plaintext highlighter-rouge">SeedSequence</code>, the paired-seed loop over $M = 100$ seeds, greedy evaluation at
$\varepsilon = 0$, and the figure code — is in the
<a href="https://github.com/jasteinberg/rl-alignment-repo/tree/main/notebooks/tabular_control">notebook</a>.</p>

<h2 id="results">Results</h2>

<p><img src="/assets/figures/cliff_greedy_paths.png" alt="Greedy rollouts after training: SARSA along the top row, Q-learning along the cliff edge" class="fig-single" /></p>

<p><em>Greedy policies after 500 episodes (seed 0). Q-learning takes the cliff-edge optimum ($-12$); SARSA backs all the way to the top row and accepts the longer route ($-16$) to stay clear of the $-100$ region. The policies differ exactly where the greedy and $\varepsilon$-soft optima disagree.</em></p>

<p><img src="/assets/figures/sarsa_learning_curves.png" alt="Online-return learning curves: SARSA settles higher than Q-learning" class="fig-single" /></p>

<p><em>Online return during training, mean over 100 CRN-paired seeds with 95% bands. SARSA settles higher (around $-28$) because it rarely falls; Q-learning sits lower (around $-50$) as it walks the edge and occasionally steps off. The dotted and dashed lines mark the two greedy-evaluation returns ($-12$ and $-16$) — far above the behavior return, which carries the cost of exploration.</em></p>

<p>Two returns are in play and they should be compared separately. The <strong>behavior return</strong> is the reward actually collected while the agent acts $\varepsilon$-greedily <em>during training</em> — exploration costs and
cliff-falls included; this is the online return plotted above (≈ $-28$ for SARSA,
≈ $-50$ for Q-learning). The <strong>greedy-evaluation return</strong> is the final learned
policy run greedily at $\varepsilon = 0$, with no exploration ($-16$ for SARSA,
$-12$ for Q-learning). SARSA earns the higher behavior return because its policy keeps a margin from the
cliff: an $\varepsilon$-greedy exploratory step from a safe row is usually just a
$-1$ detour, so few episodes carry a $-100$ fall. Q-learning walks the cliff edge,
where a single random step often drops into the $-100$ region — so more episodes
include a fall, and those penalties drag its episode-averaged return down (the
$\approx \tfrac{\varepsilon}{4}\times 100 = 2.5$ per-step cliff-adjacency cost from
before, now paid along the whole edge). The greedy-evaluation return flips the
ranking: Q-learning wins it because its target values the greedy policy regardless
of how the agent behaves. Both are correct; they answer different questions.</p>

<p><img src="/assets/figures/sarsa_delta.png" alt="Paired difference Delta between SARSA and Q-learning online return" class="fig-single" /></p>

<p><em>The paired gap $\Delta = R_{\mathrm{S}} - R_{\mathrm{Q}}$ in online return, mean over 100 seeds with a 95% band. It sits above zero through training — SARSA’s behavior-return advantage — but the band stays wide: pairing fixes the exploration confound without buying much variance reduction on this quantity, since the per-episode coupling washes out of the run total.</em></p>

<h2 id="takeaways">Takeaways</h2>

<p>The algorithmic gap between SARSA and Q-learning is one symbol — <code class="language-plaintext highlighter-rouge">max</code> versus the
on-policy $a_{t+1}$ — and it maps cleanly onto “value the greedy policy” versus
“value the policy you actually run, exploration and all.” The empirical gap on the
cliff is real but it is a statement about learning under sustained
exploration, and it dissolves as $\varepsilon \to 0$.</p>

<p>In general, for a clean comparison of two RL algorithms: first isolate
randomness into named streams; then pair the algorithms with common random
numbers so the comparison is about the update rule and not the random number
drawn; then average over many seeds and report intervals; and finally, keep
behavior return and greedy-evaluation return as separate columns. “Same seed” is table stakes, not a guarantee — and pairing’s real dividend is a
clean comparison at matched exploration, not the variance it happens to save.</p>

<p>The full analysis — the common-random-numbers coupling modes, the first-divergence
diagnostic, the annealing/GLIE study of when SARSA’s greedy policy actually
shortens, and the minimum-exposure solution set SARSA samples from — is in the
<a href="https://github.com/jasteinberg/rl-alignment-repo/tree/main/notebooks/tabular_control">tabular-control notebook</a>.</p>

<p><em><small>Prose edited with the assistance of Claude (Anthropic), which also checked the numbers against the analysis artifacts and ran supplementary experiments. I ran the main experiments and vouch for every claim here.</small></em></p>]]></content><author><name></name></author><summary type="html"><![CDATA[I treat the textbook cliff-walking comparison of SARSA and Q-learning as a controlled experiment — randomness isolated into named streams, the two algorithms paired with common random numbers — and separate behavior return from greedy-evaluation return.]]></summary></entry><entry><title type="html">Emergence as metric composition in LLMs (Part II): the training axis</title><link href="https://jasteinberg.github.io/blog/2026/training-axis-composition/" rel="alternate" type="text/html" title="Emergence as metric composition in LLMs (Part II): the training axis" /><published>2026-06-16T00:00:00+00:00</published><updated>2026-06-16T00:00:00+00:00</updated><id>https://jasteinberg.github.io/blog/2026/training-axis-composition</id><content type="html" xml:base="https://jasteinberg.github.io/blog/2026/training-axis-composition/"><![CDATA[<h2 id="introduction">Introduction</h2>

<p>In <a href="/blog/2026/parameter-axis-emergence/">Part I</a>, I studied a benchmark’s
accuracy as a function of parameter count $N$ using eight Pythia models of
increasing size, each scored at the end of training and asked whether the
sharp increase in accuracy survives a finite-size-scaling analysis in the number of model parameters $N$. This post runs the same kind of analysis on the <em>training-data</em> axis: I select a single model at fixed parameter count $N$ and record how both the per-token error rate and the benchmark built from it evolve as a function of $D$, the number of tokens seen.</p>

<p><a href="https://arxiv.org/abs/2206.07682">Wei et al. (2022)</a> did not discuss training-set size in their
emergence story, since most model families fix the number of
training tokens across sizes. As a result, the emergence debate has mostly played out along
scale. In <a href="https://arxiv.org/abs/2304.15004">Schaeffer, Miranda &amp;
Koyejo (2023)</a>, the authors made the following statement about
<em>composition</em>: a smooth per-token accuracy $p$, pushed through a hard
sequence-level metric, gives $A \approx p^{\ell}$, and that product turns a
featureless underlying curve into a sharp one. As arithmetic this is
indifferent to how the given value of $p$ evolves, but the evidence for it is read off
model <em>scale</em>, where per-token loss falls as a power law in parameters and
compute. In this post, I look at the training-dynamics of a single model and analyze how the sequence-level metric is assembled out of per-token competence. Specifically I ask: how far do the per-token errors entering the composition deviate from complete independence?
The composition law
$A \approx \prod_j p_j$ assumes the per-token errors are independent across
answer positions. This assumption can be directly probed from the data. So along training I follow two quantities: the per-token error rate
$1 - p(D)$ as it falls over the training run, and the <em>correlation</em> between those
errors across positions. Specifically I consider the gap between the product of the marginals and the observed joint accuracy as a measure of the deviation from independence.</p>

<div class="tldr gray">

  <p>In this post I examine whether the benchmark’s sharp turn-on is a transition
or an artifact of the metric, along the <em>training</em> axis for a
single model held fixed and the token budget $D$ swept across checkpoints, so
that accuracy and loss can be tracked together along one trajectory. Rather than looking at the peak in the response $\chi = dA/d\log D$ as the signature of a
transition, I will look at the behavior of the
fluctuation susceptibility and the correlation length under finite-size
scaling, i.e. the criterion that actually classifies a transition. I find that composition, $A \approx p^{\ell}$, accounts for
essentially all of the steepness. The only genuine error correlation is a
small, <em>positive</em>, short-ranged carry cascade, visible at the digit level and
opposite in sign to the negative token-level correlation a naive null reports;
and that the connected correlation stays small and non-diverging through the
peak. Along the way I reconcile the deflationary reading with the
loss-perspective account of <a href="https://arxiv.org/abs/2403.15796">Du et al. (2024)</a>,
mapping accuracy directly onto pre-training loss and finding the response peak
to sit on a smooth, monotone stretch of the loss curve. I conclude that along this axis, on this task, the turn-on is a crossover — smooth competence read through a hard metric — not a transition of any order.</p>

</div>

<p><em>Code and executed notebooks: <a href="https://github.com/jasteinberg/scaling-experiments-repo">scaling-experiments</a>.</em></p>

<h2 id="the-composition-mechanism-and-what-the-training-axis-adds">The composition mechanism, and what the training axis adds</h2>

<p>For exact match on a length-$\ell$ target a smooth per-token accuracy $p$ composes to</p>

\[A \approx p^{\ell}\]

<p>which rises abruptly even when $p$ is a featureless power law. Rescoring the
BIG-Bench catalogue with locally linear metrics — token edit distance, Brier
score — dissolves most of the emergence. This fixes the <em>full shape</em> of the sharpening of the accuracy curve, given $p$.</p>

<p>In Part I, I looked at this quantity for models of different parameter counts. Two things open up once the model is held fixed and one considers the training axis. The
first is <em>resolution</em>. For one model one can obtain a dense grid of checkpoints through training which is fine enough to measure a susceptibility in seen tokens $D$ defined as</p>

\[\chi = \frac{dA}{d\log D}\]

<p>By contrast, for model parameters $N$ there are only a few coarse points on the same curve. The second is a check on the
assumption in $A \approx \prod_j p_j$ that the per-token errors
are independent across positions. Composition fixes the joint exact-match
exactly, given the marginals, so the gap between that prediction and the
measured joint accuracy is the <em>deviation from independence</em> — a connected correlation for multi-token answers, but only when the marginals are not also pooling over heterogeneous items, as the results below make concrete.</p>

<h2 id="setup">Setup</h2>

<p>The model suite, the synthetic few-shot $d$-digit addition task, the readout, and the baselines are exactly as in <a href="/blog/2026/parameter-axis-emergence/">Part I</a>; I consider the same two measured quantities — teacher-forced per-token accuracy $p$ and exact match $A$, with the teacher-forced/free-running equivalence used as a computational shortcut — as is the $d = 2$ tokenizer coincidence that gives $\ell = 1$ and $A = p$ <em>identically</em> (composition structurally absent), against $d = 3$ with $\ell \approx 1.85$ (one- or two-token answers, composition nontrivial).</p>

<p><strong>From the parameter axis to the training axis</strong> Part I scored each Pythia size once, at the end of training, and scaled the parameter count $N$. Here I hold the model size fixed and scale $D$, the number of tokens seen, using the 154 public intermediate checkpoints released for every size. The batch schedule is identical across sizes, so checkpoint step $t$ corresponds to a fixed token count ($D \approx 2.1\times10^{6}\,t$) and the checkpoint index is a clean proxy for $D$. The sweep below walks a single fixed-size model through 11 log-spaced checkpoints — the dense $D$-coverage a handful of model sizes cannot give — which is what lets me measure a susceptibility $\chi = dA/d\log D$ in seen tokens, rather than the coarse finite difference in $N$ that Part I was limited to.</p>

<p><strong>A checkpoint-integrity caveat</strong> The training axis needs checkpoints whose weights genuinely differ, and <code class="language-plaintext highlighter-rouge">pythia-2.8b</code> fails that test: its Hub intermediate checkpoints serve the <em>final-step</em> weights at every revision, identical to the bit by weight checksum. I therefore verify each model by checksum before use and run the sweep on <code class="language-plaintext highlighter-rouge">pythia-2.8b-deduped</code>, which passes (distinct, monotonically evolving checksums).</p>

<h2 id="experiment-susceptibility-along-training">Experiment: susceptibility along training</h2>

<p><strong>The susceptibility test</strong> I fit the smooth microscopic
quantity $p(D)$ from the well-resolved regime — $D$ being the number of
tokens seen at fixed model size — compose it
through the per-item target lengths to obtain
$A_{\mathrm{pred}}(D)$, and compare</p>

\[\chi_{\mathrm{pred}} = \frac{dA_{\mathrm{pred}}}{d\log D}\]

<p>against $\chi_{\mathrm{obs}}$: peak location, height, and shape. Agreement
means the composition mechanism suffices for this task and the errors can be treated as independent; any discrepancy represents a deviation from that assumption. The composition law assumes teacher forcing, so I record
both teacher-forced and free-running accuracies and report their gap.</p>

<h2 id="results">Results</h2>

<h3 id="susceptibility-test-along-the-training-axis">Susceptibility test along the training axis</h3>

<p>To test the composition mechanism as a function of $D$ I take a single fixed-size model (<code class="language-plaintext highlighter-rouge">pythia-2.8b-deduped</code>, per the integrity check in the setup) evaluated at $11$ log-spaced
checkpoints, which walks through its training transition densely enough
to measure a response function. This is a transition in <em>one</em> model as
it learns, not a comparison across model scales — so it speaks to how
the sequence-level metric is built from per-token accuracies, not to
emergence with parameters.</p>

<p>On $d = 2$ the model traces a full sigmoid over training,
$A = 0 \to 0.85$, with the steep rise between $D \approx 7\times10^{10}$
and $1.3\times10^{11}$ tokens. Since $d = 2$ has $\ell = 1$, this curve <em>is</em>
the per-token accuracy $p(D)$ — what the figure shows is the per-token error
rate $1 - p(D)$ falling over training, with no metric standing between the
microscopic quantity and the benchmark. The susceptibility
$\chi = dA/d\log D$ has a clean peak of $0.535$ at $D = 1.3\times10^{11}$
— a clean response peak, measured along training rather than asserted. A
response peak alone is not a transition; every sigmoid has one, and what
a genuine transition would additionally require is set out in the
discussion.</p>

<p><img src="/assets/figures/susceptibility_dedup.png" alt="susceptibility" /></p>

<p>Because $d = 2$ has $\ell = 1$, the composition model’s prediction
$A_{\mathrm{pred}} = \prod_j p_j$ reduces to $p$ itself, and indeed
$A_{\mathrm{pred}}$ and $A_{\mathrm{obs}}$ agree to machine precision at
every checkpoint, with $\chi_{\mathrm{pred}}$ peaking at $0.539$ against
$\chi_{\mathrm{obs}} = 0.535$ which is shown by the same peak, in the same place. This
is the consistency check: where composition is trivially exact, it is
exactly exact.</p>

<p>For $d = 3$ the answer runs to $\ell \approx 1.85$ tokens on average and
composition is non-trivial. A clean directional gap opens: the
independence prediction <em>overpredicts</em> the observed joint at every
checkpoint with signal, and the connected residual
$\phi \equiv A_{\mathrm{obs}} - A_{\mathrm{pred}}$ is <em>negative</em>, peaking
at $\phi = -0.032$ in the steep part of the rise
($A_{\mathrm{pred}} = 0.057$ against $A_{\mathrm{obs}} = 0.025$ at
$D = 1.34\times10^{11}$). The fast reading is that this is the failure of
the independence assumption, with the sign saying per-token correctness
is <em>anti</em>-correlated. That reading is too fast — most of the residual is
not a correlation at all. The rest of this section takes it apart, and
what is left is the opposite sign.</p>

<h3 id="the-residual-is-mostly-two-answer-length-classes-pooled">The residual is mostly two answer-length classes, pooled</h3>

<p>The NeoX tokenizer assigns a <em>single</em> token to integers up to three
digits, so the $d = 3$ test set is really two answer-length classes:
three-digit sums ($\ell = 1$, one token) and four-digit sums
($\ell = 2$, two tokens), with fractions $f_1 = 0.15$, $f_2 = 0.85$ at
the residual peak. The pooled null estimates one position-0 marginal
across <em>both</em> classes and multiplies — and that pooling has an exact
cost. Writing $a_S$ for the single-token accuracy, $b_1, b_2$ for the
two-token leading/trailing accuracies, and $\phi_2 = \mathrm{Cov}(C_1,C_2)$
for the genuine within-item covariance, the residual factorizes:</p>

\[\phi_{\mathrm{pooled}}
= \underbrace{f_1 f_2\,(1-b_2)\,(a_S - b_1)}_{\text{length pooling (Simpson)}}
\;+\; \underbrace{f_2\,\phi_2}_{\text{genuine within-item}} .\]

<p>The first term is not an error correlation at all — it is a Simpson
effect from mixing two populations of different per-token difficulty. And
they differ sharply: at the peak $a_S = 0.054$ but $b_1 = 0.219$.
Producing an entire three-digit answer as one token (one draw from
$\sim 800$ values) is four times harder than producing the leading token
of a four-digit answer, whose value is nearly pinned. With $a_S \ll b_1$
the pooling term is large and negative; it dominates early and decays as
the two accuracies equalize over training:</p>

<table>
  <thead>
    <tr>
      <th>$D$</th>
      <th>$\phi_{\mathrm{pooled}}$</th>
      <th>pooling</th>
      <th>within $f_2\phi_2$</th>
      <th>$\phi_2\pm\mathrm{se}$</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td>$1.7\times10^{10}$</td>
      <td>$-0.011$</td>
      <td>$-0.009$</td>
      <td>$-0.002$</td>
      <td>$-0.002\pm.001$</td>
    </tr>
    <tr>
      <td>$6.7\times10^{10}$</td>
      <td>$-0.024$</td>
      <td>$-0.018$</td>
      <td>$-0.006$</td>
      <td>$-0.007\pm.001$</td>
    </tr>
    <tr>
      <td>$1.3\times10^{11}$</td>
      <td>$-0.032$</td>
      <td>$-0.017$</td>
      <td>$-0.015$</td>
      <td>$-0.017\pm.004$</td>
    </tr>
    <tr>
      <td>$2.1\times10^{11}$</td>
      <td>$-0.014$</td>
      <td>$-0.003$</td>
      <td>$-0.011$</td>
      <td>$-0.013\pm.007$</td>
    </tr>
    <tr>
      <td>$3.0\times10^{11}$</td>
      <td>$-0.018$</td>
      <td>$+0.003$</td>
      <td>$-0.020$</td>
      <td>$-0.024\pm.008$</td>
    </tr>
  </tbody>
</table>

<p><img src="/assets/figures/phi_decomposition_d3.png" alt="composition residual decomposed" class="fig-single" /></p>

<h3 id="conditioning-the-null-on-length-and-the-sharpness-claim">Conditioning the null on length, and the sharpness claim</h3>

<p>The fix is to condition the null on answer length — compose within each
class and recombine, $A_{\mathrm{pred}}^{\mathrm{lc}} = f_1 a_S + f_2 b_1 b_2$
— so the Simpson term is removed by construction and the residual is
exactly $f_2\phi_2$. This shifts the susceptibility comparison the test
turns on. With one finite-difference estimator throughout
($\chi = dA/d\log D$ on the 11-point log-$D$ grid, $\chi_{\mathrm{obs}} = 0.285$):</p>

<table>
  <thead>
    <tr>
      <th>null</th>
      <th>peak $\chi_{\mathrm{pred}}$</th>
      <th>$\chi_{\mathrm{obs}} - \chi_{\mathrm{pred}}$</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td>pooled</td>
      <td>$0.243$</td>
      <td>$+0.042$</td>
    </tr>
    <tr>
      <td>length-conditioned</td>
      <td>$0.265$</td>
      <td>$+0.020$</td>
    </tr>
  </tbody>
</table>

<p>Roughly half of the apparent excess sharpness was the pooling artifact:
removing it lifts the predicted susceptibility from $0.243$ to $0.265$,
and the surviving margin of $+0.020$, on eleven coarse checkpoints with a
slightly non-monotone tail, sits within finite-difference noise. So the
strong reading I was tempted by — <em>the transition is sharper than
composition predicts</em> — does not survive. What survives is the level
statement: even length-conditioned, composition overpredicts the joint,
because $\phi_2$ is genuinely negative ($-0.017 \pm 0.0044$, $3.9\sigma$ at
the peak). Whether <em>that</em> is a correlation is the last question.</p>

<p><img src="/assets/figures/susceptibility_lengthcond_d3.png" alt="susceptibility, pooled vs length-conditioned null" class="fig-single" /></p>

<h3 id="the-carry-cascade-is-real--at-the-digit-level">The carry cascade is real — at the digit level</h3>

<p>First, what a carry cascade is and why it is the error structure to expect.
Addition is computed place by place, a carry propagating from each column into
the one above whenever the column sums to ten or more. Any procedure that
respects place value couples adjacent digits: a wrong digit, or a mishandled
carry, at one place corrupts the place above it. The signature such an algorithm
must produce is a <em>positive</em> correlation between adjacent-digit errors,
concentrated where carries originate (units–tens) — so given that the model
emits the answer as a digit sequence, this is the structure to expect a priori
if it is doing arithmetic at all; a <em>negative</em> correlation would be the anomaly.
That is why the token-level $\phi_2$ below, though negative, is the wrong
quantity to interpret: the unit it is computed on is the tokenizer’s, not the
place-value alphabet’s.</p>

<p>$\phi_2$ is a covariance between two <em>tokens</em>, and the token boundary
floats: among four-digit answers the first token holds ${1,2,3,4}$
digits with counts ${11, 385, 202, 17}$. To see the arithmetic instead
of the tokenizer, I re-score the residual-peak checkpoint
($D = 1.34\times10^{11}$) and read out each answer <em>digit</em>, right-aligned
by place value. The per-place error covariance is <em>positive</em> almost
everywhere (units…thousands):</p>

\[\mathrm{Cov}(E_i, E_j) =
\begin{pmatrix}
\cdot &amp; {+}.026 &amp; {-}.006 &amp; {+}.002\\
{+}.026 &amp; \cdot &amp; {+}.006 &amp; {+}.021\\
{-}.006 &amp; {+}.006 &amp; \cdot &amp; {+}.023\\
{+}.002 &amp; {+}.021 &amp; {+}.023 &amp; \cdot
\end{pmatrix},\qquad \overline{\text{off-diag}} = +0.012,\]

<p>with the adjacent units–tens pair — where a mishandled carry actually
propagates — the largest entry at $+0.026$. Per-place accuracy runs
$[0.37, 0.17, 0.25, 0.48]$ from units to thousands: the units digit needs
no incoming carry and is the easiest low place, the carry-laden middle is
hardest, and the thousands digit (almost always $1$) is easiest of all.
This is the carry cascade the metric-artifact story always implied —
<em>positive</em> error correlation, errors clustering across digit positions.</p>

<p>The token-level sign is the opposite only because the floating split
pools over heterogeneous partitions. Conditioning $\phi_2$ on (answer
length, split point) collapses it almost entirely:</p>

<table>
  <thead>
    <tr>
      <th>answer digits, split $m$</th>
      <th>$n$</th>
      <th>token cov</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td>3, split 1</td>
      <td>277</td>
      <td>$+0.006$</td>
    </tr>
    <tr>
      <td>4, split 2</td>
      <td>385</td>
      <td>$-0.010$</td>
    </tr>
    <tr>
      <td>4, split 3</td>
      <td>202</td>
      <td>$\;\,0.000$</td>
    </tr>
    <tr>
      <td>weighted within-group</td>
      <td>864</td>
      <td>$-0.003$</td>
    </tr>
    <tr>
      <td>pooled $\phi_2$</td>
      <td>875</td>
      <td>$-0.017$</td>
    </tr>
  </tbody>
</table>

<p>Eighty-five percent of $\phi_2$ is between-group pooling; the genuine
within-(length, split) token correlation is $-0.003$, indistinguishable
from zero and sign-mixed across groups. With positive digit correlation a
<em>fixed</em> partition gives a non-negative token covariance, so the negative
$\phi_2$ is not the model’s errors anti-correlating — it is the BPE
boundary, floating against the place-value structure, mixing populations
one more time.</p>

<p>Run across all eleven checkpoints, the digit-level cascade is not a
peak artifact but a feature that <em>grows</em> with training: the units–tens
covariance climbs through the transition, from noise to $+0.026$ at the
residual peak and on to $+0.063$ by the end of training, exactly as the
model consolidates the carry algorithm — while the token-level $\phi_2$
stays negative at every checkpoint. Read mechanistically, that growth is what
it means for the model to be <em>adding</em>: competence organized per digit and place,
the residual error riding the carry chain — units into tens most strongly, the
carry-laden middle places failing together — rather than distributed holistically
over the answer or localized to single tokens.</p>

<p><img src="/assets/figures/digit_cascade_d3.png" alt="digit-level carry cascade" /></p>

<h2 id="discussion">Discussion</h2>

<p>Stripped of artifacts, the $d = 3$ training axis says something simple.
The metric-composition mechanism accounts for essentially all of the
sequence-level sharpness on this task: a smooth per-digit competence,
pushed through a hard exact-match metric, manufactures the steep
benchmark curve, exactly as the deflationary account argues. The one
piece of genuine correlation structure is the carry cascade, and it is
<em>positive</em> — visible only when the residual is computed on the
place-value alphabet rather than the tokenizer’s. Everything negative in
the token-level residual is pooling: first over answer-length classes
($\ell \in {1,2}$, the Simpson term that was half the peak gap), then,
inside the two-token class, over the floating token split (most of the
rest). There is no anti-correlation to explain. There is a carry cascade,
plus two layers of tokenization bookkeeping that disguise it and flip its
sign.</p>

<p>That is the cautionary point worth keeping, and the reason the analysis
is here in full rather than compressed to its conclusion. The connected
residual $A_{\mathrm{obs}} - A_{\mathrm{pred}}$ is a tempting order
parameter — it is exactly the deviation from independence — but its sign
and magnitude are contaminated by any heterogeneity the null pools over.
Answer length and tokenization split are two such axes, and on this task
they dominate the genuine signal and even reverse it. A composition test
is only trustworthy after conditioning on whatever the marginals are
pooled across, or — cleaner — scored at the task-natural granularity,
here the digit. The lesson inverts the <a href="/blog/2026/dyck-circuits/">Dyck-$(k,m)$
post</a>: there the structure lived in
correlations invisible to any single per-position statistic; here a
per-position null <em>invents</em> structure that the finer description
dissolves.</p>

<p><strong>The phase-transition dictionary, made explicit.</strong> The rest of this
discussion leans on the critical-phenomena analogy, and emergence is the
place where it is easiest to overclaim, so it is worth fixing the
dictionary together with the one fact it makes unavoidable: here the free
energy is the loss, and <em>continuity of the free energy classifies
nothing</em>. A first-order transition, a second-order transition, and a
smooth crossover all have a continuous free energy; at any finite system
size none of them has a true singularity at all. The loss $L(D)$ — a
per-token cross-entropy, an averaged negative log-likelihood, the
free-energy density of the model’s predictive distribution — is smooth
and monotone in $D$, and that smoothness is shared by every case, a
genuine transition included. It is not, on its own, evidence that nothing
is happening.</p>

<p><strong>The order parameter is the accuracy $A$</strong>, the mean of a ${0,1}$
correctness indicator, in the role of a magnetization $m = \langle s
\rangle$. This buys one clean statement and withholds another. A
<em>first-order</em> transition is a discontinuous jump in the order parameter,
and $A$ rises continuously, so nothing here is first-order-like. But a
continuous order parameter is equally consistent with a second-order
transition and with a crossover, so continuity alone settles nothing. The
Part I caveat also still holds: $A$ is tied to no broken symmetry and is
not $\partial L/\partial h$ for any field $h$ — this is an analogy of
role, not an identity.</p>

<p><strong>The susceptibility is two objects, and only one diagnoses
criticality.</strong> What this post measures, $\chi = dA/d\log D$, is a
<em>response</em> — how fast the order parameter turns on as the control $\log D$
is dialed — and a response peak is generic: every sigmoid has one, so its
peak marks the steepest part of the turn-on and nothing more. The
susceptibility that actually <em>diverges</em> at a continuous transition is a
different object, the <em>fluctuation</em> susceptibility $\chi_{\mathrm{fl}}
\propto \sum_{ij}\phi_{ij}$ — the variance of the order parameter, the
integrated connected correlation (here the digit-level error covariance),
tied by fluctuation–dissipation to a correlation length and divergent
only because correlations grow long-ranged. What would separate a genuine
second-order transition from composition dressing a smooth curve is
therefore not a sharper response peak, which is free, but the fluctuation
susceptibility growing without bound: $\phi$ acquiring a range that
<em>grows</em> with system size, and the response peak sharpening and rising
toward a divergence under finite-size scaling with definite exponents.
The discriminator is the singularity structure of the derivatives as the
system grows, never the continuity of the free energy. Held to that
standard the rise here is a crossover: the genuine connected correlation
is small and short-ranged — nearest-neighbor carry coupling, no growth
in range — and the response peak is finite.</p>

<p><strong>The loss-perspective rebuttal.</strong> The strongest counter to the deflationary
reading is not Schaeffer’s but <a href="https://arxiv.org/abs/2403.15796">Du et al. (2024)</a>,
who argue emergence is real once accuracy is plotted against <em>pre-training loss</em>
rather than size or compute, and that the loss threshold survives even under
continuous metrics. Nothing here conflicts with it, and with the dictionary above
the reconciliation is immediate. The loss is the free energy, and it is a smooth,
strictly monotone function of $D$; plotting $A$ against $L$ rather than against
$\log D$ is a smooth invertible reparameterization of the control axis, and a
smooth reparameterization can neither create a singularity nor change the order of
one. Du et al.’s axis cannot manufacture a transition the $D$ axis lacks. Mapped
explicitly — held-out Pile loss for the <code class="language-plaintext highlighter-rouge">pythia-2.8b-deduped</code> trajectory at the
same checkpoints — the loss has no feature of its own where the accuracy turns on:
the response peak at $D^* = 2.1\times10^{11}$ sits at $L \approx 2.0$ nats on a
smooth, monotone stretch, the loss gliding by only about $0.15$ nats across the
interval where $A$ climbs from $0.001$ to $0.205$. That shows the loss axis is a
benign reparameterization here — <em>not</em> that smoothness rules out a transition,
which, per the dictionary, it never could.</p>

<p><img src="/assets/figures/loss_mapping_d3.png" alt="accuracy and loss vs tokens" /></p>

<p>Worse for the rebuttal, the collapse onto a loss curve is <em>predicted</em> by
composition: if per-token competence $p$ is a function of how well-trained the
model is — same loss, same $p$ — then $A \approx p^{\ell}$ is automatically a
function of loss alone, so the cross-size collapse Du et al. report is a
consequence of the deflationary account, not evidence against it.</p>

<p>On this task the collapse is only partly visible, and worth showing as it is.
Along the training axis the $d = 3$ accuracy of three sizes does lie on a common
$A(L)$ curve, but only <code class="language-plaintext highlighter-rouge">pythia-2.8b-deduped</code> reaches low enough loss to climb it;
<code class="language-plaintext highlighter-rouge">pythia-1b</code> and <code class="language-plaintext highlighter-rouge">pythia-1.4b</code> never clear $A \approx 0.01$ and so only pin the
high-loss foot.</p>

<p><img src="/assets/figures/loss_collapse_d3.png" alt="accuracy vs Pile loss, three training trajectories" class="fig-single" /></p>

<p><em>Accuracy versus Pile loss along training ($d = 3$), one curve per size, training
running left as loss falls. Consistent with a single $A(L)$ curve, but only the
largest size reaches low enough loss to enter the rise.</em></p>

<p>The parameter axis traces the same relation more fully: the eight final
checkpoints, scored once each, fall on one rough $A(L)$ curve with the larger,
lower-loss models carrying the turn-on — though the fit is only as clean as the
suite, with <code class="language-plaintext highlighter-rouge">pythia-6.9b</code> sitting below trend, the documented anomaly in that
model.</p>

<p><img src="/assets/figures/loss_collapse_paramaxis.png" alt="accuracy vs final Pile loss across eight sizes" class="fig-single" /></p>

<p><em>Teacher-forced accuracy versus final-checkpoint Pile loss across the eight sizes
(70M–12B). Against pre-training loss rather than parameter count the sizes
collapse onto one rough curve — what composition predicts, since
$A \approx p^{\ell}$ inherits its loss-dependence from $p$.</em></p>

<p>The collapse is real but loose; what matters here is only that it is <em>consistent</em>
with $A$ being a function of loss, as composition requires, and not a separate
phenomenon demanding a transition. The two
analyses then answer different questions. Du et al. rebut the
<em>metric-discontinuity</em> claim — that sharpness is a hard-metric artifact that
vanishes under a continuous metric; I grant the loss dependence, even in the
continuous $p$, and ask about the <em>order</em> of the transition. By the dictionary
that question is decided by the fluctuation susceptibility and the correlation
length, on which the loss-collapse is silent: the connected correlation $\phi$
stays small, short-ranged, and non-diverging through $D^*$. A finite response
peak on a flat, short-ranged correlation is a crossover dressed sharp by exact
match — located consistently on every axis, collective on none.</p>

<p>Because the measurement runs along training in a single model, all of
this speaks to the <em>composition mechanism</em>, not to emergence with
parameter scale — that is <a href="/blog/2026/parameter-axis-emergence/">Part I</a>.
The same caution applies there: $\phi(N)$ in Part I is the same $d = 3$
residual read across model sizes, so it carries the same length-pooling
contamination and deserves the length-conditioned, digit-level treatment
applied here before its small negative value is read as a correlation.</p>

<h2 id="reproducibility">Reproducibility</h2>

<p>Hardware, precision, and the evaluation pipeline are exactly as in <a href="/blog/2026/parameter-axis-emergence/">Part I</a>. The checkpoint sweep, the length-conditioned and digit-level error analyses, and the script regenerating every figure here are in <a href="https://github.com/jasteinberg/scaling-experiments-repo">scaling-experiments</a>, with the executed notebook alongside.</p>

<hr />

<p><em><small>Prose edited with the assistance of Claude (Anthropic), which also checked the numbers against the analysis artifacts and ran supplementary experiments. I ran the main experiments and vouch for every claim here.</small></em></p>]]></content><author><name></name></author><summary type="html"><![CDATA[Part II of two. I hold the model fixed and watch the per-token error rate, and the benchmark built from it, evolve with the number of tokens seen during training.]]></summary></entry><entry><title type="html">Emergence as metric composition in LLMs (Part I): the parameter axis</title><link href="https://jasteinberg.github.io/blog/2026/parameter-axis-emergence/" rel="alternate" type="text/html" title="Emergence as metric composition in LLMs (Part I): the parameter axis" /><published>2026-06-15T00:00:00+00:00</published><updated>2026-06-15T00:00:00+00:00</updated><id>https://jasteinberg.github.io/blog/2026/parameter-axis-emergence</id><content type="html" xml:base="https://jasteinberg.github.io/blog/2026/parameter-axis-emergence/"><![CDATA[<h2 id="introduction">Introduction</h2>
<p>There have been many widely regarded studies of large language models that have demonstrated that increasing the scale of language models (i.e., training compute, model parameters, etc.) can reliably lead to better performance and sample efficiency on a range of downstream NLP tasks. In this spirit, <a href="https://arxiv.org/abs/2206.07682">Wei et al. (2022)</a> reported on a phenomenon they referred to as <em>emergent abilities</em> of large language models. They considered an ability to be emergent if it was not present in smaller models and its appearance in larger models could not be predicted by extrapolating the performance of smaller models. In the introduction, they refer to the following general definition rooted in Philip Anderson’s essay <a href="https://www.science.org/doi/10.1126/science.177.4047.393">“More is Different”</a>: <em>Emergence is when quantitative changes in a system result in qualitative changes in behavior</em> to provide further context. They do not consider training data set size because many language model families used a fixed number of training examples for all model sizes. The paper mainly focuses on few-shot prompted tasks such as modular arithmetic, TruthfulQA, and Word unscramble among others. In the second figure of their paper, they plot Accuracy as a percentage against log scale model size and show that all of the tasks they study exhibit the same trend: the accuracy stays at zero until a <em>critical</em> model size after which it sharply increases.</p>

<p>In a subsequent paper <a href="https://arxiv.org/abs/2304.15004">Schaeffer, Miranda &amp;
Koyejo (2023)</a> revisited the conclusions of <a href="https://arxiv.org/abs/2206.07682">Wei et al. (2022)</a> and challenged the presence of emergence as it had been defined. They offered an alternative interpretation: most such curves are measured with discontinuous or sharply nonlinear metrics such as accuracy, and a smooth underlying quantity composed through such a metric produces exactly
this shape. The disagreement between the two readings —
“qualitative novelty at scale” versus “measurement artifact” — has
nonetheless persisted, in part because the question is usually posed
in a binary form that the available data cannot answer.</p>

<p>Going back to the original definition of emergence from statistical physics suggests a different form for the question. In the strict sense, phase transitions do not occur in finite systems — the partition function remains analytic and the susceptibility finite — and yet transitions are established in finite systems routinely, by measuring how the <em>signatures</em> of the transition scale as the resolution of the experiment changes.</p>

<div class="tldr gray">
  <p>In this post, I probe for signatures of emergence via a scaling measurement, which I carry out for one well-controlled family of capability curves in three steps. First, I directly confirm <a href="https://arxiv.org/abs/2304.15004">Schaeffer et al.</a>’s “Prediction 2” — that small models sit at a low nonzero accuracy that small test sets censor to zero. Second, I perform a finite-size-scaling analysis of the apparent transition sharpness $s_{m}$ as a function of test-set size $m$, which cleanly separates a resolution-dependent regime in the low-accuracy foot of the accuracy vs. size curve from a resolution-invariant main transition. I show that finite statistics actively <em>manufacture</em> sharpness with a bias exponent I measure and explain via count statistics. Third, I conduct a susceptibility test comparing the observed transition shape against the one the metric-composition model predicts. Both headline mirage predictions hold where the mechanisms operate. At the main transition, though, composition <em>under</em>-predicts: the curve is sharper than independent per-token errors allow, and the shortfall is a correlation between those errors — real, but it does not sharpen with scale, a correction on top of smooth marginals rather than a transition. By the strict, finite-size-scaling definition, there is no emergence here.</p>

</div>

<p><em>Code and executed notebooks:
<a href="https://github.com/jasteinberg/scaling-experiments-repo">scaling-experiments</a>.</em></p>

<h2 id="emergence-free-energies-and-the-thermodynamic-limit">Emergence, free energies, and the thermodynamic limit</h2>

<p>While the “More is Different” phrase is thrown around in many disparate contexts, the actual
argument is quite precise. It is encapsulated in the following quote from the paper.</p>

<p><em>The behavior of large and complex aggregates of elementary particles, it turns
out, is not to be understood in terms
of a simple extrapolation of the properties of a few particles. Instead, at
each level of complexity entirely new properties appear, and the understanding of the new behaviors requires research which I think is as fundamental in its nature as any other.</em></p>

<p>The ground state of a many-body system need not
share the symmetries of its Hamiltonian: a ferromagnet’s Hamiltonian
is rotation-invariant, yet below $T_{c}$ the magnetization picks a
direction. Nothing microscopic is violated, and no amount of
single-spin analysis predicts the collective state, because the
relevant property — the broken symmetry — belongs to the state, not
the constituents.</p>

<p>In second-order phase transitions, the free energy $F(T)$ and its first derivative are continuous. The discontinuity lives in the <em>response functions</em>. For example, in the ferromagnetic transition, the susceptibility</p>

\[\chi = -\frac{\partial^2 F}{\partial h^2}\]

<p>diverges as $|T - T_c|^{-\gamma}$. 
So while the microscopic object ($F(T)$) is smooth, the macroscopic susceptibility is not. Thus, smoothness of an underlying quantity does not, by itself, settle whether sharpness in a derived
quantity is an artifact and the sharpness in a derived quantity does
not, by itself, establish a collective phenomenon. Instead, one must consider
whether a calculable smooth model reproduces the observed sharpness, and how the sharpness behaves as the resolution of the measurement changes.</p>

<p>A true singularity requires the thermodynamic limit. However, for a finite system, the partition function
is a finite sum of analytic terms. Strictly, <em>no
finite system has a phase transition</em>. However, we can still establish the existence of a transition in a finite system by looking at how the peak in the susceptibility scales with the system size. For a system of size $L$, the divergence of $\chi$ is rounded into a peak of height $\sim L^{\gamma/\nu}$ whose location drifts as
$L^{-1/\nu}$. One establishes the transition not by observing a
singularity (which is impossible) but by measuring how the rounded signature
<em>scales</em>. An LLM is always finite — finite parameters, finite data,
finite test sets — so one should never be able to observe a true discontinuity. However, one can still ask whether the finite-size signatures of a capability
transition scale like those of a resolution artifact or of something different. This is explored in the rest of the post.</p>

<h2 id="three-relationships-between-microscopics-and-collective-behavior-and-their-llm-analogues">Three relationships between microscopics and collective behavior, and their LLM analogues</h2>

<p>Condensed matter offers three reference points for how a microscopic
description can relate to collective degrees of freedom.</p>

<p><strong>Single-particle DOFs</strong> In systems like the Landau Fermi liquid, the original microscopic description survives. The low energy excitations of the system are quasiparticles that are adiabatically connected to bare electrons, but have a renormalized mass $m^*$. The LLM analogue of this is that the per-token error is the right variable,
and task-level sharpness is fully accounted for by how a smooth
per-token quantity composes through the metric.</p>

<p><strong>Non-local DOFs</strong> Single particle excitations exist but are nonlocal. An example of this is the transverse-field
Ising model in 1D. In the spin basis, this system looks strongly interacting and undergoes a genuine quantum phase transition at $g = 1$. However, a Jordan–Wigner transformation of the spins takes it to a <em>free</em> fermion theory. Yet the map is nonlocal — each fermion carries a half-infinite string of spin
operators — and although the theory is free, the physically natural
observables (spin correlators) remain nontrivial. In this situation, simplicity in the
right basis coexists with apparent complexity in the basis one measures. For LLMs this suggests that while capabilities may be simple functions of collective variables (circuits, features), the features are distributed and
nonlocal in the per-token or per-neuron description.</p>

<p><strong>No single-particle DOFs</strong> Many strongly correlated electron systems contain phases in which no long-lived single-particle excitations exist. One such phase is the strange metal phase of the cuprates. The electrons are
manifestly present — they carry the current — but spectral functions
are broad and the scattering rate saturates the Planckian scale</p>

\[\tau \sim \hbar / k_B T\]

<p>and resistivity is linear in $T$.
This prohibits any description of the system’s degrees of freedom in terms of single quasiparticle excitations: no change of basis rescues a single-particle picture, and only collective
descriptions survive. The Sachdev–Ye–Kitaev model is the solvable
representative of this physics. The LLM analogues here are capabilities with no clean
description in any accessible variable set, where only coarse-grained
scaling statements are honest.</p>

<h2 id="the-metric-composition-analysis-and-what-scaling-adds">The metric-composition analysis, and what scaling adds</h2>

<p><a href="https://arxiv.org/abs/2304.15004">Schaeffer et al.</a> observed that most claimed emergent abilities are
measured with discontinuous or sharply nonlinear metrics. For exact
match on a length-$\ell$ target, a smooth per-token accuracy $p$
composes to</p>

\[A \approx p^{\ell}\]

<p>assuming errors are independent of each other. This rises abruptly even when $p$ is a featureless power law in
scale. Empirically, re-scoring with locally linear metrics (token
edit distance, Brier score) removes most of the BIG-Bench emergence
catalogue. <a href="https://arxiv.org/abs/2304.15004">Schaeffer et al.</a>’s analysis converts
“emergence” from a narrative about accuracy curves into two calculable
mechanisms — metric composition and test-set resolution — which each make their own prediction.</p>

<p>The composition mechanism predicts, for a given per-token accuracy $p$, how a
nonlinear metric sharpens the benchmark curve; the resolution mechanism
predicts how a small test set censors low-but-nonzero accuracies to zero, and
how that censoring should weaken as the test set grows. I take both
predictions and ask how the <em>signatures</em> of the transition change as I vary
the resolution of the experiment — test-set size $m$, and system size, i.e.
parameter count $N$. Both mechanisms operate, and along the way I find a
regime in the low-accuracy foot where finite statistics actively <em>manufacture</em>
sharpness, with a bias exponent I measure and explain via count statistics.</p>

<p>Subtract those two artifacts and what is left is the actual content of the
emergence question. The composition law above — $A \approx p^{\ell}$, or
$\prod_j p_j$ once the per-position accuracies are allowed to differ — is
exact only if the per-token errors are <em>independent</em>, so the gap between it
and the observed joint accuracy is precisely the correlation between those
errors. The question of emergence then takes a concrete form: does that
correlation structure <em>reorganize</em> with scale — sharpening as the model
grows, the way a correlation length diverges at a critical point — or does it
merely strengthen smoothly? Across the Pythia sizes I observe it is the
latter. The per-token marginals are smooth, the errors are genuinely
correlated, but the correlation structure does not change character with
model size; the benchmark’s apparent threshold is the two metric artifacts
riding on smooth marginals, plus a correlation correction that is real but
analytic in $N$ — a correction, not a transition.</p>

<h2 id="setup">Setup</h2>

<p><strong>The models, and two axes of scale</strong> All experiments use the Pythia
suite (<a href="https://arxiv.org/abs/2304.01373">Biderman et al. (2023)</a>), a set
of decoder-only transformers released specifically to make training
dynamics reproducible: every model size was trained on the <em>same data in
the same order</em>, and 154 intermediate checkpoints are public for each. This gives two independent
scaling axes from one artifact family:</p>

<ul>
  <li><em>Parameter axis.</em> The eight Pythia sizes, 70M–12B, with non-embedding
parameter counts $N$ from $1.9\times10^7$ to $1.1\times10^{10}$ — about
2.8 orders of magnitude. The $d = 2$ resolution analysis uses the first
six (through 2.8B), where the transition lies; the $d = 3$ susceptibility
test uses all eight.</li>
  <li><em>Training axis.</em> For a fixed model size, the 154 public checkpoints trace
accuracy as a function of the number of tokens seen $D$. This axis is the
subject of <a href="/blog/2026/training-axis-composition/">Part II</a>.</li>
</ul>

<p><strong>The task</strong> Few-shot $d$-digit addition, generated synthetically. Each
prompt shows four worked examples and asks for a fifth, e.g. for $d = 2$:</p>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>59 + 63 = 122
15 + 43 = 58
75 + 72 = 147
61 + 48 = 109
71 + 55 =
</code></pre></div></div>

<p>with target ` 126`. The examples fix the format; the model must
complete the final sum. Operands are drawn uniformly in the $d$-digit
range with a fixed seed, so the task is fully reproducible and — the
property that matters for Experiment 1 — the test set can be extended to
<em>any</em> size, unlike a fixed benchmark. I use $n = 4096$ items per
(model, checkpoint, task) unless stated otherwise, and report $d = 2$
and $d = 3$.</p>

<p><strong>Measured output</strong> At the prediction position I read the
model’s output over the vocabulary, and I extract <em>two</em> distinct
quantities, because the emergence debate turns on the difference between
them:</p>

<ol>
  <li><em>Teacher-forced per-token accuracy</em> $p$. I feed the correct full
sequence (prompt $+$ target) and, at each answer-token position, read
the probability the model assigned to the <em>correct</em> token. This is
<em>teacher forcing</em> in the standard sequence-modeling sense
(<a href="https://doi.org/10.1162/neco.1989.1.2.270">Williams &amp; Zipser, 1989</a>):
each token is scored on the ground-truth prefix $y^{\ast}_{&lt;t}$ rather
than on the model’s own previous outputs, which is exactly the rollout
the maximum-likelihood training objective uses and the assumption the
composition law $A \approx p^{\ell}$ is built on. This is
the microscopic quantity — the smooth per-token competence that the
metric-composition argument says gets amplified into apparent
sharpness.</li>
  <li><em>Exact match</em> $A$. Whether the model produces the whole target
answer correctly. The natural definition is free-running — greedily
decode the answer tokens and compare the decoded string to the
target — but there is an equivalent and far cheaper definition that
reads off the teacher-forced pass: an item is exact iff <em>every</em>
target token is the argmax of the model’s distribution given the
correct prefix. The two coincide because greedy decoding first
departs from the target at exactly the first position whose
teacher-forced argmax is wrong. This is the macroscopic benchmark
metric — the quantity that “looks emergent.”</li>
</ol>

<p>The composition model predicts $A \approx p^{\ell}$ for an
$\ell$-token target, and that prediction <em>assumes teacher forcing</em>, so I
record both accuracies and verify the equivalence rather than assume it.
On every model small enough to decode the full test set, free-running
greedy exact match and teacher-forced argmax exact match agree to
$\lesssim 10^{-3}$ — the only items on which they can disagree are those
whose answer string admits more than one tokenization, a sub-percent
effect — so teacher forcing costs nothing here and the composition model
is tested on its own terms. I therefore report the teacher-forced argmax
match throughout: it costs a single forward pass per item, where
free-running decode costs one pass <em>per generated token</em> and becomes the
dominant expense on the largest models by orders of magnitude, and it is
also the quantity the composition law $A \approx p^{\ell}$ refers to most
directly.</p>

<p>The number of answer tokens $\ell$
is set by the tokenizer, and the NeoX tokenizer has dedicated tokens for
small integers: a two-digit sum’s answer (up to three digits) is a
<em>single</em> token. So for $d = 2$, $\ell = 1$ and $A = p$ <em>identically</em> —
the macroscopic metric and the microscopic quantity are the same number.
This is the cleanest possible setting, because it makes the composition
mechanism <em>structurally impossible</em>: there is no exponent $\ell &gt; 1$ for
a nonlinear metric to manufacture sharpness from. Any sharpness seen on
the $d = 2$ task is therefore sharpness in the per-token quantity itself.
The $d = 3$ task is the complement: its answers span one or two tokens
($\ell \approx 1.85$), so composition is nontrivial there — which is why
the susceptibility test (Experiment 2) lives on $d = 3$, where it has
something to test.</p>

<p><strong>Baselines</strong> Chance accuracy is small but nonzero: a $d$-digit sum has
on the order of a few hundred possible answers (e.g. $\sim 180$ distinct
values for $d = 2$), so chance is $\approx 5\times10^{-3}$ — a number
that matters when I claim the small models sit <em>above</em> chance rather
than at zero. All runs are on a single Apple M3 Pro chip (its 18-core integrated GPU, fp16 via the Metal Performance Shaders backend; full spec under Reproducibility), and
every figure regenerates from the linked repository.</p>

<h2 id="experiments">Experiments</h2>

<p><strong>Experiment 1 (finite-size scaling in the test set)</strong> I subsample the
test set at sizes $m = M, M/2, M/4 \ldots$ and extract the apparent
sharpness</p>

\[s_m = \max_x \, \frac{dA_m}{d\log x}\]

<p>over bootstrap resamples,
and examine $s_m$ versus $m$. The scaling of $s_m$ — its sign, its
exponent, where it saturates — is the finite-size-scaling signature
that distinguishes resolution-limited sharpness from
resolution-stable sharpness. Synthetic task generators make $m$
extensible over orders of magnitude, which fixed benchmarks cannot
offer.</p>

<p><strong>Experiment 2 (parameter-axis susceptibility — the actual emergence test)</strong> Across the eight sizes on $d = 3$, I fit smooth $p(N)$ from the well-resolved regime, compose to $A_{\mathrm{pred}}(N)$, and compare</p>

\[\chi_{\mathrm{obs}} = \frac{dA_{\mathrm{obs}}}{d\log N}\]

<p>against $\chi_{\mathrm{pred}}$ — peak location, height, shape — the training-axis test from Part II, but on the axis the emergence claim is actually about.</p>

<h2 id="results">Results</h2>

<h3 id="prediction-2-confirmed">Prediction 2 confirmed</h3>

<p>At $n = 256$, $d = 2$ accuracy reads as flat (“zero”) through 1B
parameters. At $n = 4096$ the foot resolves into a smooth,
above-chance rise: $A = 0.0078,\ 0.0085,\ 0.0129,\ 0.0210$ across
70M–1B (chance is $\sim 0.005$ against the $\sim 180$-value answer
space). This directly confirms Schaeffer et al.’s second prediction:
small models do not have zero accuracy; they have small accuracy that
a small test set censors to zero, since one cannot measure
$A \ll 1/m$. The “appears from nothing” reading of such curves is a
resolution artifact, exactly as they argued.</p>

<p><img src="/assets/figures/accuracy_vs_N.png" alt="accuracy vs N" /></p>

<h3 id="the-transitions-sharpness-is-resolution-stable">The transition’s sharpness is resolution-stable</h3>

<p>The same curves end in a steep rise: $0.114 \to 0.759$ between 1.4B
and 2.8B. For $d = 2$ the composition mechanism is structurally
absent ($\ell = 1$): the macroscopic metric and the microscopic
per-token quantity are the same number, and that number itself jumps
(mean target-token log-probability $-4.18 \to -0.78$) across the
interval. The remaining artifact candidate is test-set resolution,
which Experiment 1 addresses directly.</p>

<p>Subsampling at $m = 64, 128, \ldots, 4096$ and extracting</p>

\[s_m = \max_N \frac{dA_m}{d\log N}\]

<p>gives a two-regime answer:</p>

<ul>
  <li><strong>At the main transition, $s_m$ is flat</strong>: $\theta = +0.002$ for
$d = 2$ ($+0.001$ to $+0.004$ for $d = 3$) over a 64-fold range in
$m$, consistent with zero within bootstrap errors. The maximum
slope sits between two accuracies (0.114 and 0.759) that even 64
items resolve; additional statistics neither soften nor sharpen it.</li>
  <li><strong>In the foot (70M–1B window), $s_m$ falls with $m$</strong>: from 0.0167
to 0.0082, with a single power-law fit giving
$\theta_{\mathrm{foot}} \approx -0.2$. Here low resolution
<em>manufactures</em> apparent sharpness and statistics dissolve it — the
resolution mechanism again, in quantitative form: the bias is an
extreme-value effect with a measurable decay exponent, derived in
Appendix A. Foot-region sharpness should not be taken at face
value.</li>
</ul>

<p><img src="/assets/figures/sharpness_vs_m.png" alt="sharpness vs m" class="fig-single" /></p>

<p>The picture from the two results above is symmetric. Both artifact mechanisms are
measurably present where they apply:
censoring hides the foot, and finite statistics manufacture sharpness
within it. Both are measurably absent at the main transition:
composition is structurally unavailable at $\ell = 1$, and the
sharpness is invariant under a 64-fold change in resolution. What
survives both controls is a property of the underlying per-token
quantity, not of the measurement. The honest scope statement: with
six sizes, the parameter-scaling transition is localized to one
interval and supports “metric-stable and resolution-stable,” not
“divergent” — and that statement, about emergence <em>with model scale</em>,
is as far as six size points can take it.</p>

<h3 id="the-actual-emergence-test-parameter-axis-susceptibility">The actual emergence test: parameter-axis susceptibility</h3>

<p>The central question of the post arrives here: is the across-$N$ jump in exact
match a real sharpening of what the model knows, or only the discontinuous
metric compressing a smoothly-improving per-token quantity? Experiment 1
cleared the two <em>measurement</em> artifacts at $d = 2$, but there the composition
mechanism is structurally absent ($\ell = 1$); to put composition itself on
trial I move to $d = 3$, where $\ell \approx 1.85$ and exact match is a genuine
product of per-token accuracies. The construction is the training-axis test of Part II,
transported to the parameter axis: build the observed exact match
$A_{\mathrm{obs}}(N)$ and the per-position marginal token accuracies
$p_j(N)$; compose the marginals under independence into</p>

\[A_{\mathrm{pred}}(N) = \Big\langle \prod_{j &lt; \ell_i} p_j \Big\rangle_i\]

<p>the exact match a model with those same marginals but <em>independent</em> token
errors would post; and compare the two susceptibilities</p>

\[\chi_{\mathrm{obs}} = \frac{dA_{\mathrm{obs}}}{d\log N}\]

<p>and</p>

\[\chi_{\mathrm{pred}} = \frac{dA_{\mathrm{pred}}}{d\log N}\]

<p>across the eight sizes.
The small sizes use $n = 4096$ to clear the censoring floor; $6.9$B and
$12$B use $n = 1024$.</p>

<table>
  <thead>
    <tr>
      <th>$N_{\mathrm{ne}}$</th>
      <th>$A_{\mathrm{obs}}$</th>
      <th>$A_{\mathrm{pred}}$</th>
      <th>$p$</th>
      <th>$\chi_{\mathrm{obs}}$</th>
      <th>$\chi_{\mathrm{pred}}$</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td>$1.9\times10^{7}$ (70M)</td>
      <td>0.0007</td>
      <td>0.0080</td>
      <td>0.038</td>
      <td>$0.001$</td>
      <td>$0.003$</td>
    </tr>
    <tr>
      <td>$8.5\times10^{7}$ (160M)</td>
      <td>0.0017</td>
      <td>0.0131</td>
      <td>0.058</td>
      <td>$-0.000$</td>
      <td>$0.004$</td>
    </tr>
    <tr>
      <td>$3.0\times10^{8}$ (410M)</td>
      <td>0.0010</td>
      <td>0.0182</td>
      <td>0.073</td>
      <td>$-0.000$</td>
      <td>$-0.002$</td>
    </tr>
    <tr>
      <td>$8.1\times10^{8}$ (1B)</td>
      <td>0.0007</td>
      <td>0.0112</td>
      <td>0.055</td>
      <td>$0.021$</td>
      <td>$0.039$</td>
    </tr>
    <tr>
      <td>$1.2\times10^{9}$ (1.4B)</td>
      <td>0.0127</td>
      <td>0.0346</td>
      <td>0.127</td>
      <td>$0.114$</td>
      <td>$0.126$</td>
    </tr>
    <tr>
      <td>$2.5\times10^{9}$ (2.8B)</td>
      <td>0.2085</td>
      <td>0.2173</td>
      <td>0.454</td>
      <td>$0.118$</td>
      <td>$0.110$</td>
    </tr>
    <tr>
      <td>$6.4\times10^{9}$ (6.9B)</td>
      <td>0.1406</td>
      <td>0.1537</td>
      <td>0.381</td>
      <td>$0.091$</td>
      <td>$0.094$</td>
    </tr>
    <tr>
      <td>$1.1\times10^{10}$ (12B)</td>
      <td>0.2471</td>
      <td>0.2611</td>
      <td>0.504</td>
      <td>$0.189$</td>
      <td>$0.191$</td>
    </tr>
  </tbody>
</table>

<p><img src="/assets/figures/susceptibility_vs_N.png" alt="susceptibility vs N" /></p>

<p>Four observations:</p>

<p><strong>The foot reappears on the parameter axis</strong> The four sizes from 70M to 1B
sit at $A_{\mathrm{obs}} \approx 0.0007$–$0.0017$ — small and above chance,
not zero — while their composed prediction runs $0.008$–$0.018$. The 1B
point reads <em>exactly</em> zero at $n = 1024$ and resolves to $0.0007$ only at
$n = 4096$: the censoring of the foot, reproduced along $N$ rather than
along test-set size.</p>

<p><strong>The composition over-predicts at every size, with one sign</strong>
$A_{\mathrm{pred}}$ exceeds $A_{\mathrm{obs}}$ at all eight sizes — signed
gap $+0.007$ to $+0.022$, mean $0.013$, never changing sign. An independence
assumption <em>over</em>-counts exact match, so the true joint is harder than the
marginals imply: the token-correctness indicators are weakly <em>anti</em>-correlated,
not positively correlated. At the token level this looks like localized error — when an answer
is wrong it is usually wrong in a single token, so the errors appear to repel
rather than cluster. That reading does not survive the dissection: conditioning
on the number of output tokens and re-scoring on digits flips the sign, exposing
a <em>positive</em> carry cascade at the digit level that the floating BPE boundary
disguises as token-level anti-correlation. I work this residual — and the
length-pooling artifact hiding inside it — out in full below.
This is the same-signed residual found on the training axis in Part II,
smaller here than there (where it peaks near $0.03$).</p>

<p><strong>The susceptibilities track; the residual is in the level, not the slope</strong>
$\chi_{\mathrm{obs}} \approx \chi_{\mathrm{pred}}$ at every size — composition
reproduces the <em>shape</em> of the across-$N$ rise. The steep interval is
$1.4$B $\to 2.8$B ($A: 0.013 \to 0.21$, $p: 0.13 \to 0.45$). What the
metric-composition account captures is how the transition is shaped along
$N$; what it misses is a small, uniform, correctly-signed level offset.</p>

<p><strong>Two caveats bound the claim</strong> First, a genuine non-monotonicity:
$A_{\mathrm{obs}}$ runs $0.21 \to 0.14 \to 0.25$ from 2.8B to 6.9B to 12B,
and the dip is present in the per-token accuracy too
($0.45 \to 0.38 \to 0.50$), so it is a real weakness of Pythia-6.9B on this
task, not a metric artifact. Second, both susceptibilities peak at the
<em>largest</em> size as a one-sided edge difference: the transition is not
resolved inside $70$M–$12$B — $A$ is still climbing at 12B — so the peak is
unlocated, and a finite-difference $\chi$ over eight uneven points
straddling a dip is not a robust estimator of it. The defensible statements
are the <em>level</em> of $A$ and the <em>sign</em> of the $A_{\mathrm{pred}} -
A_{\mathrm{obs}}$ gap; the peak location is not one of them.</p>

<p>The honest reading is that along the parameter axis, the $d = 3$ transition sits on
the same metric-composition rung the training axis identifies — the
deflationary account mostly holds with scale, carrying the same small
correlated-error residual — but eight uneven Pythia points, with a real
6.9B dip and an unresolved peak, are too coarse for this to stand as an
independent confirmation. The obvious objection is that all of this is a
property of <em>Pythia</em>, not of the phenomenon. The next section answers that
objection directly.</p>

<h3 id="cross-family-replication-does-the-picture-survive-a-change-of-family">Cross-family replication: does the picture survive a change of family?</h3>

<p>The single sharpest objection to everything above is that it is a fact about
<em>Pythia</em> — one suite, one corpus (the Pile), one tokenizer, one
architecture, with a visible idiosyncrasy (the 6.9B dip) to prove the suite
has idiosyncrasies. A composition result that reproduces on an independent
family, trained on different data with a different tokenizer, is worth far
more than a tighter fit on Pythia alone.</p>

<p>BLOOM is close to the ideal control, for four reasons, only the first of
which is obvious:</p>

<ol>
  <li><strong>It is not math-tuned</strong> BLOOM (2022, the multilingual ROOTS corpus)
predates the instruction- and math-saturated pretraining of the current
generation. Three-digit addition is genuinely hard for it at small scale,
so there is a transition to see; modern suites (Qwen2.5, Llama-3,
Gemma-2) ace the task at every size and show no curve at all.</li>
  <li><strong>It holds the data axis fixed across sizes</strong> Every BLOOM size is
trained on the same ROOTS corpus <em>and</em> a matched token budget — about
$341$ billion tokens — so moving along $N$ changes neither the training
distribution nor the training duration. This is the property that makes
Pythia a clean $N$-scan, and the one that compute-optimal suites
(Cerebras-GPT, anything Chinchilla-scaled) deliberately violate by
growing the token count with $N$.</li>
  <li><strong>It is maximally independent of Pythia</strong> Different corpus (ROOTS vs the
Pile), different tokenizer (a 250k multilingual vocabulary vs the 50k
GPT-NeoX one, hence a different number tokenization and a different
$\ell$ — measured at $1.65$ here against Pythia’s $1.85$), different
positional scheme (ALiBi vs rotary). If the composition picture survives
all of that, it is not a property of any one of those choices.</li>
  <li><strong>Its non-embedding sizes nearly coincide with Pythia’s</strong> This is an
accident of BLOOM’s large vocabulary, which inflates the embedding and
leaves the non-embedding counts at
${0.30,\, 0.68,\, 1.21,\, 2.36,\, 6.04}\times 10^{9}$ — essentially
Pythia’s 410M, 1B, 1.4B, 2.8B, and 6.9B. The comparison is therefore
nearly <em>paired</em>: at matched non-embedding $N$ the two families can be read
against each other, and in particular BLOOM places a point at
$6.0\times10^{9}$, exactly where Pythia-6.9B dips — a direct test of
whether that non-monotonicity belongs to the phenomenon or to Pythia.</li>
</ol>

<p>What BLOOM does <em>not</em> do is densify the transition: in non-embedding units it
carries the same $1.2$–$2.4\times10^{9}$ gap Pythia does. This is a
robustness-and-independence check, not a finer localization.</p>

<p>The sweep is $d = 3$, 4-shot, BLOOM 560M–3B at $n = 4096$ and 7.1B at
$n = 1024$ (offloaded). All five sizes, with the same composition
construction:</p>

<table>
  <thead>
    <tr>
      <th>$N_{\mathrm{ne}}$</th>
      <th>$A_{\mathrm{obs}}$</th>
      <th>$A_{\mathrm{pred}}$</th>
      <th>$p$</th>
      <th>gap</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td>$3.0\times10^{8}$ (560M)</td>
      <td>0.0015</td>
      <td>0.0234</td>
      <td>0.052</td>
      <td>$+0.022$</td>
    </tr>
    <tr>
      <td>$6.8\times10^{8}$ (1.1B)</td>
      <td>0.0002</td>
      <td>0.0266</td>
      <td>0.060</td>
      <td>$+0.026$</td>
    </tr>
    <tr>
      <td>$1.2\times10^{9}$ (1.7B)</td>
      <td>0.0007</td>
      <td>0.0306</td>
      <td>0.064</td>
      <td>$+0.030$</td>
    </tr>
    <tr>
      <td>$2.4\times10^{9}$ (3B)</td>
      <td>0.0017</td>
      <td>0.0358</td>
      <td>0.074</td>
      <td>$+0.034$</td>
    </tr>
    <tr>
      <td>$6.0\times10^{9}$ (7.1B)</td>
      <td>0.0029</td>
      <td>0.0455</td>
      <td>0.094</td>
      <td>$+0.043$</td>
    </tr>
  </tbody>
</table>

<p>Two things reproduce; one differs, and the difference is the point.</p>

<p><strong>The foot of the accuracy curve reproduces, on a corpus and tokenizer that share nothing with the
Pile</strong> Every BLOOM size from 560M to 7.1B sits at a small, censored
$A_{\mathrm{obs}}$ (0.0002–0.0029) over above-zero per-token accuracy
($p = 0.05$–$0.09$), with a positive composition gap throughout. The
censoring-of-the-foot mechanism is not a Pythia artifact.</p>

<p><strong>The composition over-predicts with the same sign</strong> $A_{\mathrm{pred}} &gt;
A_{\mathrm{obs}}$ at all five sizes, as in Pythia. One honest qualification is that
in the foot this gap is dominated by the censoring of $A_{\mathrm{obs}}$ near
zero, not by a clean correlated-error residual at resolved accuracy — the
latter is what Pythia’s 2.8B–12B points demonstrate ($+0.009$ to $+0.014$ at
$A \sim 0.2$–$0.5$), and BLOOM has not reached resolved accuracy within its
released range. So BLOOM confirms the <em>sign</em> of the composition bias; it does
not yet independently pin its magnitude at resolved $A$.</p>

<p><strong>The transition sits roughly an order of magnitude later in $N$</strong> This is
the substantive cross-family difference. At matched non-embedding $N$, Pythia
has already climbed where BLOOM is still in the foot: at $1.2\times10^{9}$,
Pythia-1.4B reads $p = 0.13,\ A = 0.013$ against BLOOM-1.7B’s
$p = 0.064,\ A = 0.0007$; at $\sim!2.4\times10^{9}$, Pythia-2.8B has reached
$A = 0.21$ while BLOOM-3B is still at $A = 0.0017,\ p = 0.074$. BLOOM’s
per-token accuracy merely increases  from $0.052 \to 0.074$ over the same decade in
which Pythia’s runs $0.073 \to 0.454$. Three-digit addition emerges about a
decade of $N$ later for BLOOM, which is unsurprising for a multilingual ROOTS
model measured against the English-dense, arithmetic-rich Pile.</p>

<p><img src="/assets/figures/susceptibility_xfamily.png" alt="cross-family overlay" /></p>

<p>The reading is that the deflationary <em>mechanism</em> — the foot, the censoring, the
positive composition bias — is not a property of Pythia. It reappears in a
family that shares almost nothing with Pythia but its non-embedding parameter
count. What is family-dependent is the transition’s <em>location</em>: BLOOM places
it beyond the reach of its released small models, so the cross-family check
vindicates the mechanism while declining to independently resolve the steep
transition. BLOOM-7.1B, at $N = 6.0\times10^{9}$ and matched to Pythia-6.9B,
makes the point concrete: it reads $A_{\mathrm{obs}} = 0.003$, $p = 0.094$ —
still in the foot. BLOOM has not emerged anywhere in its released range, so the
dip-adjudication is moot here: there is no risen accuracy at which a dip could
appear.</p>

<h3 id="beyond-independence-does-the-correlation-reorganize-at-the-transition">Beyond independence: does the correlation reorganize at the transition?</h3>

<p>A metric artifact and a genuine collective transition can look identical in the
marginals; they need not look identical in the <em>correlations</em>. Independence is a
mean-field step — it multiplies the per-position accuracies and discards the
connected correlation — so if anything escapes the deflationary account, this is
where it hides: in the structure the composition throws away, not in the $p_j$ it
keeps. The residual is small, but its sign is a measurement. Within the two-token
answers (and at $d=3$ the answer is one or two tokens, never more),</p>

\[A_{\mathrm{obs}} - A_{\mathrm{pred}} = \mathrm{Cov}(c_1,c_2)\]

<p>exactly, with no
higher cumulant: the residual <em>is</em> that discarded correlation. I consider two
things — what sign it takes, and whether its structure <em>changes</em> across the
transition, since a genuine collective reorganization would manifest itself as
this correlation swelling at the critical $N$, exactly where the marginals (all
independence can see) show nothing. Per size:</p>

<table>
  <thead>
    <tr>
      <th>size</th>
      <th>$\phi(c_1,c_2)$</th>
      <th>$P(c_2{=}1\mid c_1{=}0)$</th>
      <th>$P(c_2{=}1\mid c_1{=}1)$</th>
      <th>$p_2$</th>
      <th>single-slip</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td>1.4B</td>
      <td>$-0.079$</td>
      <td>0.117</td>
      <td>0.051</td>
      <td>0.11</td>
      <td>26%</td>
    </tr>
    <tr>
      <td>2.8B</td>
      <td>$-0.069$</td>
      <td>0.623</td>
      <td>0.550</td>
      <td>0.60</td>
      <td>69%</td>
    </tr>
    <tr>
      <td>6.9B</td>
      <td>$-0.111$</td>
      <td>0.554</td>
      <td>0.425</td>
      <td>0.52</td>
      <td>62%</td>
    </tr>
    <tr>
      <td>12B</td>
      <td>$-0.102$</td>
      <td>0.703</td>
      <td>0.603</td>
      <td>0.67</td>
      <td>76%</td>
    </tr>
  </tbody>
</table>

<p><img src="/assets/figures/token_correlation_vs_N.png" alt="token correlation vs N" /></p>

<p><strong>The token errors are weakly anti-correlated — but the sign is the tokenizer’s, not the model’s</strong>
$\phi$ is negative at all eight sizes. At 12B, knowing the first token is
<em>wrong</em> raises the second’s hit rate to $0.70$ — above its $0.67$ marginal —
while knowing the first is <em>right</em> drops it to $0.60$. Taken at face value this
reads as localized error, failures concentrating in the harder high-order first
token ($p_1 &lt; p_2$ at every size; at 12B “first wrong, second right” outnumbers
“first right, second wrong” $403$ to $120$). But the token is the wrong unit.
Conditioning the covariance on (answer length, token split) collapses it almost
to zero, and re-scoring on the digit alphabet flips the sign: the per-digit
error covariance is <em>positive</em>, largest for the adjacent units–tens pair. This
is a <em>carry cascade</em>: place-value addition propagates a carry from each digit
into the next, so an error at one place corrupts the place above it, and
positive correlation between adjacent-digit errors is the signature any
carry-respecting algorithm must produce. The token-level $\phi$ is the wrong
quantity to interpret — its negative sign is only what that cascade looks like
after the floating BPE boundary pools over heterogeneous digit partitions. The mechanism is worked out in full on the
training axis in <a href="/blog/2026/training-axis-composition/">Part II</a>, where
conditioning kills $85\%$ of the token covariance and the digit-level units–tens
coupling grows to $+0.06$ through training; the same correction applies here,
size for size. Mechanistically, the model adds place by place with carry
propagation, its errors riding the carry chain; the apparent single-token
localization is the tokenizer’s signature, not the model’s algorithm.</p>

<p><img src="/assets/figures/digit_cascade_vs_N.png" alt="digit cascade vs N" class="fig-single" /></p>

<p><strong>The correlation does not reorganize at the transition</strong> This is the
emergence-relevant result. $\phi(N)$ runs $-0.02,\,-0.05,\,-0.07,\,-0.05,\,
-0.08,\,-0.07,\,-0.11,\,-0.10$ across the eight sizes: small, negative, drifting
only slightly, with no peak or sign-change through the steep interval
($1.4$B$\,\to\,2.8$B: $-0.079 \to -0.069$). To the resolution of eight points,
the correlation independence throws away is <em>scale-stable</em> — it does not swell
at the transition the way a critical susceptibility would. The failure mode
<em>appears</em> to reorganize — single-token failures climb from $8\%$ to $76\%$
across the rise — but that is the marginals moving: as $p_1,p_2$ rise,
single-token failures must come to dominate both-wrong even under independence,
and $\phi$, which divides that out, stays flat. The defensible statement is
“failures appear to localize as accuracy rises,” not “the model reorganizes
into a localized-error regime at the transition.”</p>

<p><strong>Most of the accuracy foot’s over-prediction is a pooling artifact, not correlation</strong>
The pooled first-token marginal mixes the easy one-token answers with the first
token of two-token answers; splitting the gap into that pooling effect plus the
genuine two-token covariance separates them. In the foot the gap is almost all
pooling — the one-token answers (merged four-digit numbers like “ 1935”) are
never correct there, so the pooled marginal over-credits them — and the
covariance is negligible. At the plateau the pooling term reverses and the
covariance takes over: at 12B the net gap $+0.014$ is a token-covariance term
$+0.019$ against pooling terms summing to $-0.005$. So the anti-correlation is
real where accuracy can resolve it, and the foot gap should not be read as
correlation at all.</p>

<p><img src="/assets/figures/phi_N_decomposition.png" alt="phi(N) decomposition and susceptibility" /></p>

<p><strong>What this says about the mirage</strong> Independence is analogous to a mean-field step —
multiply the marginals, discard the connected correlation — and recovering that
correlation does not rescue real emergence; it deepens the deflationary reading.
The sign means the exact-match readout is, if anything, slightly <em>more</em>
suppressed than independence assumes, so $p^{\ell}$ understates the apparent
sharpness rather than inventing it. However, the flatness means the discarded object
carries no transition of its own. The one place genuine collective behavior
could have hidden from the marginals — a correlation that diverges at the
critical $N$ — is, to this resolution, simply not there. BLOOM shows the same
negative $\phi$ across its foot ($-0.02$ to $-0.07$), so the sign is not a
Pythia artifact.</p>

<p>The emergence claim of Wei and Schaeffer is intrinsically about the
<strong>parameter axis</strong>: an ability is “emergent” if it is absent in small
models and present in large ones, which one can only assess by
comparing <em>different models of different sizes</em> — exactly what the
susceptibility test above does. <a href="/blog/2026/training-axis-composition/">Part II</a>
turns to the <em>training</em> axis instead, tracing a single fixed-size model
through its checkpoints. That is no longer a statement about
emergence-with-scale; it is a test of the metric-<strong>composition</strong>
mechanism itself — how per-token accuracies multiply into a
sequence-level metric, true or false independently of which axis one
moves along — and it buys the dense coverage in a single control
variable that eight size points here cannot.</p>

<h2 id="discussion-emergence-with-model-size">Discussion: emergence with model size</h2>

<p>Putting the pieces together along the parameter axis, the per-token marginals
$p_j(N)$ are smooth across the eight Pythia sizes; the exact-match curve is
sharp; and the composition law, which simply multiplies those smooth
marginals, recovers most of that sharpness on its own. What it misses is
small and has a definite sign: the connected correlation $\phi(N)$ is a few
percent, negative, and flat through the steep interval, with no peak where a critical susceptibility would diverge.
The one object that could have carried a genuine transition past the smooth
marginals is, to the resolution of eight sizes, simply not there. That value is
token-level, though, and <a href="/blog/2026/training-axis-composition/">Part II</a> shows
the token-level $d = 3$ residual to carry an answer-length pooling artifact that
flips its sign; $\phi(N)$ here inherits the same caveat, with the digit-level
read — which turns the residual into a small <em>positive</em> carry cascade — feasible
only on the dense training axis.</p>

<p>So for $d = 3$ the across-$N$ transition sits on the metric-stable rung:
the apparent emergence is a hard metric composing smooth per-token competence,
plus a small residual that is mostly tokenization pooling over a smooth,
positive digit-level carry cascade — real, but analytic in $N$. By the
strict, finite-size-scaling reading of emergence this is a negative result —
no qualitative change of state with scale — and it is a <em>quantitative</em>
negative, not a failure to find one: the would-be order parameter is measured,
and it does not move. Whether the same holds along the training axis — where
the correlation is concrete (carry propagation), the checkpoint coverage is
dense, and the loss-perspective reading of emergence
(<a href="https://arxiv.org/abs/2403.15796">Du et al. (2024)</a>) can be met head-on by
mapping accuracy onto the loss curve — is the question of
<a href="/blog/2026/training-axis-composition/">Part II</a>.</p>

<h2 id="where-the-analogy-breaks">Where the analogy breaks</h2>

<p>Phase transitions have an order parameter tied to spontaneous
symmetry breaking; LLM capabilities have no obvious one. Critical
behavior requires a thermodynamic limit; LLM scaling is always
finite, and the “sizes” being scaled — parameters, data, test items —
are not a single $L$. The analogy licenses a methodology, measuring
how signatures scale with resolution; it does not license a
structural identification, and nothing above should be read as a
claim that language models <em>are</em> near-critical systems.</p>

<p><a href="/blog/2026/training-axis-composition/">Part II</a> makes the dictionary behind
this methodology explicit for the training axis — the loss as the free energy,
the accuracy as the order parameter, and the distinction between the <em>response</em>
susceptibility $dA/d\log D$, which peaks for any sigmoid, and the <em>fluctuation</em>
susceptibility that alone diverges at a continuous transition — and uses it to
meet the loss-perspective reading of emergence head-on.</p>

<h2 id="appendix-a-count-statistics-of-the-sharpness-estimator">Appendix A: count statistics of the sharpness estimator</h2>

<p>The accuracy foot’s $\theta_{\mathrm{foot}} \approx -0.2$ from a single
power-law fit hides cleaner structure. Decomposing
$s_m = s_\infty + \delta_m$, the full-set value is
$s_\infty = 0.0082$ and the excess $\delta_m$ falls as a power law
with exponent $\alpha \approx 0.8$–$1.0$ over $m = 64$–$1024$ before
vanishing into the plateau; the shallow $-0.2$ is what a single
power law fitted across the plateau-contaminated range produces.</p>

<p>The exponent is diagnostic of the noise regime, and the controlling
parameter is the expected success count $\lambda = mA$, not $m$. The
estimator $s_m$ takes a maximum over intervals of (true slope $+$
fluctuation), and the maximum of fluctuating quantities is biased
upward — finite test sets manufacture sharpness wherever the curve is
noisy. In the Gaussian regime, $\lambda \gg 1$, accuracy fluctuations
scale as $\sqrt{A/m}$ and the bias decays with $\alpha = 1/2$. In the
count regime, $\lambda \lesssim O(10)$, a fluctuation is one discrete
extra success moving $\hat{A}$ by exactly $1/m$, and the bias decays
with $\alpha \approx 1$. With $A \sim 10^{-2}$ in the foot, the
crossover $\lambda \sim 1$ sits near $m \sim 10^2$ and the Gaussian
regime arrives only at $m \sim 10^3$ — where the excess has already
vanished. The measured $\alpha \approx 0.8$–$1.0$ is the count-regime
exponent with the expected crossover.</p>

<p>The two decay rates follow from one calculation. The estimator</p>

\[s_m = \max_i (\beta_i + \xi_i)\]

<p>maximizes over $K$ adjacent intervals,
each carrying the true slope $\beta_{i}$ plus a mean-zero fluctuation
$\xi_{i}$ of scale $\sigma_m$; the upward bias is</p>

\[\mathbb{E}[\max_i \xi_i] \sim c_K\,\sigma_m\]

<p>with $c_K$ a slowly
varying ($\sqrt{\log K}$) order-unity factor. Everything is in how
$\sigma_m$, the standard deviation of the slope estimate, scales with
$m$. The slope is a finite difference of accuracies</p>

\[\hat A = (1/m)\sum
\mathbf{1}[\text{success}]\]

<p>so $\sigma_m \propto \mathrm{sd}(\hat A)$.</p>

<p>In the <strong>Gaussian regime</strong> ($\lambda = mA \gg 1$) the success count
$\sum_i \mathbf{1}[\text{success}_i]$ is a sum of $m$ i.i.d. Bernoulli($A$)
indicators, so by the central limit theorem it is asymptotically
$\mathcal{N}!\big(mA,\, mA(1-A)\big)$ and the binomial gives</p>

\[\mathrm{sd}(\hat A) = \sqrt{A(1-A)/m} \sim m^{-1/2}\]

<p>hence</p>

<p>\(\delta_m \sim m^{-1/2}\) and 
\(\alpha = \tfrac12\).</p>

<p>In the <strong>count
regime</strong> ($\lambda \lesssim 1$) the central limit theorem has not yet taken
hold — for rare events its approach to the Gaussian is governed by the expected
count $\lambda$, not by $m$ — so successes are still Poisson with mean
$\lambda = mA$, and the slope estimate is quantized: each additional
success moves $\hat A$ by exactly $1/m$, so the smallest non-zero
fluctuation — and hence the bias floor of the maximum — is of order
$1/m$ itself. The upward bias therefore tracks $\sigma_m \sim 1/m$,
giving $\delta_m \sim m^{-1}$ and $\alpha = 1$. The crossover between
the two is at $\lambda \sim 1$, i.e. $m \sim 1/A$, exactly where the
data turn over.</p>

<p>The practical moral: a 4096-item benchmark sounds large, but for a
task at 1% accuracy it holds $\sim 40$ successes. The statistics of
capability feet are count statistics at any test-set size in common
use, and error analysis of emergence claims is naturally done in
$\lambda = mA$, not $m$.</p>

<h2 id="reproducibility">Reproducibility</h2>

<p>All experiments run on a single Apple M3 Pro chip — an Arm system-on-chip with a 12-core CPU (6 performance + 6 efficiency cores) and an 18-core integrated GPU sharing 18 GB of unified memory, with PyTorch on the Metal Performance Shaders (MPS) backend rather than CUDA. All evaluations are fp16, and the unified-memory budget sets where offloading becomes necessary rather than a hard ceiling: the 70M–2.8B models fit in the 18 GB and run on the GPU directly, 6.9B still fits in fp16, and 12B exceeds the budget and is run in fp16 with CPU offload (<code class="language-plaintext highlighter-rouge">device_map=auto</code>) — slower, but enough to evaluate it at all. Task
generators, evaluation code, sweep configurations, and the analysis
producing every number and figure above are in
<a href="https://github.com/jasteinberg/scaling-experiments-repo">scaling-experiments</a>, with the executed analysis
notebook alongside.</p>

<hr />

<p><em><small>Prose edited with the assistance of Claude (Anthropic), which also checked the numbers against the analysis artifacts and ran supplementary experiments. I ran the main experiments and vouch for every claim here.</small></em></p>]]></content><author><name></name></author><summary type="html"><![CDATA[Part I of two. Borrowing the finite-size-scaling methodology from statistical physics, I ask whether emergent abilities postulated to appear in larger models are a measurement artifact across eight Pythia sizes. Specifically, can a benchmark metric's sharpness be explained entirely from composing a smooth per-token accuracy through a hard metric or are there other signatures of emergence hidden in the per-token accuracy?]]></summary></entry><entry><title type="html">The Grokking phase diagram from a single layer transformer learning modular addition</title><link href="https://jasteinberg.github.io/blog/2026/grokking-phase-diagram/" rel="alternate" type="text/html" title="The Grokking phase diagram from a single layer transformer learning modular addition" /><published>2026-06-13T00:00:00+00:00</published><updated>2026-06-13T00:00:00+00:00</updated><id>https://jasteinberg.github.io/blog/2026/grokking-phase-diagram</id><content type="html" xml:base="https://jasteinberg.github.io/blog/2026/grokking-phase-diagram/"><![CDATA[<h2 id="introduction">Introduction</h2>

<p>Grokking is a well known phenomenon in neural networks where models exhibit an abrupt transition to a generalizing solution despite initially overfitting. This occurs long after the training loss has flatlined and is one of the cleanest examples of a phase transition occurring during training.
Grokking was first reported by <a href="https://arxiv.org/abs/2201.02177">Power et al. (2022)</a> when training a decoder only transformer on modular arithmetic. Subsequently, 
<a href="https://arxiv.org/abs/2301.05217">Nanda et al. (2023)</a> reverse-engineered
the Fourier-multiplication circuit behind the generalizable solution. Since then, many additional studies have attempted to explore the conditions in which grokking occurs and uncover mechanisms to explain why it happens.</p>

<p>A great deal of grokking work fixes the architecture and the train/test
split at a single fixed fraction and studies the learning dynamics there. However, the learning dynamics
depend strongly on whether the train/test split sits relative to the capacity of the network.
The canonical “memorize-first, generalize-late” phenomenon is only one of several regimes. The simplest way
to see this is to train with different train test fractions set by scaling the size of the network, and the value of the modulus $p$.</p>

<p>In prior work, <a href="https://arxiv.org/abs/2205.10343">Liu et al. (2022)</a> mapped a four-phase
diagram (comprehension, grokking, memorization, confusion) on modular
arithmetic and noted the existence of a small-decoder regime that does not grok.
<a href="https://arxiv.org/abs/2309.02390">Varma et al. (2023)</a> explained the
late switch through circuit efficiency — weight decay favoring a norm-efficient generalizing circuit over a memorizing lookup. <a href="https://arxiv.org/abs/2402.15175">Huang et al. (2024)</a> drew the 2D phase
diagram over model size and dataset size, and found a small-model regime
of immediate generalization.</p>

<div class="tldr gray">
  <p>In this post I reproduce grokking in the task of modular addition from scratch across a grid of model widths
and training fractions. I recover the phase structure that has appeared in several parts of the literature and analyze it through the same finite-size-scaling lens I have been applying to circuit formation and to emergence
elsewhere on this blog. I treat the critical
training requirement as a finite-size-scaling problem in the modulus
$p$ and ask which rescaling of the data axis — fraction, absolute
count, or coverage — collapses the onset across moduli, attempting to isolate a single critical
<em>fraction</em>.</p>
</div>

<p><em>Code and executed notebooks:
<a href="https://github.com/jasteinberg/interp-repo">interp-repo</a>.</em></p>

<h2 id="setup">Setup</h2>

<p>The task I study is modular addition $a + b \bmod p$ in a one-layer transformer with 4 heads, ReLU MLP trained with full-batch AdamW, and weight decay $1.0$. This is the same setup used by <a href="https://arxiv.org/abs/2301.05217">Nanda et al. (2023)</a> except that I train for up to $40{,}000$ epochs. I chose this number by previously observing the timing of the grokking transition for 10 seeds. During training, the model sees a fixed fraction of the $p^2 = 12{,}769$ possible pairs as its training set and is evaluated on the held-out remainder. 
In the initial experiments, I varied the width and train fraction, keeping other parameters fixed. I defined the width of the network to be the embedding/residual dimension and the train fraction as the share of pairs in the training set. In this setup the width is a proxy for model capacity: networks with larger $d$ can memorize more pairs as a lookup table. 
Likewise, the size of the training set determines the amount of data the model has to memorize to fit the training set.
I chose $d \in {16, 32, 64, 128, 256}$ and $f \in {0.2, 0.3, 0.4, 0.5, 0.6}$.</p>

<p>I run three seeds for each pair of values of $d$ and $f$ for a total of $5 \times 5 \times 3 = 75$ runs. For each run I
log the full train/test accuracy trajectories and read off
three quantities: the presence of a grokking transition (using the working definition of the test accuracy crossing $0.95$ within the budget), the epoch at which grokking occurs (if it does),
and the <strong>memorization delay</strong> defined as the gap in epochs
between the <em>train</em> accuracy crossing $0.95$ and the <em>test</em> accuracy reaching that value. The delay is the quantity that distinguishes the regimes: a large delay is the classic memorize-then-generalize picture; a near-zero delay
means train and test rose together, with no separate memorization phase
to grok out of.</p>

<p>Everything below regenerates from the training script
(<code class="language-plaintext highlighter-rouge">scripts/train_grokking.py</code>) and the committed per-run statistics; no
GPU is needed to reproduce the analysis.</p>

<h2 id="the-phase-diagram">The phase diagram</h2>

<p><img src="/assets/figures/phase_grok_delay.png" alt="grok time and memorization delay" /></p>

<p>In the heat maps above, the right panel shows the memorization delay.</p>

<p><strong>The data-fraction axis dominates, and there is a sharp lower edge</strong>
At $f = 0.2$ nothing groks, at any width: the model memorizes the
training set (train accuracy reaches $1$) but test accuracy never
follows within the budget. This is the data-poor regime — below a
task-and-model-dependent critical fraction, delayed generalization
simply does not occur, consistent with the critical-data-size picture of
<a href="https://arxiv.org/abs/2401.10463">Zhu et al. (2024)</a>. As $f$ increases
the delay collapses by more than two
orders of magnitude, from tens of thousands of epochs down to a few
hundred.</p>

<p><strong>The capacity axis sets the boundary at the critical fraction</strong> The
$f = 0.3$ column is the transition zone, and it is where width matters
most. At small width ($d = 16, 32, 64$) only a minority of seeds grok;
at large width ($d = 128, 256$) all of them do. Capacity controls
<em>whether</em> grokking occurs at the critical data fraction and how long it takes to happen.</p>

<p><strong>Small capacity compresses the delay</strong>
For $d = 16, f = 0.3$: train accuracy there does
not cross $0.95$ until $\sim 11{,}000$ epochs, against $100$–$700$ for
every larger model. The smallest network at the critical fraction can
barely memorize the training set — it is near its capacity limit. When the smallest models <em>do</em> grok (e.g. $d = 16, f = 0.4$), they show
the <em>smallest</em> delay in their column. This is because the memorization shortcut
is not cheaply available. When networks are near capacity, memorization is slow and it no longer wins the race by a wide margin.</p>

<h2 id="four-regimes-one-trajectory-each">Four regimes, one trajectory each</h2>

<p>I plot one representative seed from each corner of the plane to show the four types of trajectories.</p>

<p><img src="/assets/figures/phase_regime_curves.png" alt="regime curves" /></p>

<ul>
  <li><strong>No grokking</strong> ($d = 128, f = 0.2$): train accuracy saturates and
test accuracy sits at chance. There is capacity but too little data —
the model memorizes and stops.</li>
  <li><strong>Memorize-then-grok</strong> ($d = 32, f = 0.3$):
Train accuracy hits $1$ early, but test accuracy languishes for tens of
thousands of epochs, then jumps with a long flat gap between the two curves.</li>
  <li><strong>Compressed delay</strong> ($d = 16, f = 0.4$): train and test rise almost
together. There is still a gap, but a small one — a capacity-limited
model cannot race ahead on memorization, so the two phases nearly
merge.</li>
  <li><strong>Fast / concurrent</strong> ($d = 256, f = 0.6$): ample capacity, ample
data; the model generalizes about as soon as it fits, with only a
short delay.</li>
</ul>

<p>The canonical grokking examples in the literature are usually located in the memorize-then-grok regime. The width of the valley between memorization and the onset of grokking is widest in a narrow band of the
(capacity, data) plane and shrinks to nothing on either side.</p>

<h2 id="grokking-as-a-finite-size-transition">Grokking as a finite-size transition</h2>

<p>I now study the memorization delay as a function of width and data fraction, and show that the emergent jump in test accuracy is not an intrinsic constant but depends strongly on both — and, as the next section shows, on the modulus $p$.</p>

<p>Under this lens the competing explanations stop reading as rivals.
Circuit efficiency (Varma et al.) explains why the generalizing solution
eventually wins; the critical-data-size picture (Zhu et al.) explains
the lower edge in $f$; the speed-competition account (Song &amp; Ye)
explains why the boundary bends with capacity and why small models
compress the delay. Each describes a different feature of the same
surface. Drawing the surface shows they are local descriptions of one global phase structure.</p>

<h2 id="how-to-scale-the-training-set">How to scale the training set</h2>
<p><a href="https://arxiv.org/abs/2201.02177">Power et al. (2022)</a> and <a href="https://arxiv.org/abs/2301.05217">Nanda et al. (2023)</a> set the absolute size of the training set as a fraction $f$ of the $p^2$ possible pairs, which I followed for my initial runs. They quote a critical fraction (for modular addition, $f_c \approx 0.25$–$0.3$) as if it were a property of the task. However, as $p$ changes, the same fraction $f$ corresponds to a different amount of data relative to the size of the model.</p>

<p><a href="https://arxiv.org/abs/2301.05217">Nanda et al. (2023)</a> showed that the generalizing solution is a Fourier-multiplication algorithm built on a handful of
key frequencies; <a href="https://arxiv.org/abs/2205.10343">Liu et al. (2022)</a> characterized the critical training set as the least data that pins down that representation. However, that is a
statement about a <strong>count of constraints</strong>, not a fraction. How the
required amount of training scales with $p$ is an empirical question. Fixing $f$ assumes that it should be $p^2$. However, if the
structure is fixed by $O(p)$ or $O(p\log p)$ relations, then holding $f$
constant across $p$ does not define the phase boundary.</p>

<p>Which rescaling of the data axis makes the threshold
$p$-independent? The most obvious candidates are</p>

<ul>
  <li><strong>fraction $f$</strong> — implying the requirement
scales as $p^2$.</li>
  <li><strong>Absolute pairs $N = f\,p^2$</strong> — implying a fixed number of examples
suffices regardless of $p$.</li>
  <li><strong>Coverage $N/p = f\,p$</strong> — how many times each residue appears in the
training set; a fixed coverage means each symbol is seen a constant
number of times, so the requirement scales as $p$.</li>
</ul>

<p>These predict different motions for $f_c(p)$. A fixed $N$ would need
$f_c &gt; 1$ at $p = 31$, which is impossible, so a fixed count is too strong and the requirement grows <em>sub</em>-quadratically. If coverage is
invariant, $f_c$ should scale as $1/p$; if the fraction is invariant,
$f_c$ should be flat.</p>

<p>I rerun the onset measurement at seven moduli, $p \in {31, 41, 47, 53,
59, 79, 113}$, sweeping the train fraction at each to locate $f_c(p)$ —
the fraction at which the median seed first groks within budget — and
ask which axis, $f$, $N$, or $N/p$, collapses the seven onset curves
onto one.</p>

<p><img src="/assets/figures/pscale_collapse.png" alt="data collapse over p" /></p>

<p>The collapse rules out both a fixed fraction and absolute pairs. The critical fraction is not
constant: it runs from $f_c = 0.5$ at $p = 31$ down to $f_c = 0.25$ at
$p = 113$, halving across the range (left panel — the onset curves fan
out, ordered by $p$). So a single quoted critical fraction is genuinely
$p$-dependent. Absolute pairs $N$ is too demanding: the middle panel
spreads the curves wider still, and a fixed $N$ would require $f_c &gt; 1$
at the smallest prime. Coverage $N/p$ (right panel) brings the curves
closest together of the three — the best simple variable — but it
slightly over-corrects, and the onsets do not perfectly stack.</p>

<p>Fitting the exponent gives the scaling:</p>

\[f_c(p) \sim p^{-0.58},\]

<p>which is close to $p^{-1/2}$, between a $p$-independent fraction-invariance ($p^{0}$) and coverage
($p^{-1}$). In terms of a constraint count, the critical number of pairs
is</p>

\[N_c = f_c\,p^2 \sim p^{2 - 0.58} \approx p^{3/2},\]

<p>$N_c$ runs $480, 840, 884, 983, 1218, 1872, 3192$ across the seven primes — faster than coverage ($\sim p$),
slower than a fixed fraction ($\sim p^2$). The data requirement for
grokking modular addition grows like $p^{3/2}$: neither “a constant
fraction of all pairs” nor “each symbol a constant number of times,” but
between them.</p>

<p><img src="/assets/figures/pscale_exponent.png" alt="fitted exponent" /></p>

<p>I plot $f_c(p)$ on log–log axes and fit the points to a line to obtain an approximate value for the scaling exponent. The left panel places
$f_c(p)$ between the flat fraction-invariant slope and the steeper
$p^{-1}$ coverage slope; the right panel puts the critical count $N_c$ on
$p^{1.42}$, bracketed by coverage ($p^{1}$) and fixed fraction ($p^{2}$).</p>

<p>I test the $p^{3/2}$ scaling by plotting
the onset against the rescaled axis $N/p^{3/2}$ and see that the seven
per-prime curves collapse onto one.</p>

<p><img src="/assets/figures/pscale_p32_collapse.png" alt="onset collapse under N over p to the three-halves" /></p>

<p>The left panel uses the clean $3/2$, the right the fitted $1.42$. Both
stack the seven onset curves into a single step crossing
$P(\mathrm{grok}) = \tfrac12$ in a narrow band near $N/p^{3/2} \approx
2.7$. The $p^{3/2}$ count is about what one expects if the Fourier
solution needs of order $p$ constraints to fix its key frequencies and a
further $\sqrt{p}$-like redundancy to win the optimization race against
memorization, but that requires further study to confirm.</p>

<p>Overall this shows that <strong>the train fraction
is the wrong control variable</strong> for the phase diagram and that reporting
grokking thresholds as fractions — without saying at which $p$ and which
capacity — bakes in a $p^2$ scaling the data do not obviously support. It
is the same move as the finite-size-scaling collapses that fold a family
of rounded transitions onto one universal curve; here the “size” is the
modulus and the control variable is how much of the group table the
model has seen.</p>

<h2 id="discussion">Discussion</h2>

<p>I mapped grokking on modular addition across the (capacity, data-fraction)
plane and recovered the phase structure that has appeared piecemeal across the
literature. The data-fraction axis dominates: below a critical fraction nothing
groks at any width, and above it the memorization delay collapses by more than
two orders of magnitude. Capacity sets the boundary at that critical fraction
and controls how the delay behaves — models with ample capacity race ahead on
memorization and grok late, while models near their capacity limit cannot take
the memorization shortcut cheaply and so generalize almost as soon as they fit.
The four canonical regimes (no grokking, memorize-then-grok, compressed delay,
fast/concurrent) are corners of this single surface, and the competing
explanations — circuit efficiency, critical data size, and the
speed-competition account — each describe a different feature of it rather than
standing as rival mechanisms.</p>

<p>The central result is that the train fraction is the wrong control variable for
the phase boundary. The critical fraction is not a property of the task: it
runs from $f_c = 0.5$ at $p = 31$ down to $f_c = 0.25$ at $p = 113$, scaling as
$f_c(p) \sim p^{-0.58} \approx p^{-1/2}$. Read as a constraint count this is
$N_c \sim p^{3/2}$ — faster than coverage ($\sim p$), slower than a fixed
fraction ($\sim p^2$) — and rescaling the data axis by $N/p^{3/2}$ collapses
the seven per-prime onset curves onto a single step. Reporting a grokking
threshold as a fraction without naming $p$ and the capacity therefore bakes in
a $p^2$ scaling the data do not support; the natural “size” for a
finite-size-scaling treatment is the modulus, and the right control variable is
how much of the group table the model has seen.</p>

<p>I read these results as suggestive rather than settled — the grid is coarse and
the $p$-collapse spans less than a decade in $p$ — but the direction is
consistent across all seven moduli.</p>

<p>A recent information-theoretic account
(<a href="https://arxiv.org/abs/2605.09724">Song &amp; Ye, 2026</a>) argues that
grokking onset is not a static capacity threshold at all but a <em>race</em>
between a memorization timescale and a generalization timescale, both
functions of model size — a refinement that complicates the simple
threshold reading. That race unfolds along the optimization-time axis,
complementary to the data-axis rescaling I study here. The critical data
requirement governs whether the generalizing circuit is reachable at all,
while the competing timescales govern how quickly it is reached. Reading
the onset from late-time accuracy, as I do, folds the two effects together.</p>

<h2 id="reproducibility">Reproducibility</h2>

<p>All experiments run on a single Apple M3 Pro chip — an Arm system-on-chip
with a 12-core CPU (6 performance + 6 efficiency cores) and an 18-core
integrated GPU sharing 18 GB of unified memory, with PyTorch on the Metal
Performance Shaders (MPS) backend rather than CUDA. The training script,
the $75$-run sweep, the committed per-run statistics, and the analysis
that produces every figure are in
<a href="https://github.com/jasteinberg/interp-repo">interp-repo</a>, with the
executed notebook alongside.</p>

<hr />

<p><em><small>Prose edited with the assistance of Claude (Anthropic), which also checked the numbers against the analysis artifacts and ran supplementary experiments. I ran the main experiments and vouch for every claim here.</small></em></p>]]></content><author><name></name></author><summary type="html"><![CDATA[I study the (capacity, data-fraction) plane of grokking on modular addition, and ask which rescaling of the data axis collapses the onset across moduli — a finite-size-scaling look at the right control variable.]]></summary></entry><entry><title type="html">What makes a transformer use both of its layers? Circuit formation in Dyck-(k, m)</title><link href="https://jasteinberg.github.io/blog/2026/dyck-circuits/" rel="alternate" type="text/html" title="What makes a transformer use both of its layers? Circuit formation in Dyck-(k, m)" /><published>2026-06-10T00:00:00+00:00</published><updated>2026-06-10T00:00:00+00:00</updated><id>https://jasteinberg.github.io/blog/2026/dyck-circuits</id><content type="html" xml:base="https://jasteinberg.github.io/blog/2026/dyck-circuits/"><![CDATA[<h2 id="introduction">Introduction</h2>
<p>Mechanistic interpretability has assembled detailed, causal accounts of specialized attention circuits in trained transformers. The cleanest is the induction head first studied by <a href="https://transformer-circuits.pub/2022/in-context-learning-and-induction-heads/index.html">Olsson et al. (2022)</a> in <em>In-context Learning and Induction Heads</em> in the <a href="https://transformer-circuits.pub">Transformer Circuits Thread</a> — a two-layer circuit that completes <code class="language-plaintext highlighter-rouge">[A][B] … [A] → [B]</code> by attending back to the token that followed the previous occurrence of the current token. Induction heads are an unusually good object of study: they are reliably present in small attention-only models, and their appearance during training is marked by a visible phase change in the loss.
Most such accounts describe circuits in models that have already been trained. I wanted instead to watch a circuit form — to follow an identifiable head through training with every head, activation, and checkpoint available to inspect. That resolution rules out natural language: at the scales where language is learnable, a fully instrumented run from scratch is far beyond a modest compute and memory budget. To study circuit formation at all, the task has to be synthetic and small.</p>

<p>The most naive synthetic task is the induction task itself: one generates sequences of repeated random tokens and trains the model on next-token prediction. Unlike in natural language, the model finds a first-layer shortcut and never builds the two-layer composition
that defines a genuine induction circuit. This is because in this task, the only structure to learn is the repetition within a sequence, which can be done in a single layer. In order for more complex heads to develop, the <em>language</em> of the task has to carry more structure than the shortcut can capture, or
the model will simply take the shortcut.</p>

<p>The Dyck-$(k, m)$ language consists of balanced bracket
sequences over $k$ types with nesting depth at most $m$ and is about the
simplest context-free grammar that still has real structure. It separates two computational sub-tasks: <em>depth tracking</em>,
a scalar running count, and <em>type matching</em>, a non-local lookup of the
most recent unmatched OPEN. At $k = 1$ only depth tracking is
non-trivial; at $k \geq 2$ both are required.</p>

<div class="tldr gray">
  <p>In this post I report what
two-layer attention-only transformers actually learn for $k \in {1, 2, 3}$, where $k$ is used to control the structural complexity of the task. For Dyck-1 I find a two-head model reaches the entropy floor of the task without forming any induction-shaped circuitry at all. Layer 1 is inert, and strongly negative layer-0 copying scores reveal a “bracket-flipping” shortcut. This is the same single-layer escape the repeated-token task fell into. For Dyck-2 and Dyck-3 ($k\geq 2$) the <strong>type matching required forces the formation of a two-stage factored circuit</strong>. In this circuit layer 0 constructs a linear representation of bracket depth, demonstrated by the first principal component of the residual-stream correlating with ground-truth depth at $r = 0.75$, and layer 1 performs type-aware matching. Activation patching confirms this factorization causally and identifies the single causally dominant readout head as the head containing the lowest type-match score in its layer but with an OV circuit implementing the OPEN→CLOSE flip the
task needs. <strong>This head would not have been flagged by attention-pattern analysis alone — it required looking at both the QK and OV circuits.</strong> Overall, I found that <strong>width controls factorization, not capability</strong>: a two-head model
eventually solves Dyck-2 too, but via a qualitatively different,
more entangled two-pathway circuit.</p>
</div>

<p><em>Code and executed notebooks:
<a href="https://github.com/jasteinberg/interp-repo">interp-repo</a>.</em></p>

<p><strong>Related work</strong> Dyck languages are a standard testbed for sequence
models. <a href="https://arxiv.org/abs/2010.07515">Hewitt et al. (2020)</a>
introduced the Dyck-$(k, m)$ family and proved that <em>RNNs</em> can generate
it with memory $O(m \log k)$ — the construction and bounded-depth setup
this post borrows. For transformers,
<a href="https://arxiv.org/abs/2105.11115">Yao et al. (2021)</a> give a two-layer
self-attention construction for bounded Dyck, with essentially the
depth-plus-type-matching decomposition observed empirically here;
<a href="https://arxiv.org/abs/2010.04303">Ebrahimi et al. (2020)</a> study how
self-attention processes Dyck-$n$;
<a href="https://arxiv.org/abs/2210.10749">Liu et al. (2022)</a> characterize
shortcut solutions to automata-like tasks; and
<a href="https://arxiv.org/abs/2312.01429">Wen et al. (2023)</a> demonstrate
solution multiplicity and the unreliability of attention-pattern
analysis on bounded Dyck — conclusions the patching results and width
comparison below independently corroborate at small scale. These
experiments were run independently as a from-scratch exercise, and I
have not systematically audited the overlap with prior results; they are
presented as a self-contained empirical study rather than a claim of
novelty.</p>

<h2 id="setup">Setup</h2>

<p>Following <a href="https://arxiv.org/abs/2105.11115">Yao et al. (2021)</a>, I fix a
vocabulary of $k$ bracket types,
$\Sigma = \bigcup_{i \in [k]} {\,\langle_i,\ \rangle_i\,}$, where
$\langle_i$ and $\rangle_i$ are the OPEN and CLOSE brackets of type $i$.
$\mathsf{Dyck}_{k}$ is defined as the language of well-nested bracket strings,
generated from the start symbol $X$ by the context-free grammar</p>

\[X \;\to\; \varepsilon \;\;\big|\;\; \langle_i\, X\, \rangle_i\, X
\qquad (i \in [k]),\]

<p>with $\varepsilon$ the empty string. The single recursive production
$\langle_i X \rangle_i X$ is what forces every CLOSE to match the most
recent unclosed OPEN <em>of the same type</em> — the defining constraint of the
language. Reading a string left to right, the unmatched OPENs form a
stack: an OPEN pushes its type, a CLOSE pops and must agree with the top.
The <strong>depth</strong> at position $i$ is the stack height, equivalently the
running excess of OPENs over CLOSEs,</p>

\[d(w_{1:i}) = \mathsf{count}(w_{1:i}, \langle) - \mathsf{count}(w_{1:i}, \rangle).\]

<p>The bounded language $\mathsf{Dyck}_{k,m}$ restricts to strings whose
depth never exceeds $m$,</p>

\[\mathsf{Dyck}_{k,m} = \Big\{\, w_{1:n} \in \mathsf{Dyck}_k \;\Big|\;
\max_{i \in [n]} d(w_{1:i}) \le m \,\Big\},\]

<p>so that a stack of size $m$ suffices to process any string — the
bounded-memory analogue of natural-language center-embedding depth, and
the property that makes the language learnable by a fixed-width network.
For example, in $\mathsf{Dyck}_{2,4}$ the strings <code class="language-plaintext highlighter-rouge">([])</code> and <code class="language-plaintext highlighter-rouge">(()[()])</code>
are valid, while <code class="language-plaintext highlighter-rouge">([)]</code> violates type matching and <code class="language-plaintext highlighter-rouge">)(</code> violates
non-negativity of the depth.</p>

<p><strong>Data generation</strong> I sample sequences left-to-right. At current
depth $d$: if $d = 0$ the generator must OPEN (type uniform over
${1,\ldots,k}$); if $d = m$ it must CLOSE (type fixed by the stack);
at intermediate depths it opens with probability $p_o$ and otherwise
closes. Throughout I use $p_o = 0.5$, $m = 8$, and context length
$n_{\mathrm{ctx}} = 64$. The vocabulary is 
\(\{\mathrm{BOS}, \mathrm{PAD}\} \cup
\{\mathrm{OPEN}_i, \mathrm{CLOSE}_i\}_{i=1}^{k}\)
— six tokens for
Dyck-2, eight for Dyck-3.</p>

<p><strong>The entropy floor</strong> Because intermediate-depth choices are genuine
coin flips, next-token cross-entropy has an irreducible lower bound
equal to the entropy of the generating distribution,</p>

\[H = -\sum_v p_v \log p_v\]

<p>averaged over positions. The per-position entropy depends only on the current depth: at depth $0$ the generator must OPEN, at depth $m$ it must CLOSE, and in between it opens or closes with probability $p_o = 0.5$. Sampling the generator at my parameters ($m = 8$, $n_{\mathrm{ctx}} = 64$, sequences started at depth $0$), the positions split as $f_0 = 0.103$ at depth $0$, $f_m = 0.041$ at depth $m$, and $f_{\mathrm{int}} = 0.856$ intermediate — so $14.4\%$ are forced. For Dyck-1 the forced positions carry no entropy (a single bracket type) and each intermediate position carries $\log 2$, giving</p>

\[H_1 = f_{\mathrm{int}} \log 2 \approx 0.595 \text{ nats}.\]

<p>For $k \geq 2$ the bracket <em>type</em> adds entropy wherever the generator emits an OPEN with a free choice of type, and there are two such places: an intermediate position opens with probability $p_o$, contributing $p_o \log k = \tfrac{1}{2}\log k$, and a forced depth-$0$ position always opens, contributing the full $\log k$. (A CLOSE never adds type entropy — its type is pinned by the stack.) The excess over Dyck-1 is therefore</p>

\[H_k - H_1 = \big(\tfrac{1}{2} f_{\mathrm{int}} + f_0\big)\log k = 0.531\,\log k,\]

<p>i.e. predicted excesses of $0.368$ and $0.583$ nats for $k = 2, 3$, against the empirical $0.366$ and $0.581$ (floors $0.961$ and $1.176$ over Dyck-1’s $0.595$). The forced depth-$0$ OPENs are what lift the coefficient above the naive $\tfrac{1}{2}$ — which counts only the intermediate term and gives $0.347$ and $0.549$, low by ~5%. At stationarity $f_0 = f_m$ and the coefficient is exactly $\tfrac{1}{2}$; the small excess is a finite-length effect, sequences starting at depth $0$ and so spending slightly more time on the floor than the ceiling ($f_0 &gt; f_m$). All results below report the <strong>gap</strong>
$\mathcal{L}_{\mathrm{eval}} - H$ rather than raw loss: absolute loss is
uninterpretable without the floor, and the gap is the quantity that
decides whether a model has solved the task.</p>

<p>Depth tracking is a scalar derived from the running sequence; type matching is a per-type pointer
lookup requiring information routed from past positions. The two
sub-tasks are exactly the kind of pair that a layered architecture can
factor — one builds shared state, the next reads from it — but nothing
forces the model to factor them this way rather than find a shortcut.
Whether it does is the empirical question, and the next section makes
the architectural stakes precise.</p>

<h3 id="the-model">The model</h3>
<p>For this task I train 2-layer transformers with either two or three attention heads per layer.
I choose these architectures based on the computations the task requires. Validating a CLOSE
requires information from an earlier position — the matching OPEN — to be routed forward to the prediction point. This composition requires the architecture to contain two layers to complete the task. A single attention layer can
move information from one position to another exactly once when each head
reads the residual stream, attends, and writes back. However, every head in
the layer reads the <em>same</em> input, so they cannot build on one another
within the layer. Type-aware matching needs two routing steps that
depend on each other. First it must work out which bracket is open (a function
of the running history), then use that to attend to the correct prior
OPEN. The second step then requires the <em>output</em> of the first as its
input. A second layer provides heads that read from a residual stream that already contains the writes of layer 0 so that layer 1 can 
attend on the basis of what layer 0 computed. This is exactly the
structure of an induction circuit (a previous-token head in layer 0
feeding an induction head in layer 1), and it is why the canonical
induction circuit is a <em>two-layer</em> object. Depth buys <em>composition</em>;
width does not.</p>

<p>Reading off the sub-tasks, I expected on the order of three specialized
heads: a head that tracks nesting depth, a head that does the
type-aware lookup, and — by analogy to induction — a head providing something
previous-token-like or copy-like at the readout. Therefore I train models with two or
three heads per layer. This is enough to host the expected roles and to let
distinct functions land on distinct heads, making the
circuit legible, but few enough that the model cannot hide the
computation across a large redundant population. The two-versus-three
comparison is itself informative: it asks whether the clean
one-role-per-head factorization is forced by the task or is a luxury of
having a spare head.</p>

<p><strong>The network</strong> All experiments use a two-layer attention-only
transformer in the residual-stream form (TransformerLens,
<code class="language-plaintext highlighter-rouge">attn_only=True</code>, no MLPs, no layer-norm folded into the analysis). A
length-$n$ token sequence $w_{1:n}$ is embedded and given learned
positional encodings,</p>

\[x^{(0)}_t = W_E\, \mathrm{onehot}(w_t) + p_t \;\in\; \mathbb{R}^{d_{\mathrm{model}}}\]

<p>and each layer $\ell \in {0, 1}$ adds the concatenated output of its
$H$ causal attention heads back into the residual stream,</p>

\[\begin{aligned}
x^{(\ell+1)}_t &amp;= x^{(\ell)}_t + \sum_{h=1}^{H} \mathrm{head}^{(\ell,h)}_t \\
\mathrm{head}^{(\ell,h)}_t &amp;= \sum_{s \le t} A^{(\ell,h)}_{t,s}\,
\big(x^{(\ell)}_s W_V^{(\ell,h)}\big) W_O^{(\ell,h)}
\end{aligned}\]

<p>where the attention weights</p>

\[A^{(\ell,h)}_{t,\cdot} =
\mathrm{softmax}\big(x^{(\ell)}_t W_Q^{(\ell,h)} (x^{(\ell)} W_K^{(\ell,h)})^\top / \sqrt{d_{\mathrm{head}}}\big)\]

<p>are causal ($A^{(\ell,h)}_{t,s} = 0$ for $s &gt; t$). Logits come from the
model’s own unembedding with no separate classifier head,</p>

\[\mathrm{logits}_t = x^{(2)}_t W_U,\]

<p>and training is pure next-token cross-entropy at every position. This is
deliberately the minimal architecture in which the two-step composition
of §Setup <em>can</em> occur — depth in layer 0, type-matching read in layer 1
— so that whatever structure appears is forced by the task, not by
MLPs or extra depth. Two width configurations are used:
<strong>2-head</strong> ($d_{\mathrm{model}} = 64$, $d_{\mathrm{head}} = 32$) and
<strong>3-head</strong> ($d_{\mathrm{model}} = 96$, $d_{\mathrm{head}} = 32$).
I train the main two- and three-head models for $5{,}000$ steps with a batch size of $256$,
AdamW with learning rate $10^{-3}$, and weight decay $0.01$.
For the head-count and depth sweeps in “Closing the causal
chain” I train for $8{,}000$ steps, and for the two-head Dyck-2 convergence run I train for $20{,}000$ steps, with all other training parameters kept fixed.</p>

<p>The schematic below fixes notation for the analysis: the residual stream
runs left to right, each attention layer reads it and writes its heads’
outputs back, and the two computations the task needs — building the
depth representation, then reading it to match types — are expected to
land in layer 0 and layer 1 respectively.</p>

<p><img src="/assets/figures/dyck_schematic.svg" alt="Two-layer attention-only transformer: embed, layer 0 builds depth, layer 1 type-matches by reading layer 0's depth, unembed to logits" /></p>

<p>Five per-head diagnostics are computed on held-out data. Let
$A^{(\ell,h)}_{t,s}$ denote the attention weight from query position $t$
to key position $s$ for head $h$ in layer $\ell$, and write $\mathcal{C}$
for the set of CLOSE positions in a sequence. Three of the diagnostics
probe the <strong>QK circuit</strong> (where a head attends), one probes the <strong>OV
circuit</strong> (what it writes toward the unembedding), and one probes the
<strong>residual representation</strong> between layers. These are defined as follows:</p>

<p><em>Bracket-match score.</em> For each CLOSE at position $t$, let $\mu(t)$ be
the position of its unique matching OPEN (the most recent unmatched OPEN,
recovered by a stack scan). The score is the mean attention a head places
on the matched OPEN,</p>

\[\mathrm{bm}(\ell,h) = \big\langle\, A^{(\ell,h)}_{t,\,\mu(t)} \,\big\rangle_{t \in \mathcal{C}}.\]

<p><em>Type-match score.</em> For $k \geq 2$, a head can attend to a same-type OPEN
without attending to the <em>correct</em> one. Writing $\tau(t)$ for the bracket
type of the token at $t$, the type-match score sums attention onto all
prior OPENs of the matching type,</p>

\[\mathrm{tm}(\ell,h) = \Big\langle \textstyle\sum_{s &lt; t}\, A^{(\ell,h)}_{t,s}\,
\mathbf{1}\!\left[\,s \in \mathrm{OPEN},\ \tau(s) = \tau(t)\,\right] \Big\rangle_{t \in \mathcal{C}}.\]

<p>A head that does true matching scores high on both $\mathrm{bm}$ and
$\mathrm{tm}$; a head that merely routes by type but not position scores
high on $\mathrm{tm}$ alone.</p>

<p><em>Previous-token score.</em> The mean attention from each position to the one
immediately before it,</p>

\[\mathrm{pt}(\ell,h) = \big\langle\, A^{(\ell,h)}_{t,\,t-1} \,\big\rangle_{t}.\]

<p><em>Copying score (OV circuit).</em> Independently of where a head attends, does
its OV circuit map a bracket token to <em>itself</em> or to the <em>opposite</em>
bracket in logit space? Form the OV-to-logit map
$M^{(\ell,h)} = W_E\, W_V^{(\ell,h)}\, W_O^{(\ell,h)}\, W_U$, and for each
bracket type $i \in {1,\dots,k}$ take a diagonal-minus-off-diagonal
contrast on that type’s $(\mathrm{O}_i, \mathrm{C}_i)$ pair, then average
over types,</p>

\[\mathrm{cp}(\ell,h) = \frac{1}{k}\sum_{i=1}^{k}
\Big[\tfrac12\big(M_{\mathrm{O}_i,\mathrm{O}_i} + M_{\mathrm{C}_i,\mathrm{C}_i}\big)
- \tfrac12\big(M_{\mathrm{O}_i,\mathrm{C}_i} + M_{\mathrm{C}_i,\mathrm{O}_i}\big)\Big].\]

<p>A positive value is an identity-copying OV circuit (the induction
signature: promote the attended token); a <em>negative</em> value is a
<strong>flip</strong> circuit (promote the opposite bracket), which is the
task-correct readout for bracket prediction — a CLOSE should be predicted
after attending to the matching OPEN. Because the readout’s job is the
flip, the diagnostic readout heads in the results below carry strongly
<em>negative</em> copying scores.</p>

<p><em>Depth correlation (residual representation).</em> Let $z^{(\ell)}_t$ be the
residual stream after layer $\ell$ at position $t$, and $\mathrm{PC}_1$
its leading principal component over all positions. The score is the
absolute Pearson correlation between the projection and the ground-truth
stack depth $d_t$,</p>

\[\rho(\ell) = \big|\,\mathrm{corr}\big(\langle z^{(\ell)}_t, \mathrm{PC}_1\rangle,\ d_t\big)\,\big|.\]

<p>The first four are correlational — they say where a head looks and what
its OV writes, not whether that computation is <em>used</em>. The patching
analysis below supplies the interventional complement.</p>

<h2 id="results">Results</h2>

<h3 id="dyck-1-a-bracket-flipping-shortcut">Dyck-1: a bracket-flipping shortcut</h3>
<p>For the Dyck-1 task, the two-head model reaches loss $0.5941$ against a floor of $0.5948$, with a gap of $-0.0007$ within finite-batch noise. It exhibits no
induction-shaped circuitry at all. Every bracket_match score is below
$0.05$ (the uniform-attention background), every prev_token score below
$0.04$, and the depth correlation is $0.029$ (layer 0) and $0.044$
(layer 1) — essentially zero. Layer 1 is inert on every measure. The
telling signature is the layer-0 copying scores, which are strongly
<em>negative</em> ($-1.06$ and $-0.74$): the OV circuits project
attended-OPEN representations onto -CLOSE directions. This means that the model has
learned a <strong>bracket-flipping shortcut</strong> — attend diffusely to recent
tokens and emit the opposite bracket type — rather than anything
resembling matching. A three-head variant shows the same qualitative
pattern but had not fully converged at 5000 steps, gap $+0.033$;
overparameterization appears to slow optimization here, not change the
solution.</p>

<p>Why does this suffice? Although the generating distribution depends on
depth at the boundaries, only 14.4% of positions are forced; the rest
are coin flips. A model can sit on the entropy floor by getting the
boundary positions right, and a position-conditioned heuristic — long
runs of recent OPENs imply proximity to the depth ceiling, and
conversely — achieves this without tracking depth. The conclusion is
not that the conjecture fails but that <strong>Dyck-1 next-token prediction
is not a structural-complexity test</strong>: it is a depth-bounded random
walk, and the model exploits the randomness. The refined requirement on
a test language is that it demand information flow that cannot be
implemented within a single layer. Dyck-1’s depth bound does not meet
this bar; type matching at $k \geq 2$ does.</p>

<h3 id="dyck-2-and-dyck-3-a-factored-two-stage-circuit">Dyck-2 and Dyck-3: a factored two-stage circuit</h3>

<p>The three-head model solves Dyck-2 to gap $+0.0003$ and shows a
qualitatively different circuit:</p>

<table>
  <thead>
    <tr>
      <th>measure</th>
      <th>L0.H0</th>
      <th>L0.H1</th>
      <th>L0.H2</th>
      <th>L1.H0</th>
      <th>L1.H1</th>
      <th>L1.H2</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td>bracket_match</td>
      <td>0.016</td>
      <td>0.079</td>
      <td>0.045</td>
      <td>0.070</td>
      <td>0.055</td>
      <td>0.001</td>
    </tr>
    <tr>
      <td><strong>type_match</strong></td>
      <td>0.27</td>
      <td>0.31</td>
      <td>0.28</td>
      <td><strong>0.49</strong></td>
      <td><strong>0.49</strong></td>
      <td>0.37</td>
    </tr>
    <tr>
      <td>prev_token</td>
      <td>0.017</td>
      <td>0.085</td>
      <td>0.051</td>
      <td>0.045</td>
      <td>0.033</td>
      <td>0.009</td>
    </tr>
    <tr>
      <td>copying</td>
      <td>−0.09</td>
      <td>−0.10</td>
      <td>−0.07</td>
      <td>0.02</td>
      <td>0.07</td>
      <td><strong>−0.91</strong></td>
    </tr>
  </tbody>
</table>

<p>The depth correlation is $0.747$ after layer 0 and $0.060$ after
layer 1. This implies that <strong>layer 0 builds the depth representation</strong> — the
leading residual-stream PC tracks running depth — while its heads attend
to same-type OPENs only diffusely (type_match $\approx 0.28$).
<strong>Layer 1 performs type-aware matching</strong>: two heads place nearly half
their CLOSE-position attention on prior same-type OPENs, and the third
(L1.H2) carries a strongly negative copying score, marking it as the
candidate “predict the CLOSE of the matched type” readout.</p>

<p>Dyck-3 confirms the pattern at gap $+0.0034$ against the floor of
$1.176$ nats. It has depth correlation $0.656$ in layer 0 and $0.060$ in
layer 1, layer-1 type_match scores of $0.28$, $0.23$, $0.20$, and the
negative-copying readout signature distributed across the layer-1 heads
(copying scores $-0.42$, $-0.15$, $-0.35$), with the flip concentrated
on L1.H0. The per-head type_match decline from Dyck-2’s $0.49$ to
Dyck-3’s $0.28$ is consistent with a fixed attention budget spread over
more bracket types; a model with $n_{\mathrm{heads}} \approx k$ in
layer 1, allowing one head per type, is a natural follow-up I have not
run.</p>

<p>While the canonical Olsson induction circuit
copies the attended token, predicting a <em>positive</em> copying score; my
layer-1 readouts are strongly negative. This is the correct OV for the
task: at a position whose next token is $\mathrm{CLOSE}_i$, the head
attends to the matching $\mathrm{OPEN}_i$ and must emit
$\mathrm{CLOSE}_i$ — <strong>lookup-then-flip</strong> rather than lookup-then-copy.
The circuit is structurally analogous to induction (look back, route
forward) with a task-appropriate OV transformation.</p>

<h3 id="causal-validation-by-activation-patching">Causal validation by activation patching</h3>

<p>To obtain the causal picture beyond correlational diagnostics, I use activation patching. This involves running the model on a clean prompt and a
corrupted twin that demands a different answer, patching clean activations
into the corrupted run one location at a time, and measuring how much of
the clean prediction is recovered. I created clean and corrupted prompts of
length 10 that differ in exactly one OPEN type at position 1, with
intervening matched filler pairs leaving that OPEN unmatched at the
prediction position: <code class="language-plaintext highlighter-rouge">^[()()()()</code> requires <code class="language-plaintext highlighter-rouge">]</code> next, while
<code class="language-plaintext highlighter-rouge">^(()()()()</code> requires <code class="language-plaintext highlighter-rouge">)</code>. Reference logit differences are
$\pm 7.32$, averaged over 32 paired examples on the three-head Dyck-2
model.</p>

<p>The table below demonstrates that residual-stream patching localizes the computation sharply:</p>

<table>
  <thead>
    <tr>
      <th> </th>
      <th>pos 1 (swap)</th>
      <th>pos 9 (pred)</th>
      <th>other positions</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td>L0 patch</td>
      <td><strong>100%</strong></td>
      <td>0%</td>
      <td>0%</td>
    </tr>
    <tr>
      <td>L1 patch</td>
      <td>20%</td>
      <td><strong>85%</strong></td>
      <td>0%</td>
    </tr>
  </tbody>
</table>

<p>Recovery is essentially binary: type information enters the residual
stream at (layer 0, swap position), arrives at (layer 1, prediction
position), and is invisible everywhere else — precisely the
build-state-then-route pattern the per-head measures suggested.</p>

<p>I did per-head patching at the prediction position to determine <em>which</em> layer-1
head carries the effect; results are in the table below:</p>

<table>
  <thead>
    <tr>
      <th>Head</th>
      <th>recovery</th>
      <th>type_match</th>
      <th>copying</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td>L0.H0</td>
      <td>+79%</td>
      <td>0.27</td>
      <td>−0.09</td>
    </tr>
    <tr>
      <td>L0.H1</td>
      <td>−2%</td>
      <td>0.31</td>
      <td>−0.10</td>
    </tr>
    <tr>
      <td>L0.H2</td>
      <td>+3%</td>
      <td>0.28</td>
      <td>−0.07</td>
    </tr>
    <tr>
      <td>L1.H0</td>
      <td>−1%</td>
      <td>0.49</td>
      <td>+0.02</td>
    </tr>
    <tr>
      <td>L1.H1</td>
      <td>−2%</td>
      <td>0.49</td>
      <td>+0.07</td>
    </tr>
    <tr>
      <td><strong>L1.H2</strong></td>
      <td><strong>+102%</strong></td>
      <td>0.37</td>
      <td><strong>−0.91</strong></td>
    </tr>
  </tbody>
</table>

<p>This data shows that L1.H2 alone fully recovers the clean prediction — and it is the head
with the <em>lowest</em> type_match in its layer ($0.37$ against $0.49$) and
the most negative copying score. The reconciliation is that attention
statistics measure where a head looks, while the copying score measures
what its OV circuit writes toward the unembedding. The functioning
readout needs both adequate attention to the matched OPEN <em>and</em> the
flip-implementing OV, and either alone is insufficient. L1.H0 and L1.H1
attend to the matched OPEN strongly but their OV does not convert that
attention into the correct CLOSE logit so they are redundant pathways, auxiliary
signals, or computations the model does not ultimately use. Pure
attention-pattern analysis would have nominated exactly the wrong
heads, patching is what surfaces the discrepancy, consistent with <a href="https://arxiv.org/abs/2312.01429">Wen et al. (2023)</a> warnings about myopic interpretation of attention on
Dyck. Patching at the swap position completes the picture: L0.H0 and
L0.H1 recover 28% and 43% respectively there, L0.H2 only 2%, and all
layer-1 heads recover exactly 0%. Layer 1 attends only backward, so
its output at position 1 cannot influence position 9. L0.H0’s 79%
recovery at the prediction position indicates a parallel layer-0
pathway carrying type information directly via its head output rather
than through the residual PC.</p>

<h3 id="width-controls-factorization-not-capability">Width controls factorization, not capability</h3>

<p>Earlier I asked whether the clean one-role-per-head factorization is
forced by the task or a luxury of having a spare head. The answer — it
is a luxury; width controls how cleanly the circuit factorizes, not
whether the task is solved — is one of the findings below. Two heads
in <em>one</em> layer are not a substitute for two layers: they add parallel
capacity at a single routing step, not a second, dependent step.</p>

<p>Trained on Dyck-2 for the standard 5,000 steps, the two-head model
reaches a gap of only $+0.0047$ — it has not converged to the floor.
Trained 4× longer (20,000 steps) it does converge, to a gap of
$+0.0010$, but to a qualitatively different circuit:</p>

<table>
  <thead>
    <tr>
      <th> </th>
      <th>2-head, 5k</th>
      <th>2-head, 20k</th>
      <th>3-head, 5k</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td>gap</td>
      <td>+0.0047</td>
      <td><strong>+0.0010</strong></td>
      <td>+0.0003</td>
    </tr>
    <tr>
      <td>L0 depth correlation</td>
      <td>0.611</td>
      <td><strong>0.385</strong></td>
      <td>0.747</td>
    </tr>
    <tr>
      <td>L1.H0 copying</td>
      <td>−0.67</td>
      <td>−0.26</td>
      <td>+0.02</td>
    </tr>
    <tr>
      <td>L1.H1 copying</td>
      <td>−0.20</td>
      <td><strong>−1.32</strong></td>
      <td>+0.07</td>
    </tr>
    <tr>
      <td>max L1 type_match</td>
      <td>0.35</td>
      <td>0.34</td>
      <td>0.49</td>
    </tr>
  </tbody>
</table>

<p>The table above shows that the converged two-head model concentrates the flip readout on a single
head (L1.H1 copying $-1.32$, against $-0.26$ for its partner), much as
the three-head model does, but it pays for having one fewer layer-0
head by having a markedly weaker depth representation ($0.385$ versus
$0.747$), with depth tracking entangled into the same heads that do type
work rather than occupying a separable principal component. The cleanly
factored two-stage circuit — a clean depth PC in layer 0, a clean single
flip readout in layer 1 — is therefore a property of sufficient width,
not an inevitability of the task. This
echoes the broader observation that larger models often exhibit
cleaner, more localized and interpretable circuits because they can afford specialized
heads independent of whether these additional heads are required to solve the task.</p>

<h2 id="discussion">Discussion</h2>

<p>The table below assembles the experiments, which share an architecture and differ only
in data structure:</p>

<table>
  <thead>
    <tr>
      <th>experiment</th>
      <th>gap</th>
      <th>L0 depth corr.</th>
      <th>max L1 type_match</th>
      <th>layer 1</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td>synthetic induction¹</td>
      <td>+0</td>
      <td>n/a</td>
      <td>n/a</td>
      <td>dead</td>
    </tr>
    <tr>
      <td>Dyck-1</td>
      <td>−0.0007</td>
      <td>0.029</td>
      <td>n/a</td>
      <td>dead</td>
    </tr>
    <tr>
      <td>Dyck-2</td>
      <td>+0.0003</td>
      <td><strong>0.747</strong></td>
      <td><strong>0.49</strong></td>
      <td>active</td>
    </tr>
    <tr>
      <td>Dyck-3</td>
      <td>+0.0034</td>
      <td>0.656</td>
      <td>0.28</td>
      <td>active</td>
    </tr>
  </tbody>
</table>

<p>¹ Deterministic given its first half, hence no entropy floor.</p>

<p>The data supports the conclusion that tasks
solvable by single-layer shortcuts are solved by single-layer shortcuts
(with layer 1 left dead), and the two-stage circuit appears exactly
when the language demands it.  Layer separation occurs when the language requires information flow that <strong>cannot be implemented in a single layer</strong>. Type matching at $k \geq 2$ is that
minimal requirement in the Dyck-$(k, m)$ family.</p>

<p>The methodological lesson generalizes beyond bracket languages.
Correlational head diagnostics — attention patterns, OV statistics —
are necessary but not sufficient: they enumerate the circuits that
<em>might</em> be present, but in the patching analysis of my models they nominated the wrong readout heads
while assigning an unremarkable profile to the causally dominant one.
Interventional methods (patching, ablation) are what distinguished the
candidate circuitry from functioning circuitry.</p>

<h2 id="closing-the-causal-chain-four-follow-up-experiments">Closing the causal chain: four follow-up experiments</h2>

<p>The core results above leave four specific questions open, each of which
can be settled with a targeted experiment on the converged models. Below I describe each question and the corresponding experiment I ran. The results of these experiments confirm the factored-circuit picture while sharpening it
in two places.</p>

<p><strong>1. Does layer 1 actually consume layer 0’s depth representation?</strong> The
patching results show type information flowing layer-0 → layer-1, but
not that layer-1’s <em>matching</em> specifically reads the depth PC. To test
this directly, I ablated a single direction of layer-0’s output — the
depth-PC1 direction — and measured the effect on layer-1 bracket-match
attention, against a control that ablates a random orthogonal direction
of equal norm. The depth ablation degrades layer-1 matching
($\mathrm{bm}$ on the two matching heads falls $0.068 \to 0.064$ and
$0.055 \to 0.042$), while the random-direction control leaves it
essentially untouched ($0.068 \to 0.068$, $0.055 \to 0.055$). The effect
is modest but specific to the depth direction: <strong>layer 1’s matching
heads do read layer 0’s depth representation</strong>, not merely some generic
feature of its output.</p>

<p><strong>2. At what width does the clean factorization set in?</strong> The two-versus
three-head comparison suggested width controls factorization; sweeping
$n_{\mathrm{heads}} \in {2,3,4,5,6}$ on Dyck-2 (three seeds each,
$d_{\mathrm{head}}$ fixed at $32$, all trained to $8{,}000$ steps so the
slower-converging widths are compared at a matched budget) shows the
relationship is not monotone. Measuring how concentrated the flip readout is on a single
head — the ratio of the largest $|\mathrm{cp}|$ to the sum across
layer-1 heads — gives:</p>

<table>
  <thead>
    <tr>
      <th>$n_{\mathrm{heads}}$</th>
      <th>2</th>
      <th>3</th>
      <th>4</th>
      <th>5</th>
      <th>6</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td>readout concentration</td>
      <td>0.66</td>
      <td><strong>0.80</strong></td>
      <td>0.56</td>
      <td>0.58</td>
      <td>0.55</td>
    </tr>
    <tr>
      <td>L0 depth correlation</td>
      <td>0.70</td>
      <td>0.65</td>
      <td>0.54</td>
      <td>0.40</td>
      <td>0.43</td>
    </tr>
    <tr>
      <td>seeds grokked</td>
      <td>3/3</td>
      <td>3/3</td>
      <td>3/3</td>
      <td>2/3</td>
      <td>3/3</td>
    </tr>
  </tbody>
</table>

<p>The cleanest single-head readout occurs at <strong>three heads</strong>, and
<em>degrades</em> on both sides — adding heads beyond three spreads the flip
across the layer rather than concentrating it, and also dilutes the
layer-0 depth signal across more heads (the correlation with a single PC
falls from $0.70$ to $0.43$). So the clean factorization is a sweet spot
matched to the task’s natural number of roles, not the top of a
monotone trend; “more width gives cleaner circuits” holds only up to the
point where width starts hosting redundant copies. The five-head models
converged less reliably — one seed of three failed to reach the floor —
but with only three seeds I draw no strong conclusion about why.</p>

<p><strong>3. Does the linear depth code survive deeper nesting?</strong> At $m = 8$ the
depth PC tracks ground-truth depth at $\rho = 0.74$. Retraining at
$m = 16$ and $m = 32$ (with context lengths scaled to $6m$) shows the
linear representation degrading as nesting deepens:</p>

<table>
  <thead>
    <tr>
      <th>$m$</th>
      <th>8</th>
      <th>16</th>
      <th>32</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td>L0 depth correlation $\rho$</td>
      <td>0.74</td>
      <td>0.57</td>
      <td>0.56</td>
    </tr>
    <tr>
      <td>gap</td>
      <td>$-0.001$</td>
      <td>$+0.001$</td>
      <td>$+0.005$</td>
    </tr>
  </tbody>
</table>

<p>The model keeps solving the task (the gap stays near zero), but it does
so with a steadily less linear depth representation — consistent with a
drift toward a compressed, logarithmic, or counter-style code once a
single linear coordinate can no longer cover the dynamic range, as one
would expect on information-packing grounds.</p>

<p><strong>4. Do the matching heads locate their target by position or by
content?</strong> Zeroing the positional embeddings at inference does <em>not</em>
collapse layer-1 matching: bracket-match attention actually rises
($0.068 \to 0.102$, $0.055 \to 0.113$) and type-match falls only
modestly (e.g. $0.50 \to 0.36$ on one head). A purely positional matcher
would break when positions are removed; instead matching largely
survives, so <strong>the matching is substantially content-based</strong> — driven by
the type and depth content routed through the residual stream rather
than by absolute position. The signal is mixed rather than clean (the
modest type-match drop indicates position carries <em>some</em> of the
matching), and the increase in raw bracket-match attention under
position ablation is itself a caution against over-reading attention
magnitude, but the dominant story is content-driven matching.</p>

<h2 id="limitations">Limitations</h2>

<p>Two caveats survive these follow-ups. The depth-saturation and
factorization sweeps are small (three seeds, a single $k$), enough to
show the qualitative trends above but not to pin the exponents or
thresholds precisely; the five-head reliability dip in particular wants
more seeds before it means anything. The related work in §1 implies
that several of these observations plausibly reproduce known results, and a systematic comparison against the
constructions of <a href="https://arxiv.org/abs/2105.11115">Yao et al. (2021)</a> and the multiplicity results of <a href="https://arxiv.org/abs/2312.01429">Wen et al. (2023)</a> is owed before any novelty claim. The results in this post are presented as a self-contained empirical study, not a claim of priority.</p>

<h2 id="reproducibility">Reproducibility</h2>

<p>All experiments run on a single Apple M3 Pro chip — an Arm system-on-chip with a 12-core CPU (6 performance + 6 efficiency cores) and an 18-core integrated GPU sharing 18 GB of unified memory, with PyTorch on the Metal Performance Shaders (MPS) backend rather than CUDA. Training scripts,
per-head diagnostic implementations, and patching utilities are in
<a href="https://github.com/jasteinberg/interp-repo">interp-repo</a>; the executed analysis notebooks
are <code class="language-plaintext highlighter-rouge">notebooks/dyck/explore.ipynb</code> (Dyck-1),
<code class="language-plaintext highlighter-rouge">notebooks/kdyck/explore.ipynb</code> (Dyck-2/3).</p>

<hr />

<p><em><small>Prose edited with the assistance of Claude (Anthropic), which also checked the numbers against the analysis artifacts and ran supplementary experiments. I ran the main experiments and vouch for every claim here.</small></em></p>]]></content><author><name></name></author><summary type="html"><![CDATA[I study two-layer transformers learning factored circuits versus single-layer shortcuts and use activation patching to find the signature of the readout head not present in the attention patterns.]]></summary></entry></feed>