Preprint · 2026

Belief-Trajectory Energy

Measuring the Path to a Prediction

Learned models as measuring instruments: BTE characterizes an input by the belief revisions it induces across a language model's layers, so the model's own forward pass becomes a measurement of the data it reads. Four settings from the paper follow, each with an interactive case study.

01What BTE measures

When a language model is used, what gets read is the endpoint of its forward pass: the answer, a judge's verdict (Zheng et al., 2023), or the likelihood it assigns to the text (Mitchell et al., 2023). The endpoint is one function of everything the model computed on the way, and how the model got there is thrown away with it: every revision along the way.

Neuroscience stopped reading only the endpoint long ago. The same choice can be reached by different neural paths, and the path shows a change of mind that the choice hides (Resulaj et al., 2009; Kiani et al., 2014). Harder stimuli take a longer, more gradual path to the same answer (Gold & Shadlen, 2007; Kar et al., 2019). Familiar stimuli evoke a smaller response than novel ones while behavior stays the same (Grill-Spector et al., 2006). Perception itself is not one feed-forward pass but a sweep followed by recurrent refinement (Lamme & Roelfsema, 2000). Difficulty, familiarity and the route are in the trajectory, not in the output.

BTE reads the model the way neuroscience reads the brain: along the path. A residual network refines one representation block by block, the counterpart in depth of refinement over time (Greff et al., 2017; Jastrzębski et al., 2018), and the state after each block can be read as the model's belief at that depth. From how that belief moves, the paper reads difficulty (02), who wrote a review (03, 04) and what the scorer has seen before (05).

The reading is also made in the model's own terms. Difficulty is a relation between an input and whoever processes it, and neuroscience measures it on the subject, as reaction time, pupil dilation or the confidence carried by neurons (Kahneman & Beatty, 1966; Kiani & Shadlen, 2009), not as an external label. BTE measures it on the scorer: no judge and no reference answer, only the scorer's own beliefs and how far they moved on the way to the answer.

Cortex · one percept, two paths over time ventral visual stream · areas and timing recorded in macaque monkeys frontalparietaloccipitaltemporal easy image hard image V1 V2 V4 IT percept easy · about 100 ms hard · after recurrent passes feed-forward sweep · the easy image is recognized at once the hard image needs recurrent passes same percept, different paths t = 0100 mstime →
Two images of the same object. The easy one is recognized by the first feed-forward sweep through V1 → V2 → V4 → IT, within about 100 ms; the hard one needs recurrent passes before the same percept settles. The output is the same for both; the path through the cortex is not.
Transformer · one answer token, two paths across depth Llama-3.1-8B-Instruct · each row is the whole belief after that layer, split by token easy · the capital of France? revision Paris Paris Paris Paris Paris Paris hard · 9.11 or 9.9, which is larger? revision 9 9 9 layer 32 layer 30 layer 28 layer 26 layer 24 layer 22 layer 20 layer 18 layer 16 Paris9other top-5 tokensevery other token up to layer 20 both beliefs are spread over the vocabulary and still shifting from layer 20 the easy belief locks on “Paris” · the hard one keeps shifting little revision on the left, a lot on the right: that size is BTE
The scorer's own readouts, from the data in the figure below. Up to layer 20 both beliefs are spread over the vocabulary. The easy one then locks on “Paris” and stops changing; the hard one swings to “9” at layer 26, slips at 28 and ends at 0.82, changing all the way up. The small bar beside each row is the change from the layer before; its size is BTE.

Reading the trajectory

Every intermediate state of a Transformer can be read through the model's own output head as a next-token distribution, the model's belief at that depth. BTE scores how much consecutive layers revise that belief (Jensen–Shannon divergence) and keeps the sequence across depth. Averaged over the tokens of a response it is a 32-number profile; the mean over the central window (layers 12–19 for a 32-layer scorer) is the training-free scalar used below.

02Difficulty

If BTE captures the predictive revisions a computation induces, more challenging inputs should elicit stronger revisions. So reference solutions are scored given their problem, and the mean BTE over layers 12–19 is read off as a scalar, with no training. Across the 500 MATH-500 solutions it rises with the human difficulty level, and the gap sits in the middle layers.

Mean depth profile per MATH-500 level, Llama-3.1-8B-Instruct, over the middle window 12–19 (switch to 8–23 to see the levels converge again after layer 19); the chips give each level's middle-window mean, the scalar used for difficulty.

Case: an easy and a hard solution of the same subject

Five pairs of official MATH-500 solutions, each a level-1 item against a level-5 item of the same subject and of comparable length. The first pair is the one whose level-5 profile lies above the level-1 profile at the most layer transitions; the other four are the median pair of their subject. Pick a row to load its two solutions below.

Case: one theory, five proofs

ProofWriter theory AttNeg-OWA-D5-524. The facts and rules are the same for all five queries; the query moves from one inference step (Harry is big) to five (Harry is green), and each official proof extends the previous one by a step. To decide: does the average next-token loss follow the number of inference steps, and does BTE?

Case: wording and length controls

03Human vs LLM reviews

A scalar discards where in depth the revisions occur. Does the full depth profile carry finer structure about an input, such as who wrote it? To find out, human and LLM reviews of the same ICLR and NeurIPS papers are scored with the paper in the context, and a logistic regression on the 32-number profile is the whole detector; nothing else is learned.

Case: the same paper, a human review and a model's

Five papers from the paper's blind human evaluation. Each pair is the human review and one model's review of the same paper.

The same comparison over 11,200 review pairs. Depth profile of generated reviews relative to human reviews of the same papers (%). Legacy generators (solid) sit below human reviews through the middle layers; newer ones (dashed) close the gap and Claude Opus 5 reverses it. Blind human preference correlates with this mid-layer shift across the seven generators at ρ = 0.857 (Spearman).

04Generator attribution

The profile tells human from LLM; can it also tell which model? An eight-way classifier (Human + seven generators) is trained on the per-layer profile, then on per-layer statistics of the token-level BTE that averaging had removed. The map is the paper's Figure 5: LDA axes fitted on calibration papers, held-out reviews laid out with t-SNE on those axes.

Case: two reviews the per-layer mean cannot tell apart

Same paper, two sources. The 32-D classifier sees one mean per layer; the 160-D one also sees how the token energies spread at each layer, and the 352-D one adds where in the review they fall. To decide: which of the eight sources wrote each review. The same five papers of the blind human evaluation, with the readers' verdicts under each pair.

Left: eight-way source accuracy (%) on generated reviews from held-out papers as token-level statistics are added to the per-layer representation, both scorers. Right: confusion between true and predicted source for the selected scorer and feature set (% of each true source's held-out reviews); the residual errors are almost entirely GPT-5.5 ↔ GPT-5.6, the two sources whose clouds overlap on the map.

05Where the signal comes from

Where does the signal come from? BTE is a property of the input and of the scorer, so the scorer's familiarity with the data should move it. To see that, the text is held fixed and only the scorer changes: Llama-3.1-8B-Instruct before and after fine-tuning on PRM800K solutions, and OLMo-2-7B at each of its released pretraining checkpoints.

Case: the same text, before and after fine-tuning

Llama-3.1-8B-Instruct is fine-tuned on 9,704 PRM800K trajectories, and the same text is read by the scorer before fine-tuning and after 1,000, 5,000 and 9,704 trajectories. Four kinds of text: a PRM800K trajectory in the training style, and official solutions of MATH-500, GSM8K and OlympiadBench, which the fine-tuning never saw. Nothing in the text changes; only the scorer does.

Case: two solutions through pretraining

All 500 MATH-500 solutions

Depth profile on MATH-500 solutions across OLMo-2-7B checkpoints, with a Gaussian-initialized model and Llama-3.1-8B for reference.
Middle-layer BTE and its Spearman correlation with MATH level across pretraining tokens (log scale). Error bars: 95% bootstrap interval. Llama-3.1-8B, pretrained on about 15T tokens, reaches +0.36 before any instruction tuning and +0.44 after it.