Explainer · Transformer computation · 4 October 2026 · 6 minute read
A small adapter helps frozen transformers follow much longer reference chains
A task-trained adapter changes how existing transformer layers pass information between program lines. The results show why extra computation can help after adaptation, and why the layer where that adaptation starts matters.
The short version
- The tested base models follow only short reference chains by default. A small trained intervention lets their unchanged layers carry information much further.
- The adapter acts separately on each token. Existing attention connections perform the communication, extending a relay through a limited middle-layer window.
- Qwen3-8B reaches 99.0% exact accuracy on 24-line chains, compared with 15.5% without adaptation. Longer-trained adapters also make repeated computation useful.
- The strongest evidence comes from synthetic programs. Question-answering results support early adapter placement, but do not establish the same internal mechanism.
Extra layers do not automatically extend reference following
Consider a reference chain: a sequence of assignments that starts with a value and then names earlier variables. Given K = apple, B = K and D = B, answering print(D) means following the references back to apple. Everything needed is already in the prompt.
The authors of the arXiv paper study why pretrained transformers struggle when these chains grow. Across thirteen standard base models, they report reliable performance through only 1.4 to 3.6 lines, with a median of 2.2. Chain length includes the assignment that stores the starting value.
Their measure, reach, is the longest chain length maintaining accuracy of at least 80%, estimated by linear interpolation at the first downward threshold crossing. The model survey measures accuracy by choosing among the chains’ starting values. This makes the task about selecting the right reference, rather than producing an elaborate answer.
Larger models and repeated passes do not reliably remove the problem. Solved examples help answer formatting but leave reach limited. The important distinction is between computation the weights cannot perform and computation the default pass fails to engage. The experiments investigate the latter possibility.
Change the representations, not the whole model
The intervention is a LoRA, a learned low-rank update. Here it changes the residual stream, the token representation entering a transformer layer, rather than updating the model’s original weights. The authors train a rank-8 update at a single layer and keep those original weights fixed.
For Qwen3-8B, this adds 65,537 trainable parameters, less than 0.01% of the model. Crucially, the adapter processes each token separately. It cannot itself move information from one program line to another. That communication must happen through the unchanged transformer layers.
The authors describe the resulting computation as a relay: program lines pass along information about which chain they belong to. A line acquires names from earlier in its chain, then uses those names to read further back. Existing attention connections carry these reads between tokens.
Without adaptation, this relay travels only a short distance. The query follows a few more references, and a later layer copies the chosen starting value. With adaptation, the program lines themselves do more of the reference following. The improvement is not simply extra work at the final query.
Where the reference-following work happens
Without adaptation
- Program lines pass chain information only a short distance.
- The query resolves a few additional references.
- A late layer copies the selected starting value.
With the adapter
- The adapter changes each token separately.
- Frozen attention reads progressively earlier chain lines.
- The program relay travels further through middle layers.
Follow the relay through a worked program
The source example mixes an apple chain with a pear chain. Its assignments are K = apple, M = pear, B = K, Q = M, D = B and R = Q. The query print(D) should return apple. Variable names, starting nouns and the queried chain change across programs.
- Arrange the prompt. In the main long-chain setting, assignments are grouped by their depth within each chain and shuffled within those groups. This keeps definitions before their uses.
- Train the adapter with answer supervision and a text penalty on WikiText-103. The standard Qwen3-8B adapter enters layer 14 and sees chains through 20 lines. Layer numbering begins at zero.
- Let the frozen layers communicate. A line first reads its parent, the line defining the variable on its right-hand side. Earlier names become available, supporting reads further up the chain.
- Score the answer. Exact accuracy requires the correct starting value to outrank every other token in the vocabulary, not just the other chains’ possible answers.
From local token changes to a longer relay
- Apple chainK = apple; B = K; D = BPear chainM = pear; Q = M; R = Q
- Change each tokenThe trained adapter modifies representations without communicating between tokens.
- Read parent linesFrozen attention brings information from earlier chain lines.
- Extend the relayEarlier names support reads further up the same chain.
- Select the starting valueprint(D) should produce apple, not pear.
To test where information actually matters, the authors use causal tracing. They change a starting noun or redirect a reference, then restore an internal token state from the original run. Recovery of the original answer preference reveals where answer-relevant information was carried.
They also remove selected attention connections. Cutting parent reads during Qwen3-8B’s relay window, layers 14 through 22, reduces accuracy to 53%, 48% and 55% on chains of six, eight and twelve lines. Those results are near chance. Cutting the same connections later leaves 100%, 100% and 98%.
This timing matters. A connection can be necessary while the relay advances but dispensable after its information has moved elsewhere. The tests therefore support a communication process, not merely a correlation between internal representations and correct answers.
The gains depend on both training and placement
The authors report a large improvement for the standard Qwen3-8B adapter on two-chain programs. At 16, 20 and 24 lines, exact accuracy is 98.5%, 98.0% and 99.0%. The unchanged model scores 15.5%, 14.5% and 15.5%, respectively. The 24-line evaluation extends beyond the standard adapter’s training lengths.
Qwen3-8B accuracy with and without the standard adapter
Unit: Exact accuracy (%)
A separately trained adapter exposed to longer programs gives Qwen3-8B a reach of 50 lines. On 40-line chains it scores 98%, compared with 4% without adaptation. At 48 lines, the comparison is 88% against 9%. These longest results use two chains grouped by depth.
Repeated computation becomes useful too. Ouro-1.4B repeats a block of layers in loops. Its longer-trained adapter reaches 60 lines with four loops. With six loops, reach is about 146 lines; with eight, it is at least 160. Unlike the Qwen exact-accuracy tests, the default looped-model evaluations choose between the chains’ starting values.
Those longest Ouro results apply the adapter in every loop. A separate test applies it only in the first loop and reaches about 25 lines with later, unmodified loops. That narrower result shows that frozen layers can continue the computation after the adapter stops acting.
Location is consequential. Moving an otherwise matched Qwen3-8B adapter from layer 20 to layer 21 cuts reach from 20.5 to 5.2 lines. The authors’ interpretation is that adaptation must happen while the layers capable of extending the relay still lie ahead.
For MuSiQue, a benchmark requiring linked facts to answer questions, the authors train new task-specific adapters. Across 900 development questions, the earliest tested placements improve exact match by 11.4 points for Qwen3-8B, 9.4 for OLMo-3-7B and 17.9 for Llama-3.1-8B, relative to their frozen versions.
The evidence has clear boundaries
The relay mechanism is established on synthetic tasks whose needed information sits inside the prompt. Detailed tracing concentrates on Qwen3-8B and Ouro-1.4B. MuSiQue mainly supplies the correct supporting paragraphs, so its results test adapter placement rather than proving that natural-language answers use the same relay.
Prompt structure also matters. Mixing assignments while preserving definition order lowers Qwen3-8B’s 24-line accuracy to 76.5%. Using three chains lowers its 16-line accuracy to 58.0%. Both still improve on the frozen model, but the strongest scores should not be read as independent of input format.
More computation is not an unlimited remedy. The looped models lose reach when run far beyond their training loop counts. The longest looped results also require the depth-grouped arrangement and adaptation in every loop. First-loop-only success does not establish that the adapter can be removed from those longer runs.
There are evaluation limits as well. MuSiQue’s best layers were selected using the same development questions used to report performance. The placement predictor is only moderately precise and has not shown superiority over relative model depth. Broader behavior, deployment reliability and the minimum sufficient adapter rank remain unresolved.
Test which computation an intervention makes available
For researchers studying transformer computation, the result changes how to interpret a weak direct answer. Failure without adaptation does not establish that the frozen weights lack useful reference-following machinery. Here, a small learned change exposes substantially longer computation without retraining those weights.
For engineers choosing adapter locations, the lesson is to examine what computation follows the intervention. Merely leaving many layers afterward is not enough. For repeated inference, the corresponding lesson is to repeat layers that advance the task, rather than assume any extra pass will help.
This is not evidence for a universal reasoning upgrade. The adapters are trained for particular task formats, and the causal tests do not uniquely identify the relay’s algorithm. The supported conclusion is narrower: default behavior can understate what unchanged transformer layers can compute under a specified, trained intervention.
Terms used here
- Reference chain
- Assignments in which a starting value is reached by following variables that name earlier variables.
- Reach
- Chain length at the first drop below 80% accuracy, interpolated between tested lengths; a lower bound if every tested length passes.
- LoRA
- A learned low-rank update. This paper’s main version changes a token representation at one layer while keeping original model weights fixed.
- Residual stream
- The token representation entering a transformer layer, where the paper applies its main adapter.
- Relay
- Program lines passing information about their chain onward, enabling references to be followed further.
- Exact accuracy
- Accuracy requiring the correct starting value to rank above every other token in the model’s vocabulary.
- Causal tracing
- Changing an input, then restoring an original internal state to test where information affecting the answer is carried.
- MuSiQue
- The question-answering benchmark used here to test task-specific adaptation and layer placement beyond synthetic programs.
The work
Transformers Stop Thinking Too Early, and a Tiny LoRA Fixes It
- Authors
- Zehao Jin, Ruixuan Deng and Junran Wang
- Published
- 29 September 2026
- Code
- github.com
- Project page
- lunamos.github.io
Related reading
This explainer was written by AI from the source text and checked against it. Read the source for the full detail. How this site works