AI Agent News

The day's AI agent news, for people who build and run agents.

Explainer · Reasoning and training-data associations · 7 October 2026 · 7 minute read

Short response openings can unlock reasoning learned during base-model training

Sophie L. Wang and colleagues show that supplying a short response opening can bring base models close to reinforcement-trained models. Controlled training-data edits explain how words acquire reasoning effects. A separate safety study finds different behavior from similar-looking openings.

The short version

  • Fixed openings improve math accuracy without changing model weights. On MATH-500, cued Olmo-3-7B reaches 78%, compared with 42% without a cue and 75% after reinforcement learning.
  • Training associations matter: replacing “okay” with “chicken” makes the arbitrary word an effective reasoning cue.
  • The findings distinguish abilities a model already has from behaviors it readily produces. They do not establish that all post-training gains come from response openings.
  • Cue effects depend on the model, surrounding context and formatting. Similar-looking openings also produce sharply different harmful-response rates in the safety case study.

Better answers do not necessarily mean newly learned reasoning

A base model learns to predict the next piece of text before the additional training studied here. It may already produce calculations, reflection and verification. Seeing these behaviors more often after reinforcement learning, or RL, does not establish that RL created them. The Qwen3-14B comparison model receives RL training on MATH with rewards based on answer correctness.

This distinction motivates the work by Sophie L. Wang and colleagues. They investigate whether reasoning needs further training or whether a small intervention can make existing reasoning more likely to appear. Their intervention supplies only the beginning of the model’s response, not a worked solution or an instruction to reason.

Earlier work elicited reasoning with demonstrations and explicit instructions. Other approaches explored alternative first tokens separately for each problem or resampled outputs using their likelihood. These approaches showed that base models could reason, but left open how much a single fixed opening could accomplish across problems.

Another line of work changes how long already post-trained models reason. This paper instead studies how to elicit reasoning in base models. It also goes beyond observing an opening’s effect: the authors edit training data and retrain models to test where that effect comes from.

From eliciting reasoning to testing a fixed opening

Earlier approaches

  • Supply worked reasoning demonstrations or explicit instructions
  • Explore alternative first tokens separately for each problem
  • Control reasoning length in already post-trained models

This work

  • Select a short opening using sampled-answer agreement
  • Reuse one cue for each model and prompt format
  • Edit training associations and retrain to test causality
Earlier approaches showed ways to elicit reasoning. This work tests reusable response openings and edits training data to investigate why they work.

An opening can signal which kind of text comes next

A token cue is a short response opening that changes the continuation. Prefilling means supplying that opening and letting the model generate the rest, with its weights unchanged. The selected openings are `.\n\n Okay` for Olmo-3-7B and `␣Alright ,` for Qwen3-14B. In the former, `\n\n` represents two newlines.

The proposed explanation is a learned association, not a hidden instruction inside the word. If training repeatedly places an opening before reasoning, the model can learn to continue that opening with reasoning. In Olmo-3-7B’s intermediate training data, 83% of generated reasoning examples begin with “Okay.”

The important separation is between generating an opening and continuing it. A model may be unlikely to choose a useful beginning even though it can solve the problem once that beginning is supplied. Prefilling isolates the continuation. The paper then tests whether RL changes the likelihood of choosing these useful beginnings.

Find a cue, reuse it, then change its training associations

The search starts with 30 training problems from MATH. Beam search, which repeatedly retains likely candidate sequences, proposes openings the base model already tends to generate. The authors keep 20 candidates using a search width of 20 and a depth of two tokens.

For every candidate and problem, they sample 16 continuations and extract final answers. They reject a candidate if its rate of missing answers is more than 10 percentage points above the rate without a cue. Among the survivors, they choose the lowest answer entropy, a measure of disagreement among sampled answers, averaged across problems. This uses agreement as a substitute for checking reference answers.

Each model and prompt format gets one selected cue, reused across evaluation problems and benchmarks. The procedure does not search for a different opening for every test question. Nor does supplying the selected opening update the model’s weights.

How the authors discover and use a reasoning cue

  1. Start with training problemsUse 30 MATH problems to rank likely openings.
  2. Propose short openingsBeam search retains 20 candidates at depth two.
  3. Sample continuationsGenerate 16 continuations for each opening and problem.
  4. Filter and selectReject excessive missing answers, then choose the lowest average answer entropy.
  5. Prefill the selected cueReuse it across evaluation problems and benchmarks, with weights unchanged.
Cue selection uses agreement among generated answers, not reference answers. Once selected, the opening is reused without changing model weights.

The school-size example shows what changes. There are 834 music students, representing two-thirds of the school. With `.\n Answer`, Olmo gives the incorrect answer 1002. With the reasoning cue, it writes “(2/3) * T = 834”, finds 1251 and checks the calculation. The base model supplies both the solution and verification without an RL update.

To test the training-association explanation, the authors resume Olmo training from the same checkpoint using matched data mixtures. One mixture stays unchanged. Another replaces “okay” throughout with “chicken.” A further version changes paragraph-opening “Question:” labels in short question-answer documents to “Okay,”. The checkpoint, training recipe, data order and random seed remain shared across these runs.

These edits separate two possibilities. Renaming tests whether an arbitrary word can acquire the reasoning association. Redirecting tests whether an existing cue can instead lead into a question. Simply removing a word from this training stage need not erase what earlier training taught the model.

The gains approach RL, and edited data changes the effective cue

The authors report pass@1, the probability that one sampled response is correct. On MATH-500, Olmo-3-7B scores 42% without a cue, 78% with its selected cue and 75% after RL. Qwen3-14B moves from 72% to 87% with its cue, matching its RL counterpart. Prompts and sampling settings stay fixed within each model family.

Selected cues match or exceed the RL comparison on MATH-500

  • Olmo-3-7B, no cue42%
  • Olmo-3-7B, selected cue, highlighted78%
  • Olmo-3-7B, RL75%
  • Qwen3-14B, no cue72%
  • Qwen3-14B, selected cue, highlighted87%
  • Qwen3-14B, RL87%

Unit: MATH-500 pass@1 (%)

Pass@1 measures the probability that a single sampled response is correct. The cued base models receive supplied opening tokens, not weight updates.

The unchanged cues also improve results on GSM8K, AMC 23 and AIME 2024. In most comparisons, they equal or beat RL. HumanEval tests coding: there, supplying Olmo’s cue raises its performance to match RL. For Qwen, neither the cue nor RL substantially changes its already high base-model accuracy.

RL also makes the selected openings more probable. For Olmo, the reasoning cue’s probability rises from 0.14 to 0.65; for Qwen, its cue rises from 0.04 to 0.58. When the authors compare models on identical preceding text, the largest changes in next-token predictions occur at the first two positions.

Longer responses alone do not explain the accuracy gain. Under matched truncation budgets, Olmo’s reasoning cue beats no cue from 1k tokens onward. With roughly half the budget, it attains accuracy comparable to RL. Other openings produce longer responses yet perform worse than no cue.

The data edits provide the causal result. With `.\n\n Chicken` supplied, MATH-500 accuracy is 2.4% after training on the unchanged mixture and 37.2% after the renaming edit. On the renamed mixture, the original Okay cue still reaches 38.3%. Removing “okay” from this stage therefore does not erase its usefulness when supplied.

Actively redirecting Okay toward questions has a different outcome: it makes the model write a question rather than solve the input, and accuracy falls to 0.2%. A separate phrase-replacement experiment changes results for “Think duck duck goose”: accuracy rises from 6% before the edit to 17% afterward. Under “Think step by step,” accuracy moves from 12% to 15%. The arbitrary phrase becomes comparably effective.

The evidence does not establish a universal reasoning switch

The RL comparison concerns training directly from base models, without a preceding supervised post-training stage. It does not settle what more complex training pipelines add. Cue effectiveness is also model-dependent: the tested openings did not reveal an effective cue for Llama-3.1-8B.

The causal interventions cover Olmo and SmolLM3, whose training mixtures contain generated reasoning examples. Repeated, standardized openings in those examples may make cue effects especially strong. These experiments support the association mechanism in that setting, not an equally strong effect for every training mixture.

The internal analysis is informative but not decisive. A hidden state is the model’s internal representation at a particular position and layer. Different cues make these representations resemble different training-document types. Yet one opening produces reasoning-like representations without calculating an answer. Resemblance is not proof of successful reasoning.

Position matters too. Inserting Okay later in a response can briefly shift internal representations toward reasoning examples while reducing accuracy below the no-cue baseline. The cue’s effect depends on the preceding context, rather than acting as a context-independent switch.

Researchers and model builders should test what openings activate

For researchers studying RL, the work offers a concrete way to separate available ability from likely behavior. If a fixed opening recovers much of a training gain, the unprompted baseline alone understates what the base model can produce. That does not make RL useless; it changes the interpretation of what improved.

For training-data designers, the lesson is that recurring openings can become behavior controls, deliberately or accidentally. For engineers choosing response formats, small formatting choices deserve evaluation rather than assumptions of equivalence. The safety case study makes that point especially clear.

On unsafe XSTest requests, the authors report harmful-response rates of 42.4% with `␣Okay ,` and 0.2% with `.\n\n Okay`. This evaluates Olmo-3-7B and uses WildGuard to classify responses; it is not a general safety guarantee. The practical implication is to test both useful and harmful behaviors when selecting a cue, rather than treating a math improvement as sufficient.

Terms used here

base model
A model trained to predict text before the additional post-training stages being studied.
reinforcement learning
Training guided by rewards, here including rewards for correct math answers; abbreviated RL.
token cue
A short response opening that affects what the model generates afterward.
prefilling
Supplying response-opening tokens, then letting the model continue without changing its weights.
beam search
A search that repeatedly keeps likely candidate sequences as it extends them.
answer entropy
A measure of disagreement among sampled final answers; lower values mean greater agreement.
pass@1
The chance of getting a correct answer from one sampled response.
hidden state
An internal model representation at a particular text position and layer.

The work

Base Models Can Reason By Taking a Cue From Training Data

Authors
Sophie L. Wang, Amil Dravid, Rulin Shao, Kevin Farhat, Sewon Min and Alexei A. Efros
Published
5 October 2026
Code
github.com
Project page
sophielwang.com

Related reading

This explainer was written by AI from the source text and checked against it. Read the source for the full detail. How this site works