Explainer · Agent training · 6 October 2026 · 7 minute read
Recursive Self-Rewrite turns controller-assisted successes into agent training data
Specialized controllers can help an agent succeed without teaching it to work independently. Recursive Self-Rewrite converts those successes into fresh, verified demonstrations under a general controller, improving terminal-task results while leaving questions about which components drive the gains.
The short version
- The method extracts procedures from successful task histories, screens the guidance, and makes the model solve each task again in a fresh environment.
- Training keeps the verified public interactions, not the private runbooks or critic discussions. Deployment uses the general harness alone.
- On Terminal-Bench 2, the authors report 74.2% success within three attempts, compared with 57.0% for the base model and 53.4% for direct training.
- The experiments combine rewriting with more training data and filtering. They do not establish which ingredient causes the gains, and long-horizon task completion remains absent.
A successful agent run can hide help from its controller
An agent's performance depends on more than its model. Its harness, the external system managing tools, observations, progress checks, recovery, and stopping, can change which tasks it solves. A structured controller may enforce intermediate checks. Another may tell the model to continue after it mistakenly declares victory.
That creates a training problem. A trajectory, the recorded actions and environment responses from an attempt, captures the model and controller working together. Copying successful trajectories into training can also copy dependence on prompts, workflow rules, or stopping conventions that will disappear when the model runs under a simpler harness.
Earlier harness engineering improves execution without changing model weights. Other approaches retain lessons as guidance for later attempts or train directly on successful histories. The distinction here is between improving the support system and teaching the model behavior that survives when that support is removed. Merely deleting controller messages does not reconstruct the missing decision process.
Cleaning a history versus generating new practice
Direct SFT
- Start with passing discovery histories
- Remove controller-specific insertions
- Train without planning or fresh execution
Recursive Self-Rewrite
- Extract and screen procedural runbooks
- Solve again through Terminus 2
- Train on verified public histories
- Also retain direct Terminus 2 successes
Turn a discovered solution into another practice attempt
Zongxia Li and colleagues propose Recursive Self-Rewrite, or RSR, a pipeline for converting assisted successes into new training demonstrations. It uses Qwen-3.8-27B throughout discovery, planning, screening, and execution. Rather than ask a stronger teacher to supply answers, the method reorganizes the same model's successful experience.
The central object is a runbook: procedural guidance covering the desired result, milestones, checks, recovery options, and pitfalls. It should explain how to approach the work without supplying the finished deliverable. A separate execution then follows that guidance in a fresh task environment under Terminus 2, the general harness used for deployment.
The runbook stays private during generation and is excluded from training. The training example contains only the task, observations, and model responses from the new attempt. This is the proposed transfer mechanism: use extra guidance to generate useful practice, then train on the behavior produced during that practice.
The word recursive refers to revising rejected runbooks using critic feedback within a retry budget. The paper does not claim an autonomous cycle of endless model improvement. Its concrete contribution is a data-generation and training pipeline with a fixed target interface.
Discover, screen, execute again, and train
Discovery holds the model fixed and changes its support. Terminus 2 offers ordinary terminal interaction. StateM organizes work into persistent states and checked transitions. Recursive Self-Reflect Terminus continues an attempt after a failed completion check, preserving the conversation and environment while withholding the checker's internal details.
These controllers expose different successful procedures. In a system-administration example, Recursive Self-Reflect Terminus resumes after rejection, removes processes left running, and fixes a required log. StateM uniquely solves a software task by listing requirements and catching an exact dependency-version constraint. Such successes become material for rewriting.
- Compact the successful history. Keep the task, model actions, and environment observations, but strip controller-specific messages. Ask the planner to produce multiple candidate runbooks.
- Screen each candidate. Rule-based checks reject malformed guidance, known artifacts, and unsupported tools. A model critic sees only the public task and proposed runbook, then supplies feedback for revision.
- Execute accepted guidance in a fresh environment. The executor must solve the task through Terminus 2 rather than inherit completed work or replay the old interaction.
- Keep passing executions that survive filtering. Remove private guidance and critic discussions, then train on the public histories alongside retained direct Terminus 2 successes.
Private guidance produces public demonstrations
- Discover successesKeep the model fixed; vary terminal controllers.
- Compact the historyRetain task, actions, and observations; remove controller messages.
- Plan candidate runbooksDescribe milestones, checks, recovery, and pitfalls without the finished answer.
- Screen and reviseRule checks and a model critic reject leakage; feedback guides another candidate.
- Execute afreshUse private guidance to solve the task through Terminus 2 in a new environment.
- Verify and filterRetain passing histories; discard suspected hidden answer transfer.
- Train, then deployTrain without private guidance; deploy with the general harness alone.
The screening targets leakage: hidden answers or private verification information reaching the executor through supposedly procedural guidance. After execution, the model also flags values that cannot be derived from the task or current environment. Demonstrations containing those values are discarded. The filters aim to preserve legitimate procedure without secretly transferring the solution.
The Markdown inline-parsing case makes this concrete. The source attempt spends substantial effort on setup, package compatibility, and parser selection. Runbook-guided attempts reuse advice about a compatible parser and the reconstruction procedure, avoiding much of that preliminary exploration. They still have to perform the reconstruction themselves.
Of 12 rewritten attempts, 7 pass. The median successful rewrite takes 32 turns, versus 64 for the source, a reported 50.0% reduction. Yet some similarly short attempts fail. The example supports a narrow lesson: guidance can remove wasted setup without eliminating the difficult part of the task.
Rewritten demonstrations improve the reported benchmarks
The authors report that combining discovery harnesses solves 759 tasks from approximately 3K tasks, a 34.3% increase over the best individual controller's coverage in the collected data. Sampling is unequal. On a subset with equal attempt counts per harness, the combination solves 352 tasks, versus 285 for the strongest individual controller.
The abstract describes expansion from 2,001 successful source trajectories to 11,094 training trajectories. The training section additionally retains 766 direct Terminus 2 successes. Its comparison uses the original model, Direct SFT, meaning supervised training on cleaned source histories without fresh execution, and RSR training.
The main metric is pass@3, success within three attempts. On Terminal-Bench 2, the original model scores 57.0%, Direct SFT scores 53.4%, and RSR scores 74.2%. Direct training therefore hurts this benchmark, whereas rewriting improves it. On the authors' Terminal-Bench Hard, the corresponding results are 39.0%, 56.0%, and 63%.
Success within three attempts on Terminal-Bench 2
Unit: Pass@3 (%)
The same ordering favors RSR on harder evaluations, although absolute success remains low. For the original model, Direct SFT, and RSR respectively, Terminal-Bench 3 scores are 0.0%, 5.4%, and 9.5%. Terminal-Bench 4 scores are 1.5%, 4.5%, and 9.1%. On the authors' SWR100 software benchmark, they are 3.0%, 3.0%, and 6.0%.
The advantage is not limited to success across multiple attempts. Averaging individual-run success rates also favors RSR over Direct SFT on every one of these benchmarks. For Terminal-Bench 2, those averages are 70.1% and 43.8%, respectively, compared with 51.7% for the original model.
Long-Horizon Terminal Bench uses process reward, a score that can reflect partial progress without completion. It rises from 0.21 for the original model and 0.25 for Direct SFT to 0.29 for RSR. All groups nevertheless complete zero of its 46 tasks. This is progress within attempts, not evidence of completed long-horizon work.
The gains do not isolate the cause
The training comparison bundles several changes: reconstructed guidance, fresh execution, filtering, more demonstrations, and retained direct data. The reported results compare the original model with direct training and the full rewriting pipeline. They support the combined pipeline, not a conclusion that planning or critic feedback alone explains the improvement.
The scope is also narrow. The pipeline uses a single base model and evaluates terminal tasks, so transfer to other models or interaction settings remains untested here. Leakage screening partly relies on that same model. The supplied text gives no measured screening accuracy, leaving the effectiveness of the safeguards uncertain.
There is an unresolved data-accounting issue. Discovery reports successes on 759 distinct tasks, while rewriting reports coverage of 975 unique tasks without explaining the additional tasks. The generation-temperature sentence is incomplete. These gaps make the reported data pipeline harder to reproduce and interpret precisely.
Benchmark reporting also needs care. Parenthesized pass@3 task counts are rounded estimates derived from percentages, rather than independently reported exact counts. More importantly, better scores do not erase the remaining failures: several benchmarks retain low absolute success, and the long-horizon evaluation records no complete solutions.
Treat controllers as sources of experience, not just deployment machinery
For researchers training agents and engineers collecting terminal demonstrations, the useful idea is to separate discovery support from the final training interface. A specialized harness can uncover a procedure worth learning even when that harness will not be available at deployment. Re-execution tests whether the procedure can produce a valid demonstration under the intended interface.
The practical takeaway is not that cleaned logs are useless: Direct SFT improves several reported benchmarks. It is that a successful assisted history and a deployment-compatible demonstration are different objects. RSR offers a concrete way to bridge them, with encouraging results for the full pipeline and important uncertainty about safeguards and generality.
Terms used here
- Harness
- The external execution system controlling tools, observations, progress checks, recovery, and stopping.
- Trajectory
- A recorded sequence of model actions and environment responses during a task attempt.
- Recursive Self-Rewrite
- A pipeline that converts assisted successes into screened guidance and fresh, verified training demonstrations.
- Runbook
- Procedural guidance describing goals, milestones, checks, recovery options, and pitfalls without giving the finished answer.
- Leakage
- Hidden answers or private verification information entering guidance or demonstrations that should contain only legitimate information.
- Direct SFT
- Supervised training on cleaned successful source histories without reconstructing runbooks or executing the tasks again.
- pass@3
- Success on a task within three attempts.
- Process reward
- An evaluation score that can reflect partial progress even when a task is not completed.
The work
Scaling Trajectories for Complex Tasks through Recursive Self-Rewrite
- Authors
- Zongxia Li, Yucheng Shi, Zhongzhi Li, Junyao Yang, Ruhan Wang, Chengsong Huang, Fuxiao Liu, Haitao Mi and 2 more
- Published
- 2 October 2026
- Venue
- Hugging Face Daily Papers
Related reading
This explainer was written by AI from the source text and checked against it. Read the source for the full detail. How this site works