Gemini 3.5 Flash / OpenCode
Harness²: Recursive Agent
Harnessing
for an Open World
Harness the agent.
Then harness its self-improvement.
+4.3 from k = 1 to 5
Learn from execution.
Then learn from improving it.
A harness enables a model to act as an agent. Harnessing is the process of optimizing this system. Harness² makes harnessing recursive: it refines the execution harness for the current task, then updates the improvement harness with reusable harnesses and lessons to guide future refinement.
Agent = Model + Harness. Harnessing becomes recursive.
Self-improving Agent
Agent + Recursive Harnessing
Stage details Example & explanation
A new task arrives.
Start a fresh execution from the base harness. Earlier task experience, when available, belongs to the improver's library.
Reconcile invoices, incidents, and posted credits into one checked workbook.
View the complete pipeline figure from the paper
Task-level recursion.
Parallel explores independent candidates. Sequential improves the previous harness.
Each composer reads all candidates and execution evidence. Highlighted arrows illustrate the edits retained.
Eight components.
One adaptive harness.
Improvement can change more than a prompt. These components shape what enters the context, how tools are used, how work is delegated, and what happens inside the tool loop.
System prompt
The agent's persona and top-level procedure. It defines how the executor approaches a task before any individual tool is called.
systemprompt.md
Example: The Appendix defines this component as the persona and top-level procedure followed on all tasks.
# System prompt
Persona and top-level procedure
followed on all tasks.
See how all eight components fit together
The harness changes.
The model stays the same.
One improvement step raises the primary benchmark score across all four model–harness configurations. Parallel recursion at k = 5 adds further gains. Explore the reported means below.
from a single
improvement step.
Gemini 3.5 Flash with OpenCode improves from 72.4 to 77.5. Parallel recursion at k = 5 reaches 80.8.
Same executor. Same benchmark.Higher is better.
| Method | Code | Web | Office | Average |
|---|---|---|---|---|
| Base harness | 73.4 | 69.9 | 74.0 | 72.4 |
| ReasoningBank | 72.6 | 70.7 | 76.8 | 73.4 |
| Meta-Harness | 75.5 | 75.1 | 78.6 | 76.4 |
| Harness Scaling | 74.0 | 74.0 | 77.9 | 75.3 |
| Harness² · k=1 | 76.5 | 76.4 | 79.5 | 77.5 |
| Harness² · k=5 | 80.8 | 78.3 | 83.4 | 80.8 |
WorkBuddy Average is the macro-average of Code, Web, and Office. Reported means across three runs; k = 5 uses parallel recursion.
Interactive results could not load. The default table is shown above; download all reported means.
From a plausible answer
to a checked deliverable.
In this JobBench case, the base agent left source files unreconciled and never reopened its outputs. The improved harness added verification guardrails, a reconciliation skill, and precise tool guidance.
Cite Harness².
If Harness² is useful in your research, please cite the paper using the BibTeX entry. The record will be updated when the archival version is released.
@misc{xu2026harness2,
title = {Harness$^2$: Recursive Agent Harnessing for an Open World},
author = {Xu, Ruiyao and Chen, Yanfei and CuiZhu, Zhongying and
Dalvi Mishra, Bhavana and Ming, Yifei and Yu, Han and
Han, Rujun and Lee, Chen-Yu and Pfister, Tomas},
year = {2026},
url = {https://harness2.github.io/}
}
