RECURSIVE AGENT HARNESSING

Harness²: Recursive Agent Harnessing
for an Open World

1Northwestern University 2Google Cloud AI Research

Harness the agent.
Then harness its self-improvement.

WORKBUDDY-BENCH · THREE DOMAINS

Gemini 3.5 Flash / OpenCode

Code
Web
Office
TASK-LEVEL SCALING · WB-CODE

+4.3 from k = 1 to 5

01
TASK-LEVEL RECURSIONImprove the execution harness
02
HARNESSING-LEVEL RECURSIONGuide how future harnesses are improved
01 THE METHOD

Learn from execution.
Then learn from improving it.

A harness enables a model to act as an agent. Harnessing is the process of optimizing this system. Harness² makes harnessing recursive: it refines the execution harness for the current task, then updates the improvement harness with reusable harnesses and lessons to guide future refinement.

Agent = Model + Harness. Harnessing becomes recursive.

Model Frozen · 𝓜
Harness Editable · 𝓗 = (c¹, …, cᴷ)
Agent 𝒜(·; 𝓜, 𝓗)
Harness² RECURSIVE HARNESSING
Improvement Harness 𝓡
Model 𝓜FROZEN
Execution Harness 𝓗EDITABLE
Self-improving Agent Agent + Recursive Harnessing
THE TWO-LEVEL LOOP

A stream of tasks. An improving improver.

TASK STREAM k = 1 · both modes
executor starts at H₀
01Refine the execution harness
TASK LEVEL · within Task 1
ImproverR₁ · empty library
Propose → Probe → Compose
Improvement record
02Guide future harnessing
HARNESSING LEVEL · across tasks
UPDATEImprovement library
R₁
Ready to learn from the first task.
filesexperience

Reuse the library in the next improvement; each task starts at H₀.

Fixed model weights
Stage details Example & explanation
TASK STREAM · 01 / 08

A new task arrives.

Start a fresh execution from the base harness. Earlier task experience, when available, belongs to the improver's library.

CURRENT TASK

Reconcile invoices, incidents, and posted credits into one checked workbook.

View the complete pipeline figure from the paper
THE HARNESS² PIPELINE
Harness squared pipeline showing task-level recursion and harnessing-level recursion through an evolving improvement harness.
Task-level recursion refines the current execution harness. Harnessing-level recursion retains reusable harnesses and improvement experience for subsequent tasks.
02 WITHIN ONE TASK

Task-level recursion.

Parallel explores independent candidates. Sequential improves the previous harness.

Parallel recursion · k = 3 Reference

Each composer reads all candidates and execution evidence. Highlighted arrows illustrate the edits retained.

Library stays fixed within this task. next task → H₀
03 THE EDITABLE HARNESS

Eight components.
One adaptive harness.

Improvement can change more than a prompt. These components shape what enters the context, how tools are used, how work is delegated, and what happens inside the tool loop.

C1Every call
CONTEXT CONSTRUCTION

System prompt

The agent's persona and top-level procedure. It defines how the executor approaches a task before any individual tool is called.

systemprompt.md

Example: The Appendix defines this component as the persona and top-level procedure followed on all tasks.

systemprompt.md MARKDOWN
# System prompt
Persona and top-level procedure
followed on all tasks.
APPENDIX · COMPONENT DEFINITION ready

See how all eight components fit together
The eight harness components grouped by context construction, tool interface, orchestration, and middleware control.
The component names and paths follow the paper's OpenCode harness.
04 THE RESULTS

The harness changes.
The model stays the same.

One improvement step raises the primary benchmark score across all four model–harness configurations. Parallel recursion at k = 5 adds further gains. Explore the reported means below.

WORKBUDDY · AVERAGE SCORE
+5.1points

from a single
improvement step.

Gemini 3.5 Flash with OpenCode improves from 72.4 to 77.5. Parallel recursion at k = 5 reaches 80.8.

Same executor. Same benchmark.
Higher is better.
Gemini 3.5 Flash / OpenCode · Reported means across three runs
Method Code Web Office Average
Base harness 73.4 69.9 74.0 72.4
ReasoningBank 72.6 70.7 76.8 73.4
Meta-Harness 75.5 75.1 78.6 76.4
Harness Scaling 74.0 74.0 77.9 75.3
Harness² · k=1 76.5 76.4 79.5 77.5
Harness² · k=5 80.8 78.3 83.4 80.8

WorkBuddy Average is the macro-average of Code, Web, and Office. Reported means across three runs; k = 5 uses parallel recursion.

05 A CLOSER LOOK

From a plausible answer
to a checked deliverable.

In this JobBench case, the base agent left source files unreconciled and never reopened its outputs. The improved harness added verification guardrails, a reconciliation skill, and precise tool guidance.

JOBBENCH CASE STUDY · GEMINI 3.5 FLASH / OPENCODE
JobBench case study showing failure analysis, four harness interventions, and a task-score improvement from 43.1 to 77.8.
The composer retained edits supported by the execution evidence. The resulting deliverable was then evaluated independently.
06 CITATION

Cite Harness².

If Harness² is useful in your research, please cite the paper using the BibTeX entry. The record will be updated when the archival version is released.

BIBTEX
@misc{xu2026harness2,
  title  = {Harness$^2$: Recursive Agent Harnessing for an Open World},
  author = {Xu, Ruiyao and Chen, Yanfei and CuiZhu, Zhongying and
            Dalvi Mishra, Bhavana and Ming, Yifei and Yu, Han and
            Han, Rujun and Lee, Chen-Yu and Pfister, Tomas},
  year   = {2026},
  url    = {https://harness2.github.io/}
}