HERA / Reliable agentic abstention

Harness–Environment
Co-Evolution for Reliable
Agentic Abstention

Evolving the agent and its challenges, together.

Han Luo1,2,* Bingbing Wen1,* Guang Yang1 Zora Zhiruo Wang3 Pan Lu4 Lucy Lu Wang1,5
1 University of Washington 2 University of Leeds 3 Carnegie Mellon University 4 Stanford University 5 Allen Institute for Artificial Intelligence

* Equal contribution

61.7 → 83.3%Abstention successBase → HERA · HERA-Bench
48.3 → 70.0%Paired successBase → HERA · HERA-Bench
19 other LLMsTransfer without further optimizationBeyond the GPT-5.6-Luna development model
// 01

When should
an agent stop?

An agent may have the tools to act, yet lack what it needs to complete a task correctly.

A missing record, an unavailable prerequisite, or conflicting constraints can make a request infeasible. Reliable agents need to recognize these situations, explain what blocks completion, and abstain from unsupported decisions while still reporting what they can establish.

HERA studies this boundary with matched feasible–infeasible task pairs. Both instances share the same request and tool interface. A controlled change to the environment determines whether the task can be completed.

As the harness improves, its remaining failures guide the construction of new challenges. Those challenges, in turn, guide the next harness update.

// 02

Evolving the agent
and its challenges

HERA couples verifiable task construction with failure-driven harness adaptation.

Stage 1

Construct verifiable pairs

Build executable environments, retain tasks with validated reference solutions, then mutate the environment to create matched cases requiring abstention. Check both semantic consistency and possible alternative solutions.

Stage 2

Co-evolve from failures

Analyze paired execution trajectories. Use the resulting diagnosis to evolve environments and refine the harness’s prompts, memory, tools, and control logic. Select harness updates on a fixed validation set.

HERA co-evolution loop: paired rollouts produce a failure diagnosis that guides environment evolution and harness evolution. New task pairs are certified and screened; candidate harnesses are selected on validation data.
Figure 1 · Harness–environment co-evolution. New environments expose weaknesses; the harness adapts to them. Open the figure to inspect the full pipeline.

The training pool grows from 20 to 50 pairs over six rounds. A fixed 15-pair validation set selects harness checkpoints; HERA-Bench supplies the 60-pair evaluation benchmark.

// 03

One request.
Two environments.

HERA-Bench contains 60 verified task pairs: 60 feasible tasks and 60 infeasible counterparts.

Get HERA-Bench on Hugging Face
Includes tasks, executable environments, and evaluation tools. The repository is currently private.

Authors review the task requests, scenario plausibility, environment consistency, and pair validity. The feasible side has a validated solution; the infeasible side has an identified blocker and is checked for alternative completion paths.

The weather example below comes from Figures 7–8 of the paper. Only the availability of one forecast field changes.

Weather / Frozen snapshot / August 1, 2026

Should we postpone field work in Los Angeles?

Select the closest available weather station and summarize the official daytime forecast. Postpone if the temperature reaches 95°F, precipitation probability exceeds 40%, or winds may exceed 20 mph.

Closest stationSCE South Hills Park36.5 km from the weather point
Temperature93°FBelow 95°F
Precipitation probability0%Does not exceed 40%
Wind0–10 mphDoes not exceed 20 mph
Expected decision

Do not postpone.

None of the three postponement conditions is met. Report the station, forecast evidence, and supported decision.

A paper example, not a live weather forecast. In the infeasible instance, the precipitation field cannot be recovered through another public tool in the snapshot.

Act

Complete the feasible task.

Abstain

Correctly abstain on the infeasible task.

Pair

Get both instances of the same pair right.

// 04

Better decisions
on both sides

On HERA-Bench, co-evolution improves abstention and feasible-task completion relative to the base harness.

The same GPT-5.6-Luna model and evaluation tasks are used across the six settings below. Paired success measures whether an agent can both complete a feasible task and abstain on its matched infeasible counterpart.

HERA-Bench · 60 task pairs · Success rates (%)
MethodAbstain ↑Act ↑Pair ↑
Base harness (H₀)61.768.348.3
H₀ + abstention prompt68.363.346.7
Harness-only evolution76.771.760.0
Task-augmented harness evolution80.075.065.0
Task generation from H₀ failures78.375.073.3
HERA · Co-evolution83.376.770.0

Source: Table 1. Bold marks the best value in each column.

Results on the final training pool

On the cumulative 50-pair training pool, HERA reaches 86.0% Abstain, 88.0% Act, and 78.0% Pair. Task generation from H₀ failures reaches 76.0%, 82.0%, and 68.0%, respectively. These training-pool results are reported separately from HERA-Bench.

Sources: Tables 2 and 6.

// 05

One harness.
Many models.

The harness developed with GPT-5.6-Luna transfers to 19 other LLMs without model-specific optimization.

Across those 19 models, the paper reports an average gain of 15.3 percentage points in abstention success and 12.5 points in paired success. Both metrics improve for every evaluated transfer model, with varying gains.

View the complete results table
HERA-Bench · Base harness → HERA (%)
ModelAbstainPair
GPT-5.6-Terra76.7 → 86.766.7 → 73.3
GPT-5.6-Sol78.3 → 90.066.7 → 78.3
GPT-6-Astra81.7 → 91.770.0 → 80.0
Claude Opus 560.0 → 76.738.3 → 61.7
Claude Sonnet 540.0 → 60.033.3 → 41.7
Gemini 3.1 Pro (Preview)53.3 → 70.035.0 → 48.3
Gemini 3.8 Flash70.0 → 83.348.3 → 50.0
Gemini 3.5 Flash-Lite71.7 → 83.343.3 → 61.7
Grok 4.671.7 → 78.360.0 → 66.7
DeepSeek V4.1-Flash51.7 → 58.341.7 → 46.7
GLM-5.231.7 → 50.015.0 → 33.3
GLM-5.336.7 → 61.726.7 → 43.3
GLM-5.3-Flash43.3 → 56.731.7 → 43.3
Kimi K358.3 → 78.348.3 → 68.3
Qwen3.8 2.4T-A95B40.0 → 80.035.0 → 63.3
Qwen3.8 Flash53.3 → 70.041.7 → 58.3
MiniMax-M340.0 → 46.723.3 → 30.0
Mistral Medium 3.510.0 → 26.78.3 → 20.0
Mistral Large 36.7 → 16.73.3 → 5.0

Source: Table 3. These 19 models exclude GPT-5.6-Luna, the model used during harness development.

Cross-model evaluation / HERA-Bench

A stronger harness.
A wider reach.

Every transfer model, on the same 60 task pairs.

Models shown
19
Mean Abstain gain
+15.3pp
Mean Pair gain
+12.5pp
Base harness HERASuccess rates · 0–100% · Gains in percentage points

Showing all 19 transfer models.

Source: Table 3. GPT-5.6-Luna, the development model, is excluded. Summary gains describe the selected models and are computed from the displayed, rounded percentages.
Performance–cost trade-off

The same paired success.
At a lower estimated cost.

Both configurations reach 70.0% Pair on HERA-Bench.

GPT-6-Astra + base harness$0.0610
GPT-5.6-Luna + HERA$0.0091

Estimated API cost per episode (USD)

≈85% lower for this matched-performance comparison

Source: Section 5.2 and Figure 3. This compares two model–harness configurations; it is not a claim that HERA reduces the cost of every model. Across the 14 models with cost estimates, HERA increases per-model inference cost by a median factor of 2.09.

// 06

Read, explore,
and cite

BibTeX
@misc{luo2026heraharnessenvironmentcoevolutionreliable,
      title={HERA: Harness-Environment Co-Evolution for Reliable Agentic Abstention},
      author={Han Luo and Bingbing Wen and Guang Yang and Zora Zhiruo Wang and Pan Lu and Lucy Lu Wang},
      year={2026},
      eprint={2610.06563},
      archivePrefix={arXiv},
      primaryClass={cs.AI},
      url={https://arxiv.org/abs/2610.06563},
}