If an agent repeats the same failure, should you replace the model first?

A coding agent fixes a failed test, then undoes the fix. A research agent looks up material it has already confirmed, and a document agent forgets conditions settled earlier by the end of a long task. When this keeps happening, attaching a more powerful model can seem like the quickest solution.

But if the cause is not a lack of knowledge or reasoning ability, but lost task state, no completion check, or gaps in post-failure recovery rules, changing the model will leave the same problems in place. What makes NVIDIA’s AVO case interesting is not Claude Opus 5 itself so much as the harness around it that keeps the model working through long tasks.

Still, you should not read the roughly 30 points and 100 points in the title as simple before-and-after scores. The two runs were not a controlled experiment that changed only the harness. The takeaway here is not a number called “a 70-percentage-point gain,” but a way to compare the model and execution system separately when evaluating an agent.

This was not a controlled experiment from 30 points to 100

AVO is not a new language model; it is an agent system for long-horizon autonomous work. A primary agent repeatedly investigates the situation, plans, implements, and evaluates, while persistent memory carries prior attempts and evaluation results into the next run. A separate supervisor intervenes when exploration stalls or falls into unproductive repetition.

25
public ARC-AGI-3 environments
183
total levels completed
6,624
actions taken in the environments

AVO, powered by Claude Opus 5, completed all 183 levels across the 25 public ARC-AGI-3 environments and recorded 100.00 RHAE. It took 6,624 actions in the environments. RHAE reflects both whether a level was completed and the efficiency of environment actions relative to the first-attempt human baseline. Internal reasoning and tool calls are not included in the action count unless they change the environment state.

The roughly 30% NVIDIA also mentioned came from a separate Claude Opus 5 evaluation by ARC Prize. Its reasoning settings, agent system, and evaluation setup differed from the AVO run. NVIDIA also noted that this difference does not directly measure AVO’s contribution alone. So you cannot calculate that “the harness produced exactly 70 percentage points” or that it “improved performance by 3.3 times.”

A score of 100 does not mean AGI has been achieved or guarantee a 100% success rate in real work.

The result came from a set with publicly available problems. The ARC Prize technical report explains that a harness tuned to public environments can overfit, so it does not view public-set scores as a valid measure of AGI progress. At the same time, it distinguishes this from the potential economic value of harness research for work automation.

What a harness manages is not an answer, but continuity of work

A harness does more than wrap a model in a long prompt. Between model calls, it manages what to remember, which tools to run, who judges the result, and whether to continue or stop after failure.

Recurring failure Harness element to check first Records to compare
Repeats research or edits already completed Persistent state that retains goals, confirmed facts, and prior attempts Duplicate tool calls and repeated-failure count
Declares an incorrect result complete External judges such as tests, schemas, and static analysis Validation failure rate after a completion declaration
Repeats only the same strategy after failure Failure records and recovery policy Post-failure recovery rate and number of human interventions
Loses the goal during a long run A supervisory layer that monitors progress and budget Goal drift, stop point, and cost per completed task

Other experiments also show that harness configuration can have a substantial effect. OpenAI reported that GPT-5.6 Sol scored 13.3% on the public set in its official harness, but reached 38.3% in a Responses API harness using both reasoning-state preservation and context compression, while output tokens fell to one-sixth. Because the two settings were applied together, their individual contributions cannot be separated, but the case shows that the same model’s results and cost can change depending on how execution state is handled.

There was no single right answer for input representation either. VISTA used 512×512 rendered images and circular visual memory that could be revisited with the same Claude Opus 5, reaching 100.00 across the 25 public environments. It took 7,542 actions. AVO reported 6,624 using 64×64 text grids, but because their backends and memory structures differ, this difference alone cannot establish which is better.

Today, divide failure cases into calibration and holdout sets

Your first move is not to build a new multi-agent system. Gather reproducible failure cases from past run logs and first divide them into a calibration set you will inspect while changing the harness and a holdout set you will not inspect until the changes are finished. Fixing one case repeatedly and then succeeding on it again is not evidence of generalization. It is the same issue behind ARC Prize’s warning that a harness tailored to public, observed environments may not transfer to unseen ones.

  1. Create a set of cases with the same work type and acceptance criteria. For payment-bug fixes, define completion conditions that apply to every case, such as passing regression tests, staying within the allowed changed-file scope, and zero duplicate charges. Separate calibration cases to use in harness design from holdout cases to use only in the final evaluation, and record which case belongs to which set.
  2. Run the current configuration as a baseline on the calibration set. Fix the model, reasoning settings, inputs, and tool permissions. If you can control variation such as seed or temperature, use the same values. If you cannot, run each configuration multiple times on the same cases so a single success or failure does not determine the conclusion.
  3. Create a baseline table with a visible denominator. Do not write only “60% completion rate”; also record the number of tasks and repetitions, such as “6 of 10 cases completed” or “9 of 15 completed after running 5 cases 3 times each.” Track repeated failures, post-failure recovery, human intervention, model-call volume, tokens, runtime, and cost in the same run unit. ARC-AGI-3 action efficiency does not include all internal reasoning and tool-call costs, so practical costs must be measured separately.
  4. Add persistent state one element at a time. Pass the goal, confirmed facts, attempted approaches, evaluation results, and remaining work into the next run, then re-evaluate the same calibration set under identical conditions. Next, add external judges, post-failure recovery policies, and a supervisory layer one at a time, comparing each with the immediately preceding configuration.
  5. Open the holdout set for the first time only after finalizing the configuration. Apply the best configuration from the calibration cases to the holdout cases without tuning it further. Evaluate the baseline under the same holdout set and run conditions, then compare completed cases, recovered cases, and costs. If you modify the harness again after seeing the holdout results, those cases are now calibration cases, so a new holdout set is needed for final validation.
  6. Replace only the model at the end. Change the model while keeping the finalized harness, the same case set, and the same permissions and acceptance criteria. That is how you can separate the effect of a model change from a harness change.

Calibration task state does not need to be elaborate. You can start by recording judgeable information and the data split like this.

{
  "task_id": "bugfix-017",
  "split": "calibration",
  "goal": "Fix retries after payment failures so no duplicate charge occurs",
  "expected_checks": [
    "Regression tests pass",
    "Changed-file scope is respected",
    "0 duplicate charges"
  ],
  "prior_attempts": [],
  "budget": {"max_model_calls": 20}
}

Success criteria need a denominator and holdout results.

For example, if you run 10 calibration cases three times each, record how many of the 30 total runs were completed and recovered. When you add one harness element, check whether those rates improve along with repeated failures, human intervention, tokens, time, and cost per completed task. Finally, you need to see changes in the same direction in holdout cases that were not inspected during design before you have grounds to expand to the next kind of work. If you have few cases, report the number of runs and observed outcomes as they are instead of claiming a definitive improvement rate.

Connecting tools does not make long-running work safe

Errors in long-running work can appear as accumulated state corruption rather than a single wrong answer. The DELEGATE-52 study evaluated 19 LLMs on long-form document editing across 52 specialized domains and reported that even frontier models corrupted an average of 25% of document content by the end of the task. Adding tool use alone did not improve performance, and corruption increased with document size, interaction length, and distracting files.

You cannot directly apply this result as an error rate for coding or game agents. But the practical warning is clear: do not treat a successful tool call as equivalent to completed work. An external judge must verify the result, and failure state must be recoverable or stoppable before it moves to the next stage.

Do not write only the model name on your next agent evaluation sheet. Include the memory used, completion judge, recovery policy, supervision conditions, and execution budget so you can compare the same system again. The most practical change NVIDIA’s 100 score demonstrates is not how to choose the top model, but how to broaden the unit of evaluation from the model to the full execution system.

If you want to dig deeper

NVIDIA AVO Reaches 100% on ARC-AGI-3 — You can review AVO’s structure, the public-set result, and the limits of comparing it with the roughly 30% baseline. developer.nvidia.com

ARC-AGI-3 Scoring Methodology — See how RHAE calculates completion and environment-action efficiency, and what it does not include in cost. docs.arcprize.org

How enabling two settings tripled our scores on the ARC-AGI-3 benchmark — Explains how scores and output tokens changed when reasoning-state preservation and context compression were applied together. openai.com

LLMs Corrupt Your Documents When You Delegate — Explore a counterexample where corruption accumulates in long document work and cannot be solved merely by connecting tools. arxiv.org