Fine Japanese Calligraphy

The Art of Master Japanese Calligrapher Eri Takase

0

2026-10-04 - ScholarEvolve: Learning from Research — a Review

Notes on Learning from Research: Toward Lifelong Agent Harness Evolution (arXiv:2609.40169v1).

What the paper does

An AI agent works inside a harness: the software around the model that decides which tools it can call, what it remembers, what goes into its context, and how it plans. The model stays fixed; the harness is what gets improved. A recent approach (Meta-Harness, arXiv:2603.28052) has a coding agent read the agent's own failures and edit the harness. This paper, ScholarEvolve, adds a step. It turns the observed failures into research questions, searches arXiv, groups the returned papers into families of mechanisms, has a coding agent build several paper-derived candidates for each harness module (four per module on AppWorld, three on Telecom; App. A.1), searches over combinations of them, and keeps a new harness only if it beats the current best on a held-out validation set by a statistical margin.

The headline: on AppWorld, Qwen3.5-27B's task completion rises from 49.6% to 63.6% on the hard split, and on τ²-Bench Telecom, GPT-5.4-mini's single-try success rises from 72.7% to 81.9%. The authors also ran Meta-Harness themselves on the same tasks, and on these two results it reached 54.6% and 66.5%. On Telecom it ended below the starting harness. (How they ran it matters; see issue 4.)

What the paper does well

The evaluation discipline is careful. The paper keeps three separate task sets: one to learn from, one to choose with, and one to report on. It names related work that optimized and reported on the same tasks. A change must clear a statistical bar before it is kept. The authors report their failed candidates rather than hiding them: three of the twenty, tested alone, collapsed performance. They publish a detailed failure analysis and release code. The issues below concern what the evidence can carry, not whether the authors were careless.

Issues, most important first

1. The search itself is not reported as repeated. The standard deviations in the main results come from re-running a finished harness three times (§4.1: "We report means and standard deviations over three independent evaluation runs"). The paper does not report running either improvement process more than once. These processes are long chains of model calls, and the paper does not measure how much their outcome varies from one search to the next, so we cannot tell how much of the gap between ScholarEvolve and Meta-Harness would survive a second search.

2. In the one failure audit reported, most of the drop in failures is in answer delivery. The authors audited every failed episode of one model (Qwen3.5-27B) before and after. On the hard split, failures fall from 631 to 455. The category submitted an answer in the wrong form falls from 407 to 177. The category finished with an accepted answer but the task's goal was not met rises from 203 to 251 (App. A.5, Fig. 8). The authors say so themselves: "Evolution therefore improves answer delivery most strongly, while requirement tracking and evidence coverage remain recurring challenges" (§4.5). Part of that rise may come from the audit's ordering: once a format check passes, a later failure can get recorded instead. The paper says the change "combines changes in execution with changes in which failure is recorded first."

3. The comparison method is never audited the same way. The failure audit covers only ScholarEvolve's harness against the starting harness. We cannot tell whether Meta-Harness also fixed the answer-delivery problem, or whether that is where the two methods differ.

4. The comparison method is re-run with a different editing agent, not the published system. Here Meta-Harness uses GPT-5.4 through Codex as its editing agent, at the same budget as ScholarEvolve (§4.1). The Meta-Harness paper itself used Claude Code on Opus 4.6 (arXiv:2603.28052, "the proposer P is Claude Code [4] with Opus-4.6"). The weak showing (it drops GPT-5.4-mini below its starting point on Telecom: 72.7% → 66.5%) describes this re-run.

5. The part credited to "research" is not separated from the rest. The ablation removes components one at a time, in one fixed order (Table 2). Each step changes Qwen's score by 1.4 to 1.8 points, about the same size as the run-to-run spread reported for those rows. The research step is never removed on its own. Some of the modules built from papers are also simplified adaptations of them. The memory module credited to one paper ranks past episodes mainly by word overlap, described as "a lightweight version of the paper's memory lifecycle" (App. D.2). In our reading, the papers may supply directions to try as much as working methods.

6. "Research instead of failures" is not quite what was run. The framing contrasts learning from the literature with reacting to observed failures. In the experiments, the literature search starts from a failure audit of the training trajectories (§3.2.2: pool construction "begins with trajectories collected on D_evo"). The method is better described as failures translated into the literature's vocabulary, then searched. The "proactive" claim, that new papers can supply fixes for failures not yet seen, is not tested. The lifelong run shows newer papers adding gains while the archived training trajectories stay fixed (App. A.2), which is a real result, but those are the same failures throughout, and no agent is deployed to meet new ones.

7. The lifelong experiment changes more than the publication window. It covers one model on one benchmark split (Table 5). The coding agent changes from GPT-5.4 to GPT-5.5 for this run (App. A.1 vs A.2), and the three publication windows are "reconstructed from the current index." Whether the models doing the research already knew of work later than each window is not discussed. That is our question, not a finding.

8. Smaller points. - "Closing approximately 74% of the gap" to GPT-5.4 compares the evolved Qwen harness with GPT-5.4 in its unimproved starting harness, not with GPT-5.4 given the same treatment. - For hard-split tasks needing more APIs than any training task, the text says gains stay positive on average. Qwen's 95% interval there is [−0.7, 19.6] points (App. A.3), which includes zero. - The finding that combined modules beat their best single part (Table 3) is measured on the same 57 tasks used to choose the modules. - The three collapsed candidates (−43 to −68 points, Table 7) are never diagnosed. They show a built candidate can be badly harmful. They do not show whether the source paper's idea or its implementation was at fault. - Larger gains on the hard split are not by themselves evidence that a mechanism transfers to unseen tools: the hard split also starts lower, with more room to rise. - There is no Limitations section. The failure categories were assigned by a language model, in two passes by the same model. The paper reports no human agreement check.

What holds up

Within its setup, the best harness the paper found beats the starting harness on average on both models and both benchmarks, on held-out tasks, by wide margins except where the starting harness was already near the ceiling (Qwen on Telecom, 96.7% → 98.1%). More tasks become reliably solved: for Qwen, tasks solved in all three runs rise from 205 to 333 across both splits. Searching structured families of ideas, and testing each candidate before keeping it, is a sound design, and Table 7 shows that the testing step matters.

Bottom line

A careful systems paper whose framing is bigger than its evidence. It shows that a failure-seeded, literature-guided, carefully tested search beat one re-run of one baseline, with neither search reported as repeated, on two benchmarks; in the one model and benchmark audited, most of the drop in failures was in how answers were delivered. It does not yet show how much of that is due to the research itself, or that the approach adapts ahead of failures. A repeated search, a failure audit of the baseline, and an ablation that removes only the research step would address the three largest gaps.

What we took from it anyway

None of the issues above stopped us using the paper. We read papers for ideas we can use, and the useful idea is often not the paper's central claim: sometimes it is a detail in an appendix, and sometimes it is an idea the paper sets up and does not follow through. This paper gave us one of each, and one we already had.

Sources