Skip to content 跳至正文
Simple Agent Lab
Home 首页 GitHub

Contents 目录

  • Abstract 摘要
  • A long-horizon loop pointed at an AI system 面向 AI 系统的长程改进循环
  • AutoML, AI4AI, RSI: where the line falls AutoML、AI4AI 与 RSI 的分界
  • Model and harness 模型与 harness 两条路线
  • Why composition is less reliable 组件组合后的可靠性断层
  • Conclusion 结论
  • How to cite 如何引用
Survey overview: component benchmarks and environments on the left, six task pressures and the model-harness-environment-evaluation stack in the middle, model-side and harness-side methods on the right.

Survey · Preprints.org · August 2026 综述 · Preprints.org · 2026 年 8 月

AI4AI Survey: From Long-Horizon Agents to Recursive Self-Improvement

A survey of AI4AI: definitions, reliable horizons, and open problems. AI4AI 综述:核心定义、长程可靠性边界与关键开放问题。

Kai Wu1,*, Hao Lyu1,*, Zhen Luo1,*, Chaofan Wang2, Siyu Ye7, Jinghao Lin7, Xiaozhong Ji7, Boyuan Jiang7, Shengzhi Wang1, Zihan Wang7, Yiwen Ye7, Hao Wang7, Zimu Wang3, Wenzhe Liu7, Ruobing Wang7, Kai Cai7, Mingliang Xiong1, Wen Fang1, Mingqing Liu1, Yifan Zhang4, Lei Yang6, Xiaobin Hu5, Qingwen Liu1,†

1Tongji University · 2Shanghai Jiao Tong University · 3UC Berkeley · 4University of Chinese Academy of Sciences · 5National University of Singapore · 6Nanyang Technological University · 7Simple Agent Lab
*Equal contribution · †Corresponding author

Paper on Preprints.org 论文(Preprints.org) ↗ AI4AI harness: start your research here AI4AI 实验框架(RSIHub 开源代码) ↗ Short version (English) 英文短版 ↗

Abstract: Frontier releases including Claude Fable 5, GPT-5.6, Kimi K3, and GLM-5.3 are converging on sustained agentic work. AI agents can now run experiments, modify code and training pipelines, and iteratively improve AI artifacts, yet the relevant information is scattered across distinct concepts such as long-horizon agents, AI4AI, self-improvement, and recursive self-improvement. This survey asks how far an AI system can reliably carry an improvement process from idea to a validated result. Across model design, agent harnesses, benchmarks, automated research, and self-modifying systems, a consistent pattern emerges: AI is becoming increasingly capable at planning, coding, experimentation, optimization, and repair. Yet humans still largely determine the goals, evaluation criteria, and what counts as progress. Strong component performance also rarely translates into reliable end-to-end improvement; we call this the composition gap. Current systems can produce impressive improvements under bounded conditions, but evidence for reliable research judgment, causal experimentation, persistent gains, and compounding improvement remains limited. 摘要:以 Claude Fable 5、GPT-5.6、Kimi K3、GLM-5.3 为代表的前沿模型,正在展现出支撑长程连续作业的潜力。AI Agent 已经能够调度实验、修改代码并优化训练管线,但相关研究目前分散在长程智能体、AI4AI、自我改进与递归自改进(RSI)等割裂概念中。本综述围绕一个核心问题展开:AI 系统能否独立且可靠地将一个改进设想推进至真实验证?通过系统梳理模型设计、Agent harness、基准评测、自动化科研与自修改系统,我们归纳出一个清晰格局:AI 在具体研发任务中已展现出极高的执行力,涵盖实验规划、代码实现、流程调度、超参优化与故障修复;但核心目标设定、评测准则裁决以及“有效进步”的界定权,依然严格受控于人类设计者。更关键的是,单项基准的高分极难转化为端到端的可靠自改进能力,我们称之为 composition gap。现有系统仅能在严格受限的边界内取得局部增益,而关于可靠科研判断、严谨因果实验、增益长期存续以及复合式自我改进的实证依据依然非常有限。

The survey covers the component capabilities that AI4AI requires and how they are measured today, gives a taxonomy for locating where the loop closes, analyzes reliable execution within the coupled model, harness, environment, and evaluation stack, unfolds in depth along the model and harness routes, and ends with the failures that bound recursive improvement. 这篇综述梳理了 AI4AI 所需的基础能力及现有度量方式,给出一套判断系统自主程度的分类框架;把长程可靠执行拆解到模型、harness、运行环境、评测体系四层耦合结构中分析,深入对比模型与 harness 两条路线,并系统总结了阻碍系统实现真正递归改进的各类失效模式。

Alongside the survey, we also prepared an open-source framework for AI4AI experiments, RSIHub: build agents that improve, and keep the evidence. It is a ready starting point for your own research. 除了综述本身,我们还开源了配套的实验框架 RSIHub:让 Agent 在持续迭代中保留完整的验证证据链,方便大家直接上手开展自进化实验。

Outline 内容大纲

  1. AI4AI is a long-horizon loop of goal, plan, execute, feedback, and repair. Frontier models already handle each stage well enough on its own, so the pieces RSI needs are mostly in place and we are on the eve of AI4AI. The hard part is keeping the loop running for a long time without breaking. AI4AI 是一套面向 AI 系统本身的长程改进循环,涵盖目标设定、规划、执行、反馈、修复五个环节。前沿模型在单项原子任务上已相对成熟,真正的难点在于维持长程科研迭代链条的全局可靠性。
  2. Where the line falls between AutoML, AI4AI, and RSI. Which stages a system runs on its own. Of the 35 systems we audited, every one runs execute itself and none sets its own goal. AutoML、AI4AI 与 RSI 的分界。核心区分标准是五个环节中系统自己承担了哪几环。我们盘点了现有的 35 个系统,执行环节几乎全由系统自主跑完,但核心目标无一例外全是人类给的。
  3. Two routes to a longer horizon. The model lowers per-step error, the harness contains and recovers from errors. Among comparable models, changing the harness can matter more than changing the model. 延长执行步数的两条路线:模型与 harness。模型负责降低单步出错率,harness 负责隔离错误并自动恢复。在底座能力相当时,优化 harness 带来的提升往往比直接换模型更显著。
  4. Why composition is less reliable than its parts. Strong pieces, unreliable whole. The paper calls this the composition gap. Gains can be bought with budget, so a success rate means little without its cost. 组件拼装后的可靠性断层(composition gap)。单点能力都很强,但串起来就容易崩溃。此外,许多看似亮眼的改进本质是用算力刷出来的,抛开推理成本只谈成功率毫无意义。
  5. Conclusion. What AI4AI lacks is not capability. It is the authority to decide. 结论。AI4AI 现在的最大瓶颈在于决策能力:系统会做任务,但无法为自己决定方向。

AI4AI is a long-horizon loop pointed at an AI system AI4AI 是面向 AI 系统的长程改进循环

Each atomic skill is already strong on its own. The bottleneck is keeping the improvement loop running stably for a long time. AI4AI 的核心是将完整的科研实验循环作用于 AI 系统自身。单项研发能力已相对完备,核心瓶颈在于维持长程改进迭代的全局可靠性。

Our reading is that frontier models already handle today's long-horizon tasks well. Treat those abilities as atomic capabilities and, in terms of coverage, the pieces RSI needs are mostly on the table: planning, writing code, running experiments, analyzing results, repairing failures. Each alone is good enough. That is why we say we are on the eve of AI4AI. The bottleneck at this stage is chaining these actions and keeping the chain running over a long stretch. Whether a system can set its own goal is a later question, and the conclusion comes back to it. 从能力边界来看,前沿模型在单点工程与实验任务上已具备相当水平。自改进循环所依赖的规划推导、代码实现、实验调度、指标分析与故障修复等原子能力,单独评估时均已达到可用标准。然而,单点任务的成功并不等同于系统具备端到端的自进化能力。当前的核心瓶颈,在于将这些离散动作串联为长程实验流时,系统极易因状态漂移与误差累积而崩溃失效。至于系统能否自主确立具有科研价值的改进方向,则是尚未解决的更深层决策障碍。

Long-horizon is often equated with long context, many turns, or many tool calls. Those are only proxies. The survey builds its definition and metrics around four questions: what is improved, why long-horizon matters, how to measure it, and when the loop starts to recurse (RSI). 在现有讨论中,长程任务常被简化为上下文长度、交互轮数或工具调用频次,但这些代理指标并不能等同于真正的长程能力。本综述围绕四个核心问题建立形式化度量体系:改进对象的具体范畴、长程执行的必要性、任务压力的量化方式,以及循环在何种条件下真正介入自身的演化机制(RSI)。

Four organizing questions: what the loop may change, why a single pass is already long-horizon, how to measure task pressure, closure and evidence, and when self-reference turns the loop inward.
Our taxonomy of the long-horizon loop. 长程改进循环的四维分析坐标系。

What separates AutoML, AI4AI, and RSI: which stages the system runs itself AutoML、AI4AI 与 RSI 的分界:核心看系统接管了哪几环

Ask which of the five stages the system runs on its own. Today, no system picks its goal. 检验自进化程度的标准是看五环接管了多少。至少目前,没有一个系统能自主选定目标。

Of the taxonomy's five dimensions, closure does most of the discriminating work. It asks which stages of (goal, plan, execute, feedback, repair) the system owns. 在分析框架的五个维度中,区分度最高的是闭合度(closure):也就是在 (目标 goal, 规划 plan, 执行 execute, 反馈 feedback, 修复 repair) 五个阶段里,究竟哪些环节真正由系统独立接管。

Paradigm 范式 Target 改进对象 Self-reference 是否改进自身(自指) Typical closure 典型自主阶段
AutoML / NAS architectures and hyperparameters 模型架构与超参数 no 否 humans fix the goal and the search space; the system closes execute only 人类固定目标与搜索空间,系统仅负责执行搜索
AI4AI any part of an AI system AI 系统的任意模块(数据、训练、代码等) not required 不强制要求 the system closes plan and execute; feedback is mix under a human-written evaluator 系统接管规划与执行;反馈依赖人类编写的评测器(记为 mix)
Recursive self-improvement the improvement mechanism itself 改进机制本身(递归自增强) required 必须具备 all four execution stages close; the goal is still given by humans 规划、执行、反馈、修复均由系统接管;但底层目标仍由人类设定

From least to most recursive: 从低到高梳理当前的几种自演化形态:

  • Fixed harness (ReAct, Reflexion, MemGPT): a longer reliable horizon, but the agent mechanism itself is unchanged. 固定 harness 框架(如 ReAct、Reflexion、MemGPT):虽然延长了可靠执行步数,但 Agent 自身的机制没有发生变化。
  • Outer agent searches scaffold code (ADAS, Meta-Harness): the improver and the improved are still two roles. 外层 Agent 搜索优化脚手架代码(如 ADAS、Meta-Harness):改进者与被改进者仍旧是解耦的两个角色。
  • Agent edits its own code (Darwin Gödel Machine, Self-Harness): still judged by a fixed evaluator. Agent 自主修改自身代码(如 Darwin Gödel Machine、Self-Harness):系统改了自己,但裁判权仍依赖人类预先写死的评测器。
  • Only when a pass feeds into the process that proposes the next improvement goal does the loop become truly recursive. 只有当单次迭代结果能够反哺并重塑“下一个改进目标如何提出”这一过程时,才算真正触及递归自改进。

Two routes to a longer reliable horizon: model and harness 延长可靠执行步数的两条路线:模型与 harness

On AI4AI's long-horizon tasks, the reliable horizon is not set by the model alone: among comparable models, changing the harness can matter more than changing the model. 在长程 AI4AI 任务中,有效步数不单由模型决定:在能力相当的基座模型之间,优化 harness 带来的提升往往远超直接换模型。

Model-side interventions mapped onto the plan, execute, feedback, and repair stages of the loop.
Model-side interventions by stage: plan yields supervision, execute grounds in tools and state, feedback becomes step-level credit, repair rewrites the training recipe. 模型侧在各阶段的干预策略:规划阶段产出监督信号,执行阶段依托工具交互与状态维护,反馈阶段做单步归因评分,修复阶段重写训练方案。
  • Model route: lower the per-step error. Each of the four stages gets its own intervention (see the figure above). The limit: end-to-end reliability still decays as tasks lengthen. 模型路线:全力压缩单步出错概率。规划、执行、反馈、修复四个阶段各有对应的技术手段(见上图)。其物理极限在于:随着任务步数拉长,哪怕极低的单步失误率在端到端复合下依然会迅速累积至失效。
  • Harness route: contain errors and recover from them. The harness is the runtime outside the weights, and its leverage is larger than usually assumed: among comparable frontier models, changing the harness can reverse rankings, and on decomposable work, decomposition with per-step correction has run more than a million steps without a final error. The limit: tightly coupled tasks cannot be decomposed this way. harness 路线:主动隔离错误并动态纠偏。harness 是模型权重之外的外部运行时框架,它的实际杠杆远超很多人预期:在实力相近的前沿模型间,光是更换 harness 架构就能直接反转榜单排序;在高度可分解的任务中,结合逐级校验机制甚至跑出过百万步无致命崩溃的记录。它的边界在于:高度紧耦合、无法切分的长程逻辑很难套用这种模式。

Why composition is less reliable than its parts 单点能力拼装后为什么会失效(Composition Gap)

Strong parts, unreliable whole. Two mechanisms open the gap, and the failures show up in three places. 单模块评测高分不等于系统整体可用。局部误差会随步骤放大,而各环节交接时的信息损耗往往直接中断整个流程。

The paper calls this the composition gap. Two mechanisms produce it. First, a pass depends on every component: one failure terminates the whole pass, and even a local error changes all later state. Second, handoffs between components lose information: summarization can drop a constraint that only a later experiment needs. 我们把这种“散装能力极强、整机一跑就垮”的落差称为 composition gap。核心痛点通常来自两方面:首先是长链条的木桶效应,端到端迭代高度依赖每一步的确定性,任何一个环节掉链子都会直接导致流程中断,即便勉强往下跑,早期的隐蔽错误也会持续污染后续的状态;其次是各环节交接时的信息衰减,例如上一环节做上下文压缩摘要时,很容易抹掉后续实验中极为关键的细微边界约束。

Where the failures show up: 在实际工程中,这种失效集中体现在三个层面:

  • Across passes. Memory is state management, not a bigger context window. Stale state or a summarized-away constraint drifts the goal, and later observations hide the drift. Judge recovery as state repair, not as another attempt. 跨迭代的记忆退化。长程记忆从来不是单纯把上下文窗口拉大,而是严密的状态管理。过期残留的状态、或者在信息压缩中被漏掉的隐性约束,都会悄悄导致核心目标发生漂移,而后续轮次的新观测又会掩盖这种偏离。评估 Agent 的自我纠错能力,重点必须看它内部状态到底修复没有,而不是盲目重试碰巧蒙对。
  • Across generations. A system that both proposes and judges an improvement produces no evidence. Weak signals invite reward hacking, and searches rarely stop at their peak: in RSIBench-Data, more than three quarters that continue past the best checkpoint end below it. 跨代际的虚假繁荣。当一个系统自己提方案、又自己当裁判打分时,拿出的所谓改进结论是站不住脚的。微弱的评测信号极易引发对评测指标的过度迎合(reward hacking),而且搜索策略很难恰好刹停在全局最优:在 RSIBench-Data 真实测试中,越过最佳 checkpoint 继续自迭代的实验里,超过 75% 最终都会越改越差,回落到峰值之下。
  • Evidence and control. Gains can be bought with budget, so report a success rate together with its cost. A credible claim also needs model–harness attribution, contamination controls, and disclosed human oversight. 算力刷分与证据归因。很多性能提升本质是靠狂砸推理预算堆出来的,脱离了计算成本孤立谈成功率没有任何实际意义。严谨的自改进报告必须把模型自身能力与 harness 工程技巧剥离归因,严格控制测试数据污染,并清晰说明是否有隐性的人工介入。

Conclusion 结论

What AI4AI lacks is not capability. It is the authority to decide. AI4AI 现在的最大瓶颈在于决策能力:系统会做任务,但无法为自己决定方向。

Plan, execute, feedback, and repair are each already strong, and across the 35 audited systems execute is fully system-owned. No system picks its own goal, and none can judge what counts as “better” without a human. 从规划、执行到反馈、修复,现有前沿系统在各个分项上展现出的执行力已经毋庸置疑;在我们深入梳理的 35 个自改进系统中,执行环节几乎 100% 能够自主跑通。但目前没有一个系统能够自主决定下一步要研究什么目标,也没有任何一个系统能在脱离人类裁判的前提下判定“到底怎样才算变得更好”。

The eve ends the first time a system, with no human-specified goal or evaluation, has successors that reliably strengthen it. We hope this taxonomy and this audit help that moment arrive a little sooner. 什么时候一个系统能脱离人类预设的目标和裁判,稳定地把下一代改得更强,AI4AI 才算真正走通。我们梳理这套分类框架和 35 个系统,就是想把当前离散的研究认知拉平,提出通往递归自进化真正的可行研究方向。

If reading leaves you wanting to build: RSIHub is our open-source framework for self-evolving agents, and a place to start your own AI4AI research. Issues and pull requests are welcome. 如果你想亲自上手验证:我们搭建了开源实验框架 RSIHub,专为自进化 Agent 的实验设计与全证据链复现打造,你的 AI4AI 研究可以从这里起步,欢迎 issue 和 PR。

How to cite this AI4AI survey 如何引用这篇 AI4AI 综述

The survey is published on Preprints.org as doi:10.20944/preprints202608.2108.v1. If it helps your work, please cite it as: 本篇综述已在 Preprints.org 正式发布(doi:10.20944/preprints202608.2108.v1)。如果本项工作对你的研究有所启发,欢迎引用:

@misc{wu2026eveai4ai,
  title   = {{AI4AI} Survey: From Long-Horizon Agents to Recursive Self-Improvement---Definitions, Reliable Horizons, and Open Problems},
  author  = {Wu, Kai and Lyu, Hao and Luo, Zhen and Wang, Chaofan and
             Ye, Siyu and Lin, Jinghao and Ji, Xiaozhong and Jiang, Boyuan and
             Wang, Shengzhi and Wang, Zihan and Ye, Yiwen and Wang, Hao and
             Wang, Zimu and Liu, Wenzhe and Wang, Ruobing and Cai, Kai and
             Xiong, Mingliang and Fang, Wen and Liu, Mingqing and
             Zhang, Yifan and Yang, Lei and Hu, Xiaobin and Liu, Qingwen},
  howpublished = {Preprints.org},
  year    = {2026},
  doi     = {10.5281/zenodo.22198847},
  url     = {https://doi.org/10.5281/zenodo.22198847}
}
← Back to Simple Agent Lab ← 返回 Simple Agent Lab
Simple Agent Lab Disclaimer 免责声明 © 2026