Evolve your
Codex
on
Terminal-Bench 2
在
Terminal-Bench 2
上进化
Codex
RSIHub is a framework focused on providing the environment for the RSI process, integrating multiple RSI methods and benchmarks. RSIHub 是一个专注于为 RSI 过程提供环境的框架,集成了多种 RSI 方法与 benchmark。
$ git clone https://github.com/simple-agent-lab/RSIHub.git
$ cd RSIHub && ./scripts/setup_terminal_bench.sh ahe
$ ./scripts/run_recipe_demo.sh ahe
gen 0 seed train 60.0% baseline tagged
gen 2 candidate train 68.0% gate ✓ admitted
gen 5 best train 74.0% archive.jsonl stamped
Seed and best train scores from the published Codex × AHE run; intermediate generations abridged. seed 和最佳 train 分数来自真实的 Codex × AHE 运行,中间几代有省略。
Results Benchmark 结果
Audited gains on real benchmarks. 在真实 benchmark 上的提升
Scores are percentages shown as seed → evolved agent. The train score is measured on the recipe’s training split; the full benchmark score is measured across the complete benchmark. All runs use a GPT-5.4-high target model and a GPT-5.4-xhigh Codex mutate operator. 分数格式为 seed → 进化后。Train 分数在 recipe 的训练集上得出,完整分数在整个 benchmark 上得出。所有实验的目标模型是 GPT-5.4-high,mutate 使用 GPT-5.4-xhigh 的 Codex。
Terminal-Bench 2
Tau³ Banking
Full results table 完整结果表格
| Benchmark Benchmark | Target agent 目标 agent | Method 方法 | Train score Train 分数 | Full benchmark score 完整 benchmark 分数 |
|---|---|---|---|---|
| Terminal-Bench 2 50 train / 19 gate / 20 sealed |
MiniSWE | AHE | 70.0% → 74.0% (+4.0%) |
55.1% → 56.2% (+1.1%) |
| Hyperagents | 58.0% → 68.0% (+10.0%) |
55.1% → 68.5% (+13.4%) |
||
| A-Evolve | 66.0% → 68.0% (+2.0%) |
55.1% → 65.2% (+10.1%) |
||
| GEPA | 58.0% → 68.0% (+10.0%) |
55.1% → 59.6% (+4.5%) |
||
| Codex | AHE | 60.0% → 74.0% (+14.0%) |
60.7% → 66.3% (+5.6%) |
|
| Hyperagents | 58.0% → 72.0% (+14.0%) |
60.7% → 70.8% (+10.1%) |
||
| A-Evolve | 62.0% → 62.0% (0.0%) |
60.7% → 64.0% (+3.3%) |
||
| GEPA | 64.0% → 64.0% (0.0%) |
60.7% → 60.7% (0.0%) |
||
| Tau³ Banking 50 train / 20 gate / 27 sealed |
MiniSWE | AHE | 34.0% → 36.0% (+2.0%) |
12.4% → 22.7% (+10.3%) |
| Hyperagents | 30.0% → 38.0% (+8.0%) |
12.4% → 28.9% (+16.5%) |
||
| A-Evolve | 30.0% → 34.0% (+4.0%) |
12.4% → 24.7% (+12.3%) |
||
| GEPA | 30.0% → 32.0% (+2.0%) |
12.4% → 22.7% (+10.3%) |
||
| Codex | AHE | 32.0% → 36.0% (+4.0%) |
11.3% → 26.8% (+15.5%) |
|
| Hyperagents | 34.0% → 36.0% (+2.0%) |
11.3% → 39.2% (+27.9%) |
||
| A-Evolve | 30.0% → 38.0% (+8.0%) |
11.3% → 13.4% (+2.1%) |
||
| GEPA | 10.0% → 16.0% (+6.0%) |
11.3% → 15.5% (+4.2%) |
How it works 工作原理
A typical iteration loop. 一个典型的迭代循环
A recipe decides how parents are selected, how traces are analyzed, what may be modified, and how evaluation runs. The framework carries those decisions: managing candidates, running evals, and archiving changes. Recipe 决定如何选择父代、如何分析轨迹、修改范围、以及如何eval。框架负责承载这些决策:管理 candidate、运行 eval、归档改动。
Five built-in recipes compose the loop for different goals: 内置的五个 recipe 各有侧重:
hill_climb
Improve one candidate from its current best parent. 在当前最优 candidate 上持续改进。
Classic loop 经典循环See the recipe guide for each strategy’s workflow and configuration. 每个策略的用法和配置见 recipe 指南。
Showcase
A Skill that learned poster design. 一个学会海报设计的 Skill。
RSIHub can improve a Skill as a complete package: instructions, references, and validation scripts evolve together while a frozen evaluator keeps the comparison honest. In this local Paper2Poster run, the same Codex model and paper prompt produced both LoRA posters below. RSIHub 可以把 Skill 当成一个整体来进化:指令、参考资料、验证脚本一起改,eval 保持冻结来保证对比公平。下面两张 LoRA 海报来自同一个 Codex 模型和同一条 prompt,只是 Skill 的代数不同。
Across the four-paper showcase, the deterministic completion pass rate moved from 1/4 at Gen 0 to 4/4 at Gen 2. This is a representative evolution run rather than a broad benchmark; see the result snapshot, frozen rubric, and minimal seed Skill. 四篇论文的测试中,通过率从 Gen 0 的 1/4 涨到 Gen 2 的 4/4。这只是一次示例运行,不是严格的 benchmark;结果快照、评分标准和种子 Skill 都在仓库里。
Build agents that improve —
and keep the evidence.
让智能体持续改进,
并保留证据。
Clone the repository and run your first generation in an afternoon. 克隆仓库,一个下午跑通第一代。
git clone https://github.com/simple-agent-lab/RSIHub.git