RSIHub

Evolve your Codex
on Terminal-Bench 2
在 Terminal-Bench 2
上进化 Codex

RSIHub is a framework focused on providing the environment for the RSI process, integrating multiple RSI methods and benchmarks. RSIHub 是一个专注于为 RSI 过程提供环境的框架,集成了多种 RSI 方法与 benchmark。

$ git clone https://github.com/simple-agent-lab/RSIHub.git
$ cd RSIHub && ./scripts/setup_terminal_bench.sh ahe
$ ./scripts/run_recipe_demo.sh ahe

gen 0  seed        train 60.0%   baseline tagged
gen 2  candidate   train 68.0%   gate ✓ admitted
gen 5  best        train 74.0%   archive.jsonl stamped

Seed and best train scores from the published Codex × AHE run; intermediate generations abridged. seed 和最佳 train 分数来自真实的 Codex × AHE 运行,中间几代有省略。

Results Benchmark 结果

Audited gains on real benchmarks. 在真实 benchmark 上的提升

Scores are percentages shown as seed → evolved agent. The train score is measured on the recipe’s training split; the full benchmark score is measured across the complete benchmark. All runs use a GPT-5.4-high target model and a GPT-5.4-xhigh Codex mutate operator. 分数格式为 seed → 进化后。Train 分数在 recipe 的训练集上得出,完整分数在整个 benchmark 上得出。所有实验的目标模型是 GPT-5.4-high,mutate 使用 GPT-5.4-xhigh 的 Codex。

Full results table 完整结果表格
Benchmark Benchmark Target agent 目标 agent Method 方法 Train score Train 分数 Full benchmark score 完整 benchmark 分数
Terminal-Bench 2
50 train / 19 gate / 20 sealed
MiniSWE AHE 70.0% → 74.0%
(+4.0%)
55.1% → 56.2%
(+1.1%)
Hyperagents 58.0% → 68.0%
(+10.0%)
55.1% → 68.5%
(+13.4%)
A-Evolve 66.0% → 68.0%
(+2.0%)
55.1% → 65.2%
(+10.1%)
GEPA 58.0% → 68.0%
(+10.0%)
55.1% → 59.6%
(+4.5%)
Codex AHE 60.0% → 74.0%
(+14.0%)
60.7% → 66.3%
(+5.6%)
Hyperagents 58.0% → 72.0%
(+14.0%)
60.7% → 70.8%
(+10.1%)
A-Evolve 62.0% → 62.0%
(0.0%)
60.7% → 64.0%
(+3.3%)
GEPA 64.0% → 64.0%
(0.0%)
60.7% → 60.7%
(0.0%)
Tau³ Banking
50 train / 20 gate / 27 sealed
MiniSWE AHE 34.0% → 36.0%
(+2.0%)
12.4% → 22.7%
(+10.3%)
Hyperagents 30.0% → 38.0%
(+8.0%)
12.4% → 28.9%
(+16.5%)
A-Evolve 30.0% → 34.0%
(+4.0%)
12.4% → 24.7%
(+12.3%)
GEPA 30.0% → 32.0%
(+2.0%)
12.4% → 22.7%
(+10.3%)
Codex AHE 32.0% → 36.0%
(+4.0%)
11.3% → 26.8%
(+15.5%)
Hyperagents 34.0% → 36.0%
(+2.0%)
11.3% → 39.2%
(+27.9%)
A-Evolve 30.0% → 38.0%
(+8.0%)
11.3% → 13.4%
(+2.1%)
GEPA 10.0% → 16.0%
(+6.0%)
11.3% → 15.5%
(+4.2%)

How it works 工作原理

A typical iteration loop. 一个典型的迭代循环

A recipe decides how parents are selected, how traces are analyzed, what may be modified, and how evaluation runs. The framework carries those decisions: managing candidates, running evals, and archiving changes. Recipe 决定如何选择父代、如何分析轨迹、修改范围、以及如何eval。框架负责承载这些决策:管理 candidate、运行 eval、归档改动。

Five built-in recipes compose the loop for different goals: 内置的五个 recipe 各有侧重:

hill_climb

Improve one candidate from its current best parent. 在当前最优 candidate 上持续改进。

Classic loop 经典循环
aevolve

Evolve prompts and reusable agent skills. 进化 prompt 和可复用的 skill。

Code代码
ahe

Engineer the agent harness against evaluator feedback. 根据 eval 反馈改进 agent harness。

Paper论文
gepa

Balance multiple objectives with minibatch validation. 用 minibatch 验证平衡多个目标。

Paper论文
hyperagents

Co-evolve the target and the selected evolution policy. 让目标和进化策略一起进化。

Paper论文

See the recipe guide for each strategy’s workflow and configuration. 每个策略的用法和配置见 recipe 指南。

Showcase

A Skill that learned poster design. 一个学会海报设计的 Skill。

RSIHub can improve a Skill as a complete package: instructions, references, and validation scripts evolve together while a frozen evaluator keeps the comparison honest. In this local Paper2Poster run, the same Codex model and paper prompt produced both LoRA posters below. RSIHub 可以把 Skill 当成一个整体来进化:指令、参考资料、验证脚本一起改,eval 保持冻结来保证对比公平。下面两张 LoRA 海报来自同一个 Codex 模型和同一条 prompt,只是 Skill 的代数不同。

Generation zero LoRA research poster with a generic dashboard-style layout
Gen 0 · minimal 12-line Skill Gen 0 · 12 行的最小 Skill Deterministic geometry gate failed: 14 text elements overflowed the SVG viewBox. 没通过几何检查:14 个文本元素溢出了 SVG viewBox。
Generation two LoRA research poster with a paper-specific editorial layout and low-rank matrix visualization
Gen 2 · evolved editorial Skill Gen 2 · 进化后的 Skill Passed deterministic renderability and geometry gates; paper fidelity remained advisory reviewer feedback. 渲染与几何检查全部通过;论文忠实度作为评审参考意见。

Across the four-paper showcase, the deterministic completion pass rate moved from 1/4 at Gen 0 to 4/4 at Gen 2. This is a representative evolution run rather than a broad benchmark; see the result snapshot, frozen rubric, and minimal seed Skill. 四篇论文的测试中,通过率从 Gen 0 的 1/4 涨到 Gen 2 的 4/4。这只是一次示例运行,不是严格的 benchmark;结果快照、评分标准和种子 Skill 都在仓库里。

Build agents that improve —
and keep the evidence.
让智能体持续改进,
并保留证据。

Clone the repository and run your first generation in an afternoon. 克隆仓库,一个下午跑通第一代。

git clone https://github.com/simple-agent-lab/RSIHub.git