指南从已验证任务出发,让模型生成更难的任务、参考解和测试,经干净容器中的机械检查后再作为下一轮种子。内容包含设计、代码骨架和示例,并提出记录各难度成功率与候选拒绝率,尚未报告实测数据。
Guides
Glossary
RU
面向终端智能体的任务工厂:逐级验证,复现任务的递归扩展。
难度:高级 · 阅读时间:75 分钟 · 更新日期:2026 年 10 月 10 日
手工为终端智能体编写高难度、可检验的任务很费时间:每个任务都需要环境、参考解法和测试,而且三者必须相互一致。完全合成的任务成本低,但任务指令、测试和容器往往会逐渐偏离。这样生成的任务,要么无法完成,要么测试在没有实际依据的情况下也能通过。本指南采用一条折中的路线。你从一个已经验证过的任务开始,让模型提出一个更难的版本,编写新的解法和测试。生成结果必须在干净的容器中通过一组自动化检查关卡,才能被接受。被接受的任务随后成为下一级的种子。最后,你将测量智能体的成功率随难度级别提升下降了多少,以及在此过程中有多少候选任务被检查关卡拒绝。
本文包含什么、不包含什么:本文提供完整的设计、可运行的代码骨架,以及一个完整演示的示例任务,但不报告实测数据。我们没有为本文使用特定模型运行这条流水线,因此下文所有结果表格都是留给你填写的空白模板。示例中出现的任何数字都只是说明,不是实验结论。
要训练或评估终端智能体,你需要由三个相互一致的部分组成的任务:智能体阅读的任务指令、运行环境(包含文件和工具的容器镜像),以及自动检查器(智能体完成任务后运行的测试)。人能编写高质量任务,但数量有限。语言模型能编写大量任务,但它的错误往往出现在几个可预见的地方:测试检查了指令从未提出的要求,镜像缺少必需的工具,参考解法只是因为测试为空才“通过”,或者解法根本没有实际运行过。递归扩展之所以有帮助,是因为每个新任务都继承了一个已经能正常工作的基础。你只需要验证改动部分,这比验证一个全新的任务容易得多。
一种脚本可以检查的任务目录格式:包含任务指令、Dockerfile、参考解法、测试和需求映射。
一种脚本可以检查的任务目录格式:包含任务指令、Dockerfile、参考解法、测试和需求映射。
一个 extend.py 步骤:通过你配置的命令调用任意模型,要求它以严格的 JSON 格式生成一个更难的子任务。
一个 extend.py 步骤:通过你配置的命令调用任意模型,要求它以严格的 JSON 格式生成一个更难的子任务。
一个包含七道检查关卡的 validate.py 步骤。这些关卡在禁用网络的沙箱容器中运行。
一个包含七道检查关卡的 validate.py 步骤。这些关卡在禁用网络的沙箱容器中运行。
一个 grow.py 循环:构建深度为 0 → 1 → 2 → 3 的任务谱系,并记录每次拒绝及其原因。
一个 grow.py 循环:构建深度为 0 → 1 → 2 → 3 的任务谱系,并记录每次拒绝及其原因。
一个 measure.py 步骤:对每个任务运行你的智能体 N 次,报告各深度的成功率及其 Wilson 置信区间,并给出每条任务谱系的配对成功率降幅。
一个 measure.py 步骤:对每个任务运行你的智能体 N 次,报告各深度的成功率及其 Wilson 置信区间,并给出每条任务谱系的配对成功率降幅。
安装了 Docker(无根模式或标准模式)的 Linux 或 macOS、Python 3.11+,以及安装在宿主机上、用于本地语法检查的 pytest。
安装了 Docker(无根模式或标准模式)的 Linux 或 macOS、Python 3.11+,以及安装在宿主机上、用于本地语法检查的 pytest。
一个从 stdin 接收提示词、向 stdout 输出模型响应的命令。我们将它称为 $EXTENDER_CMD。它可以封装任意提供商的 CLI 或 SDK。本文不限定具体使用哪一种。
一个从 stdin 接收提示词、向 stdout 输出模型响应的命令。我们将它称为 $EXTENDER_CMD。它可以封装任意提供商的 CLI 或 SDK。本文不限定具体使用哪一种。
一个在任务容器中运行你的智能体的命令。我们将它称为 $AGENT_CMD。第 9 节会定义它的接口约定。
一个在任务容器中运行你的智能体的命令。我们将它称为 $AGENT_CMD。第 9 节会定义它的接口约定。
一份预算。每次扩展尝试都需要一次模型调用和若干次容器运行,而测量需要对每个任务执行 N 次智能体运行。在扩大规模之前,先估算这些成本。
一份预算。每次扩展尝试都需要一次模型调用和若干次容器运行,而测量需要对每个任务执行 N 次智能体运行。在扩大规模之前,先估算这些成本。
每个任务都是一个目录。智能体只能看到 instruction.md,以及 Dockerfile 放入镜像中的内容。测试和参考解法会在智能体完成任务后才挂载进去,绝不会预先打包进镜像。
tasks/
logreport-d0/
task.json # id, parent, depth, requirement ids, superseded tests
instruction.md # what the agent reads
Dockerfile # environment; must not COPY tests/ or solution.sh
env/ # files copied into the image (fixtures)
solution.sh # reference ("oracle") solution
tests/
test_outputs.py # pytest; each test tagged with a requirement id
种子任务的 task.json:
{
"id": "logreport-d0",
"parent": null,
"depth": 0,
"requirements": {
"R1": "Read every *.log file in /app/logs",
"R2": "Write /app/report.json mapping HTTP status code (string) to count (int)"
},
"superseded_tests": []
}
需求映射有实际作用。检查关卡 G5 用它来检查任务指令和测试是否描述了同一件事,而这恰恰是合成任务容易出现偏差的地方。
我们特意使用一个小型种子任务,确保每一级任务都能由人读懂。这个案例的重点是机制,种子任务本身的难度在这里并不重要。
instruction.md(深度 0)
Access logs in combined format are in /app/logs/*.log.
Produce /app/report.json: a JSON object whose keys are HTTP status
codes as strings and whose values are the number of requests with
that status, summed across all files. [R1] [R2]
FROM python:3.12-slim@sha256:<pin-a-digest-you-have-pulled>
RUN pip install --no-cache-dir pytest==8.3.3
WORKDIR /app
COPY env/logs /app/logs
使用摘要固定基础镜像。否则,多次运行之间,“干净环境”会悄悄发生变化,你观察到的成功率随深度下降,实际上可能是镜像变化造成的。
#!/usr/bin/env bash
set -euo pipefail
python3 - <<'PY'
import json, glob, re, collections
pat = re.compile(r'"\S+ \S+ \S+" (\d{3}) ')
c = collections.Counter()
for f in sorted(glob.glob("/app/logs/*.log")):
with open(f, encoding="utf-8", errors="replace") as fh:
for line in fh:
m = pat.search(line)
if m:
c[m.group(1)] += 1
json.dump(dict(c), open("/app/report.json", "w"), sort_keys=True)
PY
tests/test_outputs.py
import json, pathlib, pytest
REPORT = pathlib.Path("/app/report.json")
@pytest.mark.req("R2")
def test_report_exists_and_is_object():
data = json.loads(REPORT.read_text())
assert isinstance(data, dict)
assert all(isinstance(k, str) and isinstance(v, int) for k, v in data.items())
@pytest.mark.req("R1", "R2")
def test_counts_match_fixture():
# Expected counts are computed once from env/logs by a human and frozen here.
expected = {"200": 7, "301": 1, "404": 2, "500": 1}
assert json.loads(REPORT.read_text()) == expected
上面的预期计数对应的是我们的测试数据。请根据你自己的 env/logs 计算计数,并将它们固定在测试文件中。如果预期值是在测试运行时用与解法相同的逻辑计算出来的,那么测试只能证明代码与自身一致。
在 tests/conftest.py 中注册标记:
def pytest_configure(config):
config.addinivalue_line("markers", "req(*ids): requirement ids covered by the test")
继续之前,请使用之后验证子任务时所用的同一个验证器(第 7 节),亲自检查种子任务。如果种子任务无法通过自身的检查关卡,它就会把缺陷传递给由它衍生出的每个任务。
扩展器会收到完整的父任务,并且必须以 JSON 格式返回一个子任务。子任务增加一项新能力,并保留父任务的每项需求,除非它明确声明替代其中某项。每一步都尽量减少改动,这样,当子任务未能通过某道检查关卡时,你就能知道是哪项新增内容导致的。
You are extending a verified terminal task. You receive the parent
instruction, Dockerfile, env file listing, reference solution, tests
and requirement map.
生成一个难度更高的子任务,要求:
只返回 JSON,包含以下键: instruction_md、dockerfile、env_files(路径 -> 内容,仅限文本)、 solution_sh、tests_py、requirements(ID -> 文本)、 superseded_tests(测试名称列表)、rationale。
import json, os, pathlib, shlex, subprocess, sys
def read_task(d: pathlib.Path) -> dict: env = {str(p.relative_to(d)): p.read_text(errors="replace") for p in (d / "env").rglob("*") if p.is_file()} return { "task": json.loads((d / "task.json").read_text()), "instruction_md": (d / "instruction.md").read_text(), "dockerfile": (d / "Dockerfile").read_text(), "env_files": env, "solution_sh": (d / "solution.sh").read_text(), "tests_py": (d / "tests" / "test_outputs.py").read_text(), }
def extend(parent: pathlib.Path, child: pathlib.Path) -> None: prompt = pathlib.Path("prompts/extend.txt").read_text() payload = prompt + "\n\nPARENT:\n" + json.dumps(read_task(parent), indent=1) out = subprocess.run(shlex.split(os.environ["EXTENDER_CMD"]), input=payload, capture_output=True, text=True, timeout=900, check=True).stdout start, end = out.find("{"), out.rfind("}") spec = json.loads(out[start:end + 1])
p = json.loads((parent / "task.json").read_text()) child.mkdir(parents=True) (child / "tests").mkdir() for rel, content in spec["env_files"].items(): dst = child / rel if not dst.resolve().is_relative_to(child.resolve()): raise ValueError(f"env path escapes task dir: {rel}") dst.parent.mkdir(parents=True, exist_ok=True) dst.write_text(content) (child / "instruction.md").write_text(spec["instruction_md"]) (child / "Dockerfile").write_text(spec["dockerfile"]) (child / "solution.sh").write_text(spec["solution_sh"]) (child / "tests" / "test_outputs.py").write_text(spec["tests_py"]) (child / "tests" / "conftest.py").write_text( (parent / "tests" / "conftest.py").read_text()) (child / "task.json").write_text(json.dumps({ "id": child.name, "parent": p["id"], "depth": p["depth"] + 1, "requirements": spec["requirements"], "superseded_tests": spec.get("superseded_tests", []), "rationale": spec.get("rationale", ""), }, indent=2))
if name == "main": extend(pathlib.Path(sys.argv[1]), pathlib.Path(sys.argv[2]))
注意:环境文件由模型重写。如果父任务的测试数据文件较大或为二进制文件,应直接从父任务复制,只让模型添加新文件。否则,即使明确要求模型保留这些测试数据,它仍可能悄悄修改它们。
只有通过所有关卡,任务扩展才会被接受。每个关卡都在根据子任务自身的 Dockerfile 构建的全新容器中运行,并使用 `--network none`。如果某个关卡需要联网才能通过,那说明任务本身有问题,而不是这次运行运气不好。
关卡检查捕获的缺陷
G0 BuildImage builds with --no-cache; Dockerfile does not COPY tests/ or solution.shMissing tools, unpinned installs, test leakage into the image G1 NoveltyParent solution fails the child tests"Extensions" that add nothing G2 SolvableChild solution passes the child tests in 3 of 3 fresh runsUnsolvable tasks, flaky tests G3 RegressionInherited parent tests (minus superseded ones) pass against the child solution in the child imageSilent loss of earlier requirements G4 Non-vacuousA no-op solution (true) fails the child testsTests that pass on an untouched environment G5 AlignmentEvery requirement id appears in the instruction and in at least one test marker; every test has a marker; no unknown idsInstruction–test drift G6 LeakageInstruction does not contain expected literal values from the tests; image filesystem contains no test fileTasks answerable by copying
G1 和 G4 是最常被遗漏的关卡,而它们能捕获合成任务中最常见的缺陷:测试无法真正区分正确和错误的解决方案。
import ast, hashlib, json, pathlib, re, shutil, subprocess, sys, tempfile
DOCKER_RUN = ["docker", "run", "--rm", "--network", "none", "--memory", "2g", "--cpus", "2", "--pids-limit", "512"]
def image_tag(task: pathlib.Path) -> str: h = hashlib.sha256() for p in sorted([task / "Dockerfile", (task / "env").rglob("")]): if p.is_file(): h.update(str(p.relative_to(task)).encode()); h.update(p.read_bytes()) return f"taskfactory/{task.name}:{h.hexdigest()[:12]}"
def build(task): df = (task / "Dockerfile").read_text() if re.search(r"^\s*(COPY|ADD)\s+.*(tests|solution.sh)", df, re.M): raise GateError("G0", "Dockerfile copies tests or solution") tag = image_tag(task) r = subprocess.run(["docker", "build", "--no-cache", "-q", "-t", tag, str(task)], capture_output=True, text=True, timeout=1800) if r.returncode: raise GateError("G0", r.stderr[-2000:]) return tag
def run(tag, solution: pathlib.Path, tests_dir: pathlib.Path, deselect=()): """Run solution, then tests, in one fresh container. Returns (passed, junit_text).""" with tempfile.TemporaryDirectory() as tmp: out = pathlib.Path(tmp) args = " ".join(f"--deselect /tests/test_outputs.py::{t}" for t in deselect) script = (f"bash /solution.sh >/out/solution.log 2>&1; " f"pytest -q -p no:cacheprovider {args} --junitxml=/out/junit.xml /tests " f">/out/pytest.log 2>&1") r = subprocess.run(DOCKER_RUN + [ "-v", f"{solution.resolve()}:/solution.sh:ro", "-v", f"{tests_dir.resolve()}:/tests:ro", "-v", f"{out}:/out", tag, "bash", "-c", script], capture_output=True, text=True, timeout=1200) junit = (out / "junit.xml").read_text() if (out / "junit.xml").exists() else "" return r.returncode == 0, junit
class GateError(Exception): def init(self, gate, detail): super().init(f"{gate}: {detail}"); self.gate = gate
def markers(tests_py: str): """Map test function name -> set of requirement ids from @pytest.mark.req(...).""" tree, found = ast.parse(tests_py), {} for node in ast.walk(tree): if isinstance(node, ast.FunctionDef) and node.name.startswith("test_"): ids = set() for d in node.decorator_list: if (isinstance(d, ast.Call) and getattr(d.func, "attr", "") == "req"): ids |= {a.value for a in d.args if isinstance(a, ast.Constant)} found[node.name] = ids return found
def validate(child: pathlib.Path, parent: pathlib.Path | None) -> dict: meta = json.loads((child / "task.json").read_text()) tests_py = (child / "tests" / "test_outputs.py").read_text() instr = (child / "instruction.md").read_text() report = {"task": meta["id"], "gates": {}}
req = set(meta["requirements"])
tmap = markers(tests_py)
in_instr = set(re.findall(r"\[(R\d+)\]", instr))
tagged = set().union(*tmap.values()) if tmap else set()
problems = []
if req - in_instr: problems.append(f"not in instruction: {sorted(req - in_instr)}")
if req - tagged: problems.append(f"no test: {sorted(req - tagged)}")
if tagged - req: problems.append(f"unknown ids in tests: {sorted(tagged - req)}")
if [t for t, ids in tmap.items() if not ids]: problems.append("untagged tests")
if problems: raise GateError("G5", "; ".join(problems))
report["gates"]["G5"] = "pass"
lits = {n.value for n in ast.walk(ast.parse(tests_py))
if isinstance(n, ast.Constant) and isinstance(n.value, str) and len(n.value) >= 12}
leaked = [l for l in lits if l in instr and not l.startswith("/")]
if leaked: raise GateError("G6", f"instruction contains test literals: {leaked[:3]}")
tag = build(child); report["gates"]["G0"] = "
关于这段代码,需要了解以下几点:
G3 会在子任务的镜像中运行父任务的测试。当子任务修改测试夹具时,父任务中某些使用固定值的测试确实会变得不再正确。superseded_tests 就是为此准备的,但你需要亲自阅读其中的每一项,因为模型也可能借此删除那些仅仅让它觉得麻烦的测试。第 10 节建议为此设置一个上限。
G3 会在子任务的镜像中运行父任务的测试。当子任务修改测试夹具时,父任务中某些使用固定值的测试确实会变得不再正确。superseded_tests 就是为此准备的,但你需要亲自阅读其中的每一项,因为模型也可能借此删除那些仅仅让它觉得麻烦的测试。第 10 节建议为此设置一个上限。
针对字面量的泄漏检查是一种启发式方法。它会漏掉数字和短字符串,也可能误报合法的文件路径,因此检查时排除了路径。把它当作预警机制,而不是证据。
针对字面量的泄漏检查是一种启发式方法。它会漏掉数字和短字符串,也可能误报合法的文件路径,因此检查时排除了路径。把它当作预警机制,而不是证据。
连续 3 次全部通过的规则只能发现明显的不稳定性。对于涉及时序、并发或随机性的任务,应增加运行次数。
连续 3 次全部通过的规则只能发现明显的不稳定性。对于涉及时序、并发或随机性的任务,应增加运行次数。
## 8. 递归增长循环
每个被接受的子任务都会成为下一次尝试的父任务。被拒绝的子任务会被记录,然后重试,每一层的尝试次数有固定上限。如果某一层耗尽了尝试次数,这条演化链就到此为止。不能从未经验证的任务直接跳到下一层。
import json, pathlib, shutil, subprocess, sys, time from extend import extend from validate import validate, GateError
def grow(seed: pathlib.Path, max_depth=3, attempts=4, log="runs/grow.jsonl"): pathlib.Path(log).parent.mkdir(exist_ok=True) parent = seed for depth in range(1, max_depth + 1): accepted = None for a in range(attempts): child = seed.parent / f"{seed.name.rsplit('-d', 1)[0]}-d{depth}-a{a}" shutil.rmtree(child, ignore_errors=True) rec = {"ts": time.time(), "parent": parent.name, "child": child.name, "depth": depth, "attempt": a} try: extend(parent, child) rec |= validate(child, parent); rec["status"] = "accepted" accepted = child except GateError as e: rec |= {"status": "rejected", "gate": e.gate, "detail": str(e)[:500]} except Exception as e: # malformed JSON, timeouts, path errors rec |= {"status": "error", "detail": repr(e)[:500]} with open(log, "a") as fh: fh.write(json.dumps(rec) + "\n") if accepted: break if not accepted: print(f"lineage stopped at depth {depth}"); return parent = accepted
if name == "main": grow(pathlib.Path(sys.argv[1]))
用多个种子任务运行,而不是只用一个。单条演化链只能反映模型的一条决策链。要对“递归扩展”得出具有普遍意义的结论,你需要多条独立的演化链,每条都有自己的种子任务。
export EXTENDER_CMD="your-model-cli --max-tokens 16000"
python validate.py tasks/logreport-d0 # the seed must pass its own gates
for s in tasks/*-d0; do python grow.py "$s"; done
jq -r 'select(.status!="accepted") | [.depth,.gate // "error"] | @tsv' runs/grow.jsonl
| sort | uniq -c
### 我们的种子任务可能形成怎样的演化链(示例)
下面展示了模型可能提出的一种扩展链。这是我们自己编写的示例,并非实际运行的输出:
d1:轮转日志 access.log.1.gz、access.log.2.gz 也必须计入统计 [R3]。测试夹具中新增 gzip 文件,父任务的解法会少算,因此 G1 会按预期触发。
d1:轮转日志 access.log.1.gz、access.log.2.gz 也必须计入统计 [R3]。测试夹具中新增 gzip 文件,父任务的解法会少算,因此 G1 会按预期触发。
d2:排除来自 /app/config/healthcheck_ips.txt 中所列 IP 的请求(每行一个 IP,允许使用 # 注释)[R4]。父任务的测试 test_counts_match_fixture 会被取代,因为其中固定的计数值发生了变化。
d2:排除来自 /app/config/healthcheck_ips.txt 中所列 IP 的请求(每行一个 IP,允许使用 # 注释)[R4]。父任务的测试 test_counts_match_fixture 会被取代,因为其中固定的计数值发生了变化。
d3:还要写入 /app/report_hourly.csv,表头为 hour_utc,status,count,将带有不同偏移量的时间戳统一转换为 UTC,并对行进行排序 [R5]。正确的解法需要同时处理时区、CSV 格式以及之前的所有过滤条件。
d3:还要写入 /app/report_hourly.csv,表头为 hour_utc,status,count,将带有不同偏移量的时间戳统一转换为 UTC,并对行进行排序 [R5]。正确的解法需要同时处理时区、CSV 格式以及之前的所有过滤条件。
注意 d2 发生的变化。取代原测试是合理的,但替代测试必须使用新的固定值重新覆盖 R1 和 R2。否则,G5 仍然可能通过(R1 可能由另一个测试覆盖),而部分基础行为却不再受到检查。请将被取代的测试列表与新的测试集放在一起审查。
## 9. 衡量成功率下降
$AGENT_CMD 接收镜像标签和指令文件路径。它必须基于该镜像启动一个容器(网络策略由你决定,但要记录下来),让智能体执行任务,然后将容器最终的文件系统保存为一个新镜像,以便对其运行测试。使用 docker commit 是一种简单且如实保留执行结果的方式:
#!/usr/bin/env bash
set -euo pipefail cid=$(docker run -d --network none "$1" sleep infinity) your-agent --container "$cid" --instruction "$2" --max-steps 60 --timeout 1800 || true docker commit "$cid" "$3" >/dev/null docker rm -f "$cid" >/dev/null
随后,使用一个不执行任何操作的解法,在 OUT_TAG 上运行测试。智能体完成的工作就是解法。在各个深度保持相同的步骤预算、超时时间和智能体模型。如果你在不同深度之间改变其中任何一项,衡量的就是这项变化,而不是任务难度。
import json, math, pathlib, subprocess, sys, collections from validate import run, image_tag
NOOP = pathlib.Path("noop.sh"); NOOP.write_text("#!/usr/bin/env bash\ntrue\n")
def wilson(k, n, z=1.96): if n == 0: return (float("nan"),) * 3 p = k / n; d = 1 + zz/n c = (p + zz/(2n)) / d h = z * math.sqrt(p(1-p)/n + zz/(4n*n)) / d return p, max(0, c - h), min(1, c + h)
def episode(task, i): tag = image_tag(task); out = f"{tag}-ep{i}".replace(":", "-") subpro