文章要求在属性规范、测试夹具、随机种子调度和失败特征一致时,对比父提交与补丁的失败结果,才允许提出测试波动冻结方案。父提交通过而补丁失败必须判为回归,反复重跑补丁不能改变这一分类。
Agent 提交的补丁,并不会因为新代码树中的某项 property 连续失败两次,就被认定为 flaky。只有父提交在同一份 contract 下也未通过同一项 property,才允许提出 flake-freeze 提案。如果父提交通过、补丁失败,就应归类为 regression。无论再对补丁重跑多少次,都不会改变这个分类。
这道 gate 与 fixture 锁定、输入缩减和断言豁免规则并行使用,并不取代它们。它回答的是那些检查尚未解决的一个问题:这个失败在父提交中是否已经存在?
只有以下七项条件全部成立,才能写入提案文件。
缺少任何一项,都必须 reject。不存在“差不多符合”的情况。
两份日志如果没有共享的 contract,就无法判定。把 contract 放进评审目录,忽略所有引用其他 hash 的运行结果。
contract_version: 1
parent_sha: "abc123"
patch_sha: "def456"
spec_hash: "spec-9f2c"
fixture_closure_hash: "fix-44ab"
seeds: [11, 29, 47]
properties:
- id: prop.order_total_non_negative
oracle: "total >= 0"
这些 hash 是前序检查提供的输入,本文不重新定义它们。如果 Agent 修改了 property 文本或 fixture 集合,hash 就会不一致,这道 gate 随即停止。Spec review 走的是另一条评审队列。
在不同的 worktree 中运行父提交和补丁。凭记忆还原父提交的结果,不能算作 baseline。
git worktree add ../parent-tree "$PARENT_SHA"
git worktree add ../patch-tree "$PATCH_SHA"
缺少一行,意味着 contract 的执行不完整,而不是通过。status 只能是 pass 或 fail。
python3 run_properties.py \
--tree ../parent-tree \
--contract differential-contract.yaml \
--out parent.jsonl
python3 run_properties.py \
--tree ../patch-tree \
--contract differential-contract.yaml \
--out patch.jsonl
把每个 fail 映射到你已经信任的 witness 类别:property id、异常类型,以及稳定的断言 id。输入最小化交给 shrinking gate。如果类别不同,即使两边都失败,也说明补丁引入了新的失败模式。
判定器读取两份 JSONL 文件和 seed 列表,写出一个决策对象。它不会编辑测试,也不会删除失败的断言。
propose_freeze 是一张评审工单。由人工补上有效期和原因代码。有效期一到,就删除工单。下一个补丁必须重新满足条件,才能获得新工单。
把这个文件直接沿用到后续补丁,是流程错误,不是续期。
seed 结果出现分化,意味着至少一侧的调度结果中,该 property 同时出现 pass 和 fail。全部失败并不是 flaky 的证据,而是稳定失败;这条流程不会豁免稳定失败。
只要出现一项 regression,就拒绝整个补丁,即使另一项 property 看起来属于双方共有的 flaky 失败。合并的单位是整个补丁。如果允许部分通过,真正的破坏就可能搭着一个无关的噪声检查一起混进来。
这个模块是一份参考实现。信任它之前,请先自己运行。它不是生产环境统计报告,也不会调用远程模型。
#!/usr/bin/env python3
"""Score parent vs patch property runs. Unexecuted until you run it."""
from __future__ import annotations
import json
from pathlib import Path
from typing import Iterable
def load_jsonl(path: Path) -> list[dict]:
rows = []
for line in path.read_text().splitlines():
if line.strip():
rows.append(json.loads(line))
return rows
def split(statuses: Iterable[str]) -> bool:
seen = set(statuses)
return "pass" in seen and "fail" in seen
def decide(parent: list[dict], patch: list[dict], seeds: list[int]) -> dict:
if not parent or not patch:
return {"decision": "reject", "reason": "incomplete_contract", "results": []}
def index(rows: list[dict]) -> dict:
return {(r["prop"], int(r["seed"])): r for r in rows}
pmap, cmap = index(parent), index(patch)
props = sorted({r["prop"] for r in parent} | {r["prop"] for r in patch})
results = []
for prop in props:
p_rows = [pmap.get((prop, s)) for s in seeds]
c_rows = [cmap.get((prop, s)) for s in seeds]
if any(r is None for r in p_rows + c_rows):
results.append({"prop": prop, "decision": "reject", "reason": "missing_seed"})
continue
def side_status(rows: list[dict]) -> str:
return "fail" if any(r["status"] == "fail" for r in rows) else "pass"
def witnesses(rows: list[dict]) -> list[str]:
return sorted({r.get("witness", "") for r in rows if r["status"] == "fail"})
p_status, c_status = side_status(p_rows), side_status(c_rows)
p_wit, c_wit = witnesses(p_rows), witnesses(c_rows)
if p_status == "pass" and c_status == "pass":
item = {"prop": prop, "decision": "accept", "reason": "both_pass"}
elif p_status == "pass" and c_status == "fail":
item = {"prop": prop, "decision": "reject", "reason": "regression"}
elif p_status == "fail" and c_status == "pass":
item = {"prop": prop, "decision": "accept", "reason": "parent_fail_cleared"}
elif p_wit != c_wit:
item = {"prop": prop, "decision": "reject", "reason": "witness_mismatch"}
elif not (split(r["status"] for r in p_rows) or split(r["status"] for r in c_rows)):
item = {"prop": prop, "decision": "reject", "reason": "uniform_fail"}
else:
item = {"prop": prop, "decision": "propose_freeze", "reason": "shared_variant_fail"}
results.append(item)
if any(r["decision"] == "reject" for r in results):
return {"decision": "reject", "results": results}
if any(r["decision"] == "propose_freeze" for r in results):
return {"decision": "propose_freeze", "results": results}
return {"decision": "accept", "results": results}
def main() -> None:
parent = load_jsonl(Path("parent.jsonl"))
patch = load_jsonl(Path("patch.jsonl"))
contract = json.loads(Path("differential-contract.json").read_text())
print(json.dumps(decide(parent, patch, contract["seeds"]), indent=2))
if __name__ == "__main__":
main()
regression 示例。父提交的记录显示通过,补丁的记录显示失败。按照代码中的分支,这组结果会被判为 regression。
printf '%s\n' '{"prop":"prop.order_total_non_negative","seed":11,"status":"pass","witness":""}' > parent.jsonl
printf '%s\n' '{"prop":"prop.order_total_non_negative","seed":11,"status":"fail","witness":"AssertionError:total"}' > patch.jsonl
printf '%s\n' '{"seeds":[11]}' > differential-contract.json
python3 differential_gate.py
双方共有的变动性失败示例。seed 11 在两边都失败,且 witness 相同。seed 29 在父提交中通过,在补丁中失败。这样就出现了结果分化,因此函数可以输出 propose_freeze。这个文件仍然只是一份提案。
cat > parent.jsonl <<'EOF'
{"prop":"prop.order_total_non_negative","seed":11,"status":"fail","witness":"AssertionError:total"}
{"prop":"prop.order_total_non_negative","seed":29,"status":"pass","witness":""}
EOF
cat > patch.jsonl <<'EOF'
{"prop":"prop.order_total_non_negative","seed":11,"status":"fail","witness":"AssertionError:total"}
{"prop":"prop.order_total_non_negative","seed":29,"status":"fail","witness":"AssertionError:total"}
EOF
printf '%s\n' '{"seeds":[11,29]}' > differential-contract.json
python3 differential_gate.py
先数清记录行数,再争论结果是绿还是红。评审说明如果拿不出这些计数,就不完整。
python3 - <<'PY'
import json
from collections import Counter
from pathlib import Path
for name in ("parent.jsonl", "patch.jsonl"):
c = Counter(
json.loads(line)["status"]
for line in Path(name).read_text().splitlines()
if line.strip()
)
print(name, dict(c), "lines", sum(c.values()))
PY
公布对象数量,不要公布你没有测量过的 flaky 比例。两个 seed 中出现两次 fail,在这份 contract 下属于全部失败。它不是整个测试套件的 flaky 比例,也不能作为其他 property 的证据。
如果要延长调度,就修改 contract 中的 seeds,并重新运行两棵代码树。只给补丁追加运行,直到出现一次通过,这是挑选结果,不是测量。
父提交中的失败被消除后,还需要进一步复核,而这个函数不会做这一步。parent_fail_cleared 的意思是不要发起 freeze,并不意味着补丁保留了原有 oracle。如果 diff 修改了边界值、比较运算符或 fixture,即使判定表给出 accept,也要交给 spec review。
披露:本文是 MonkeyCode 产品推广活动的一部分。
在进入 gate 之前,可以把 MonkeyCode 的免费模型访问作为起草辅助。让它根据补丁摘要提出候选 property,然后编辑这些文本,直到它们成为你愿意实际执行的 oracle。将编辑后的 spec 提交到版本库。判定器从不把模型给出的标签当作输入。
一段把失败日志解释成时序噪声的文字,不能推翻父提交通过这一事实。
只有本地 CPU 成为瓶颈时,才把免费服务器选项作为第 2 步的额外计算资源。任务 spec 必须携带 parent_sha、patch_sha、spec_hash、fixture_closure_hash 和 seed 列表,并返回两份 JSONL 文件。在你控制的机器上,用本文的模块判定这些文件。响应如果为空、不完整,或关联的 contract hash 并非你提交的那个,就丢弃它。
传输成功不等于测试结果。
本文不声明任何模型名称、配额、硬件规格或截止日期。免费访问和免费服务器是运营方提供的选项,在某一天,两者中的任何一个都可能不可用。如果不可用,就在其他地方运行同一份 contract。判定表不会因为服务中断而放宽条件。
把分类器保存在你正在评审的仓库里,将远程运行环境视为可以替换的 worker。
较短的调度可能漏掉父提交中罕见的失败。这时,面对一个原本可能是双方共有的补丁失败,gate 会将其报告为 regression。这种误拦属于保守方向的错误,可以接受。相反,豁免真正的破坏则不可接受。
不要为了解除阻拦而移除父提交的运行。
witness 相等判断的准确程度,取决于归一化器。时间戳、内存地址和临时路径会把同一种失败拆成多个类别,导致 witness_mismatch。先去掉这些字段。
共享的外部服务可能让两棵代码树都失败,看起来像双方共有的变动性失败。如果 property 会访问网络、读取你无法控制的时钟,或调用其他团队的 staging 主机,就不要使用这道 gate。与外界隔离的 fixture 是前提条件,不是附带说明。
这个函数也不考虑覆盖率。补丁可能删除某个调用点,而一项已不再执行高风险分支的 property 仍然显示 both_pass。如果这种风险对你的代码库很重要,就为这张判定表配套覆盖率检查或 mutation check。本文没有提供第二项检查。
没有父提交时,例如首次导入生成的测试,就跳过这道 gate,因为没有可供差分比较的 baseline。spec hash 或 fixture closure hash 发生变化时,也要跳过。
对于无法将失败归纳为稳定 witness 类别的截图测试和像素测试,跳过这道 gate。如果你的 merge bot 把 freeze 文件当作自动放行依据,也要跳过。这套设计不接受这种工作流。
无法保存 JSONL 文件时,同样要跳过。一条声称“父提交也失败了”的文字评论,不能替代这些文件。没有文件,后来的读者就无法核查全部条件是否同时成立,提案也就无效。
保存 contract hash、顶层 decision、每项 property 的原因代码,以及 JSONL 行数。这四行就是证据包。缺少任何一行,都不要凭一个 flaky 标签就合并。
父提交的代码树是对照组,补丁的代码树是实验组。freeze 是针对双方共有的变动性失败、设有有效期的工单,不是用来静音某个对照组从未出现过的 regression 的按钮。
如需进一步采取行动,你可以考虑屏蔽此人和/或举报滥用行为。