feat(quality): ⑨升通用底座第12条「核心操作非无脑」落 RICHNESS_CHECKLIST(L2 6→7,前11条一字不动,prompt 计数改动态 len 同源防漂)+ 测试分母随升;图说预算两段式「待实施/CLI仍hard」翻已落实态(cheap_budget 同源工厂+真M3实证) (质量SoT §4 规范一 12条口径,ba55508e 配套;金标复验=C波)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
This commit is contained in:
parent
b78a8dfa4f
commit
3495470d89
@ -44,9 +44,14 @@ RICHNESS_CHECKLIST = [
|
||||
("首3分钟脚本", "L3", "首 3 分钟脚本成立:0–10s 零阅读上手 / 10–60s 首次升级 / 1–3min 露出 2–3 个后续锁"),
|
||||
("炫耀时刻", "L4", "有可截图分享的结算/成就画面(打完有结算展示场景,不是直接黑屏)"),
|
||||
("同款钩子", "L4", "有清晰的品类原型 + 主题标签元数据,可供 remix 预填做同款"),
|
||||
# 第 12 条(2026-07-03 增补,质量 SoT §4 规范一「通用底座 12 条口径」;创始人授权默认项拍板采纳):
|
||||
# 源自便宜档主 prompt 自检第 ⑨ 条(策划知识包 v2 否决项判定句式),升为评分尺条目。
|
||||
# 锚定纪律:追加于末尾、前 11 条一字不动;L2 分母随之 6→7,档位观测线百分比不变、分母随升;
|
||||
# 落地与金标复验同批(复验漂移超线按规范三复采样定性再判回退)。
|
||||
("核心操作非无脑", "L2", "每次主操作含真实决策或技巧含量、判错有真代价(如补错货压库存),不是点哪都加分/点了就自动结算"),
|
||||
]
|
||||
|
||||
_MAX = len(RICHNESS_CHECKLIST) # = 11
|
||||
_MAX = len(RICHNESS_CHECKLIST) # = 12(2026-07-03 ⑨升第 12 条;此前 v2 = 11)
|
||||
_LAYERS = ("L2", "L3", "L4") # 三分组固定顺序(消费面按分组小计各自取阈,见 §3.6 档位表)
|
||||
|
||||
# ── 品类扩展 rubric(W-GENRE 品类件④;质量模型 SoT §4 规范二)──
|
||||
@ -101,7 +106,7 @@ def load_genre_checklist(genre) -> list:
|
||||
|
||||
|
||||
def _layer_maxes(checklist=RICHNESS_CHECKLIST) -> dict:
|
||||
"""各层条目数(分母)。通用底座 v2 = L2:6 / L3:3 / L4:2。"""
|
||||
"""各层条目数(分母)。通用底座 12 条口径 = L2:7 / L3:3 / L4:2(2026-07-03 第 12 条入 L2)。"""
|
||||
m = {ly: 0 for ly in _LAYERS}
|
||||
for _name, layer, _meaning in checklist:
|
||||
m[layer] = m.get(layer, 0) + 1
|
||||
@ -129,19 +134,21 @@ _MAX_SRC_BYTES = 24000 # 喂 judge 的源码总量上限(控 token / 成本
|
||||
_JUDGE_MAX_TOKENS = 4000 # judge 输出小(11 条短理由 + 一句点评),4000 足够、M3 无 thinking。
|
||||
|
||||
# judge 的 system / user 框定——只评分、不改码、要看源码真实现、输出严格 JSON。
|
||||
# 计数用 len(RICHNESS_CHECKLIST) 动态拼(2026-07-03 ⑨升第 12 条时改:此前「11 条」写死在文案里,
|
||||
# 清单变更即漂;动态取数后条数永远与评分尺同源。评分尺变更本身仍受锚定纪律约束=金标复验)。
|
||||
_JUDGE_SYSTEM = (
|
||||
"你是轻量小游戏的【丰富度评审 agent】。任务 = 读一款便宜档 LittleJS 小游戏的源码,"
|
||||
"只评判它「作为一款游戏够不够丰富、好玩、耐玩、想传出去」,逐条给出 11 条丰富度清单的命中与否 + 一句中文理由,"
|
||||
f"只评判它「作为一款游戏够不够丰富、好玩、耐玩、想传出去」,逐条给出 {len(RICHNESS_CHECKLIST)} 条丰富度清单的命中与否 + 一句中文理由,"
|
||||
"最后输出严格 JSON。你只评分、不修改代码、不阻断发布——这是非阻塞的质量信号。"
|
||||
"这 11 条分三层:L2 内容丰富(有料耐玩)、L3 留存结构(想再玩的结构前提)、L4 传播钩子(想让别人看/做同款的前提);"
|
||||
f"这 {len(RICHNESS_CHECKLIST)} 条分三层:L2 内容丰富(有料耐玩)、L3 留存结构(想再玩的结构前提)、L4 传播钩子(想让别人看/做同款的前提);"
|
||||
"评判要看源码里**真实实现**的玩法、数值成长、解锁阶梯、即时反馈、音效、结算画面与品类元数据,别被空壳或注释骗。"
|
||||
)
|
||||
|
||||
# 品类扩展评分时的 system 变体(W-GENRE 件④):仅把「11 条」的量词放宽到「通用 11 条 + 品类扩展若干条」,
|
||||
# 其余判据措辞一字不动(锚定纪律:无品类路走上面原版 _JUDGE_SYSTEM,逐字节不变)。
|
||||
# 品类扩展评分时的 system 变体(W-GENRE 件④):仅把量词放宽到「通用 N 条 + 品类扩展若干条」,
|
||||
# 其余判据措辞一字不动(锚定纪律:无品类路走上面原版 _JUDGE_SYSTEM)。
|
||||
_JUDGE_SYSTEM_GENRE = (
|
||||
"你是轻量小游戏的【丰富度评审 agent】。任务 = 读一款便宜档 LittleJS 小游戏的源码,"
|
||||
"只评判它「作为一款游戏够不够丰富、好玩、耐玩、想传出去」,逐条给出丰富度清单(通用 11 条 + 品类扩展若干条)的命中与否 + 一句中文理由,"
|
||||
f"只评判它「作为一款游戏够不够丰富、好玩、耐玩、想传出去」,逐条给出丰富度清单(通用 {len(RICHNESS_CHECKLIST)} 条 + 品类扩展若干条)的命中与否 + 一句中文理由,"
|
||||
"最后输出严格 JSON。你只评分、不修改代码、不阻断发布——这是非阻塞的质量信号。"
|
||||
"条目分三层:L2 内容丰富(有料耐玩)、L3 留存结构(想再玩的结构前提)、L4 传播钩子(想让别人看/做同款的前提);"
|
||||
"评判要看源码里**真实实现**的玩法、数值成长、解锁阶梯、即时反馈、音效、结算画面与品类元数据,别被空壳或注释骗。"
|
||||
@ -283,18 +290,18 @@ def _build_judge_user(src_text: str, brief: str = "", genre_key: str = None, gen
|
||||
for i, (name, layer, meaning) in enumerate(genre_checklist)
|
||||
)
|
||||
genre_block = (
|
||||
f"\n另有 {n_genre} 条【{genre_key} 品类扩展清单】(分母独立、不与上面 11 条混合,同样逐条独立判命中):\n"
|
||||
f"\n另有 {n_genre} 条【{genre_key} 品类扩展清单】(分母独立、不与上面 {len(RICHNESS_CHECKLIST)} 条混合,同样逐条独立判命中):\n"
|
||||
f"{genre_lines}\n"
|
||||
)
|
||||
total_note = f"共 {len(RICHNESS_CHECKLIST) + n_genre} 条(通用 11 条在前、品类 {n_genre} 条紧随其后)"
|
||||
total_note = f"共 {len(RICHNESS_CHECKLIST) + n_genre} 条(通用 {len(RICHNESS_CHECKLIST)} 条在前、品类 {n_genre} 条紧随其后)"
|
||||
name_note = "name 用上面各条的中文名"
|
||||
else:
|
||||
genre_block = ""
|
||||
total_note = "共 11 条"
|
||||
name_note = "name 用上面 11 条的中文名"
|
||||
total_note = f"共 {len(RICHNESS_CHECKLIST)} 条"
|
||||
name_note = f"name 用上面 {len(RICHNESS_CHECKLIST)} 条的中文名"
|
||||
return (
|
||||
f"{brief_block}"
|
||||
"下面是这款便宜档小游戏的 L3 源码。请逐条评判它是否命中这 11 条丰富度清单(每条前的 [L2]/[L3]/[L4] 是分层标注——"
|
||||
f"下面是这款便宜档小游戏的 L3 源码。请逐条评判它是否命中这 {len(RICHNESS_CHECKLIST)} 条丰富度清单(每条前的 [L2]/[L3]/[L4] 是分层标注——"
|
||||
"L2 内容丰富、L3 留存结构、L4 传播钩子——仅用于分组,不改变你对每条的独立判命中),"
|
||||
"据**源码里真实实现了的玩法/数值/音效/解锁/结算画面/品类元数据**判断(别被注释或空壳骗:比如只 import 了 audioMusic "
|
||||
"但收益处没真调 playSfx,则「音反馈」不算命中):\n\n"
|
||||
|
||||
@ -33,16 +33,18 @@ def _full_checks(hits):
|
||||
return [{"name": n, "hit": h, "why": "理由"} for (n, _l, _m), h in zip(cheap_verify.RICHNESS_CHECKLIST, hits)]
|
||||
|
||||
|
||||
# 通用底座 v2 层索引(据 RICHNESS_CHECKLIST 顺序):L2={0,1,3,4,5,7} L3={2,6,8} L4={9,10}
|
||||
# 通用底座层索引(据 RICHNESS_CHECKLIST 顺序):L2={0,1,3,4,5,7,11} L3={2,6,8} L4={9,10}
|
||||
def test_checklist_shape_v2():
|
||||
"""底座 = 11 条,分层 L2×6 / L3×3 / L4×2(分组归属对齐 §3.3 第 1 条)。"""
|
||||
"""底座 = 12 条(2026-07-03 ⑨升第 12 条「核心操作非无脑」),分层 L2×7 / L3×3 / L4×2(质量 SoT §4 规范一 12 条口径)。"""
|
||||
cl = cheap_verify.RICHNESS_CHECKLIST
|
||||
assert len(cl) == 11 and cheap_verify._MAX == 11
|
||||
assert len(cl) == 12 and cheap_verify._MAX == 12
|
||||
layers = [layer for _n, layer, _m in cl]
|
||||
assert layers.count("L2") == 6 and layers.count("L3") == 3 and layers.count("L4") == 2
|
||||
assert layers.count("L2") == 7 and layers.count("L3") == 3 and layers.count("L4") == 2
|
||||
# 原 8 条 name/顺序锁定(锚定纪律:评分尺变更不得动原条目)
|
||||
assert [n for n, _l, _m in cl][:8] == [
|
||||
"即时反馈", "可见成长", "下一个解锁", "30秒爽点", "数值滚雪球", "情感锚", "放置回归", "音反馈"]
|
||||
# 第 12 条钉名+钉位(追加于末尾、层 L2;锚定纪律=前 11 条一字不动)
|
||||
assert cl[11][0] == "核心操作非无脑" and cl[11][1] == "L2"
|
||||
|
||||
|
||||
# ───────────────────────── ① parse_judge_output 纯函数(无 agentscope/无网络)─────────────────────────
|
||||
@ -53,18 +55,18 @@ def test_parse_valid_full():
|
||||
"notes": "还行"}, ensure_ascii=False)
|
||||
r = cheap_verify.parse_judge_output(text)
|
||||
assert r["degraded"] is False
|
||||
assert r["score"] == 4 and r["max"] == 11
|
||||
assert len(r["hits"]) == 11 and r["hits"][0]["name"] == "即时反馈" and r["hits"][0]["layer"] == "L2"
|
||||
# 前 8 条命中态 [T,T,F,T,F,F,F,T]:L2(0,1,3,4,5,7)=4;L3(2,6,8=F,F,未喂F)=0;L4(9,10 未喂)=0
|
||||
assert r["groups"] == {"L2": {"score": 4, "max": 6}, "L3": {"score": 0, "max": 3}, "L4": {"score": 0, "max": 2}}
|
||||
assert r["score"] == 4 and r["max"] == 12
|
||||
assert len(r["hits"]) == 12 and r["hits"][0]["name"] == "即时反馈" and r["hits"][0]["layer"] == "L2"
|
||||
# 前 8 条命中态 [T,T,F,T,F,F,F,T]:L2(0,1,3,4,5,7=4 中,11 未喂 F)=4;L3(2,6,8=F,F,未喂F)=0;L4(9,10 未喂)=0
|
||||
assert r["groups"] == {"L2": {"score": 4, "max": 7}, "L3": {"score": 0, "max": 3}, "L4": {"score": 0, "max": 2}}
|
||||
assert r["notes"] == "还行"
|
||||
|
||||
|
||||
def test_parse_11_full_and_groups():
|
||||
"""11 条全命中 → score=11、三分组小计满分(L2 6/6 · L3 3/3 · L4 2/2)。"""
|
||||
r = cheap_verify.parse_judge_output(json.dumps({"checks": _full_checks([True] * 11), "notes": "满"}))
|
||||
assert r["degraded"] is False and r["score"] == 11 and r["max"] == 11
|
||||
assert r["groups"] == {"L2": {"score": 6, "max": 6}, "L3": {"score": 3, "max": 3}, "L4": {"score": 2, "max": 2}}
|
||||
"""12 条全命中 → score=12、三分组小计满分(L2 7/7 · L3 3/3 · L4 2/2)。"""
|
||||
r = cheap_verify.parse_judge_output(json.dumps({"checks": _full_checks([True] * 12), "notes": "满"}))
|
||||
assert r["degraded"] is False and r["score"] == 12 and r["max"] == 12
|
||||
assert r["groups"] == {"L2": {"score": 7, "max": 7}, "L3": {"score": 3, "max": 3}, "L4": {"score": 2, "max": 2}}
|
||||
|
||||
|
||||
def test_parse_groups_by_layer():
|
||||
@ -72,23 +74,23 @@ def test_parse_groups_by_layer():
|
||||
pat = [True, True, False, True, False, False, False, True, False, True, True]
|
||||
r = cheap_verify.parse_judge_output(json.dumps({"checks": _full_checks(pat)}))
|
||||
assert r["score"] == 6
|
||||
assert r["groups"]["L2"] == {"score": 4, "max": 6}
|
||||
assert r["groups"]["L2"] == {"score": 4, "max": 7}
|
||||
assert r["groups"]["L3"] == {"score": 0, "max": 3}
|
||||
assert r["groups"]["L4"] == {"score": 2, "max": 2}
|
||||
|
||||
|
||||
def test_parse_fenced_json():
|
||||
"""markdown 围栏 + 前后赘语 → json_repair 兜得住、正确解析。"""
|
||||
body = json.dumps({"checks": _full_checks([True] * 11), "notes": "满"}, ensure_ascii=False)
|
||||
body = json.dumps({"checks": _full_checks([True] * 12), "notes": "满"}, ensure_ascii=False)
|
||||
text = "这是我的评分:\n```json\n" + body + "\n```\n以上。"
|
||||
r = cheap_verify.parse_judge_output(text)
|
||||
assert r["degraded"] is False and r["score"] == 11
|
||||
assert r["degraded"] is False and r["score"] == 12
|
||||
|
||||
|
||||
def test_parse_garbage_degraded():
|
||||
"""完全非 JSON → degraded(score=None, groups=None),不抛。"""
|
||||
r = cheap_verify.parse_judge_output("完全不是 JSON 的一段话")
|
||||
assert r["degraded"] is True and r["score"] is None and r["max"] == 11 and r["groups"] is None
|
||||
assert r["degraded"] is True and r["score"] is None and r["max"] == 12 and r["groups"] is None
|
||||
|
||||
|
||||
def test_parse_none_and_empty_degraded():
|
||||
@ -115,7 +117,7 @@ def test_parse_coerce_and_positional():
|
||||
r = cheap_verify.parse_judge_output(json.dumps({"checks": checks}))
|
||||
assert r["degraded"] is False
|
||||
assert r["score"] == 2 # 前两条 "true"/1 命中,第三条「否」不命中,其余 False
|
||||
assert r["max"] == 11 and r["hits"][0]["layer"] == "L2" # 位置兜底也带 checklist 权威 layer
|
||||
assert r["max"] == 12 and r["hits"][0]["layer"] == "L2" # 位置兜底也带 checklist 权威 layer
|
||||
|
||||
|
||||
# ───────────────────────── ② verify_richness 非阻塞契约(fake model,零真网络)─────────────────────────
|
||||
@ -238,7 +240,7 @@ def test_parse_with_genre_checklist_independent_denominator():
|
||||
checks += [{"name": n, "hit": (i != len(gc) - 1), "why": "w"} for i, (n, _l, _m) in enumerate(gc)]
|
||||
r = cheap_verify.parse_judge_output(json.dumps({"checks": checks, "notes": "n"}),
|
||||
genre_checklist=gc, genre_key="trpg")
|
||||
assert r["score"] == 11 and r["max"] == 11, "通用分母被品类条目污染"
|
||||
assert r["score"] == 12 and r["max"] == 12, "通用分母被品类条目污染"
|
||||
g = r["genre"]
|
||||
assert g["key"] == "trpg" and g["max"] == len(gc) and g["score"] == len(gc) - 1
|
||||
total = sum(v["max"] for v in g["groups"].values())
|
||||
@ -248,18 +250,18 @@ def test_parse_with_genre_checklist_independent_denominator():
|
||||
def test_parse_genre_positional_fallback_offset():
|
||||
"""品类段位置兜底:name 全对不上时按「通用在前、品类紧随」的偏移对位。"""
|
||||
gc = cheap_verify.load_genre_checklist("trpg")
|
||||
n = 11 + len(gc)
|
||||
n = len(cheap_verify.RICHNESS_CHECKLIST) + len(gc)
|
||||
checks = [{"name": f"未知{i}", "hit": True, "why": ""} for i in range(n)]
|
||||
r = cheap_verify.parse_judge_output(json.dumps({"checks": checks}),
|
||||
genre_checklist=gc, genre_key="trpg")
|
||||
assert r["score"] == 11 and r["genre"]["score"] == len(gc)
|
||||
assert r["score"] == 12 and r["genre"]["score"] == len(gc)
|
||||
|
||||
|
||||
def test_parse_no_genre_backcompat_no_genre_key():
|
||||
"""不传品类:输出形状与既有完全一致(无 genre 键——老消费面零感知)。"""
|
||||
checks = [{"name": n, "hit": True, "why": "w"} for (n, _l, _m) in cheap_verify.RICHNESS_CHECKLIST]
|
||||
r = cheap_verify.parse_judge_output(json.dumps({"checks": checks}))
|
||||
assert "genre" not in r and r["score"] == 11
|
||||
assert "genre" not in r and r["score"] == 12
|
||||
|
||||
|
||||
def test_build_judge_user_genre_block_additive():
|
||||
@ -267,8 +269,8 @@ def test_build_judge_user_genre_block_additive():
|
||||
base = cheap_verify._build_judge_user("SRC", "B")
|
||||
gc = cheap_verify.load_genre_checklist("trpg")
|
||||
withg = cheap_verify._build_judge_user("SRC", "B", genre_key="trpg", genre_checklist=gc)
|
||||
assert "品类扩展清单" not in base and "共 11 条" in base
|
||||
assert "trpg 品类扩展清单" in withg and f"共 {11 + len(gc)} 条" in withg
|
||||
assert "品类扩展清单" not in base and "共 12 条" in base
|
||||
assert "trpg 品类扩展清单" in withg and f"共 {12 + len(gc)} 条" in withg
|
||||
for name, _l, _m in gc:
|
||||
assert name in withg, f"品类条目「{name}」应出现在 judge prompt"
|
||||
|
||||
|
||||
@ -109,9 +109,10 @@ def test_genre_loader_and_none_path_identity():
|
||||
src = "// demo src"
|
||||
assert cv._build_judge_user(src, "b") == cv._build_judge_user(src, "b", genre_key="puzzle", genre_checklist=None), \
|
||||
"无品类条目时 judge prompt 必须与既有版本逐字节一致(锚定纪律)"
|
||||
# 品类段真的进 prompt(有品类时;主干契约=单 checks 数组、通用 11 条在前、品类条目紧随)。
|
||||
# 品类段真的进 prompt(有品类时;主干契约=单 checks 数组、通用底座在前、品类条目紧随)。
|
||||
p = cv._build_judge_user(src, "b", genre_key="puzzle", genre_checklist=tuple(items))
|
||||
assert "品类扩展清单" in p and items[0][0] in p and f"共 {11 + len(items)} 条" in p
|
||||
n_base = len(cv.RICHNESS_CHECKLIST) # 12(2026-07-03 ⑨升第 12 条)
|
||||
assert "品类扩展清单" in p and items[0][0] in p and f"共 {n_base + len(items)} 条" in p
|
||||
|
||||
|
||||
def test_genre_parse_separate_subtotals():
|
||||
@ -126,17 +127,17 @@ def test_genre_parse_separate_subtotals():
|
||||
genre_checks = [{"name": n, "hit": (i % 2 == 0), "why": "ok"} for i, (n, _l, _m) in enumerate(items)]
|
||||
text = json.dumps({"checks": base_checks + genre_checks, "notes": "n"}, ensure_ascii=False)
|
||||
r = cv.parse_judge_output(text, genre_checklist=items, genre_key="puzzle")
|
||||
assert r["degraded"] is False and r["score"] == 11
|
||||
assert r["degraded"] is False and r["score"] == len(cv.RICHNESS_CHECKLIST)
|
||||
g = r["genre"]
|
||||
assert g["key"] == "puzzle" and g["max"] == len(items) and len(g["hits"]) == len(items)
|
||||
gg = g["groups"]
|
||||
layer_max = {ly: sum(1 for (_n, l, _m) in items if l == ly) for ly in ("L2", "L3", "L4")}
|
||||
assert {ly: gg[ly]["max"] for ly in gg} == layer_max, "品类 groups 分母应来自品类清单(独立小计)"
|
||||
assert g["score"] == sum(1 for i in range(len(items)) if i % 2 == 0)
|
||||
# 模型漏答品类段(checks 只有通用 11 条)→ 扩展全 0、底座不受影响。
|
||||
# 模型漏答品类段(checks 只有通用底座条目)→ 扩展全 0、底座不受影响。
|
||||
r2 = cv.parse_judge_output(json.dumps({"checks": base_checks, "notes": "n"}),
|
||||
genre_checklist=items, genre_key="puzzle")
|
||||
assert r2["score"] == 11 and all(not h["hit"] for h in r2["genre"]["hits"])
|
||||
assert r2["score"] == len(cv.RICHNESS_CHECKLIST) and all(not h["hit"] for h in r2["genre"]["hits"])
|
||||
# 无品类路:输出键集与既有版本一致(不带 genre 键)。
|
||||
r3 = cv.parse_judge_output(json.dumps({"checks": base_checks, "notes": "n"}))
|
||||
assert "genre" not in r3
|
||||
|
||||
@ -367,11 +367,11 @@ tier2(最高深度档)的运行时按 AgentScope 2.0.2 的真实对象结构落
|
||||
|
||||
**续修原语的落点在 2026-07-01/02 又前进了一步(护城河 middleware 线)。** 上面讲的"有界外层 resume"是 spike 期的形态——编排层(`studio.py`)在 agent 停下后,判断没真 finish、门没绿、还有预算,就带 verdict 失败反馈再 reply 一次。护城河续修把它从编排层外挂**迁进了 `on_reasoning` 的 RepairMiddleware**:整局在**单次 POST 内**于 finish 点拦截——agent 想 finish 时,middleware 独立重跑门,没绿就压制这次 finish、把结构化反馈注入同一会话让它接着修,不再每轮重开 POST。这条是**生产主路**,便宜档 Service 与富档共用同一份 RepairMiddleware。tier2 本地 runner 的 `studio.py` 外层 resume 仍然并存(收敛环就走它),作为 fallback 形态——两套语义一致(都是"门没绿加有预算就带反馈续修、门始终是 judge 纯代码判"),差别只在拦截点:一个在 POST 内的 finish 处,一个在 POST 外的编排层。同文件头注释曾写"不再外层 repair",与它自己 400 行开外的外层 resume 代码相左,随本轮一并纠正。
|
||||
|
||||
**四道熔断加预算闸**叠在中间件洋葱上,任一先触发即停本局:步数硬顶(给每系统的构建-修复定上限、给整局定 max_iters,防模型靠多轮反复试错把门擦边混过去)、预算闸、卡死探测(语义层判 agent 是否原地打转、空转换汤不换药,而非单纯计步;还要防 agent 改 driver 来绕过它)、双层超时(单步钉死一次工具调用、整局钉死本局总时长)。这套熔断 tier2 是**自建的、且实际比"软刹"强**——要纠一处早稿错:早稿把"官方 `ReplyBudgetControlMiddleware` 软刹"当现成件,但 2.0.2 源码核验**该类根本不存在**;tier2 用自建的 `on_system_prompt` 变换钩子 + 中间件洋葱上订阅模型调用结束事件的硬熔断替代,把 token 按 new-api 计费口径折成 ¥ 累进。**¥ 越限的行为在 2026-07-02/03 定为两段式(创始人裁定,起因是护城河续修不该被预算硬闸旁路)**:先是**软停线**(便宜档 ¥10 / 富档 ¥50)——越线不断链,设软停标记、经 `on_system_prompt` 强提醒 agent 基于当前工程状态尽快 finish 交尽力产物,只许收尾类动作(finish、构建、跑门),禁新增大额生成调用,给续修留活路;再是**硬地板**(软停线 ×1.5 = 便宜档 ¥15 / 富档 ¥75)——它与轮数、墙钟两道闸任一先到即停,保证软停之后 ¥ 上界数学上仍封得死("软停不等于无界")。**档位行为分叉**:soft 档(tier2 与便宜档 Service 路)走上面这套软停;hard 档(便宜档 CLI 及默认)仍是越限即 fail-closed 抛熔断、守 ¥ 硬地板(spike 实测 60/80 轮硬熔断兜底已落)。生产 Service、CLI、本图说三处取值统一。这道预算闸的三层强制架构见 §5.2。
|
||||
**四道熔断加预算闸**叠在中间件洋葱上,任一先触发即停本局:步数硬顶(给每系统的构建-修复定上限、给整局定 max_iters,防模型靠多轮反复试错把门擦边混过去)、预算闸、卡死探测(语义层判 agent 是否原地打转、空转换汤不换药,而非单纯计步;还要防 agent 改 driver 来绕过它)、双层超时(单步钉死一次工具调用、整局钉死本局总时长)。这套熔断 tier2 是**自建的、且实际比"软刹"强**——要纠一处早稿错:早稿把"官方 `ReplyBudgetControlMiddleware` 软刹"当现成件,但 2.0.2 源码核验**该类根本不存在**;tier2 用自建的 `on_system_prompt` 变换钩子 + 中间件洋葱上订阅模型调用结束事件的硬熔断替代,把 token 按 new-api 计费口径折成 ¥ 累进。**¥ 越限的行为在 2026-07-02/03 定为两段式(创始人裁定,起因是护城河续修不该被预算硬闸旁路)**:先是**软停线**(便宜档 ¥10 / 富档 ¥50)——越线不断链,设软停标记、经 `on_system_prompt` 强提醒 agent 基于当前工程状态尽快 finish 交尽力产物,只许收尾类动作(finish、构建、跑门),禁新增大额生成调用,给续修留活路;再是**硬地板**(软停线 ×1.5 = 便宜档 ¥15 / 富档 ¥75)——它与轮数、墙钟两道闸任一先到即停,保证软停之后 ¥ 上界数学上仍封得死("软停不等于无界")。**生产两档已统一走两段式**(2026-07-03 W-ARCH② 实施):便宜档 CLI 与生产 Service 经 `cheap_budget.build_cheap_breaker` 单一装配点组 breaker、数值同源 genconfig(generation.yaml budget 区),软停后 `on_acting` 拦生成面工具(write_file/write_source/scaffold_init,收尾类恒放行),硬地板 fail-closed 数学封顶;middleware 的 hard 单段语义原样保留为默认档(未显式装配面与实验路的兜底,spike 实测 60/80 轮硬熔断已落)。这道预算闸的三层强制架构见 §5.2。
|
||||
|
||||
**M3 的接法**是这条线能不能成的根因之一。tier2 agent 经官方 `AnthropicChatModel` 走 `AnthropicCredential.base_url`(指向 new-api 的 Anthropic 端点),到 MiniMax-M3,计费统一从 new-api 一个平面走。"用对 M3"是三件事:走 Anthropic 原生协议加 agentic 工具循环(循环调工具写源 / build / 跑门 / 读 verdict / 改,而不是 OpenAI 式单次 JSON 填空);thinking 分离(开 thinking,且 max_tokens 须严格大于 thinking_budget);完整 response 与历史保留(每轮的 thinking / text / tool_use 块原样回传入历史,否则 M3 的交错思维失效)。旧用法 60% 失败的最深根因正是反着来——关 thinking、单次 JSON 填空、失败从头重生成,是用法错、不是模型天花板。
|
||||
|
||||
> **现 / 建(2026-07-03 更新)**:本面核心**已落并经 feie-005 accept**——M3 Anthropic 接法链路、有界外层 resume + finish 门、自建硬熔断(非官方软刹,该类 2.0.2 不存在)、写源/改/快检/跑门内循环都已跑通;`max_tokens > thinking_budget` 启动校验、`RecordingChatModel(Anthropic)` 成本取证已落。**续修原语已迁进 `RepairMiddleware`**(护城河 middleware 线,2026-07-01/02 已代码落地,单 POST 内 finish 点拦截续修,便宜档 Service 与富档共用),外层 resume 在本地 runner 路并存。**预算两段式**:软停(`soft_budget` 档,tier2 与便宜档 Service 走)已落;硬地板 ×1.5 的显式公式为 2026-07-03 裁定、待实施(W-ARCH②);便宜档 CLI 仍走 hard fail-closed。**收敛环已真跑判 conditional**(见 §4.2)。待:agent 层 win-balance 攻坚、把偏脆的文本契约检查(play-scene 工厂结构)做成更稳的结构化校验。
|
||||
> **现 / 建(2026-07-03 更新)**:本面核心**已落并经 feie-005 accept**——M3 Anthropic 接法链路、有界外层 resume + finish 门、自建硬熔断(非官方软刹,该类 2.0.2 不存在)、写源/改/快检/跑门内循环都已跑通;`max_tokens > thinking_budget` 启动校验、`RecordingChatModel(Anthropic)` 成本取证已落。**续修原语已迁进 `RepairMiddleware`**(护城河 middleware 线,2026-07-01/02 已代码落地,单 POST 内 finish 点拦截续修,便宜档 Service 与富档共用),外层 resume 在本地 runner 路并存。**预算两段式已全落**(W-ARCH② 2026-07-03 实施):软停线越线只许收尾(`on_acting` 拦生成面工具)+ 硬地板 ×1.5 fail-closed;便宜档 CLI 与生产 Service 经 `cheap_budget.build_cheap_breaker` 同源,真 M3 smoke 实证软停触发与硬地板 fail-closed 都真转。**收敛环已真跑判 conditional**(见 §4.2)。待:agent 层 win-balance 攻坚、把偏脆的文本契约检查(play-scene 工厂结构)做成更稳的结构化校验。
|
||||

|
||||
|
||||

|
||||
|
||||
Loading…
x
Reference in New Issue
Block a user