按 tier2 图说目标做缺口分析(8族逐元素比对),补齐 0号 spike 为过门收窄掉、 但图说明确要求的「核心引擎待补」项。全部加性/observe-only:金标冒烟仍 ACCEPT (九门9/9+富游戏三门3/3,门一道没放松),真依赖下全链 import+自测+一款真 M3 跑验证通过。 G族(n≥30 runbook 执行基建): - worker/config.py: build_model_openai 便宜档 client(deepseek 经 new-api OpenAI 兼容路,与 M3 Anthropic 路并存) - worker/run_record.py: G4 采集字段表 → 可序列化 RunRecord(含退路树分流键 fail_system) - worker/fallback_tree.py: 退路树五出口判定器(Q1–Q4 数字触发线,★阈值常量区待校准) - batch_run.py / aggregate.py: model×variant×n 批跑(断点续跑/失败隔离)+ 矩阵聚合三图喂判定器 H族(观测/成本接线,把孤儿件缝进 run 主链): - observability/newapi_pricing.py: 活读 new-api /api/pricing 倍率(取不到回落显式参数+告警) - middleware.py: Tier2TraceMiddleware 挂 writer agent 最外层洋葱,ReAct 全事件旁路 ingest - agent_loop/studio.py 收口: records→cost_for_run 折¥;真跑实测 cost_rmb=1.29(newapi-live)、trace 647事件 dropped=0 - contracts/trace/: additive trace 事件契约位(忠实 trace.py 落 sink 形状) D族(L3 视觉软检接线,observe-only): - agent_loop/studio.py: 收口调一次 M3 多模态(真截图+真玩取证→fun映射0-100),只写 verdict.L3,绝不参与 decision - 真跑实测 L3 score=25 准确指出空心表现层;decision=fix 仍由 L1硬门/熔断裁、与 L3 无关(防 Goodhart 成立) 留后(不投机抢建):工作室 Agent Team/第二装载落库/控制面/Agent Service 等按 plan 决策②⑤ gate 到 B门后; n≥30 等统计相是「跑」非「写」(批跑底座已就位);A-model 4插件复用待合并对账;4处图说 spec-drift 待 doc 线回写。 详见 tier2/HANDOFF.md「图说对账补全」节。 Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
334 lines
18 KiB
Python
334 lines
18 KiB
Python
#!/usr/bin/env python3
|
||
"""aggregate.py —— tier2 0号 spike · JSONL run-records → 矩阵级三图聚合 + 退路树判定。
|
||
|
||
【这份在 spike runbook 里的位置(权威 = G 族图说)】
|
||
docs/architecture/架构/生成引擎/tier2细节图说-G-spike-runbook.md 图 G4 底部那条汇总带:
|
||
run 级采集字段(每款一行,batch_run.py 落的 JSONL)卷成【矩阵级三张图】——
|
||
① 过门率 = pass 款数 / n;
|
||
② ¥/成功款 = Σcost / pass 款数;
|
||
③ 收敛中位数 = median(repairs);
|
||
外加 fail_system 分布(经营品类特有,退路树 Q2/Q3 分流的关键依据)。
|
||
这三图按 model × brief_variant 分组(G3:过门率/成本/墙钟各有归属对象,「60% 是哪个模型的 60%」要答得上),
|
||
并卷成 fallback_tree.decide 吃的矩阵级 by_model 聚合,直接喂退路树判 go/no-go(图 G5)。
|
||
|
||
【职责边界:纯读 + 纯算 + 输出报告,零生成副作用】
|
||
- 只读 JSONL 台账(batch_run 落的 RunRecord 行)、纯算聚合、调 fallback_tree.decide(纯函数判定器)。
|
||
- 不发网络、不跑 chrome、不改任何 run 记录;输出 = 一个聚合 JSON + 一段人读文本摘要(写文件 + 打印)。
|
||
- 不依赖 agentscope(只 import worker.run_record + worker.fallback_tree,二者皆纯数据/纯函数),6c6g 可跑。
|
||
|
||
【observe-only / 防 Goodhart】
|
||
本模块只把 run 记录里 judge 纯代码判出的 pass/fail 聚合呈现,绝不改裁决;退路树出口也只是「据数字裁一步」的
|
||
建议,人锚软门(创始人试玩判肥鹅味)不可被本聚合替代——GO 出口的文本会显式提示「人锚仍须过」。
|
||
|
||
CLI 用法:
|
||
python aggregate.py --in results/spike-runs.jsonl
|
||
python aggregate.py --in results/spike-runs.jsonl --out-json results/spike-agg.json --out-txt results/spike-agg.txt
|
||
|
||
可 import 用法:
|
||
from aggregate import load_records, aggregate, build_decide_stats
|
||
recs = load_records("results/spike-runs.jsonl")
|
||
agg = aggregate(recs) # 矩阵级三图 + fail_system 分布
|
||
stats = build_decide_stats(agg) # 卷成 fallback_tree.decide 的 by_model schema
|
||
verdict = fallback_tree.decide(stats) # 五出口判定
|
||
"""
|
||
|
||
from __future__ import annotations
|
||
|
||
import argparse
|
||
import json
|
||
import statistics
|
||
import sys
|
||
from pathlib import Path
|
||
from typing import Any
|
||
|
||
# 包内/直跑兼容:把 gen-worker/ 加进 sys.path 使顶层包 `worker` 可解析。
|
||
sys.path.insert(0, str(Path(__file__).resolve().parent))
|
||
|
||
from worker import fallback_tree # noqa: E402 —— A4 退路树判定器(纯函数)
|
||
from worker.run_record import RunRecord # noqa: E402 —— A3 采集记录(纯数据)
|
||
|
||
# fail_system 桶名(对齐 run_record.FailSystem Literal + fallback_tree 的桶名常量,集成接缝唯一口径)。
|
||
_FAIL_SYSTEM_BUCKETS = ("resource", "merge", "order", "presentation")
|
||
|
||
|
||
# ──────────────────────────────────────────────────────────────────────────
|
||
# 1) 读 JSONL 台账 → RunRecord 列表
|
||
# ──────────────────────────────────────────────────────────────────────────
|
||
def load_records(path: str) -> list[RunRecord]:
|
||
"""逐行读 JSONL 台账,反序列化成 RunRecord 列表。
|
||
|
||
坏行(JSON 解析失败 / 字段缺失到无法构造)跳过不计、不中断(best-effort:一行坏不该让整份聚合崩),
|
||
但打印告警让人看见(采集数据有坏行是要查的)。空文件 / 文件不存在 → 空列表。
|
||
"""
|
||
p = Path(path)
|
||
if not p.exists():
|
||
print(f"[aggregate] ⚠ 台账不存在:{path}(无数据可聚合)", file=sys.stderr)
|
||
return []
|
||
out: list[RunRecord] = []
|
||
for ln, line in enumerate(p.read_text(encoding="utf-8").splitlines(), 1):
|
||
line = line.strip()
|
||
if not line:
|
||
continue
|
||
try:
|
||
out.append(RunRecord.from_jsonl_line(line))
|
||
except Exception as e: # noqa: BLE001 —— 坏行跳过、告警、不中断
|
||
print(f"[aggregate] ⚠ 第 {ln} 行解析失败已跳过:{type(e).__name__}: {e}", file=sys.stderr)
|
||
return out
|
||
|
||
|
||
# ──────────────────────────────────────────────────────────────────────────
|
||
# 2) 分组聚合工具(纯函数:一组 RunRecord → 三图指标 + fail_system 分布)
|
||
# ──────────────────────────────────────────────────────────────────────────
|
||
def _agg_cell(recs: list[RunRecord]) -> dict[str, Any]:
|
||
"""把一组 RunRecord(同一分组,如某 model×variant 格)聚成三图指标 + fail_system 分布。
|
||
|
||
产出字段(对齐 G4 矩阵级三图 + fallback_tree.decide 的 by_model schema 字段名):
|
||
- n = 这组的总跑次数(分母);
|
||
- pass_count = 过门款数(pass_gate=True);
|
||
- pass_rate = pass_count / n(图① 过门率;n=0 → 0.0);
|
||
- cost_rmb_sum = Σcost_yuan(全部款,含失败款——失败也烧了钱);
|
||
- cost_rmb_per_pass = Σcost_yuan / pass_count(图② ¥/成功款;无过门款 → None,避免除零误导);
|
||
- wall_s_median = median(wall_seconds)(墙钟图;空 → 0.0);
|
||
- repairs_median = median(repairs)(图③ 收敛中位数;空 → 0);
|
||
- fail_system = {桶: 失败款数}(仅 fail_gate 款计入;退路树 Q2/Q3 读它);
|
||
- fail_stage = {段: 失败款数}(失败定位的段分布,人读摘要用)。
|
||
"""
|
||
n = len(recs)
|
||
passes = [r for r in recs if r.pass_gate]
|
||
pass_count = len(passes)
|
||
cost_sum = round(sum(r.cost_yuan for r in recs), 5)
|
||
fail_system: dict[str, int] = {}
|
||
fail_stage: dict[str, int] = {}
|
||
for r in recs:
|
||
if r.pass_gate:
|
||
continue
|
||
if r.fail_system:
|
||
fail_system[r.fail_system] = fail_system.get(r.fail_system, 0) + 1
|
||
if r.fail_stage:
|
||
fail_stage[r.fail_stage] = fail_stage.get(r.fail_stage, 0) + 1
|
||
return {
|
||
"n": n,
|
||
"pass_count": pass_count,
|
||
"pass_rate": round(pass_count / n, 4) if n else 0.0,
|
||
"cost_rmb_sum": cost_sum,
|
||
# ¥/成功款:无过门款时 None(不写 0,0 会被误读成「免费过门」;退路树读 by_model 时也不用它判 GO)。
|
||
"cost_rmb_per_pass": round(cost_sum / pass_count, 5) if pass_count else None,
|
||
"wall_s_median": round(statistics.median([r.wall_seconds for r in recs]), 2) if recs else 0.0,
|
||
"repairs_median": int(statistics.median([r.repairs for r in recs])) if recs else 0,
|
||
"fail_system": fail_system,
|
||
"fail_stage": fail_stage,
|
||
}
|
||
|
||
|
||
def _group_by(recs: list[RunRecord], key) -> dict[str, list[RunRecord]]:
|
||
"""按 key(rec → 分组键)把记录分桶,保持首次出现顺序。"""
|
||
out: dict[str, list[RunRecord]] = {}
|
||
for r in recs:
|
||
out.setdefault(key(r), []).append(r)
|
||
return out
|
||
|
||
|
||
# ──────────────────────────────────────────────────────────────────────────
|
||
# 3) 主聚合:aggregate(recs) → 矩阵级三图(by_model / by_variant / by_model_variant + overall)
|
||
# ──────────────────────────────────────────────────────────────────────────
|
||
def aggregate(recs: list[RunRecord]) -> dict[str, Any]:
|
||
"""把 run-records 卷成矩阵级聚合(三图按 model / variant / model×variant 三种分组 + 全局)。
|
||
|
||
Returns(聚合报告主体):
|
||
{
|
||
"total_runs": int, # 总跑次数
|
||
"overall": {三图指标 + fail_system 分布}, # 不分组的全局聚合
|
||
"by_model": {model: {三图...}}, # 按模型档分组(G3:过门率/成本归属到档)
|
||
"by_variant":{variant: {三图...}}, # 按题面变体分组(变体维度方差)
|
||
"by_model_variant": {"model | variant": {三图...}}, # 5×6 矩阵每格(最细粒度)
|
||
"fail_system_overall": {桶: 失败款数}, # 跨全部便宜档/全档的 fail_system 分布(退路树兜底用)
|
||
}
|
||
"""
|
||
by_model = {m: _agg_cell(g) for m, g in _group_by(recs, lambda r: r.model).items()}
|
||
by_variant = {v: _agg_cell(g) for v, g in _group_by(recs, lambda r: r.brief_variant).items()}
|
||
by_mv = {f"{r_m} | {r_v}": _agg_cell(g)
|
||
for (r_m, r_v), g in _group_by(recs, lambda r: (r.model, r.brief_variant)).items()}
|
||
|
||
# 全局 fail_system 分布(退路树未给 fail_system_overall 时的兜底来源;这里直接全档合)。
|
||
fail_overall: dict[str, int] = {}
|
||
for r in recs:
|
||
if not r.pass_gate and r.fail_system:
|
||
fail_overall[r.fail_system] = fail_overall.get(r.fail_system, 0) + 1
|
||
|
||
return {
|
||
"total_runs": len(recs),
|
||
"overall": _agg_cell(recs),
|
||
"by_model": by_model,
|
||
"by_variant": by_variant,
|
||
"by_model_variant": by_mv,
|
||
"fail_system_overall": fail_overall,
|
||
}
|
||
|
||
|
||
# ──────────────────────────────────────────────────────────────────────────
|
||
# 4) 卷成 fallback_tree.decide 吃的 by_model schema(集成接缝:聚合产物 → 判定器)
|
||
# ──────────────────────────────────────────────────────────────────────────
|
||
def build_decide_stats(agg: dict[str, Any]) -> dict[str, Any]:
|
||
"""把 aggregate() 的 by_model 转成 fallback_tree.decide 的 stats schema。
|
||
|
||
fallback_tree.decide 读的字段(见其 docstring 的 stats schema):
|
||
by_model[档] = {pass_rate, n, pass_count, fail_system, cost_rmb_per_pass, repairs_median}
|
||
+ 顶层 fail_system_overall(Q2 判失败是否集中表现层)。
|
||
aggregate 的 by_model 每格已含这些字段(字段名刻意对齐),这里只做「挑字段 + 透传」,不重算。
|
||
"""
|
||
by_model_in = agg.get("by_model") or {}
|
||
by_model_out: dict[str, Any] = {}
|
||
for model, cell in by_model_in.items():
|
||
by_model_out[model] = {
|
||
"pass_rate": cell.get("pass_rate", 0.0),
|
||
"n": cell.get("n", 0),
|
||
"pass_count": cell.get("pass_count", 0),
|
||
"fail_system": cell.get("fail_system", {}),
|
||
"cost_rmb_per_pass": cell.get("cost_rmb_per_pass"),
|
||
"repairs_median": cell.get("repairs_median", 0),
|
||
}
|
||
return {
|
||
"by_model": by_model_out,
|
||
# 顶层跨档 fail_system 分布(退路树 Q2/Q3 优先读它;缺则它内部由各便宜档现合)。
|
||
"fail_system_overall": agg.get("fail_system_overall") or {},
|
||
}
|
||
|
||
|
||
# ──────────────────────────────────────────────────────────────────────────
|
||
# 5) 人读文本摘要(资深工程师看的散文式报告,不堆电报体)
|
||
# ──────────────────────────────────────────────────────────────────────────
|
||
def _fmt_pct(x: float | None) -> str:
|
||
return f"{x:.0%}" if isinstance(x, (int, float)) else "—"
|
||
|
||
|
||
def _fmt_yuan(x: float | None) -> str:
|
||
return f"¥{x:.3f}" if isinstance(x, (int, float)) else "—(无过门款)"
|
||
|
||
|
||
def render_text_summary(agg: dict[str, Any], decide_out: dict[str, Any]) -> str:
|
||
"""把聚合 + 退路树判定渲染成一段人读文本摘要(供终端打印 + 落 .txt)。
|
||
|
||
结构:总览 → 按模型档三图(过门率/¥每成功款/墙钟中位/收敛中位/失败分布)→ 按变体 → 退路树裁决。
|
||
阈值口径不在这里硬编码,引用 fallback_tree 的常量(单一事实源),避免摘要与判定器漂移。
|
||
"""
|
||
lines: list[str] = []
|
||
ov = agg.get("overall") or {}
|
||
lines.append("=" * 72)
|
||
lines.append("tier2 0号 spike · 矩阵级聚合报告(过门率 / 成本 / 墙钟 by model×variant + 退路树)")
|
||
lines.append("=" * 72)
|
||
lines.append(f"总跑 {agg.get('total_runs', 0)} 款;"
|
||
f"全局过门率 {_fmt_pct(ov.get('pass_rate'))}"
|
||
f"({ov.get('pass_count', 0)}/{ov.get('n', 0)});"
|
||
f"全局 ¥/成功款 {_fmt_yuan(ov.get('cost_rmb_per_pass'))};"
|
||
f"总成本 ¥{ov.get('cost_rmb_sum', 0)}。")
|
||
|
||
# ── 按模型档(G3:三图归属到档,这是 go/no-go 的主分组)──
|
||
lines.append("")
|
||
lines.append("【按模型档】(过门率 / ¥每成功款 / 墙钟中位 / 收敛中位 / 失败系统分布)")
|
||
by_model = agg.get("by_model") or {}
|
||
if not by_model:
|
||
lines.append(" (无数据)")
|
||
for model, c in by_model.items():
|
||
fs = c.get("fail_system") or {}
|
||
fs_desc = ", ".join(f"{k}:{v}" for k, v in fs.items()) or "无失败或系统不明"
|
||
lines.append(
|
||
f" {model:<20} 过门率 {_fmt_pct(c.get('pass_rate')):>5}"
|
||
f"({c.get('pass_count', 0)}/{c.get('n', 0)}) "
|
||
f"¥/成功款 {_fmt_yuan(c.get('cost_rmb_per_pass'))} "
|
||
f"墙钟中位 {c.get('wall_s_median', 0)}s "
|
||
f"收敛中位 {c.get('repairs_median', 0)} 轮 "
|
||
f"失败系统[{fs_desc}]")
|
||
|
||
# ── 按题面变体(变体维度方差;G3 要 5 变体各跑,看变体间过门率差异)──
|
||
lines.append("")
|
||
lines.append("【按题面变体】(过门率 / ¥每成功款 / 墙钟中位)")
|
||
by_variant = agg.get("by_variant") or {}
|
||
if not by_variant:
|
||
lines.append(" (无数据)")
|
||
for v, c in by_variant.items():
|
||
lines.append(
|
||
f" {v:<8} 过门率 {_fmt_pct(c.get('pass_rate')):>5}"
|
||
f"({c.get('pass_count', 0)}/{c.get('n', 0)}) "
|
||
f"¥/成功款 {_fmt_yuan(c.get('cost_rmb_per_pass'))} "
|
||
f"墙钟中位 {c.get('wall_s_median', 0)}s")
|
||
|
||
# ── 全局 fail_system 分布(退路树分流的关键依据)──
|
||
lines.append("")
|
||
fso = agg.get("fail_system_overall") or {}
|
||
fso_desc = ", ".join(f"{k}:{v}" for k, v in fso.items()) or "无失败或系统不明"
|
||
lines.append(f"【全局 fail_system 分布】{fso_desc}")
|
||
|
||
# ── 退路树裁决(fallback_tree.decide 的五出口 + 逐条触发线 reasons)──
|
||
lines.append("")
|
||
lines.append("【退路树裁决(图 G5 五出口)】")
|
||
lines.append(f" 出口 = {decide_out.get('exit')}")
|
||
for r in (decide_out.get("reasons") or []):
|
||
lines.append(f" · {r}")
|
||
lines.append("")
|
||
lines.append("注:退路树只裁数字面;人锚软门(创始人试玩判肥鹅味)不可被替代,GO 仍须另行过人锚。")
|
||
lines.append(" ★ 阈值(过门率门 / ¥3每款 / 集中比例)= directional v1,需创始人和实测校准"
|
||
"(集中在 worker/fallback_tree.py 常量区)。")
|
||
lines.append("=" * 72)
|
||
return "\n".join(lines)
|
||
|
||
|
||
# ──────────────────────────────────────────────────────────────────────────
|
||
# 6) 顶层:从 JSONL 直接产完整报告(聚合 + decide stats + 退路树 + 文本摘要)
|
||
# ──────────────────────────────────────────────────────────────────────────
|
||
def report_from_jsonl(path: str) -> dict[str, Any]:
|
||
"""读 JSONL → 聚合 → 卷 decide stats → 退路树判定 → 组装完整报告 dict(含人读文本)。
|
||
|
||
Returns:
|
||
{
|
||
"aggregate": {...}, # aggregate() 的矩阵级三图
|
||
"decide_stats": {...}, # 喂 fallback_tree 的 by_model schema
|
||
"fallback_decision": {exit, reasons}, # 退路树五出口判定
|
||
"text_summary": "...", # 人读文本摘要
|
||
}
|
||
"""
|
||
recs = load_records(path)
|
||
agg = aggregate(recs)
|
||
stats = build_decide_stats(agg)
|
||
decision = fallback_tree.decide(stats)
|
||
text = render_text_summary(agg, decision)
|
||
return {
|
||
"aggregate": agg,
|
||
"decide_stats": stats,
|
||
"fallback_decision": decision,
|
||
"text_summary": text,
|
||
}
|
||
|
||
|
||
# ──────────────────────────────────────────────────────────────────────────
|
||
# CLI
|
||
# ──────────────────────────────────────────────────────────────────────────
|
||
def main() -> None:
|
||
ap = argparse.ArgumentParser(
|
||
description="tier2 0号 spike 聚合(JSONL run-records → 矩阵级三图 + 退路树判定)")
|
||
ap.add_argument("--in", dest="in_path", required=True, help="输入 JSONL 台账(batch_run 落的)")
|
||
ap.add_argument("--out-json", default=None,
|
||
help="聚合报告 JSON 落点(默认 = 输入同目录 <stem>-agg.json)")
|
||
ap.add_argument("--out-txt", default=None,
|
||
help="人读文本摘要落点(默认 = 输入同目录 <stem>-agg.txt)")
|
||
args = ap.parse_args()
|
||
|
||
rep = report_from_jsonl(args.in_path)
|
||
|
||
in_p = Path(args.in_path)
|
||
out_json = Path(args.out_json) if args.out_json else in_p.with_name(in_p.stem + "-agg.json")
|
||
out_txt = Path(args.out_txt) if args.out_txt else in_p.with_name(in_p.stem + "-agg.txt")
|
||
|
||
# 落 JSON(完整聚合 + decide stats + 判定;text_summary 也一并落,便于单文件回看)。
|
||
out_json.parent.mkdir(parents=True, exist_ok=True)
|
||
out_json.write_text(json.dumps(rep, ensure_ascii=False, indent=2), encoding="utf-8")
|
||
out_txt.write_text(rep["text_summary"], encoding="utf-8")
|
||
|
||
# 终端打印人读摘要(创始人要紧凑;细节落文件)。
|
||
print(rep["text_summary"])
|
||
print(f"\n[aggregate] 报告 → JSON {out_json} / TXT {out_txt}")
|
||
sys.exit(0)
|
||
|
||
|
||
if __name__ == "__main__":
|
||
main()
|