lili 63f178d615 feat(prompts): 四道闸CI接线段B——真模型四闸判定器 eval_gate + prompt-eval workflow + safety 基线
薄版重建四道闸(不复活已删的 evalflow/judge 整套编排),对改动的 live prompt 拿金标集真跑 M3:
- 闸1 Schema(产物合 output_schema)/ 闸2 成功率 ≥0.8 / 闸3 回归 diff(vs baseline)/ 闸4 成本延迟增量;
- 只焊 live 面:非 live SKIP 豁免、live 无金标 fail-closed(绝不假绿);单跑预算 cap ¥5 硬编码、超即 fail;
- 真模型 = MiniMax-M3 经 new-api(thinking 关、实测 reasoning_tokens=0),urllib ProxyHandler({}) 绕代理,key 从 env 读、不硬编码不设长期 env;
- 触发 = workflow_dispatch(prompt-eval.yml,与 push 门分工);台账 append-only 落 eval/<id>/runs/;
- 调用失败(网关 500/限流)排除出判定分母、建议复跑(SoT §7:eval 跑不通不阻塞、不假绿)。

safety.prompt-check 真跑证据:10 条违规负例经 M3 判定,9/9 responded 全绿(neg:5 稳定触发 new-api 上游 500 已按基础设施失败排除),本跑 ¥0.0071,建立 baseline 供闸3/4 后续真判。

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-04 09:10:00 -07:00

600 lines
29 KiB
Python
Executable File
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

#!/usr/bin/env python3
"""
eval_gate.py — Prompt 四道闸 · 真模型闸(段 B,薄版重建)
这是什么:对一条改动的 live prompt,拿它的 Golden 金标集真跑一遍模型,过「四道闸」才放行。
历史判定实现(agent-loop v1 的 orchestrator/evalflow.py + judge.py)已随 spike 存档删除、只在 git;
本脚本按 registry 契约重建一个薄版四闸判定,不复活整套批跑编排——只做「一条 prompt × 它的金标集 →
四闸裁决」这一件事,触发方式改成 CI 手动 workflow_dispatch(见 .gitea/workflows/prompt-eval.yml)。
四道闸(对齐 prompt治理 SoT §4):
闸1 Schema —— 模型对每条输入的产物必须能解析成 JSON 且符合该 prompt 声明的 output_schema。
闸2 成功率 —— 拿金标集跑一遍,判对率必须 ≥0.8(对齐 MVP 顶层指标「AI 生成成功率 ≥80%」)。
闸3 回归 diff —— 与基线快照(baseline.json)逐条比关键裁决,关键字段相对基线翻转即拦。
闸4 成本/延迟 —— 与基线比单跑成本/延迟增量,超阈值即拦(变贵/变慢)。
另有前置闸0(version 升号)由段 A check_version_bump.py 管,本脚本不重复。
只焊 live 面(step 0 消费面对账,见 registry.yaml 头部三态总表):
· 非 live(fossil / batch-relic)条目 → 显式 SKIP 豁免,绝不花真模型钱焊死面;
· live 但无金标集(当前 cheap-system / 8 条 tier2,eval 待建)→ fail-closed 报「金标缺失」,绝不假绿;
· live 且有金标集(当前唯一 = safety.prompt-check)→ 真跑四闸。
LIVE_PROMPT_IDS 是段 B 的判定白名单,逐条依据见 registry.yaml 头部对账总表;启动自检它们都在 registry 内。
真模型:MiniMax-M3 经 new-api 网关(thinking 关:裸 OpenAI 兼容端点,实测 reasoning_tokens=0)。
凭据从环境变量读(NEWAPI_KEY 必需 / NEWAPI_BASE_URL 可选),绝不硬编码 key、不设长期 env——
key 权威源 = docs/内网凭据与端点.md。内网直连一律绕系统代理(urllib ProxyHandler({}))。
单跑预算 cap:硬编码 ≤¥5,累计估算成本超即 fail(防真模型闸自身失控烧钱)。
成本按 new-api 权威计费倍率折算(M3: model_ratio=0.15 / completion_ratio=4;qpu=500000;usd_rate=7.3,
实测自 /api/pricing + /api/status,2026-07-04);台账同时记原始 token,权威成本以 new-api 对账为准。
台账:每次真跑 append 一行到 eval/<id>/runs/(只追加,不回改),记 token/成本/延迟/四闸裁决,可审计可复现。
用法:
NEWAPI_KEY=sk-xxx python3 contracts/prompts/eval_gate.py --id safety.prompt-check
NEWAPI_KEY=sk-xxx python3 contracts/prompts/eval_gate.py --changed --base HEAD~1 # 自动对改动的 live 条目跑
# 附加:--update-baseline 建/更新基线快照;--ledger-dir <dir> 台账写别处(演示用,避免污染仓)
# exit 0 = 四闸全绿(或 SKIP 豁免);exit 1 = 有闸未过 / 金标缺失 fail-closed;exit 2 = 模型跑不通(人工兜底)
零三方依赖:只用标准库(urllib / json / re)+ 调 git。
"""
import argparse
import json
import os
import re
import subprocess
import sys
import time
import urllib.request
from datetime import datetime, timezone
from pathlib import Path
PROMPTS_DIR = Path(__file__).resolve().parent
REGISTRY_FILE = PROMPTS_DIR / "registry.yaml"
EVAL_DIR = PROMPTS_DIR / "eval"
# ── 段 B 判定白名单:仅这些 live 条目才跑真模型闸(逐条依据见 registry.yaml 头部消费面三态对账总表)──
# live = 现行生产生成路径运行时读它、渲染后发给 LLM。非白名单条目一律 SKIP 豁免(fossil / batch-relic)。
LIVE_PROMPT_IDS = {
"safety.prompt-check", # SafetyCheckClient GP9 合规门(有 eval 金标集 → 真跑)
"config.cheap-system", # cheap_roles.py 便宜档 A-model(无金标集 → fail-closed)
"tier2.design-system", # 以下 8 条 tier2:roles.py prompts.load 运行时读(无金标集 → fail-closed)
"tier2.design-systems-designer",
"tier2.design-economy-designer",
"tier2.design-level-designer",
"tier2.design-presentation-designer",
"tier2.design-leader-system",
"tier2.writer-system",
"tier2.player-system",
}
# ── 硬约束参数 ────────────────────────────────────────────────────────────────
PER_RUN_BUDGET_CAP_RMB = 5.0 # 单跑预算 cap(硬编码,累计估算成本超即 fail)
MIN_SUCCESS_RATE = 0.8 # 闸2 成功率阈值(对齐 MVP「AI 生成成功率 ≥80%」)
MAX_COST_DELTA_RATIO = 0.5 # 闸4 成本增量上限(相对基线 +50%);prompt 契约未声明时的缺省
MAX_LATENCY_DELTA_RATIO = 1.0 # 闸4 延迟增量上限(相对基线 +100%);缺省
MODEL_CALL_TIMEOUT_S = 60 # 单次模型调用超时
MODEL_CALL_RETRIES = 2 # 单条最多重试次数(限流/超时 best-effort)
# M3 new-api 权威计费倍率(源:GET /api/pricing 的 MiniMax-M3 + /api/status,2026-07-04 实测)。
# quota = (prompt_tokens + completion_tokens × completion_ratio) × model_ratio(保守不计 cache 折扣);
# rmb = quota / quota_per_unit × usd_rate。权威成本以 new-api quota 对账为准,此处为 cap/闸4 用的一致口径估算。
M3_MODEL_RATIO = 0.15
M3_COMPLETION_RATIO = 4.0
NEWAPI_QUOTA_PER_UNIT = 500000.0
NEWAPI_USD_RATE = 7.3
DEFAULT_MODEL = "MiniMax-M3"
DEFAULT_BASE_URL = "http://100.64.0.8:3000" # new-api host 根(NEWAPI_BASE_URL 覆盖);Python 侧调用补 /v1
# ══════════════════════ registry / frontmatter 解析(零依赖,复用 check_registry 口径)══════════════════════
def _strip_inline_comment(value):
out, in_quote = [], None
for ch in value:
if in_quote:
if ch == in_quote:
in_quote = None
out.append(ch)
elif ch in ("'", '"'):
in_quote = ch
out.append(ch)
elif ch == '#':
break
else:
out.append(ch)
s = "".join(out).strip()
if len(s) >= 2 and s[0] == s[-1] and s[0] in ("'", '"'):
s = s[1:-1]
return s
def parse_registry(path):
"""解析 registry.yaml → {id: {file, version, eval, ...}}(最小 YAML 子集,无三方依赖)。"""
text = path.read_text(encoding="utf-8")
entries = {}
current_id = None
in_prompts = False
for line in text.splitlines():
stripped = line.strip()
if not stripped or stripped.startswith('#'):
continue
if re.match(r'^prompts\s*:\s*(#.*)?$', line):
in_prompts = True
continue
if not in_prompts:
continue
m = re.match(r'^\s*-\s+id\s*:\s*(.+)$', line)
if m:
current_id = _strip_inline_comment(m.group(1))
entries[current_id] = {"id": current_id}
continue
m = re.match(r'^\s+([A-Za-z_][A-Za-z0-9_]*)\s*:\s*(.*)$', line)
if m and current_id is not None:
entries[current_id][m.group(1)] = _strip_inline_comment(m.group(2))
continue
if re.match(r'^[A-Za-z_]', line):
in_prompts = False
current_id = None
return entries
def split_frontmatter(text):
"""返回 (frontmatter_text, body);无合法 frontmatter → ("", 原文)。"""
lines = text.split("\n")
if not lines or lines[0].strip() != "---":
return "", text
for i in range(1, len(lines)):
if lines[i].strip() == "---":
fm = "\n".join(lines[1:i])
body = "\n".join(lines[i + 1:])
return fm, body
return "", text
def parse_output_schema(frontmatter_text):
"""从 frontmatter 提取 output_schema 的 required 字段与各字段 type(JSON Schema 子集)。
只解析本项目 prompt 契约实际用到的形态(safety 为例):
output_schema:
type: object
required: [safe, reason]
properties:
safe: { type: boolean }
reason: { type: string }
返回 {"required": ["safe","reason"], "types": {"safe":"boolean","reason":"string"}};
无 output_schema 块 → None(闸1 退化为「仅要求是合法 JSON 对象」)。
更复杂的嵌套 schema 需扩展本解析(诚实边界,非静默放过)。
"""
lines = frontmatter_text.split("\n")
# 定位 output_schema: 块的缩进范围
start = None
base_indent = None
for i, line in enumerate(lines):
if re.match(r'^output_schema\s*:', line):
start = i
base_indent = len(line) - len(line.lstrip())
break
if start is None:
return None
block = []
for line in lines[start + 1:]:
if not line.strip():
block.append(line)
continue
indent = len(line) - len(line.lstrip())
if indent <= base_indent:
break # 回到同级/更外层,块结束
block.append(line)
block_text = "\n".join(block)
required = []
m = re.search(r'required\s*:\s*\[([^\]]*)\]', block_text)
if m:
required = [x.strip().strip('"\'') for x in m.group(1).split(",") if x.strip()]
types = {}
# properties 下形如 ` safe: { type: boolean }`
for line in block:
pm = re.match(r'^\s+([A-Za-z_][A-Za-z0-9_]*)\s*:\s*\{\s*type\s*:\s*([A-Za-z]+)\s*\}', line)
if pm:
types[pm.group(1)] = pm.group(2)
return {"required": required, "types": types}
# ══════════════════════ 金标集加载 ══════════════════════
def load_golden(eval_id):
"""加载 eval/<eval_id>/{inputs,labels}.jsonl,按 evalKey 关联成 [{evalKey, input, label}]。
返回 (samples, err):err 非 None 表示金标缺失/不可用(上层 fail-closed)。
"""
d = EVAL_DIR / eval_id
inputs_f = d / "inputs.jsonl"
labels_f = d / "labels.jsonl"
if not inputs_f.exists() or not labels_f.exists():
return None, f"金标集缺失:{d} 下无 inputs.jsonl / labels.jsonl"
def _read_jsonl(p):
rows = []
for line in p.read_text(encoding="utf-8").splitlines():
line = line.strip()
if not line:
continue
try:
rows.append(json.loads(line))
except json.JSONDecodeError:
continue # 坏行跳过(不阻断),不计入
return rows
inputs = {r.get("evalKey"): r for r in _read_jsonl(inputs_f) if r.get("evalKey")}
labels = {r.get("evalKey"): r for r in _read_jsonl(labels_f) if r.get("evalKey")}
keys = [k for k in inputs if k in labels]
if not keys:
return None, f"金标集为空或 inputs/labels 无对齐 evalKey:{d}"
samples = [{"evalKey": k, "input": inputs[k], "label": labels[k]} for k in keys]
return samples, None
# ══════════════════════ 模型调用(M3 via new-api,thinking 关,绕代理)══════════════════════
def _no_proxy_opener():
"""构造禁用系统代理的 urllib opener(内网直连必须绕本机 fake-ip 代理,否则栽 502/连接关闭)。"""
return urllib.request.build_opener(urllib.request.ProxyHandler({}))
def call_model(base_url, api_key, model, user_content):
"""调 M3(OpenAI 兼容 /v1/chat/completions,temperature=0,thinking 关)。
返回 (content_str, prompt_tokens, completion_tokens, latency_ms, err);
err 非 None 表示调用失败(限流/超时/HTTP 错),content 为 None。best-effort 重试 MODEL_CALL_RETRIES 次。
"""
url = base_url.rstrip("/") + "/v1/chat/completions"
payload = json.dumps({
"model": model,
"temperature": 0,
"messages": [{"role": "user", "content": user_content}],
}).encode("utf-8")
opener = _no_proxy_opener()
last_err = None
for attempt in range(1, MODEL_CALL_RETRIES + 2): # 首次 + 重试
t0 = time.time()
try:
req = urllib.request.Request(
url, data=payload,
headers={"Authorization": f"Bearer {api_key}", "Content-Type": "application/json"},
method="POST",
)
with opener.open(req, timeout=MODEL_CALL_TIMEOUT_S) as resp:
body = json.loads(resp.read().decode("utf-8"))
latency_ms = int((time.time() - t0) * 1000)
content = body["choices"][0]["message"]["content"]
usage = body.get("usage") or {}
pt = int(usage.get("prompt_tokens") or 0)
ct = int(usage.get("completion_tokens") or 0)
return content, pt, ct, latency_ms, None
except Exception as exc: # noqa: BLE001 —— best-effort:网络/超时/解析异常都进重试,不中断整轮
last_err = f"{type(exc).__name__}: {exc}"
time.sleep(1.0 * attempt) # 退避
return None, 0, 0, 0, last_err
def estimate_cost_rmb(prompt_tokens, completion_tokens):
"""按 new-api M3 权威倍率估算单次调用成本(¥);口径见文件头计费说明。"""
quota = (prompt_tokens + completion_tokens * M3_COMPLETION_RATIO) * M3_MODEL_RATIO
return quota / NEWAPI_QUOTA_PER_UNIT * NEWAPI_USD_RATE
# ══════════════════════ 产物解析 + 四闸判定 ══════════════════════
def extract_json_obj(content):
"""从模型输出抽取第一个 JSON 对象;裸 JSON 直接解析,带 markdown/前后文则正则兜。失败返回 None。"""
if content is None:
return None
try:
obj = json.loads(content.strip())
return obj if isinstance(obj, dict) else None
except json.JSONDecodeError:
pass
m = re.search(r'\{.*\}', content, re.DOTALL)
if m:
try:
obj = json.loads(m.group(0))
return obj if isinstance(obj, dict) else None
except json.JSONDecodeError:
return None
return None
_TYPE_PY = {"boolean": bool, "string": str, "number": (int, float),
"integer": int, "object": dict, "array": list}
def gate1_schema_ok(obj, schema):
"""闸1:产物合 output_schema(是对象 + 含 required 字段 + 字段类型匹配)。返回 (ok, 失败字段列表)。"""
if not isinstance(obj, dict):
return False, ["<非 JSON 对象>"]
if schema is None:
return True, [] # 无 schema 声明 → 只要求是对象(上面已保证)
bad = []
for field in schema.get("required", []):
if field not in obj:
bad.append(f"缺字段 {field}")
for field, tname in schema.get("types", {}).items():
if field in obj:
pytype = _TYPE_PY.get(tname)
if pytype and not isinstance(obj[field], pytype):
bad.append(f"字段 {field} 类型应为 {tname}")
return (len(bad) == 0), bad
def judge_correct(prompt_id, obj, label):
"""闸2 单条判对:产物是否与金标 label 一致。按 prompt 语义适配。
当前仅 safety.prompt-check 真跑(唯一 live+有金标):产物 safe 必须 == 金标 expectedSafe。
其它 live 条目补金标时在此加对应判定分支(诚实边界:未适配的 id 直接判「无判定口径」→ 不假绿)。
"""
if prompt_id == "safety.prompt-check":
return obj.get("safe") == label.get("expectedSafe")
return None # 无判定口径(上层视为未过 + 报告标注需补判定适配)
def key_verdict(prompt_id, obj):
"""闸3 关键裁决快照(回归 diff 的比对键)。safety = safe 字段。"""
if prompt_id == "safety.prompt-check":
return {"safe": obj.get("safe")} if isinstance(obj, dict) else {"safe": None}
return {}
# ══════════════════════ 基线 / 台账 ══════════════════════
def baseline_path(eval_id):
return EVAL_DIR / eval_id / "baseline.json"
def load_baseline(eval_id):
p = baseline_path(eval_id)
if not p.exists():
return None
try:
return json.loads(p.read_text(encoding="utf-8"))
except (OSError, json.JSONDecodeError):
return None
def write_ledger(ledger_dir, eval_id, record):
"""台账 append-only:eval/<id>/runs/run-<date>.jsonl 追加一行运行记录。"""
d = Path(ledger_dir) if ledger_dir else (EVAL_DIR / eval_id / "runs")
d.mkdir(parents=True, exist_ok=True)
day = datetime.now(timezone.utc).strftime("%Y%m%d")
f = d / f"run-{day}.jsonl"
with f.open("a", encoding="utf-8") as fh:
fh.write(json.dumps(record, ensure_ascii=False) + "\n")
return f
# ══════════════════════ 单条目四闸主流程 ══════════════════════
def run_gate_for_id(prompt_id, registry, base_url, api_key, model, *, update_baseline, ledger_dir):
"""对单个 prompt 条目跑段 B。返回 exit code(0 绿/SKIP,1 未过/金标缺失,2 跑不通)。"""
print(f"\n{'='*70}\n[段B·真模型闸] 目标条目:{prompt_id}\n{'='*70}")
# ── 三态门:非 live 一律 SKIP 豁免(不花真模型钱焊死面)──
if prompt_id not in LIVE_PROMPT_IDS:
print(f"[SKIP] {prompt_id} 非 live(fossil / batch-relic,见 registry 头部三态对账)→ 真模型闸豁免,放行。")
return 0
entry = registry.get(prompt_id)
if not entry:
print(f"[FAIL] {prompt_id} 不在 registry.yaml(白名单与 registry 漂移?)", file=sys.stderr)
return 1
# ── 金标集:live 但无金标 → fail-closed(绝不假绿)──
eval_id = prompt_id # eval 目录名 = prompt id(registry eval 字段亦指向 eval/<id>/)
samples, gerr = load_golden(eval_id)
if gerr:
print(f"[FAIL·fail-closed] {prompt_id} 是 live 面但金标集不可用:{gerr}")
print(" → 按纪律 fail-closed(绝不假绿)。补齐 eval/<id>/inputs.jsonl+labels.jsonl 后本闸即真跑。")
return 1
# ── 加载 prompt 正文 + output_schema ──
rel_file = entry.get("file")
prompt_text = (PROMPTS_DIR / rel_file).read_text(encoding="utf-8")
fm, body = split_frontmatter(prompt_text)
schema = parse_output_schema(fm)
version = entry.get("version")
print(f"[信息] version={version} 金标集={len(samples)} 条 output_schema={'有' if schema else '无(仅校验JSON对象)'}")
# ── 逐条真跑模型 → 收集裁决 ──
results = []
total_pt = total_ct = 0
total_latency = 0
acc_cost = 0.0
n_error = 0
for s in samples:
user_content = body.replace("{{input.prompt}}", str(s["input"].get("prompt", "")))
content, pt, ct, latency_ms, cerr = call_model(base_url, api_key, model, user_content)
total_pt += pt
total_ct += ct
total_latency += latency_ms
acc_cost += estimate_cost_rmb(pt, ct)
# 预算 cap 硬闸:累计估算成本超 ¥5 立即中止
if acc_cost > PER_RUN_BUDGET_CAP_RMB:
print(f"[FAIL·预算] 累计估算成本 ¥{acc_cost:.4f} 超单跑 cap ¥{PER_RUN_BUDGET_CAP_RMB} → 中止。")
return 1
if cerr:
n_error += 1
results.append({"evalKey": s["evalKey"], "error": cerr, "schema_ok": False, "correct": False})
print(f" · {s['evalKey']}: [调用失败] {cerr}")
continue
obj = extract_json_obj(content)
s_ok, bad = gate1_schema_ok(obj, schema)
correct = judge_correct(prompt_id, obj or {}, s["label"])
results.append({
"evalKey": s["evalKey"], "schema_ok": s_ok,
"correct": bool(correct), "verdict": key_verdict(prompt_id, obj or {}),
"schema_bad": bad,
})
flag = "OK" if (s_ok and correct) else ("schema✘" if not s_ok else "判错✘")
print(f" · {s['evalKey']}: {flag} 产物={json.dumps(obj, ensure_ascii=False) if obj else content!r}")
# ── 调用失败(网关 500 / 限流 / 超时 = 基础设施问题,非 prompt 质量)从判定分母排除,不误判 prompt ──
# 多数失败 或 全失败 → 判定不可信,交人工兜底复跑(SoT §7:eval 跑不通不阻塞、不假绿)。
responded = [r for r in results if not r.get("error")]
if not responded or n_error > len(samples) * 0.2:
print(f"[SKIP·人工兜底] {n_error}/{len(samples)} 条模型调用失败(网关/限流/超时)→ 判定不可信,交人工复跑。")
return 2
n_judged = len(responded)
# ── 闸1 Schema(只对成功响应的条目判)──
n_schema_ok = sum(1 for r in responded if r["schema_ok"])
gate1 = (n_schema_ok == n_judged)
# ── 闸2 成功率(分母=成功响应数,不惩罚基础设施失败)──
n_correct = sum(1 for r in responded if r["correct"])
rate = n_correct / n_judged
gate2 = rate >= MIN_SUCCESS_RATE
# ── 闸3 回归 diff(有基线才判;无基线 advisory)──
baseline = load_baseline(eval_id)
gate3, gate3_note = True, "无基线,首次跑(advisory,未拦)"
if baseline and baseline.get("perKey"):
flipped = []
for r in responded:
bk = baseline["perKey"].get(r["evalKey"])
if bk is not None and bk != r.get("verdict"):
flipped.append(r["evalKey"])
gate3 = (len(flipped) == 0)
gate3_note = "无关键裁决翻转" if gate3 else f"关键裁决翻转:{flipped}"
# ── 闸4 成本/延迟(有基线才判增量;无基线 advisory)──
avg_cost = acc_cost / n_judged
avg_latency = total_latency / n_judged
gate4, gate4_note = True, "无基线,首次跑(advisory,未拦)"
if baseline and baseline.get("cost") is not None:
base_cost = baseline["cost"] or 1e-9
base_lat = baseline.get("latency") or 1e-9
cost_delta = (avg_cost - base_cost) / base_cost
lat_delta = (avg_latency - base_lat) / base_lat
cost_ok = cost_delta <= MAX_COST_DELTA_RATIO
lat_ok = lat_delta <= MAX_LATENCY_DELTA_RATIO
gate4 = cost_ok and lat_ok
gate4_note = f"成本Δ={cost_delta:+.1%}(限+{MAX_COST_DELTA_RATIO:.0%}) 延迟Δ={lat_delta:+.1%}(限+{MAX_LATENCY_DELTA_RATIO:.0%})"
all_green = gate1 and gate2 and gate3 and gate4
# ── 报告 ──
print(f"\n{'-'*70}")
if n_error:
print(f"[注] {n_error}/{len(samples)} 条模型调用失败(网关/限流,非 prompt 质量),已排除出判定分母;建议复跑补齐。")
print(f"[闸1 Schema ] {'绿' if gate1 else '红'} {n_schema_ok}/{n_judged} 条产物合 output_schema")
print(f"[闸2 成功率 ] {'绿' if gate2 else '红'} 判对率 {rate:.0%}(阈值 ≥{MIN_SUCCESS_RATE:.0%},{n_correct}/{n_judged})")
print(f"[闸3 回归diff] {'绿' if gate3 else '红'} {gate3_note}")
print(f"[闸4 成本延迟] {'绿' if gate4 else '红'} {gate4_note}")
print(f"[成本] 本跑估算 ¥{acc_cost:.4f}(tokens: prompt={total_pt} completion={total_ct};cap ¥{PER_RUN_BUDGET_CAP_RMB})")
print(f"[延迟] 累计 {total_latency}ms(均 {avg_latency:.0f}ms/条)")
print(f"{'-'*70}\n[裁决] {'四闸全绿 → 放行' if all_green else '四闸未全绿 → 拦截'}")
# ── 台账 append ──
record = {
"ts": datetime.now(timezone.utc).isoformat(),
"promptId": prompt_id, "version": version, "model": model,
"nSamples": len(samples), "nJudged": n_judged, "nSchemaOk": n_schema_ok, "nCorrect": n_correct,
"successRate": round(rate, 4),
"gate1_schema": gate1, "gate2_success": gate2, "gate3_regression": gate3, "gate4_cost_latency": gate4,
"allGreen": all_green,
"estCostRmb": round(acc_cost, 6), "promptTokens": total_pt, "completionTokens": total_ct,
"totalLatencyMs": total_latency, "nError": n_error,
"gate3_note": gate3_note, "gate4_note": gate4_note,
}
ledger_f = write_ledger(ledger_dir, eval_id, record)
print(f"[台账] 已追加 → {ledger_f}")
# ── 更新基线(显式请求时)──
if update_baseline:
bl = {
"promptId": prompt_id, "version": version, "updatedAt": record["ts"],
"perKey": {r["evalKey"]: r.get("verdict") for r in results if not r.get("error")},
"cost": round(avg_cost, 6), "latency": round(avg_latency, 1),
"note": "基线快照:关键裁决 + 单条均成本/延迟;由 eval_gate --update-baseline 建立",
}
baseline_path(eval_id).write_text(json.dumps(bl, ensure_ascii=False, indent=2) + "\n", encoding="utf-8")
print(f"[基线] 已更新 → {baseline_path(eval_id)}")
return 0 if all_green else 1
# ══════════════════════ 改动条目发现(--changed)══════════════════════
def changed_prompt_ids(registry, base):
"""git diff base 找改动的 prompt .md,映射回 registry id。"""
try:
root = subprocess.run(["git", "rev-parse", "--show-toplevel"], cwd=str(PROMPTS_DIR),
capture_output=True, text=True, check=True).stdout.strip()
out = subprocess.run(["git", "diff", "--name-only", base, "--", "contracts/prompts/"],
cwd=root, capture_output=True, text=True, check=True).stdout
except (subprocess.CalledProcessError, OSError) as e:
print(f"[错误] git diff 失败:{e}", file=sys.stderr)
return []
changed_files = {ln.strip() for ln in out.splitlines() if ln.strip().endswith(".md")}
ids = []
for pid, entry in registry.items():
rel = entry.get("file")
if rel and f"contracts/prompts/{rel}" in changed_files:
ids.append(pid)
return ids
def main():
parser = argparse.ArgumentParser(description="Prompt 四道闸 · 真模型闸(段 B)")
g = parser.add_mutually_exclusive_group(required=True)
g.add_argument("--id", help="对指定 prompt id 跑真模型闸")
g.add_argument("--changed", action="store_true", help="对相对基线改动的 live 条目跑(配 --base)")
parser.add_argument("--base", default="HEAD~1", help="--changed 的对比基线(默认 HEAD~1)")
parser.add_argument("--model", default=os.environ.get("EVAL_MODEL", DEFAULT_MODEL), help="模型名(默认 MiniMax-M3)")
parser.add_argument("--update-baseline", action="store_true", help="真跑后建/更新基线快照(闸3/4 后续据此判)")
parser.add_argument("--ledger-dir", default=None, help="台账目录(默认 eval/<id>/runs/;演示可指临时目录避免污染仓)")
args = parser.parse_args()
# 凭据:从 env 读,绝不硬编码 key(权威源 docs/内网凭据与端点.md)。
api_key = os.environ.get("NEWAPI_KEY", "").strip()
base_url = os.environ.get("NEWAPI_BASE_URL", DEFAULT_BASE_URL).strip()
registry = parse_registry(REGISTRY_FILE)
# 白名单自检:LIVE_PROMPT_IDS 必须都在 registry(防漂移/typo)
drift = [pid for pid in LIVE_PROMPT_IDS if pid not in registry]
if drift:
print(f"[错误] 段B白名单与 registry 漂移,以下 id 不在 registry:{drift}", file=sys.stderr)
sys.exit(1)
if args.changed:
targets = changed_prompt_ids(registry, args.base)
if not targets:
print(f"[OK] 相对基线 {args.base} 无改动的 prompt 条目 → 真模型闸无需跑,放行。")
sys.exit(0)
print(f"[信息] 改动条目:{targets}")
else:
targets = [args.id]
# 真跑需要 key;仅当存在「live 且有金标」目标(真会调模型)时才强制 key。
def _will_call_model(pid):
if pid not in LIVE_PROMPT_IDS:
return False
samples, err = load_golden(pid)
return err is None
if any(_will_call_model(t) for t in targets) and not api_key:
print("[错误] 需真调模型但 NEWAPI_KEY 未设(读 docs/内网凭据与端点.md,经 env 传入,勿设长期 env)。",
file=sys.stderr)
sys.exit(1)
worst = 0
for t in targets:
code = run_gate_for_id(t, registry, base_url, api_key, args.model,
update_baseline=args.update_baseline, ledger_dir=args.ledger_dir)
worst = max(worst, code) if code != 2 else (worst if worst == 1 else 2)
sys.exit(worst)
if __name__ == "__main__":
main()