框架: 建立细纲回放评测冻结与评分合同
This commit is contained in:
parent
1a3204827d
commit
fbb262cb57
@ -1,22 +1,25 @@
|
||||
---
|
||||
name: judge
|
||||
description: 质量评委——二层保护节点(质量门控与 LLM-Judge)的打分器,不可被装配替换;维度体系与 rubric 以 quality-gate skill 为合同。
|
||||
description: 质量评委——二层保护节点(质量门控与 LLM-Judge)的打分器,不可被装配替换;正文与细纲回放的维度体系、rubric 以 quality-gate skill 为合同。
|
||||
tools: Read, Grep, Glob, Write
|
||||
model: opus
|
||||
---
|
||||
|
||||
你是质量评委,保护节点角色(不可被装配替换)。**功能合同=`quality-gate` skill**——维度体系(3 叙事关键+5 非关键)、达标口径、rubric 全在那里,你只执行不自造维度。报告落 `works/<书>/评审/第NNN章-评分.md`,除此不写任何文件。
|
||||
你是质量评委,保护节点角色(不可被装配替换)。**功能合同=`quality-gate` skill**——正文任务使用 `quality_gate` profile;细纲回放任务使用 `fine_outline_replay` profile。维度、证据要求和稳定性口径全在那里,你只执行不自造维度。报告按当前运行合同落临时评测目录或 `works/<书>/评审/第NNN章-评分.md`,除此不写正式创作文件。
|
||||
|
||||
## 元数据纪律(怎么用元数据)
|
||||
|
||||
- 维度体系是元数据(专题-04 质量策略族):策略加维度,quality-gate skill 更新表格,你一字不改;
|
||||
- 上下文=writer 基线包(**同证独立**):评的是"在写手所知条件下写得好不好",不索取额外资料,不拿包外信息扣分。
|
||||
- 细纲回放只使用冻结快照和结构化 `reference scaffold proxy`;不得读取目标章全文或完整目标章细纲。
|
||||
- 细纲回放每个分数必须有 `evidence` 字段;不得使用正文文风、文笔或可读性作为评分维度。
|
||||
|
||||
## 打分纪律
|
||||
|
||||
- 每维给分必附一句引文证据(好在哪/差在哪,引原句);没有证据的分数无效。
|
||||
- 同一维度复评同一章分差应 ≤0.5;严格度不因收敛压力改变,不放水不加戏。
|
||||
- 末尾「最值得改的三点」按提升空间排序:问题→根因层猜测(prompt/上下文/设定卡)→具体改法。
|
||||
- 细纲回放时,把末尾建议替换为“最值得补齐的三项结构缺口”,并标注它属于候选结构、公共大纲、卡注入、原文检索还是标准事实不确定;若两次同维分差大于 0.5,只写稳定性警告,不强行裁决。
|
||||
|
||||
## 禁区
|
||||
|
||||
|
||||
@ -1,11 +1,16 @@
|
||||
---
|
||||
name: planner
|
||||
description: 规划师——规划槽位默认绑定件,承接 planning 功能;产出结构=schema 字段清单本身,功能细节以 planning skill 为合同;产出全为草稿。
|
||||
description: 规划师——规划槽位默认绑定件,承接 planning 与 fine_outline 功能;产出结构=schema 字段清单本身,功能细节以对应 skill 为合同;产出全为草稿。
|
||||
tools: Read, Write, Grep, Glob
|
||||
model: opus
|
||||
---
|
||||
|
||||
你是这部书的总规划,规划槽位的默认绑定件。**功能合同=`planning` skill**(立项与修订同用)。产出全部不提交;未确认的规划不进生成上下文。
|
||||
你是这部书的总规划,规划槽位的默认绑定件。功能合同按本次任务二选一:
|
||||
|
||||
- `planning`:遵守 `planning` skill,负责立项与规划修订;
|
||||
- `fine_outline`:遵守 `fine-outline` skill,只产结构细纲,不写正文。
|
||||
|
||||
产出全部不提交;未确认的规划不进生成上下文。回放任务中,`fine-outline` skill 的冻结边界优先于本身份段里面向正式创作的全局规划能力。
|
||||
|
||||
## 元数据纪律(怎么用元数据)
|
||||
|
||||
|
||||
@ -1,6 +1,6 @@
|
||||
---
|
||||
name: detect
|
||||
description: 检测的功能合同(scenario: validation/consistency_check,检测槽位)。检查清单由 schema 字段自动生成——凡 aiContext 含 detection 的字段即检查项。
|
||||
description: 检测的功能合同(scenario: validation/consistency_check/fine_outline_replay,检测槽位)。检查清单由 schema 字段自动生成——凡 aiContext 含 detection 的字段即检查项。
|
||||
disable-model-invocation: true
|
||||
---
|
||||
|
||||
@ -35,6 +35,23 @@ schema 给字段加上 detection 用途,检查项自动+1,本 skill 与 detector
|
||||
|
||||
末尾:阻塞性 N 条(不修不建议采纳)/建议性 N 条 + 「本次检查用不上但设定卡缺失的字段」(设计发现)。
|
||||
|
||||
## 细纲回放分支(`fine_outline_replay`)
|
||||
|
||||
何时用:`fine-outline` 候选进入独立评分前。检测器只看冻结到 `as_of` 的规划上下文和候选,不看目标章标准事实,不把评测答案倒灌回规划侧。
|
||||
|
||||
检查对象和证据格式:
|
||||
|
||||
| 检查对象 | 检查项 | 高严重度条件 | 证据必须包含 |
|
||||
|---|---|---|---|
|
||||
| 候选结构 | 必填字段、目标章号、事件 ID、事件顺序 | 缺字段、重复事件或目标章错误导致无法评估 | 字段路径 + 候选值摘要 |
|
||||
| 因果链 | 事件触发、行动、结果方向是否自洽 | 后一事件无前置、结果与前置状态矛盾 | 事件 ID 对 + 冲突原因 |
|
||||
| 实体状态 | 角色/势力/地点/能力是否违反冻结事实 | 引用未来事实或越过 N 时点能力边界 | `sourceId` + 冻结字段 |
|
||||
| 伏笔动作 | 埋、推进、回收和不确定标记 | 将未知或未来回收写成确定事实 | 伏笔标识 + 当前台账定位 |
|
||||
| 来源引用 | `sourceRefs` 是否来自快照、是否包含目标章及以后 | 目标章/未来来源出现在候选引用中 | 来源 ID + 章号范围 |
|
||||
| 未知项纪律 | `unknowns` / `assumptions` 是否显式承载缺口 | 用无来源断言替代未知项 | 候选字段路径 |
|
||||
|
||||
报告仍然只产审查结果,不修改候选。回放中任一高严重度问题阻断该臂进入 judge;“卡里缺少目标新角色”要单列为资料覆盖发现,不冒充规划器错误。
|
||||
|
||||
## 红线
|
||||
|
||||
只产报告,不改任何创作文件;证据先行——无依据的观感问题归「建议」并标明主观;底牌信息仅用于检测判断。
|
||||
|
||||
72
.claude/skills/fine-outline/SKILL.md
Normal file
72
.claude/skills/fine-outline/SKILL.md
Normal file
@ -0,0 +1,72 @@
|
||||
---
|
||||
name: fine-outline
|
||||
description: 细纲规划合同。根据规划上下文和已授权的冻结事实,产出下一章的结构细纲草稿;不写正文、不读取目标章答案。
|
||||
disable-model-invocation: true
|
||||
---
|
||||
|
||||
# 细纲规划(`scenario: fine_outline`)
|
||||
|
||||
何时用:在正文生成前,把作品大纲、当前状态和已授权设定组织成单章结构细纲。评测回放使用同一合同,但把上下文固定为 `as_of` 冻结快照;正式创作可以按产品授权扩展到连续章节或全书规划。
|
||||
|
||||
## 输入合同
|
||||
|
||||
上下文必须由 `read-context` 组装,至少包含四层中的以下部分:
|
||||
|
||||
- **L0 当前任务**:目标章号、章节输出合同、用户意图和未知项纪律;
|
||||
- **L1 叙事现在时**:截至当前章的状态、活动线程和必要的近章结构摘要;
|
||||
- **L2 作品事实**:已确认的大纲、设定、知识卡和伏笔台账;
|
||||
- **L3 授权资料**:仅使用装配层已绑定、且来源授权允许本次用途的资料。
|
||||
|
||||
回放模式额外要求:所有来源的绝对章号或完整窗口上界必须 `<= as_of`。目标章正文、目标章细纲、目标章出场清单、未来里程碑和终态摘要不得进入规划上下文。无法证明时间边界的来源按未知处理,不凭名称或窗口号猜测。
|
||||
|
||||
## 规划步骤
|
||||
|
||||
1. 先确认章节在当前大纲弧线中的位置,写出本章戏剧目标和承接关系。
|
||||
2. 再按因果顺序拆关键事件:触发、参与者、行动、结果方向和不可逆变化必须能互相解释。
|
||||
3. 将实体只列为本章确实需要的角色、势力、地点、物品或规则,并标明本章作用;不能把目标章答案反推成实体清单。
|
||||
4. 对伏笔明确写 `埋`、`推进`、`回收` 或 `不确定`,不得把冻结资料没有证明的结果静默写成确定事实。
|
||||
5. 写出章末状态变化和下一步钩子;无法由来源支持的细节放入 `unknowns` 或 `assumptions`。
|
||||
6. 输出前逐字段自查,确认没有正文段落、对白、原文复述、目标章引用或未来来源引用。
|
||||
|
||||
## 输出合同
|
||||
|
||||
候选必须是一个结构化对象,字段完整且顺序稳定:
|
||||
|
||||
```yaml
|
||||
targetChapter: 489
|
||||
chapterGoal: "本章要完成的戏剧任务"
|
||||
keyEvents:
|
||||
- id: event-1
|
||||
order: 1
|
||||
event: "事件"
|
||||
participants: ["实体"]
|
||||
trigger: "触发因果"
|
||||
resultDirection: "结果方向"
|
||||
entities:
|
||||
- name: "实体"
|
||||
type: character
|
||||
role: "本章作用"
|
||||
foreshadowing:
|
||||
- action: 推进
|
||||
subject: "伏笔"
|
||||
evidence: "来源 ID 或 unknown"
|
||||
stateChanges: ["可观察的状态变化"]
|
||||
hook: "章末钩子"
|
||||
unknowns: ["无法由当前资料证明的内容"]
|
||||
assumptions: ["为组织结构暂时采用的假设"]
|
||||
sourceRefs: ["冻结快照中的 sourceId"]
|
||||
```
|
||||
|
||||
`sourceRefs` 是可选的回溯字段,但存在时只能引用快照登记的 `sourceId`,不能粘贴原文。`keyEvents`、`entities`、`foreshadowing`、`stateChanges`、`unknowns`、`assumptions` 必须是数组;`chapterGoal` 和 `hook` 必须是字符串。
|
||||
|
||||
## 回放专用红线
|
||||
|
||||
- 不得使用参考作品的目标章细纲作为生成输入;它只能在独立评审侧作为 `reference scaffold proxy`;
|
||||
- 不得输出正文、场景对白、完整原文摘要或“我猜原书下一章是……”之类的答案复述;
|
||||
- 不得因为知识卡存在而减少对公共大纲和叙事现在时的使用;卡是检索索引和事实补充,不替代公共上下文;
|
||||
- 正确卡、无卡、错配卡三臂只能改变卡注入分区,其他任务、模型、预算和输出合同保持一致;
|
||||
- 不确定事实显式留在 `unknowns`,审查智能体据此区分资料缺失和规划错误。
|
||||
|
||||
## 产物边界
|
||||
|
||||
候选细纲属于 Shadow 草稿,只进入临时评测目录或 `works/<书>/评审/` 的运行噪音,不直接写入正式大纲、知识库或正文。评测最终只保留结构评分、摘要、章节定位、阻断类别和哈希。
|
||||
46
.claude/skills/fine-outline/scripts/test_contract.py
Normal file
46
.claude/skills/fine-outline/scripts/test_contract.py
Normal file
@ -0,0 +1,46 @@
|
||||
#!/usr/bin/env python3
|
||||
"""细纲功能合同的离线回归测试,不启动模型、不访问作品数据。"""
|
||||
|
||||
from pathlib import Path
|
||||
import unittest
|
||||
|
||||
|
||||
ROOT = Path(__file__).resolve().parents[4]
|
||||
SKILL = (ROOT / ".claude/skills/fine-outline/SKILL.md").read_text(encoding="utf-8")
|
||||
PLANNER = (ROOT / ".claude/agents/planner.md").read_text(encoding="utf-8")
|
||||
CHAINS = (ROOT / "meta/chains/README.md").read_text(encoding="utf-8")
|
||||
|
||||
|
||||
class FineOutlineContractTest(unittest.TestCase):
|
||||
def test_contract_is_bound_to_existing_planner_slot(self):
|
||||
self.assertIn("scenario: fine_outline", SKILL)
|
||||
self.assertIn("fine_outline", PLANNER)
|
||||
self.assertIn("fine_outline", CHAINS)
|
||||
self.assertIn("规划→planner", CHAINS)
|
||||
|
||||
def test_candidate_fields_and_shadow_boundary_are_explicit(self):
|
||||
for field in (
|
||||
"targetChapter",
|
||||
"chapterGoal",
|
||||
"keyEvents",
|
||||
"entities",
|
||||
"foreshadowing",
|
||||
"stateChanges",
|
||||
"hook",
|
||||
"unknowns",
|
||||
"assumptions",
|
||||
):
|
||||
self.assertIn(field, SKILL)
|
||||
self.assertIn("不得输出正文", SKILL)
|
||||
self.assertIn("Shadow", SKILL)
|
||||
self.assertIn("目标章细纲作为生成输入", SKILL)
|
||||
|
||||
def test_freeze_and_two_line_retrieval_rules_are_explicit(self):
|
||||
self.assertIn("<= as_of", SKILL)
|
||||
self.assertIn("卡是检索索引和事实补充", SKILL)
|
||||
self.assertIn("目标章正文", SKILL)
|
||||
self.assertIn("unknowns", SKILL)
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
unittest.main()
|
||||
@ -1,6 +1,6 @@
|
||||
---
|
||||
name: quality-gate
|
||||
description: 质量评分的功能合同(scenario: quality_gate,保护节点)。维度体系对齐专题-04:3 叙事关键定达标线+5 非关键出建议;rubric 与达标口径的单一来源。
|
||||
description: 质量评分的功能合同(scenario: quality_gate/fine_outline_replay,保护节点)。正文质量维度和细纲回放 rubric 分开管理;rubric 与达标口径的单一来源。
|
||||
disable-model-invocation: true
|
||||
---
|
||||
|
||||
@ -33,3 +33,28 @@ disable-model-invocation: true
|
||||
## 输出合同
|
||||
|
||||
报告落 `works/<书>/评审/第NNN章-评分.md`:逐维分数+引文证据;达标结论(关键三维口径);「最值得改的三点」按提升空间排序,每点:问题→根因层猜测(prompt/上下文/设定卡)→具体改法。
|
||||
|
||||
## 细纲回放 profile(`fine_outline_replay`)
|
||||
|
||||
细纲评分是结构事实评估,不复用正文 `style_fit`、`readability`、`文风一致性`、`文笔` 等维度。每维 1–5 分,分数必须有结构化证据摘要;评委只使用冻结快照和评测侧的 `reference scaffold proxy`,不读取目标章全文。
|
||||
|
||||
| 维度 ID | 评分问题 |
|
||||
|---|---|
|
||||
| `structure_completeness` | 目标、关键事件、实体、伏笔、状态变化和钩子是否齐全 |
|
||||
| `direction_causality` | 关键冲突、结果方向和事件因果是否命中且成立 |
|
||||
| `order_pacing` | 事件先后和章内推进节拍是否接近结构化标准 |
|
||||
| `entity_state` | 实体身份、阵营、能力层级和 N 时点状态是否一致 |
|
||||
| `foreshadowing_action` | 伏笔埋设、推进、回收和时机是否正确 |
|
||||
| `handoff_hook` | 能否从 N 时点自然承接,并立住下一步钩子 |
|
||||
|
||||
评审报告使用以下最小形状:
|
||||
|
||||
```yaml
|
||||
profile: fine_outline_replay
|
||||
scores:
|
||||
direction_causality:
|
||||
score: 4
|
||||
evidence: "事件-2 的触发和结果方向与标准事实摘要一致"
|
||||
```
|
||||
|
||||
两次独立评审的同维差异大于 `0.5` 时只输出稳定性警告,不据此下卡效用结论。三臂顺序、arm 名称和卡 manifest 对评委隐藏;汇总报告分别给出每维分数、结构门结果、`B-A` 和 `C-A`,不压成没有统计意义的单一“卡质量总分”。
|
||||
|
||||
@ -7,7 +7,7 @@ description: 组装智能体上下文包——统一创作数据读取器(专题
|
||||
|
||||
SoT 对齐:包结构=专题-03 §4.2 **四层上下文**;字段级 aiContext 裁剪=专题-06 §7 统一读取器。
|
||||
|
||||
**输入**:作品名、功能(scenario,专题-03 §4.1:continuation/rewrite/expansion/polish/extraction/full_parse/planning/validation/consistency_check/quality_gate)、目标(如"写第7章")。**用途(purpose)由 scenario 映射**——生成类→`generation`,抽取类→`extraction`,规划→`planning`,检测类→`detection`,quality_gate→基线包即 writer 视图;purpose 定字段可见集(aiContext),scenario 定功能指令段(`meta/chains/`)与 L0 形态,两者不混。
|
||||
**输入**:作品名、功能(scenario,专题-03 §4.1:continuation/rewrite/expansion/polish/extraction/full_parse/planning/fine_outline/validation/consistency_check/quality_gate)、目标(如"写第7章")。**用途(purpose)由 scenario 映射**——生成类→`generation`,抽取类→`extraction`,规划类(`planning`/`fine_outline`)→`planning`,检测→`detection`,quality_gate→基线包即 writer 视图;purpose 定字段可见集(aiContext),scenario 定功能指令段(`meta/chains/`)与 L0 形态,两者不混。
|
||||
|
||||
## 步骤
|
||||
|
||||
@ -46,6 +46,7 @@ SoT 对齐:包结构=专题-03 §4.2 **四层上下文**;字段级 aiContext 裁
|
||||
|---|---|---|---|---|
|
||||
| writer(generation) | 本章细纲+任务 | **必**:上章尾 1–2 场景原文+前章摘要+状态 | 必:设定(裁底牌)+近三章细纲+**仅出场卡** | 按本章场景类型选绑定范式卡 |
|
||||
| planner(planning) | 规划任务+范围 | 可省:近章摘要即可 | **必**:可见度最高——底牌字段开+未来卷粗纲开+知识卡**全量索引** | trope/公式类范式优先 |
|
||||
| planner(fine_outline) | 目标章细纲任务+未知项纪律 | 可省:只读冻结视图中的近章结构摘要 | **必**:当前大纲、叙事现在时、截至冻结点的安全设定/卡索引;回放时关闭目标章及未来来源 | 仅使用已授权且通过冻结审计的资料 |
|
||||
| extractor(extraction) | 抽取任务+目标型清单 | 待处理章**全文**(是处理对象,不截尾) | schema 字段合同+既有知识**全集**(判重/判冲突) | **关**:防外部范式诱导脑补 |
|
||||
| detector(detection) | 检测任务+待检候选全文 | 必:近邻正文(连续性比对) | **必**:知识全集+状态台账+**底牌开**(查提前泄底必须知道谜底) | 范式卡「失效风险」字段可选 |
|
||||
| judge(quality_gate) | 待评候选+rubric | 随基线包 | **=writer 基线**(共享前缀;评设定一致性需同一基准,不另加底牌) | **关**:不拿范式当标准答案 |
|
||||
|
||||
38
.claude/skills/replay-eval/SKILL.md
Normal file
38
.claude/skills/replay-eval/SKILL.md
Normal file
@ -0,0 +1,38 @@
|
||||
---
|
||||
name: replay-eval
|
||||
description: 回放评测的冻结与结果边界合同。把参考作品冻结到 as_of 章号,生成可审计的输入清单,并在生成/评分前阻断未来信息、未授权来源和全文留存。
|
||||
disable-model-invocation: true
|
||||
---
|
||||
|
||||
# 回放评测(`next_fine_outline_replay_v0`)
|
||||
|
||||
本 skill 只负责确定性的评测编排边界,不调用模型、不替代统一读取器,也不写正式规划或知识。评测目的、三臂定义和细纲评分合同见 `docs/2026-07-19-回放评测-细纲首跑设计与计划.md`。
|
||||
|
||||
## 输入合同
|
||||
|
||||
- `reference_work`:参考作品标识和版本。
|
||||
- `as_of`:冻结章号,必须是正整数;所有历史证据只能来自绝对章号 `<= as_of`。
|
||||
- `snapshot_version`:不可变的快照版本。
|
||||
- 作品大纲窗口:使用 `from_order` / `to_order`,只允许完整窗口 `to_order <= as_of`。
|
||||
- 实体卡和里程碑:必须带可验证的绝对章号;无明确上界的区间不进入快照。
|
||||
- 来源合同:每个来源要有 `sourceId`、`sourceVersion`、用途授权快照和来源状态。
|
||||
|
||||
## 冻结规则
|
||||
|
||||
1. 章号支持整数和数字字符串,拒绝布尔值、浮点猜测、空值和无法证明上界的字符串。
|
||||
2. 里程碑按绝对章号过滤;未来条目、跨过 `as_of` 的区间和无章号条目进入 `omittedSources`,不改写成当前事实。
|
||||
3. 窗级大纲按 `to_order` 过滤并按 `from_order` 排序;`window_no` 只是展示字段,不能作为冻结键。
|
||||
4. 终态摘要、未来弧线、原始当前态等存储字段不能直接透传;只能保留安全历史并由后续消费方明确标记推导状态。
|
||||
5. 任何目标章事实、目标实体清单或目标章标签不得参与规划器侧来源选择。
|
||||
|
||||
## 输出合同
|
||||
|
||||
输出是临时 `snapshot_manifest`:记录来源 ID/版本、SHA-256、字符数、字段裁剪、来源省略和快照版本,不保存原书正文。三臂的 manifest 除卡注入分区外必须字节一致;不一致时整组作废。
|
||||
|
||||
授权、来源版本、目标章禁读或内容级泄露检查任一失败,都必须 fail-closed,返回明确的阻断状态,不靠重试绕过。
|
||||
|
||||
## 产物边界
|
||||
|
||||
- 原始候选、标准事实摘要和完整输入只能留在临时运行目录或 `/tmp`。
|
||||
- 最终报告只允许评分、摘要、章节定位、失败类别和哈希。
|
||||
- 禁止写入原书正文、完整目标细纲、完整 Prompt/Response、供应商原始响应、token、密钥或未脱敏授权资料。
|
||||
340
.claude/skills/replay-eval/scripts/build_snapshot.py
Normal file
340
.claude/skills/replay-eval/scripts/build_snapshot.py
Normal file
@ -0,0 +1,340 @@
|
||||
#!/usr/bin/env python3
|
||||
"""回放评测的冻结快照纯函数。
|
||||
|
||||
本模块只处理内存中的结构化数据,不访问数据库、不调用模型、不写正式数据。
|
||||
它把「截至第 N 章」变成可重复、可审计的来源清单,避免把未来事实误当成卡的价值。
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import argparse
|
||||
import copy
|
||||
import hashlib
|
||||
import json
|
||||
import re
|
||||
from pathlib import Path
|
||||
from typing import Any, Iterable, Mapping, Sequence
|
||||
|
||||
|
||||
class SnapshotError(ValueError):
|
||||
"""快照输入违反冻结或全文留存边界。"""
|
||||
|
||||
|
||||
TERMINAL_FIELDS = frozenset(
|
||||
{
|
||||
"final_summary",
|
||||
"terminal_summary",
|
||||
"final_state",
|
||||
"current_state",
|
||||
"future_arc",
|
||||
"future_plan",
|
||||
"成长弧线",
|
||||
"当前态",
|
||||
"终态摘要",
|
||||
"终局状态",
|
||||
"未来弧线",
|
||||
"未来计划",
|
||||
}
|
||||
)
|
||||
|
||||
FINAL_REPORT_FORBIDDEN_FIELDS = frozenset(
|
||||
{
|
||||
"raw",
|
||||
"raw_text",
|
||||
"body",
|
||||
"content",
|
||||
"full_text",
|
||||
"full_body",
|
||||
"原文",
|
||||
"正文",
|
||||
"正文全文",
|
||||
"完整目标细纲",
|
||||
"target_chapter_text",
|
||||
"prompt",
|
||||
"response",
|
||||
}
|
||||
)
|
||||
|
||||
_INTEGER_RE = re.compile(r"^\s*(\d+)\s*$")
|
||||
_RANGE_RE = re.compile(r"^\s*(?:第\s*)?(\d+)\s*(?:-|–|—|~|至|到)\s*(?:第\s*)?(\d+)\s*(?:章)?\s*$")
|
||||
|
||||
|
||||
def normalize_chapter(value: Any) -> int | None:
|
||||
"""只接受明确的正整数章号;不把 bool、浮点或模糊文本猜成章号。"""
|
||||
|
||||
if isinstance(value, bool):
|
||||
return None
|
||||
if isinstance(value, int):
|
||||
return value if value > 0 else None
|
||||
if not isinstance(value, str):
|
||||
return None
|
||||
matched = _INTEGER_RE.fullmatch(value)
|
||||
if not matched:
|
||||
return None
|
||||
chapter = int(matched.group(1))
|
||||
return chapter if chapter > 0 else None
|
||||
|
||||
|
||||
def normalize_chapter_range(value: Any) -> tuple[int, int] | None:
|
||||
"""解析单章或有明确起止上界的章区间。"""
|
||||
|
||||
chapter = normalize_chapter(value)
|
||||
if chapter is not None:
|
||||
return chapter, chapter
|
||||
if not isinstance(value, str):
|
||||
return None
|
||||
matched = _RANGE_RE.fullmatch(value)
|
||||
if not matched:
|
||||
return None
|
||||
start, end = int(matched.group(1)), int(matched.group(2))
|
||||
if start <= 0 or end < start:
|
||||
return None
|
||||
return start, end
|
||||
|
||||
|
||||
def _chapter_value(record: Mapping[str, Any]) -> Any:
|
||||
"""兼容实验台和 schema 中的中英文章号键。"""
|
||||
|
||||
for key in ("chapter", "chapter_no", "order_no", "章", "章号"):
|
||||
if key in record:
|
||||
return record[key]
|
||||
if "from_order" in record or "to_order" in record:
|
||||
start = record.get("from_order")
|
||||
end = record.get("to_order")
|
||||
if start is not None and end is not None:
|
||||
return f"{start}-{end}"
|
||||
return None
|
||||
|
||||
|
||||
def _safe_json(value: Any) -> str:
|
||||
"""用固定格式序列化,确保相同输入产生相同哈希。"""
|
||||
|
||||
return json.dumps(value, ensure_ascii=False, sort_keys=True, separators=(",", ":"))
|
||||
|
||||
|
||||
def sha256_value(value: Any) -> str:
|
||||
"""返回结构化值的稳定 SHA-256。"""
|
||||
|
||||
return hashlib.sha256(_safe_json(value).encode("utf-8")).hexdigest()
|
||||
|
||||
|
||||
def _omitted(index: int, reason: str, record: Any) -> dict[str, Any]:
|
||||
"""只记录定位和原因,不把被排除的原文复制到回显。"""
|
||||
|
||||
return {"index": index, "reason": reason, "recordHash": sha256_value(record)}
|
||||
|
||||
|
||||
def filter_milestones(
|
||||
milestones: Iterable[Mapping[str, Any]], as_of: int
|
||||
) -> tuple[list[dict[str, Any]], list[dict[str, Any]]]:
|
||||
"""保留完整落在 as_of 以前的里程碑,返回(保留项、排除项)。"""
|
||||
|
||||
as_of = normalize_chapter(as_of)
|
||||
if as_of is None:
|
||||
raise SnapshotError("as_of 必须是正整数章号")
|
||||
|
||||
kept: list[dict[str, Any]] = []
|
||||
omitted: list[dict[str, Any]] = []
|
||||
for index, item in enumerate(milestones):
|
||||
if not isinstance(item, Mapping):
|
||||
omitted.append(_omitted(index, "not_an_object", item))
|
||||
continue
|
||||
bounds = normalize_chapter_range(_chapter_value(item))
|
||||
if bounds is None:
|
||||
omitted.append(_omitted(index, "missing_or_unbounded_chapter", item))
|
||||
continue
|
||||
if bounds[1] > as_of:
|
||||
omitted.append(_omitted(index, "future_or_crosses_as_of", item))
|
||||
continue
|
||||
kept.append(copy.deepcopy(dict(item)))
|
||||
|
||||
kept.sort(key=lambda item: normalize_chapter_range(_chapter_value(item)) or (0, 0))
|
||||
return kept, omitted
|
||||
|
||||
|
||||
def filter_outline_windows(
|
||||
windows: Iterable[Mapping[str, Any]], as_of: int
|
||||
) -> tuple[list[dict[str, Any]], list[dict[str, Any]]]:
|
||||
"""只保留 to_order<=as_of 的完整窗口,并按 from_order 排序。"""
|
||||
|
||||
as_of = normalize_chapter(as_of)
|
||||
if as_of is None:
|
||||
raise SnapshotError("as_of 必须是正整数章号")
|
||||
|
||||
kept: list[dict[str, Any]] = []
|
||||
omitted: list[dict[str, Any]] = []
|
||||
for index, item in enumerate(windows):
|
||||
if not isinstance(item, Mapping):
|
||||
omitted.append(_omitted(index, "not_an_object", item))
|
||||
continue
|
||||
start = normalize_chapter(item.get("from_order"))
|
||||
end = normalize_chapter(item.get("to_order"))
|
||||
if start is None or end is None or start > end:
|
||||
omitted.append(_omitted(index, "missing_or_invalid_window_bounds", item))
|
||||
continue
|
||||
if end > as_of:
|
||||
omitted.append(_omitted(index, "future_or_crosses_as_of", item))
|
||||
continue
|
||||
kept.append(copy.deepcopy(dict(item)))
|
||||
|
||||
kept.sort(key=lambda item: (normalize_chapter(item["from_order"]) or 0, normalize_chapter(item["to_order"]) or 0))
|
||||
return kept, omitted
|
||||
|
||||
|
||||
def remove_terminal_fields(value: Any) -> Any:
|
||||
"""递归删除不能直接作为 as_of 事实使用的终态/未来字段。"""
|
||||
|
||||
if isinstance(value, list):
|
||||
return [remove_terminal_fields(item) for item in value]
|
||||
if not isinstance(value, Mapping):
|
||||
return copy.deepcopy(value)
|
||||
|
||||
projected: dict[str, Any] = {}
|
||||
for key, item in value.items():
|
||||
if str(key) in TERMINAL_FIELDS:
|
||||
continue
|
||||
projected[str(key)] = remove_terminal_fields(item)
|
||||
return projected
|
||||
|
||||
|
||||
def _validate_final_report(value: Any, path: str = "finalReport") -> None:
|
||||
"""禁止最终报告携带原书全文、完整响应或完整 prompt。"""
|
||||
|
||||
if isinstance(value, Mapping):
|
||||
for key, item in value.items():
|
||||
if str(key) in FINAL_REPORT_FORBIDDEN_FIELDS:
|
||||
raise SnapshotError(f"{path}.{key} 不得进入最终报告")
|
||||
_validate_final_report(item, f"{path}.{key}")
|
||||
elif isinstance(value, list):
|
||||
for index, item in enumerate(value):
|
||||
_validate_final_report(item, f"{path}[{index}]")
|
||||
|
||||
|
||||
def _source_record(section: str, value: Any) -> dict[str, Any]:
|
||||
"""将来源内容收敛成元数据,绝不把 payload 写入 manifest。"""
|
||||
|
||||
if isinstance(value, Mapping) and "payload" in value:
|
||||
source_id = str(value.get("sourceId") or section)
|
||||
source_version = str(value.get("sourceVersion") or "unknown")
|
||||
payload = value["payload"]
|
||||
omitted_fields = list(value.get("omittedFields") or [])
|
||||
omitted_sources = list(value.get("omittedSources") or [])
|
||||
else:
|
||||
source_id = section
|
||||
source_version = "unknown"
|
||||
payload = value
|
||||
omitted_fields = []
|
||||
omitted_sources = []
|
||||
|
||||
serialized = _safe_json(payload)
|
||||
return {
|
||||
"section": section,
|
||||
"sourceId": source_id,
|
||||
"sourceVersion": source_version,
|
||||
"sha256": hashlib.sha256(serialized.encode("utf-8")).hexdigest(),
|
||||
"charCount": len(serialized),
|
||||
"omittedFields": omitted_fields,
|
||||
"omittedSources": omitted_sources,
|
||||
}
|
||||
|
||||
|
||||
def build_snapshot_manifest(
|
||||
*,
|
||||
as_of: int,
|
||||
snapshot_version: str,
|
||||
sections: Mapping[str, Any],
|
||||
omitted_fields: Sequence[Any] | None = None,
|
||||
omitted_sources: Sequence[Any] | None = None,
|
||||
final_report: Mapping[str, Any] | None = None,
|
||||
) -> dict[str, Any]:
|
||||
"""构造稳定 manifest;最终报告只允许摘要、定位、评分和哈希。"""
|
||||
|
||||
normalized_as_of = normalize_chapter(as_of)
|
||||
if normalized_as_of is None:
|
||||
raise SnapshotError("as_of 必须是正整数章号")
|
||||
if not snapshot_version.strip():
|
||||
raise SnapshotError("snapshot_version 不能为空")
|
||||
if final_report is not None:
|
||||
_validate_final_report(final_report)
|
||||
|
||||
manifest: dict[str, Any] = {
|
||||
"snapshotVersion": snapshot_version,
|
||||
"asOfChapter": normalized_as_of,
|
||||
"sections": {
|
||||
str(section): _source_record(str(section), value)
|
||||
for section, value in sorted(sections.items(), key=lambda pair: str(pair[0]))
|
||||
},
|
||||
"omittedFields": list(omitted_fields or []),
|
||||
"omittedSources": list(omitted_sources or []),
|
||||
"finalReport": copy.deepcopy(final_report or {}),
|
||||
}
|
||||
manifest["manifestSha256"] = sha256_value(manifest)
|
||||
return manifest
|
||||
|
||||
|
||||
def build_snapshot(data: Mapping[str, Any], as_of: int, snapshot_version: str) -> dict[str, Any]:
|
||||
"""从最小 JSON 输入生成冻结后的结构化快照和 manifest。"""
|
||||
|
||||
safe = copy.deepcopy(dict(data))
|
||||
all_omitted: list[dict[str, Any]] = []
|
||||
|
||||
if isinstance(safe.get("milestones"), list):
|
||||
milestones, omitted = filter_milestones(safe["milestones"], as_of)
|
||||
safe["milestones"] = milestones
|
||||
all_omitted.extend({"source": "milestones", **item} for item in omitted)
|
||||
|
||||
if isinstance(safe.get("outlineWindows"), list):
|
||||
windows, omitted = filter_outline_windows(safe["outlineWindows"], as_of)
|
||||
safe["outlineWindows"] = windows
|
||||
all_omitted.extend({"source": "outlineWindows", **item} for item in omitted)
|
||||
|
||||
if isinstance(safe.get("cards"), list):
|
||||
frozen_cards: list[dict[str, Any]] = []
|
||||
for card_index, card in enumerate(safe["cards"]):
|
||||
if not isinstance(card, Mapping):
|
||||
all_omitted.append(_omitted(card_index, "card_not_an_object", card))
|
||||
continue
|
||||
projected = copy.deepcopy(dict(card))
|
||||
history_key = next(
|
||||
(key for key in ("milestones", "演变历程") if isinstance(projected.get(key), list)),
|
||||
None,
|
||||
)
|
||||
if history_key is not None:
|
||||
history, omitted = filter_milestones(projected[history_key], as_of)
|
||||
projected[history_key] = history
|
||||
all_omitted.extend(
|
||||
{"source": f"cards[{card_index}].{history_key}", **item}
|
||||
for item in omitted
|
||||
)
|
||||
frozen_cards.append(projected)
|
||||
safe["cards"] = frozen_cards
|
||||
|
||||
safe = remove_terminal_fields(safe)
|
||||
manifest = build_snapshot_manifest(
|
||||
as_of=as_of,
|
||||
snapshot_version=snapshot_version,
|
||||
sections={"snapshot": {"sourceId": "frozen-snapshot", "sourceVersion": snapshot_version, "payload": safe}},
|
||||
omitted_sources=all_omitted,
|
||||
)
|
||||
return {"snapshot": safe, "manifest": manifest}
|
||||
|
||||
|
||||
def _parse_args() -> argparse.Namespace:
|
||||
parser = argparse.ArgumentParser(description="构造 as_of 章号冻结快照")
|
||||
parser.add_argument("--input", type=Path, required=True, help="结构化 JSON 输入")
|
||||
parser.add_argument("--output", type=Path, required=True, help="输出 JSON 路径")
|
||||
parser.add_argument("--as-of", type=int, required=True, dest="as_of")
|
||||
parser.add_argument("--snapshot-version", default="next_fine_outline_replay_v0")
|
||||
return parser.parse_args()
|
||||
|
||||
|
||||
def main() -> int:
|
||||
args = _parse_args()
|
||||
data = json.loads(args.input.read_text(encoding="utf-8"))
|
||||
result = build_snapshot(data, args.as_of, args.snapshot_version)
|
||||
args.output.write_text(_safe_json(result) + "\n", encoding="utf-8")
|
||||
return 0
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
raise SystemExit(main())
|
||||
315
.claude/skills/replay-eval/scripts/check_snapshot.py
Normal file
315
.claude/skills/replay-eval/scripts/check_snapshot.py
Normal file
@ -0,0 +1,315 @@
|
||||
#!/usr/bin/env python3
|
||||
"""回放快照的 fail-closed 校验。
|
||||
|
||||
输入是冻结脚本产出的结构化 manifest 和运行时登记信息;本模块不访问数据库、不调用模型。
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import argparse
|
||||
import json
|
||||
from pathlib import Path
|
||||
from typing import Any, Mapping, Sequence
|
||||
|
||||
from build_snapshot import normalize_chapter, normalize_chapter_range
|
||||
|
||||
|
||||
STATUS_READY = "ready"
|
||||
STATUS_BLOCKED_AUTHORIZATION = "blocked_authorization"
|
||||
STATUS_INVALID_SNAPSHOT = "invalid_snapshot"
|
||||
STATUS_INVALID_ARM_DIFF = "invalid_arm_diff"
|
||||
STATUS_TARGET_SOURCE_FORBIDDEN = "target_source_forbidden"
|
||||
STATUS_SCHEMA_INVALID = "schema_invalid"
|
||||
|
||||
FORBIDDEN_SOURCE_STATUSES = frozenset(
|
||||
{"revoked", "delisted", "recalled", "blocked", "owner_missing", "unauthorized"}
|
||||
)
|
||||
ALLOWED_SOURCE_STATUSES = frozenset({"active", "approved", "authorized", "licensed"})
|
||||
FORBIDDEN_COPYRIGHT_STATUSES = frozenset(
|
||||
{"unauthorized", "unlicensed", "revoked", "expired", "blocked"}
|
||||
)
|
||||
CARD_KEYS = frozenset(
|
||||
{
|
||||
"arm",
|
||||
"card",
|
||||
"cards",
|
||||
"cardInjection",
|
||||
"cardSection",
|
||||
"cardSections",
|
||||
"cardManifest",
|
||||
"cardIds",
|
||||
"cardSourceIds",
|
||||
"cardStrategy",
|
||||
"l2-card",
|
||||
"l2-placebo",
|
||||
"card_injection",
|
||||
"card_section",
|
||||
"card_sections",
|
||||
"card_manifest",
|
||||
"card_ids",
|
||||
"card_source_ids",
|
||||
"card_strategy",
|
||||
"l2_card",
|
||||
"l2_placebo",
|
||||
}
|
||||
)
|
||||
NORMALIZED_CARD_KEYS = frozenset(key.lower() for key in CARD_KEYS)
|
||||
REQUIRED_CANDIDATE_FIELDS = (
|
||||
"targetChapter",
|
||||
"chapterGoal",
|
||||
"keyEvents",
|
||||
"entities",
|
||||
"foreshadowing",
|
||||
"stateChanges",
|
||||
"hook",
|
||||
"unknowns",
|
||||
"assumptions",
|
||||
)
|
||||
FORBIDDEN_CANDIDATE_FIELDS = frozenset(
|
||||
{"body", "raw", "rawText", "content", "正文", "原文", "正文全文", "完整目标细纲"}
|
||||
)
|
||||
|
||||
|
||||
def _result(status: str, errors: Sequence[str] = (), warnings: Sequence[str] = ()) -> dict[str, Any]:
|
||||
return {"status": status, "ok": status == STATUS_READY, "errors": list(errors), "warnings": list(warnings)}
|
||||
|
||||
|
||||
def _field(value: Mapping[str, Any], *keys: str) -> Any:
|
||||
for key in keys:
|
||||
if key in value:
|
||||
return value[key]
|
||||
return None
|
||||
|
||||
|
||||
def check_authorization(authorization: Mapping[str, Any] | None) -> dict[str, Any]:
|
||||
"""授权信息缺失、用途不符或来源进入危险状态时关闭评测。"""
|
||||
|
||||
if not isinstance(authorization, Mapping):
|
||||
return _result(STATUS_BLOCKED_AUTHORIZATION, ["缺少不可变授权快照"])
|
||||
snapshot = _field(authorization, "authorizationSnapshot", "authorization_snapshot")
|
||||
if not isinstance(snapshot, Mapping) or not snapshot:
|
||||
return _result(STATUS_BLOCKED_AUTHORIZATION, ["授权快照为空"])
|
||||
|
||||
source_status = str(_field(authorization, "sourceStatus", "source_status") or "").lower()
|
||||
if not source_status:
|
||||
return _result(STATUS_BLOCKED_AUTHORIZATION, ["缺少 sourceStatus"])
|
||||
if source_status in FORBIDDEN_SOURCE_STATUSES:
|
||||
return _result(STATUS_BLOCKED_AUTHORIZATION, [f"来源状态禁止评测: {source_status}"])
|
||||
if source_status not in ALLOWED_SOURCE_STATUSES:
|
||||
return _result(STATUS_BLOCKED_AUTHORIZATION, [f"来源状态未登记,拒绝评测: {source_status}"])
|
||||
|
||||
copyright_status = str(
|
||||
_field(authorization, "copyrightStatus", "copyright_status") or ""
|
||||
).lower()
|
||||
if copyright_status in FORBIDDEN_COPYRIGHT_STATUSES:
|
||||
return _result(STATUS_BLOCKED_AUTHORIZATION, [f"版权状态禁止评测: {copyright_status}"])
|
||||
|
||||
allowed = _field(authorization, "allowedPurpose", "allowed_purpose")
|
||||
if isinstance(allowed, str):
|
||||
allowed = [allowed]
|
||||
if not isinstance(allowed, Sequence) or isinstance(allowed, (str, bytes)):
|
||||
return _result(STATUS_BLOCKED_AUTHORIZATION, ["allowedPurpose 不是用途列表"])
|
||||
if "offline_evaluation" not in allowed:
|
||||
return _result(STATUS_BLOCKED_AUTHORIZATION, ["allowedPurpose 不包含 offline_evaluation"])
|
||||
return _result(STATUS_READY)
|
||||
|
||||
|
||||
def check_target_sources(target_chapter: int, sources: Sequence[Mapping[str, Any]]) -> dict[str, Any]:
|
||||
"""规划器侧来源不得包含目标章或更晚章号。"""
|
||||
|
||||
target = normalize_chapter(target_chapter)
|
||||
if target is None:
|
||||
return _result(STATUS_INVALID_SNAPSHOT, ["目标章号无效"])
|
||||
errors: list[str] = []
|
||||
for index, source in enumerate(sources):
|
||||
if not isinstance(source, Mapping):
|
||||
errors.append(f"source[{index}] 不是对象")
|
||||
continue
|
||||
value = _field(source, "chapter", "chapterNo", "chapter_no", "章", "章号")
|
||||
bounds = normalize_chapter_range(value)
|
||||
if bounds is None:
|
||||
value = _field(
|
||||
source,
|
||||
"chapterRange",
|
||||
"chapter_range",
|
||||
"range",
|
||||
"章节范围",
|
||||
)
|
||||
bounds = normalize_chapter_range(value)
|
||||
if bounds is None and "from_order" in source:
|
||||
value = f"{source.get('from_order')}-{source.get('to_order')}"
|
||||
bounds = normalize_chapter_range(value)
|
||||
if bounds is not None and bounds[1] >= target:
|
||||
errors.append(f"source[{index}] 包含目标章或未来章: {bounds[0]}-{bounds[1]}")
|
||||
return _result(STATUS_TARGET_SOURCE_FORBIDDEN if errors else STATUS_READY, errors)
|
||||
|
||||
|
||||
def _without_card_fields(value: Any) -> Any:
|
||||
"""去掉三臂允许变化的 arm/card 分区,保留所有公共输入用于字节级比较。"""
|
||||
|
||||
if isinstance(value, list):
|
||||
return [_without_card_fields(item) for item in value]
|
||||
if not isinstance(value, Mapping):
|
||||
return value
|
||||
result: dict[str, Any] = {}
|
||||
for key, item in value.items():
|
||||
key_text = str(key)
|
||||
if key_text.lower() in NORMALIZED_CARD_KEYS or key_text in {"卡", "知识卡", "卡片"}:
|
||||
continue
|
||||
result[key_text] = _without_card_fields(item)
|
||||
return result
|
||||
|
||||
|
||||
def check_arm_manifests(
|
||||
manifests: Mapping[str, Mapping[str, Any]],
|
||||
required_arms: Sequence[str] | None = None,
|
||||
) -> dict[str, Any]:
|
||||
"""确保三臂除卡注入区外完全一致。"""
|
||||
|
||||
required = set(required_arms or ("outline_only", "outline_plus_cards", "outline_plus_placebo_cards"))
|
||||
actual = set(manifests)
|
||||
missing = sorted(required - actual)
|
||||
if missing:
|
||||
return _result(STATUS_INVALID_ARM_DIFF, [f"缺少评测臂: {','.join(missing)}"])
|
||||
unexpected = sorted(actual - required)
|
||||
if unexpected:
|
||||
return _result(STATUS_INVALID_ARM_DIFF, [f"存在未登记评测臂: {','.join(unexpected)}"])
|
||||
if any(not isinstance(manifest, Mapping) for manifest in manifests.values()):
|
||||
return _result(STATUS_INVALID_ARM_DIFF, ["评测臂 manifest 必须是对象"])
|
||||
|
||||
names = sorted(required)
|
||||
baseline = _without_card_fields(manifests[names[0]])
|
||||
differences = [name for name in names[1:] if _without_card_fields(manifests[name]) != baseline]
|
||||
if differences:
|
||||
return _result(STATUS_INVALID_ARM_DIFF, [f"公共输入区不一致: {','.join(differences)}"])
|
||||
return _result(STATUS_READY)
|
||||
|
||||
|
||||
def check_candidate_output(
|
||||
candidate: Mapping[str, Any],
|
||||
target_chapter: int,
|
||||
source_catalog: Sequence[Mapping[str, Any]] = (),
|
||||
) -> dict[str, Any]:
|
||||
"""校验细纲候选的结构边界,不判断内容是否命中标准答案。"""
|
||||
|
||||
if not isinstance(candidate, Mapping):
|
||||
return _result(STATUS_SCHEMA_INVALID, ["候选不是对象"])
|
||||
expected_target = normalize_chapter(target_chapter)
|
||||
if expected_target is None:
|
||||
return _result(STATUS_INVALID_SNAPSHOT, ["目标章号无效"])
|
||||
errors = [f"缺少字段: {field}" for field in REQUIRED_CANDIDATE_FIELDS if field not in candidate]
|
||||
forbidden = sorted(set(candidate) & FORBIDDEN_CANDIDATE_FIELDS)
|
||||
if forbidden:
|
||||
errors.append(f"候选包含正文/原文字段: {','.join(forbidden)}")
|
||||
for field in ("keyEvents", "entities", "foreshadowing", "stateChanges", "unknowns", "assumptions"):
|
||||
if field in candidate and not isinstance(candidate[field], list):
|
||||
errors.append(f"字段必须是数组: {field}")
|
||||
for field in ("chapterGoal", "hook"):
|
||||
if field in candidate and not isinstance(candidate[field], str):
|
||||
errors.append(f"字段必须是字符串: {field}")
|
||||
events = candidate.get("keyEvents")
|
||||
if isinstance(events, list):
|
||||
event_ids = [item.get("id") for item in events if isinstance(item, Mapping) and item.get("id")]
|
||||
if len(event_ids) != len(set(event_ids)):
|
||||
errors.append("keyEvents 包含重复事件 id")
|
||||
if any(not isinstance(item, Mapping) for item in events):
|
||||
errors.append("keyEvents 每项必须是对象")
|
||||
if any(isinstance(item, Mapping) and not item.get("id") for item in events):
|
||||
errors.append("keyEvents 每项必须有 id")
|
||||
source_errors: list[str] = []
|
||||
source_refs = candidate.get("sourceRefs")
|
||||
if source_refs is not None:
|
||||
if not isinstance(source_refs, list):
|
||||
errors.append("sourceRefs 必须是数组")
|
||||
else:
|
||||
catalog = {
|
||||
str(source.get("sourceId")): source
|
||||
for source in source_catalog
|
||||
if isinstance(source, Mapping) and source.get("sourceId")
|
||||
}
|
||||
for index, reference in enumerate(source_refs):
|
||||
source = reference if isinstance(reference, Mapping) else catalog.get(str(reference))
|
||||
if source is None:
|
||||
if source_catalog:
|
||||
errors.append(f"sourceRefs[{index}] 未登记来源")
|
||||
continue
|
||||
source_value = _field(
|
||||
source,
|
||||
"chapter",
|
||||
"chapterNo",
|
||||
"chapter_no",
|
||||
"章",
|
||||
"章号",
|
||||
"chapterRange",
|
||||
"chapter_range",
|
||||
"range",
|
||||
)
|
||||
bounds = normalize_chapter_range(source_value)
|
||||
if bounds is None and "from_order" in source:
|
||||
bounds = normalize_chapter_range(f"{source.get('from_order')}-{source.get('to_order')}")
|
||||
if bounds is not None and bounds[1] >= expected_target:
|
||||
source_errors.append(
|
||||
f"sourceRefs[{index}] 包含目标章或未来章: {bounds[0]}-{bounds[1]}"
|
||||
)
|
||||
actual_target = normalize_chapter(candidate.get("targetChapter"))
|
||||
if actual_target != expected_target:
|
||||
errors.append(f"候选目标章错误: expected={expected_target}, actual={actual_target}")
|
||||
if source_errors:
|
||||
return _result(STATUS_TARGET_SOURCE_FORBIDDEN, source_errors)
|
||||
if errors:
|
||||
return _result(STATUS_SCHEMA_INVALID, errors)
|
||||
return _result(STATUS_READY)
|
||||
|
||||
|
||||
def check_replay(
|
||||
*,
|
||||
authorization: Mapping[str, Any] | None,
|
||||
target_chapter: int,
|
||||
planner_sources: Sequence[Mapping[str, Any]],
|
||||
arm_manifests: Mapping[str, Mapping[str, Any]],
|
||||
required_arms: Sequence[str] | None = None,
|
||||
) -> dict[str, Any]:
|
||||
"""执行回放前置门,任何一项失败都不允许进入模型调用。"""
|
||||
|
||||
checks = [
|
||||
check_authorization(authorization),
|
||||
check_target_sources(target_chapter, planner_sources),
|
||||
check_arm_manifests(arm_manifests, required_arms),
|
||||
]
|
||||
failures = [check for check in checks if not check["ok"]]
|
||||
if failures:
|
||||
return _result(
|
||||
failures[0]["status"],
|
||||
[error for check in failures for error in check["errors"]],
|
||||
[warning for check in failures for warning in check["warnings"]],
|
||||
)
|
||||
return _result(STATUS_READY)
|
||||
|
||||
|
||||
def _parse_args() -> argparse.Namespace:
|
||||
parser = argparse.ArgumentParser(description="校验回放快照和三臂输入")
|
||||
parser.add_argument("--authorization", type=Path, required=True)
|
||||
parser.add_argument("--sources", type=Path, required=True)
|
||||
parser.add_argument("--arms", type=Path, required=True)
|
||||
parser.add_argument("--target-chapter", type=int, required=True)
|
||||
parser.add_argument("--output", type=Path, required=True)
|
||||
return parser.parse_args()
|
||||
|
||||
|
||||
def main() -> int:
|
||||
args = _parse_args()
|
||||
authorization = json.loads(args.authorization.read_text(encoding="utf-8"))
|
||||
sources = json.loads(args.sources.read_text(encoding="utf-8"))
|
||||
arms = json.loads(args.arms.read_text(encoding="utf-8"))
|
||||
result = check_replay(
|
||||
authorization=authorization,
|
||||
target_chapter=args.target_chapter,
|
||||
planner_sources=sources,
|
||||
arm_manifests=arms,
|
||||
)
|
||||
args.output.write_text(json.dumps(result, ensure_ascii=False, sort_keys=True) + "\n", encoding="utf-8")
|
||||
return 0 if result["ok"] else 2
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
raise SystemExit(main())
|
||||
77
.claude/skills/replay-eval/scripts/fine_outline_rubric.py
Normal file
77
.claude/skills/replay-eval/scripts/fine_outline_rubric.py
Normal file
@ -0,0 +1,77 @@
|
||||
#!/usr/bin/env python3
|
||||
"""细纲回放 rubric 的确定性校验器。
|
||||
|
||||
它不替代 judge 打分,只检查评分报告是否使用正确维度、分数范围和证据字段。
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
from typing import Any, Mapping
|
||||
|
||||
|
||||
RUBRIC_PROFILE = "fine_outline_replay"
|
||||
DIMENSIONS = (
|
||||
"structure_completeness",
|
||||
"direction_causality",
|
||||
"order_pacing",
|
||||
"entity_state",
|
||||
"foreshadowing_action",
|
||||
"handoff_hook",
|
||||
)
|
||||
PROSE_DIMENSIONS = frozenset(
|
||||
{"style_fit", "readability", "文风一致性", "文笔", "pacing_tension", "information_density"}
|
||||
)
|
||||
|
||||
|
||||
def validate_scores(scores: Mapping[str, Any]) -> list[str]:
|
||||
"""返回报告问题;每个维度必须有 1-5 分和非空证据。"""
|
||||
|
||||
errors: list[str] = []
|
||||
if not isinstance(scores, Mapping):
|
||||
return ["scores 必须是对象"]
|
||||
missing = [dimension for dimension in DIMENSIONS if dimension not in scores]
|
||||
if missing:
|
||||
errors.append(f"缺少 rubric 维度: {','.join(missing)}")
|
||||
unexpected = sorted(set(scores) - set(DIMENSIONS))
|
||||
if unexpected:
|
||||
errors.append(f"存在未登记 rubric 维度: {','.join(unexpected)}")
|
||||
prose = sorted(set(unexpected) & PROSE_DIMENSIONS)
|
||||
if prose:
|
||||
errors.append(f"细纲 rubric 禁止正文质量维度: {','.join(prose)}")
|
||||
for dimension in DIMENSIONS:
|
||||
value = scores.get(dimension)
|
||||
if not isinstance(value, Mapping):
|
||||
errors.append(f"维度必须包含 score/evidence 对象: {dimension}")
|
||||
continue
|
||||
score = value.get("score")
|
||||
if isinstance(score, bool) or not isinstance(score, (int, float)) or not 1 <= score <= 5:
|
||||
errors.append(f"分数必须在 1-5: {dimension}")
|
||||
evidence = value.get("evidence")
|
||||
if not isinstance(evidence, str) or not evidence.strip():
|
||||
errors.append(f"分数缺少证据: {dimension}")
|
||||
return errors
|
||||
|
||||
|
||||
def stability_warning(
|
||||
first: Mapping[str, float],
|
||||
second: Mapping[str, float],
|
||||
threshold: float = 0.5,
|
||||
) -> dict[str, Any]:
|
||||
"""比较两次评审,返回差异和是否需要人工复核。"""
|
||||
|
||||
gaps = {
|
||||
dimension: abs(float(first[dimension]) - float(second[dimension]))
|
||||
for dimension in DIMENSIONS
|
||||
if dimension in first and dimension in second
|
||||
}
|
||||
return {"stable": all(gap <= threshold for gap in gaps.values()), "gaps": gaps}
|
||||
|
||||
|
||||
def validate_report(report: Mapping[str, Any]) -> list[str]:
|
||||
"""校验一个评委报告的 profile 和评分结构。"""
|
||||
|
||||
errors: list[str] = []
|
||||
if report.get("profile") != RUBRIC_PROFILE:
|
||||
errors.append(f"profile 必须是 {RUBRIC_PROFILE}")
|
||||
errors.extend(validate_scores(report.get("scores", {})))
|
||||
return errors
|
||||
114
.claude/skills/replay-eval/scripts/test_build_snapshot.py
Normal file
114
.claude/skills/replay-eval/scripts/test_build_snapshot.py
Normal file
@ -0,0 +1,114 @@
|
||||
#!/usr/bin/env python3
|
||||
"""build_snapshot.py 的无网络离线测试。"""
|
||||
|
||||
import pathlib
|
||||
import sys
|
||||
import unittest
|
||||
|
||||
sys.path.insert(0, str(pathlib.Path(__file__).resolve().parent))
|
||||
from build_snapshot import ( # noqa: E402
|
||||
SnapshotError,
|
||||
build_snapshot,
|
||||
build_snapshot_manifest,
|
||||
filter_milestones,
|
||||
filter_outline_windows,
|
||||
normalize_chapter,
|
||||
remove_terminal_fields,
|
||||
)
|
||||
|
||||
|
||||
class BuildSnapshotTest(unittest.TestCase):
|
||||
def test_normalize_chapter_rejects_guessing(self):
|
||||
self.assertEqual(normalize_chapter(488), 488)
|
||||
self.assertEqual(normalize_chapter(" 488 "), 488)
|
||||
self.assertIsNone(normalize_chapter(True))
|
||||
self.assertIsNone(normalize_chapter(488.0))
|
||||
self.assertIsNone(normalize_chapter("第488章"))
|
||||
|
||||
def test_milestones_are_frozen_by_absolute_chapter(self):
|
||||
kept, omitted = filter_milestones(
|
||||
[
|
||||
{"章": 487, "台阶": "已知"},
|
||||
{"章": "488", "台阶": "已知字符串"},
|
||||
{"章": "489", "台阶": "未来"},
|
||||
{"章": "487-490", "台阶": "跨冻结点"},
|
||||
{"台阶": "没有章号"},
|
||||
],
|
||||
488,
|
||||
)
|
||||
self.assertEqual([item["台阶"] for item in kept], ["已知", "已知字符串"])
|
||||
self.assertEqual(
|
||||
[item["reason"] for item in omitted],
|
||||
["future_or_crosses_as_of", "future_or_crosses_as_of", "missing_or_unbounded_chapter"],
|
||||
)
|
||||
self.assertTrue(all("台阶" not in item for item in omitted))
|
||||
|
||||
def test_outline_windows_require_complete_to_order(self):
|
||||
kept, omitted = filter_outline_windows(
|
||||
[
|
||||
{"window_no": 2, "from_order": 401, "to_order": 430},
|
||||
{"window_no": 1, "from_order": 431, "to_order": 490},
|
||||
{"window_no": 3, "from_order": "bad", "to_order": 500},
|
||||
],
|
||||
488,
|
||||
)
|
||||
self.assertEqual([item["from_order"] for item in kept], [401])
|
||||
self.assertEqual([item["reason"] for item in omitted], ["future_or_crosses_as_of", "missing_or_invalid_window_bounds"])
|
||||
|
||||
def test_terminal_fields_are_removed_recursively(self):
|
||||
value = {
|
||||
"name": "苏铭",
|
||||
"current_state": "终局状态",
|
||||
"nested": {"future_arc": "未来计划", "safe": "保留"},
|
||||
"items": [{"终态摘要": "不应出现"}, {"safe": "保留"}],
|
||||
}
|
||||
self.assertEqual(remove_terminal_fields(value), {"name": "苏铭", "nested": {"safe": "保留"}, "items": [{}, {"safe": "保留"}]})
|
||||
|
||||
def test_nested_card_history_is_frozen(self):
|
||||
result = build_snapshot(
|
||||
{
|
||||
"cards": [
|
||||
{
|
||||
"name": "生物机甲",
|
||||
"milestones": [
|
||||
{"chapter": 488, "step": "可见"},
|
||||
{"chapter": 489, "step": "未来"},
|
||||
],
|
||||
"current_state": "终态字段不得透传",
|
||||
}
|
||||
]
|
||||
},
|
||||
488,
|
||||
"v0",
|
||||
)
|
||||
card = result["snapshot"]["cards"][0]
|
||||
self.assertEqual(card["milestones"], [{"chapter": 488, "step": "可见"}])
|
||||
self.assertNotIn("current_state", card)
|
||||
self.assertEqual(result["manifest"]["omittedSources"][0]["reason"], "future_or_crosses_as_of")
|
||||
|
||||
def test_manifest_is_stable_and_does_not_store_payload(self):
|
||||
kwargs = {
|
||||
"as_of": 488,
|
||||
"snapshot_version": "v0",
|
||||
"sections": {"l0": {"sourceId": "task", "sourceVersion": "1", "payload": {"target": 489}}},
|
||||
"final_report": {"status": "ready", "sourceHash": "abc"},
|
||||
}
|
||||
first = build_snapshot_manifest(**kwargs)
|
||||
second = build_snapshot_manifest(**kwargs)
|
||||
self.assertEqual(first, second)
|
||||
self.assertNotIn("payload", first["sections"]["l0"])
|
||||
self.assertEqual(len(first["manifestSha256"]), 64)
|
||||
self.assertNotIn("原书", str(first["finalReport"]))
|
||||
|
||||
def test_final_report_rejects_original_text_fields(self):
|
||||
with self.assertRaises(SnapshotError):
|
||||
build_snapshot_manifest(
|
||||
as_of=488,
|
||||
snapshot_version="v0",
|
||||
sections={},
|
||||
final_report={"raw_text": "原书全文"},
|
||||
)
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
unittest.main()
|
||||
119
.claude/skills/replay-eval/scripts/test_check_snapshot.py
Normal file
119
.claude/skills/replay-eval/scripts/test_check_snapshot.py
Normal file
@ -0,0 +1,119 @@
|
||||
#!/usr/bin/env python3
|
||||
"""check_snapshot.py 的无网络离线测试。"""
|
||||
|
||||
import pathlib
|
||||
import sys
|
||||
import unittest
|
||||
|
||||
sys.path.insert(0, str(pathlib.Path(__file__).resolve().parent))
|
||||
from check_snapshot import ( # noqa: E402
|
||||
STATUS_BLOCKED_AUTHORIZATION,
|
||||
STATUS_INVALID_ARM_DIFF,
|
||||
STATUS_READY,
|
||||
STATUS_SCHEMA_INVALID,
|
||||
STATUS_TARGET_SOURCE_FORBIDDEN,
|
||||
check_arm_manifests,
|
||||
check_authorization,
|
||||
check_candidate_output,
|
||||
check_replay,
|
||||
check_target_sources,
|
||||
)
|
||||
|
||||
|
||||
AUTH = {
|
||||
"sourceStatus": "active",
|
||||
"allowedPurpose": ["offline_evaluation"],
|
||||
"authorizationSnapshot": {"id": "auth-1", "version": "v1"},
|
||||
}
|
||||
|
||||
|
||||
def arm(common, cards):
|
||||
return {**common, "cardInjection": cards}
|
||||
|
||||
|
||||
class CheckSnapshotTest(unittest.TestCase):
|
||||
def test_authorization_is_fail_closed(self):
|
||||
self.assertEqual(check_authorization(None)["status"], STATUS_BLOCKED_AUTHORIZATION)
|
||||
self.assertTrue(check_authorization(AUTH)["ok"])
|
||||
denied = {**AUTH, "allowedPurpose": ["read"]}
|
||||
self.assertEqual(check_authorization(denied)["status"], STATUS_BLOCKED_AUTHORIZATION)
|
||||
unknown_status = {**AUTH, "sourceStatus": "temporary"}
|
||||
self.assertEqual(check_authorization(unknown_status)["status"], STATUS_BLOCKED_AUTHORIZATION)
|
||||
unlicensed = {**AUTH, "copyrightStatus": "unlicensed"}
|
||||
self.assertEqual(check_authorization(unlicensed)["status"], STATUS_BLOCKED_AUTHORIZATION)
|
||||
|
||||
def test_target_source_is_blocked(self):
|
||||
allowed = check_target_sources(489, [{"chapter": 488}, {"from_order": 450, "to_order": 482}])
|
||||
self.assertEqual(allowed["status"], STATUS_READY)
|
||||
blocked = check_target_sources(489, [{"chapter": 489}])
|
||||
self.assertEqual(blocked["status"], STATUS_TARGET_SOURCE_FORBIDDEN)
|
||||
|
||||
def test_arm_common_input_must_match(self):
|
||||
common = {"snapshotVersion": "v0", "asOfChapter": 488, "l0": {"target": 489}}
|
||||
manifests = {
|
||||
"outline_only": arm(common, []),
|
||||
"outline_plus_cards": arm(common, [{"id": 1}]),
|
||||
"outline_plus_placebo_cards": arm(common, [{"id": 2}]),
|
||||
}
|
||||
self.assertTrue(check_arm_manifests(manifests)["ok"])
|
||||
extra = {**manifests, "unregistered_arm": arm(common, [])}
|
||||
self.assertEqual(check_arm_manifests(extra)["status"], STATUS_INVALID_ARM_DIFF)
|
||||
changed = dict(manifests)
|
||||
changed["outline_plus_cards"] = arm({**common, "l0": {"target": 490}}, [{"id": 1}])
|
||||
self.assertEqual(check_arm_manifests(changed)["status"], STATUS_INVALID_ARM_DIFF)
|
||||
|
||||
def test_two_arm_smoke_can_be_explicit(self):
|
||||
common = {"snapshotVersion": "v0", "asOfChapter": 488}
|
||||
manifests = {"outline_only": arm(common, []), "outline_plus_cards": arm(common, [{"id": 1}])}
|
||||
self.assertTrue(check_arm_manifests(manifests, ["outline_only", "outline_plus_cards"])["ok"])
|
||||
|
||||
def test_candidate_contract_is_structural(self):
|
||||
candidate = {
|
||||
"targetChapter": 489,
|
||||
"chapterGoal": "突破",
|
||||
"keyEvents": [],
|
||||
"entities": [],
|
||||
"foreshadowing": [],
|
||||
"stateChanges": [],
|
||||
"hook": "悬念",
|
||||
"unknowns": [],
|
||||
"assumptions": [],
|
||||
}
|
||||
self.assertTrue(check_candidate_output(candidate, 489)["ok"])
|
||||
bad = {**candidate, "正文全文": "原文"}
|
||||
self.assertEqual(check_candidate_output(bad, 489)["status"], STATUS_SCHEMA_INVALID)
|
||||
bad_type = {**candidate, "keyEvents": "事件"}
|
||||
self.assertEqual(check_candidate_output(bad_type, 489)["status"], STATUS_SCHEMA_INVALID)
|
||||
duplicate = {**candidate, "keyEvents": [{"id": "event-1"}, {"id": "event-1"}]}
|
||||
self.assertEqual(check_candidate_output(duplicate, 489)["status"], STATUS_SCHEMA_INVALID)
|
||||
missing_id = {**candidate, "keyEvents": [{"event": "没有 id"}]}
|
||||
self.assertEqual(check_candidate_output(missing_id, 489)["status"], STATUS_SCHEMA_INVALID)
|
||||
future_ref = {**candidate, "sourceRefs": ["target-scaffold"]}
|
||||
self.assertEqual(
|
||||
check_candidate_output(
|
||||
future_ref,
|
||||
489,
|
||||
[{"sourceId": "target-scaffold", "chapter": 489}],
|
||||
)["status"],
|
||||
STATUS_TARGET_SOURCE_FORBIDDEN,
|
||||
)
|
||||
|
||||
def test_replay_fails_closed_before_model(self):
|
||||
common = {"snapshotVersion": "v0", "asOfChapter": 488}
|
||||
manifests = {
|
||||
"outline_only": arm(common, []),
|
||||
"outline_plus_cards": arm(common, [{"id": 1}]),
|
||||
"outline_plus_placebo_cards": arm(common, [{"id": 2}]),
|
||||
}
|
||||
result = check_replay(
|
||||
authorization=AUTH,
|
||||
target_chapter=489,
|
||||
planner_sources=[{"chapter": 489}],
|
||||
arm_manifests=manifests,
|
||||
)
|
||||
self.assertFalse(result["ok"])
|
||||
self.assertEqual(result["status"], STATUS_TARGET_SOURCE_FORBIDDEN)
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
unittest.main()
|
||||
54
.claude/skills/replay-eval/scripts/test_rubric.py
Normal file
54
.claude/skills/replay-eval/scripts/test_rubric.py
Normal file
@ -0,0 +1,54 @@
|
||||
#!/usr/bin/env python3
|
||||
"""细纲 rubric 的离线回归测试。"""
|
||||
|
||||
import pathlib
|
||||
import sys
|
||||
import unittest
|
||||
|
||||
sys.path.insert(0, str(pathlib.Path(__file__).resolve().parent))
|
||||
from fine_outline_rubric import ( # noqa: E402
|
||||
DIMENSIONS,
|
||||
RUBRIC_PROFILE,
|
||||
stability_warning,
|
||||
validate_report,
|
||||
validate_scores,
|
||||
)
|
||||
|
||||
|
||||
def valid_scores():
|
||||
return {
|
||||
dimension: {"score": 4, "evidence": f"证据-{dimension}"}
|
||||
for dimension in DIMENSIONS
|
||||
}
|
||||
|
||||
|
||||
class FineOutlineRubricTest(unittest.TestCase):
|
||||
def test_every_dimension_requires_evidence(self):
|
||||
self.assertEqual(validate_scores(valid_scores()), [])
|
||||
missing_evidence = valid_scores()
|
||||
missing_evidence[DIMENSIONS[0]] = {"score": 4}
|
||||
self.assertTrue(any("缺少证据" in error for error in validate_scores(missing_evidence)))
|
||||
|
||||
def test_prose_dimensions_are_rejected(self):
|
||||
scores = valid_scores()
|
||||
scores["style_fit"] = {"score": 5, "evidence": "不应出现"}
|
||||
self.assertTrue(any("禁止正文质量维度" in error for error in validate_scores(scores)))
|
||||
|
||||
def test_profile_and_score_range_are_checked(self):
|
||||
report = {"profile": RUBRIC_PROFILE, "scores": valid_scores()}
|
||||
self.assertEqual(validate_report(report), [])
|
||||
bad = {"profile": "quality_gate", "scores": valid_scores()}
|
||||
bad["scores"][DIMENSIONS[1]] = {"score": 6, "evidence": "超范围"}
|
||||
self.assertEqual(len(validate_report(bad)), 2)
|
||||
|
||||
def test_large_reviewer_gap_warns(self):
|
||||
first = {dimension: 4 for dimension in DIMENSIONS}
|
||||
second = {dimension: 4 for dimension in DIMENSIONS}
|
||||
second[DIMENSIONS[2]] = 5
|
||||
result = stability_warning(first, second)
|
||||
self.assertFalse(result["stable"])
|
||||
self.assertEqual(result["gaps"][DIMENSIONS[2]], 1.0)
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
unittest.main()
|
||||
@ -21,6 +21,7 @@ agent 的提示词按**变化轴**拆三段,不做"一个 agent 一个大 prom
|
||||
| expansion 扩写 | `expansion` | 写作→writer | generation | 同 continuation | 已建 |
|
||||
| polish 润色 | `polish` | 写作→writer | generation | 同 continuation(只动表达层) | 已建 |
|
||||
| planning 规划 | `planning` | 规划→planner | planning | read-context(planning 视图)→槽位→用户确认(confirm) | 已建 |
|
||||
| fine_outline 细纲规划 | `fine-outline` | 规划→planner | planning | read-context(fine_outline 冻结视图)→槽位→detect→quality-gate | 评测版首建(2026-07-19) |
|
||||
| full_parse 拆书 | `parse-book` | 分析→extractor | extraction | import 分章→逐章槽位→db 落 draft→管理员确认(G3 门) | 已建;B2 PG 版首验 |
|
||||
| extraction 章后抽取 | `extract-knowledge` | 分析→extractor | extraction | 采纳后触发→槽位→草稿/冲突队列→confirm | 已建;C5 首验 |
|
||||
| validation / consistency_check 检测 | `detect` | 检测→detector | detection | read-context(detection 视图)→槽位→报告落评审/ | 已建 |
|
||||
|
||||
Loading…
x
Reference in New Issue
Block a user