diff --git a/.claude/agents/judge.md b/.claude/agents/judge.md index ab4e43b..19eebe2 100644 --- a/.claude/agents/judge.md +++ b/.claude/agents/judge.md @@ -1,22 +1,25 @@ --- name: judge -description: 质量评委——二层保护节点(质量门控与 LLM-Judge)的打分器,不可被装配替换;维度体系与 rubric 以 quality-gate skill 为合同。 +description: 质量评委——二层保护节点(质量门控与 LLM-Judge)的打分器,不可被装配替换;正文与细纲回放的维度体系、rubric 以 quality-gate skill 为合同。 tools: Read, Grep, Glob, Write model: opus --- -你是质量评委,保护节点角色(不可被装配替换)。**功能合同=`quality-gate` skill**——维度体系(3 叙事关键+5 非关键)、达标口径、rubric 全在那里,你只执行不自造维度。报告落 `works/<书>/评审/第NNN章-评分.md`,除此不写任何文件。 +你是质量评委,保护节点角色(不可被装配替换)。**功能合同=`quality-gate` skill**——正文任务使用 `quality_gate` profile;细纲回放任务使用 `fine_outline_replay` profile。维度、证据要求和稳定性口径全在那里,你只执行不自造维度。报告按当前运行合同落临时评测目录或 `works/<书>/评审/第NNN章-评分.md`,除此不写正式创作文件。 ## 元数据纪律(怎么用元数据) - 维度体系是元数据(专题-04 质量策略族):策略加维度,quality-gate skill 更新表格,你一字不改; - 上下文=writer 基线包(**同证独立**):评的是"在写手所知条件下写得好不好",不索取额外资料,不拿包外信息扣分。 +- 细纲回放只使用冻结快照和结构化 `reference scaffold proxy`;不得读取目标章全文或完整目标章细纲。 +- 细纲回放每个分数必须有 `evidence` 字段;不得使用正文文风、文笔或可读性作为评分维度。 ## 打分纪律 - 每维给分必附一句引文证据(好在哪/差在哪,引原句);没有证据的分数无效。 - 同一维度复评同一章分差应 ≤0.5;严格度不因收敛压力改变,不放水不加戏。 - 末尾「最值得改的三点」按提升空间排序:问题→根因层猜测(prompt/上下文/设定卡)→具体改法。 +- 细纲回放时,把末尾建议替换为“最值得补齐的三项结构缺口”,并标注它属于候选结构、公共大纲、卡注入、原文检索还是标准事实不确定;若两次同维分差大于 0.5,只写稳定性警告,不强行裁决。 ## 禁区 diff --git a/.claude/agents/planner.md b/.claude/agents/planner.md index 6eb4882..e6f7af5 100644 --- a/.claude/agents/planner.md +++ b/.claude/agents/planner.md @@ -1,11 +1,16 @@ --- name: planner -description: 规划师——规划槽位默认绑定件,承接 planning 功能;产出结构=schema 字段清单本身,功能细节以 planning skill 为合同;产出全为草稿。 +description: 规划师——规划槽位默认绑定件,承接 planning 与 fine_outline 功能;产出结构=schema 字段清单本身,功能细节以对应 skill 为合同;产出全为草稿。 tools: Read, Write, Grep, Glob model: opus --- -你是这部书的总规划,规划槽位的默认绑定件。**功能合同=`planning` skill**(立项与修订同用)。产出全部不提交;未确认的规划不进生成上下文。 +你是这部书的总规划,规划槽位的默认绑定件。功能合同按本次任务二选一: + +- `planning`:遵守 `planning` skill,负责立项与规划修订; +- `fine_outline`:遵守 `fine-outline` skill,只产结构细纲,不写正文。 + +产出全部不提交;未确认的规划不进生成上下文。回放任务中,`fine-outline` skill 的冻结边界优先于本身份段里面向正式创作的全局规划能力。 ## 元数据纪律(怎么用元数据) diff --git a/.claude/skills/detect/SKILL.md b/.claude/skills/detect/SKILL.md index 637b18f..8225d0b 100644 --- a/.claude/skills/detect/SKILL.md +++ b/.claude/skills/detect/SKILL.md @@ -1,6 +1,6 @@ --- name: detect -description: 检测的功能合同(scenario: validation/consistency_check,检测槽位)。检查清单由 schema 字段自动生成——凡 aiContext 含 detection 的字段即检查项。 +description: 检测的功能合同(scenario: validation/consistency_check/fine_outline_replay,检测槽位)。检查清单由 schema 字段自动生成——凡 aiContext 含 detection 的字段即检查项。 disable-model-invocation: true --- @@ -35,6 +35,23 @@ schema 给字段加上 detection 用途,检查项自动+1,本 skill 与 detector 末尾:阻塞性 N 条(不修不建议采纳)/建议性 N 条 + 「本次检查用不上但设定卡缺失的字段」(设计发现)。 +## 细纲回放分支(`fine_outline_replay`) + +何时用:`fine-outline` 候选进入独立评分前。检测器只看冻结到 `as_of` 的规划上下文和候选,不看目标章标准事实,不把评测答案倒灌回规划侧。 + +检查对象和证据格式: + +| 检查对象 | 检查项 | 高严重度条件 | 证据必须包含 | +|---|---|---|---| +| 候选结构 | 必填字段、目标章号、事件 ID、事件顺序 | 缺字段、重复事件或目标章错误导致无法评估 | 字段路径 + 候选值摘要 | +| 因果链 | 事件触发、行动、结果方向是否自洽 | 后一事件无前置、结果与前置状态矛盾 | 事件 ID 对 + 冲突原因 | +| 实体状态 | 角色/势力/地点/能力是否违反冻结事实 | 引用未来事实或越过 N 时点能力边界 | `sourceId` + 冻结字段 | +| 伏笔动作 | 埋、推进、回收和不确定标记 | 将未知或未来回收写成确定事实 | 伏笔标识 + 当前台账定位 | +| 来源引用 | `sourceRefs` 是否来自快照、是否包含目标章及以后 | 目标章/未来来源出现在候选引用中 | 来源 ID + 章号范围 | +| 未知项纪律 | `unknowns` / `assumptions` 是否显式承载缺口 | 用无来源断言替代未知项 | 候选字段路径 | + +报告仍然只产审查结果,不修改候选。回放中任一高严重度问题阻断该臂进入 judge;“卡里缺少目标新角色”要单列为资料覆盖发现,不冒充规划器错误。 + ## 红线 只产报告,不改任何创作文件;证据先行——无依据的观感问题归「建议」并标明主观;底牌信息仅用于检测判断。 diff --git a/.claude/skills/fine-outline/SKILL.md b/.claude/skills/fine-outline/SKILL.md new file mode 100644 index 0000000..071bc72 --- /dev/null +++ b/.claude/skills/fine-outline/SKILL.md @@ -0,0 +1,72 @@ +--- +name: fine-outline +description: 细纲规划合同。根据规划上下文和已授权的冻结事实,产出下一章的结构细纲草稿;不写正文、不读取目标章答案。 +disable-model-invocation: true +--- + +# 细纲规划(`scenario: fine_outline`) + +何时用:在正文生成前,把作品大纲、当前状态和已授权设定组织成单章结构细纲。评测回放使用同一合同,但把上下文固定为 `as_of` 冻结快照;正式创作可以按产品授权扩展到连续章节或全书规划。 + +## 输入合同 + +上下文必须由 `read-context` 组装,至少包含四层中的以下部分: + +- **L0 当前任务**:目标章号、章节输出合同、用户意图和未知项纪律; +- **L1 叙事现在时**:截至当前章的状态、活动线程和必要的近章结构摘要; +- **L2 作品事实**:已确认的大纲、设定、知识卡和伏笔台账; +- **L3 授权资料**:仅使用装配层已绑定、且来源授权允许本次用途的资料。 + +回放模式额外要求:所有来源的绝对章号或完整窗口上界必须 `<= as_of`。目标章正文、目标章细纲、目标章出场清单、未来里程碑和终态摘要不得进入规划上下文。无法证明时间边界的来源按未知处理,不凭名称或窗口号猜测。 + +## 规划步骤 + +1. 先确认章节在当前大纲弧线中的位置,写出本章戏剧目标和承接关系。 +2. 再按因果顺序拆关键事件:触发、参与者、行动、结果方向和不可逆变化必须能互相解释。 +3. 将实体只列为本章确实需要的角色、势力、地点、物品或规则,并标明本章作用;不能把目标章答案反推成实体清单。 +4. 对伏笔明确写 `埋`、`推进`、`回收` 或 `不确定`,不得把冻结资料没有证明的结果静默写成确定事实。 +5. 写出章末状态变化和下一步钩子;无法由来源支持的细节放入 `unknowns` 或 `assumptions`。 +6. 输出前逐字段自查,确认没有正文段落、对白、原文复述、目标章引用或未来来源引用。 + +## 输出合同 + +候选必须是一个结构化对象,字段完整且顺序稳定: + +```yaml +targetChapter: 489 +chapterGoal: "本章要完成的戏剧任务" +keyEvents: + - id: event-1 + order: 1 + event: "事件" + participants: ["实体"] + trigger: "触发因果" + resultDirection: "结果方向" +entities: + - name: "实体" + type: character + role: "本章作用" +foreshadowing: + - action: 推进 + subject: "伏笔" + evidence: "来源 ID 或 unknown" +stateChanges: ["可观察的状态变化"] +hook: "章末钩子" +unknowns: ["无法由当前资料证明的内容"] +assumptions: ["为组织结构暂时采用的假设"] +sourceRefs: ["冻结快照中的 sourceId"] +``` + +`sourceRefs` 是可选的回溯字段,但存在时只能引用快照登记的 `sourceId`,不能粘贴原文。`keyEvents`、`entities`、`foreshadowing`、`stateChanges`、`unknowns`、`assumptions` 必须是数组;`chapterGoal` 和 `hook` 必须是字符串。 + +## 回放专用红线 + +- 不得使用参考作品的目标章细纲作为生成输入;它只能在独立评审侧作为 `reference scaffold proxy`; +- 不得输出正文、场景对白、完整原文摘要或“我猜原书下一章是……”之类的答案复述; +- 不得因为知识卡存在而减少对公共大纲和叙事现在时的使用;卡是检索索引和事实补充,不替代公共上下文; +- 正确卡、无卡、错配卡三臂只能改变卡注入分区,其他任务、模型、预算和输出合同保持一致; +- 不确定事实显式留在 `unknowns`,审查智能体据此区分资料缺失和规划错误。 + +## 产物边界 + +候选细纲属于 Shadow 草稿,只进入临时评测目录或 `works/<书>/评审/` 的运行噪音,不直接写入正式大纲、知识库或正文。评测最终只保留结构评分、摘要、章节定位、阻断类别和哈希。 diff --git a/.claude/skills/fine-outline/scripts/test_contract.py b/.claude/skills/fine-outline/scripts/test_contract.py new file mode 100644 index 0000000..26d1224 --- /dev/null +++ b/.claude/skills/fine-outline/scripts/test_contract.py @@ -0,0 +1,46 @@ +#!/usr/bin/env python3 +"""细纲功能合同的离线回归测试,不启动模型、不访问作品数据。""" + +from pathlib import Path +import unittest + + +ROOT = Path(__file__).resolve().parents[4] +SKILL = (ROOT / ".claude/skills/fine-outline/SKILL.md").read_text(encoding="utf-8") +PLANNER = (ROOT / ".claude/agents/planner.md").read_text(encoding="utf-8") +CHAINS = (ROOT / "meta/chains/README.md").read_text(encoding="utf-8") + + +class FineOutlineContractTest(unittest.TestCase): + def test_contract_is_bound_to_existing_planner_slot(self): + self.assertIn("scenario: fine_outline", SKILL) + self.assertIn("fine_outline", PLANNER) + self.assertIn("fine_outline", CHAINS) + self.assertIn("规划→planner", CHAINS) + + def test_candidate_fields_and_shadow_boundary_are_explicit(self): + for field in ( + "targetChapter", + "chapterGoal", + "keyEvents", + "entities", + "foreshadowing", + "stateChanges", + "hook", + "unknowns", + "assumptions", + ): + self.assertIn(field, SKILL) + self.assertIn("不得输出正文", SKILL) + self.assertIn("Shadow", SKILL) + self.assertIn("目标章细纲作为生成输入", SKILL) + + def test_freeze_and_two_line_retrieval_rules_are_explicit(self): + self.assertIn("<= as_of", SKILL) + self.assertIn("卡是检索索引和事实补充", SKILL) + self.assertIn("目标章正文", SKILL) + self.assertIn("unknowns", SKILL) + + +if __name__ == "__main__": + unittest.main() diff --git a/.claude/skills/quality-gate/SKILL.md b/.claude/skills/quality-gate/SKILL.md index b1e2b60..a979f68 100644 --- a/.claude/skills/quality-gate/SKILL.md +++ b/.claude/skills/quality-gate/SKILL.md @@ -1,6 +1,6 @@ --- name: quality-gate -description: 质量评分的功能合同(scenario: quality_gate,保护节点)。维度体系对齐专题-04:3 叙事关键定达标线+5 非关键出建议;rubric 与达标口径的单一来源。 +description: 质量评分的功能合同(scenario: quality_gate/fine_outline_replay,保护节点)。正文质量维度和细纲回放 rubric 分开管理;rubric 与达标口径的单一来源。 disable-model-invocation: true --- @@ -33,3 +33,28 @@ disable-model-invocation: true ## 输出合同 报告落 `works/<书>/评审/第NNN章-评分.md`:逐维分数+引文证据;达标结论(关键三维口径);「最值得改的三点」按提升空间排序,每点:问题→根因层猜测(prompt/上下文/设定卡)→具体改法。 + +## 细纲回放 profile(`fine_outline_replay`) + +细纲评分是结构事实评估,不复用正文 `style_fit`、`readability`、`文风一致性`、`文笔` 等维度。每维 1–5 分,分数必须有结构化证据摘要;评委只使用冻结快照和评测侧的 `reference scaffold proxy`,不读取目标章全文。 + +| 维度 ID | 评分问题 | +|---|---| +| `structure_completeness` | 目标、关键事件、实体、伏笔、状态变化和钩子是否齐全 | +| `direction_causality` | 关键冲突、结果方向和事件因果是否命中且成立 | +| `order_pacing` | 事件先后和章内推进节拍是否接近结构化标准 | +| `entity_state` | 实体身份、阵营、能力层级和 N 时点状态是否一致 | +| `foreshadowing_action` | 伏笔埋设、推进、回收和时机是否正确 | +| `handoff_hook` | 能否从 N 时点自然承接,并立住下一步钩子 | + +评审报告使用以下最小形状: + +```yaml +profile: fine_outline_replay +scores: + direction_causality: + score: 4 + evidence: "事件-2 的触发和结果方向与标准事实摘要一致" +``` + +两次独立评审的同维差异大于 `0.5` 时只输出稳定性警告,不据此下卡效用结论。三臂顺序、arm 名称和卡 manifest 对评委隐藏;汇总报告分别给出每维分数、结构门结果、`B-A` 和 `C-A`,不压成没有统计意义的单一“卡质量总分”。 diff --git a/.claude/skills/read-context/SKILL.md b/.claude/skills/read-context/SKILL.md index 371dc3d..dba9ddc 100644 --- a/.claude/skills/read-context/SKILL.md +++ b/.claude/skills/read-context/SKILL.md @@ -7,7 +7,7 @@ description: 组装智能体上下文包——统一创作数据读取器(专题 SoT 对齐:包结构=专题-03 §4.2 **四层上下文**;字段级 aiContext 裁剪=专题-06 §7 统一读取器。 -**输入**:作品名、功能(scenario,专题-03 §4.1:continuation/rewrite/expansion/polish/extraction/full_parse/planning/validation/consistency_check/quality_gate)、目标(如"写第7章")。**用途(purpose)由 scenario 映射**——生成类→`generation`,抽取类→`extraction`,规划→`planning`,检测类→`detection`,quality_gate→基线包即 writer 视图;purpose 定字段可见集(aiContext),scenario 定功能指令段(`meta/chains/`)与 L0 形态,两者不混。 +**输入**:作品名、功能(scenario,专题-03 §4.1:continuation/rewrite/expansion/polish/extraction/full_parse/planning/fine_outline/validation/consistency_check/quality_gate)、目标(如"写第7章")。**用途(purpose)由 scenario 映射**——生成类→`generation`,抽取类→`extraction`,规划类(`planning`/`fine_outline`)→`planning`,检测→`detection`,quality_gate→基线包即 writer 视图;purpose 定字段可见集(aiContext),scenario 定功能指令段(`meta/chains/`)与 L0 形态,两者不混。 ## 步骤 @@ -46,6 +46,7 @@ SoT 对齐:包结构=专题-03 §4.2 **四层上下文**;字段级 aiContext 裁 |---|---|---|---|---| | writer(generation) | 本章细纲+任务 | **必**:上章尾 1–2 场景原文+前章摘要+状态 | 必:设定(裁底牌)+近三章细纲+**仅出场卡** | 按本章场景类型选绑定范式卡 | | planner(planning) | 规划任务+范围 | 可省:近章摘要即可 | **必**:可见度最高——底牌字段开+未来卷粗纲开+知识卡**全量索引** | trope/公式类范式优先 | +| planner(fine_outline) | 目标章细纲任务+未知项纪律 | 可省:只读冻结视图中的近章结构摘要 | **必**:当前大纲、叙事现在时、截至冻结点的安全设定/卡索引;回放时关闭目标章及未来来源 | 仅使用已授权且通过冻结审计的资料 | | extractor(extraction) | 抽取任务+目标型清单 | 待处理章**全文**(是处理对象,不截尾) | schema 字段合同+既有知识**全集**(判重/判冲突) | **关**:防外部范式诱导脑补 | | detector(detection) | 检测任务+待检候选全文 | 必:近邻正文(连续性比对) | **必**:知识全集+状态台账+**底牌开**(查提前泄底必须知道谜底) | 范式卡「失效风险」字段可选 | | judge(quality_gate) | 待评候选+rubric | 随基线包 | **=writer 基线**(共享前缀;评设定一致性需同一基准,不另加底牌) | **关**:不拿范式当标准答案 | diff --git a/.claude/skills/replay-eval/SKILL.md b/.claude/skills/replay-eval/SKILL.md new file mode 100644 index 0000000..79475ad --- /dev/null +++ b/.claude/skills/replay-eval/SKILL.md @@ -0,0 +1,38 @@ +--- +name: replay-eval +description: 回放评测的冻结与结果边界合同。把参考作品冻结到 as_of 章号,生成可审计的输入清单,并在生成/评分前阻断未来信息、未授权来源和全文留存。 +disable-model-invocation: true +--- + +# 回放评测(`next_fine_outline_replay_v0`) + +本 skill 只负责确定性的评测编排边界,不调用模型、不替代统一读取器,也不写正式规划或知识。评测目的、三臂定义和细纲评分合同见 `docs/2026-07-19-回放评测-细纲首跑设计与计划.md`。 + +## 输入合同 + +- `reference_work`:参考作品标识和版本。 +- `as_of`:冻结章号,必须是正整数;所有历史证据只能来自绝对章号 `<= as_of`。 +- `snapshot_version`:不可变的快照版本。 +- 作品大纲窗口:使用 `from_order` / `to_order`,只允许完整窗口 `to_order <= as_of`。 +- 实体卡和里程碑:必须带可验证的绝对章号;无明确上界的区间不进入快照。 +- 来源合同:每个来源要有 `sourceId`、`sourceVersion`、用途授权快照和来源状态。 + +## 冻结规则 + +1. 章号支持整数和数字字符串,拒绝布尔值、浮点猜测、空值和无法证明上界的字符串。 +2. 里程碑按绝对章号过滤;未来条目、跨过 `as_of` 的区间和无章号条目进入 `omittedSources`,不改写成当前事实。 +3. 窗级大纲按 `to_order` 过滤并按 `from_order` 排序;`window_no` 只是展示字段,不能作为冻结键。 +4. 终态摘要、未来弧线、原始当前态等存储字段不能直接透传;只能保留安全历史并由后续消费方明确标记推导状态。 +5. 任何目标章事实、目标实体清单或目标章标签不得参与规划器侧来源选择。 + +## 输出合同 + +输出是临时 `snapshot_manifest`:记录来源 ID/版本、SHA-256、字符数、字段裁剪、来源省略和快照版本,不保存原书正文。三臂的 manifest 除卡注入分区外必须字节一致;不一致时整组作废。 + +授权、来源版本、目标章禁读或内容级泄露检查任一失败,都必须 fail-closed,返回明确的阻断状态,不靠重试绕过。 + +## 产物边界 + +- 原始候选、标准事实摘要和完整输入只能留在临时运行目录或 `/tmp`。 +- 最终报告只允许评分、摘要、章节定位、失败类别和哈希。 +- 禁止写入原书正文、完整目标细纲、完整 Prompt/Response、供应商原始响应、token、密钥或未脱敏授权资料。 diff --git a/.claude/skills/replay-eval/scripts/build_snapshot.py b/.claude/skills/replay-eval/scripts/build_snapshot.py new file mode 100644 index 0000000..a11ad33 --- /dev/null +++ b/.claude/skills/replay-eval/scripts/build_snapshot.py @@ -0,0 +1,340 @@ +#!/usr/bin/env python3 +"""回放评测的冻结快照纯函数。 + +本模块只处理内存中的结构化数据,不访问数据库、不调用模型、不写正式数据。 +它把「截至第 N 章」变成可重复、可审计的来源清单,避免把未来事实误当成卡的价值。 +""" + +from __future__ import annotations + +import argparse +import copy +import hashlib +import json +import re +from pathlib import Path +from typing import Any, Iterable, Mapping, Sequence + + +class SnapshotError(ValueError): + """快照输入违反冻结或全文留存边界。""" + + +TERMINAL_FIELDS = frozenset( + { + "final_summary", + "terminal_summary", + "final_state", + "current_state", + "future_arc", + "future_plan", + "成长弧线", + "当前态", + "终态摘要", + "终局状态", + "未来弧线", + "未来计划", + } +) + +FINAL_REPORT_FORBIDDEN_FIELDS = frozenset( + { + "raw", + "raw_text", + "body", + "content", + "full_text", + "full_body", + "原文", + "正文", + "正文全文", + "完整目标细纲", + "target_chapter_text", + "prompt", + "response", + } +) + +_INTEGER_RE = re.compile(r"^\s*(\d+)\s*$") +_RANGE_RE = re.compile(r"^\s*(?:第\s*)?(\d+)\s*(?:-|–|—|~|至|到)\s*(?:第\s*)?(\d+)\s*(?:章)?\s*$") + + +def normalize_chapter(value: Any) -> int | None: + """只接受明确的正整数章号;不把 bool、浮点或模糊文本猜成章号。""" + + if isinstance(value, bool): + return None + if isinstance(value, int): + return value if value > 0 else None + if not isinstance(value, str): + return None + matched = _INTEGER_RE.fullmatch(value) + if not matched: + return None + chapter = int(matched.group(1)) + return chapter if chapter > 0 else None + + +def normalize_chapter_range(value: Any) -> tuple[int, int] | None: + """解析单章或有明确起止上界的章区间。""" + + chapter = normalize_chapter(value) + if chapter is not None: + return chapter, chapter + if not isinstance(value, str): + return None + matched = _RANGE_RE.fullmatch(value) + if not matched: + return None + start, end = int(matched.group(1)), int(matched.group(2)) + if start <= 0 or end < start: + return None + return start, end + + +def _chapter_value(record: Mapping[str, Any]) -> Any: + """兼容实验台和 schema 中的中英文章号键。""" + + for key in ("chapter", "chapter_no", "order_no", "章", "章号"): + if key in record: + return record[key] + if "from_order" in record or "to_order" in record: + start = record.get("from_order") + end = record.get("to_order") + if start is not None and end is not None: + return f"{start}-{end}" + return None + + +def _safe_json(value: Any) -> str: + """用固定格式序列化,确保相同输入产生相同哈希。""" + + return json.dumps(value, ensure_ascii=False, sort_keys=True, separators=(",", ":")) + + +def sha256_value(value: Any) -> str: + """返回结构化值的稳定 SHA-256。""" + + return hashlib.sha256(_safe_json(value).encode("utf-8")).hexdigest() + + +def _omitted(index: int, reason: str, record: Any) -> dict[str, Any]: + """只记录定位和原因,不把被排除的原文复制到回显。""" + + return {"index": index, "reason": reason, "recordHash": sha256_value(record)} + + +def filter_milestones( + milestones: Iterable[Mapping[str, Any]], as_of: int +) -> tuple[list[dict[str, Any]], list[dict[str, Any]]]: + """保留完整落在 as_of 以前的里程碑,返回(保留项、排除项)。""" + + as_of = normalize_chapter(as_of) + if as_of is None: + raise SnapshotError("as_of 必须是正整数章号") + + kept: list[dict[str, Any]] = [] + omitted: list[dict[str, Any]] = [] + for index, item in enumerate(milestones): + if not isinstance(item, Mapping): + omitted.append(_omitted(index, "not_an_object", item)) + continue + bounds = normalize_chapter_range(_chapter_value(item)) + if bounds is None: + omitted.append(_omitted(index, "missing_or_unbounded_chapter", item)) + continue + if bounds[1] > as_of: + omitted.append(_omitted(index, "future_or_crosses_as_of", item)) + continue + kept.append(copy.deepcopy(dict(item))) + + kept.sort(key=lambda item: normalize_chapter_range(_chapter_value(item)) or (0, 0)) + return kept, omitted + + +def filter_outline_windows( + windows: Iterable[Mapping[str, Any]], as_of: int +) -> tuple[list[dict[str, Any]], list[dict[str, Any]]]: + """只保留 to_order<=as_of 的完整窗口,并按 from_order 排序。""" + + as_of = normalize_chapter(as_of) + if as_of is None: + raise SnapshotError("as_of 必须是正整数章号") + + kept: list[dict[str, Any]] = [] + omitted: list[dict[str, Any]] = [] + for index, item in enumerate(windows): + if not isinstance(item, Mapping): + omitted.append(_omitted(index, "not_an_object", item)) + continue + start = normalize_chapter(item.get("from_order")) + end = normalize_chapter(item.get("to_order")) + if start is None or end is None or start > end: + omitted.append(_omitted(index, "missing_or_invalid_window_bounds", item)) + continue + if end > as_of: + omitted.append(_omitted(index, "future_or_crosses_as_of", item)) + continue + kept.append(copy.deepcopy(dict(item))) + + kept.sort(key=lambda item: (normalize_chapter(item["from_order"]) or 0, normalize_chapter(item["to_order"]) or 0)) + return kept, omitted + + +def remove_terminal_fields(value: Any) -> Any: + """递归删除不能直接作为 as_of 事实使用的终态/未来字段。""" + + if isinstance(value, list): + return [remove_terminal_fields(item) for item in value] + if not isinstance(value, Mapping): + return copy.deepcopy(value) + + projected: dict[str, Any] = {} + for key, item in value.items(): + if str(key) in TERMINAL_FIELDS: + continue + projected[str(key)] = remove_terminal_fields(item) + return projected + + +def _validate_final_report(value: Any, path: str = "finalReport") -> None: + """禁止最终报告携带原书全文、完整响应或完整 prompt。""" + + if isinstance(value, Mapping): + for key, item in value.items(): + if str(key) in FINAL_REPORT_FORBIDDEN_FIELDS: + raise SnapshotError(f"{path}.{key} 不得进入最终报告") + _validate_final_report(item, f"{path}.{key}") + elif isinstance(value, list): + for index, item in enumerate(value): + _validate_final_report(item, f"{path}[{index}]") + + +def _source_record(section: str, value: Any) -> dict[str, Any]: + """将来源内容收敛成元数据,绝不把 payload 写入 manifest。""" + + if isinstance(value, Mapping) and "payload" in value: + source_id = str(value.get("sourceId") or section) + source_version = str(value.get("sourceVersion") or "unknown") + payload = value["payload"] + omitted_fields = list(value.get("omittedFields") or []) + omitted_sources = list(value.get("omittedSources") or []) + else: + source_id = section + source_version = "unknown" + payload = value + omitted_fields = [] + omitted_sources = [] + + serialized = _safe_json(payload) + return { + "section": section, + "sourceId": source_id, + "sourceVersion": source_version, + "sha256": hashlib.sha256(serialized.encode("utf-8")).hexdigest(), + "charCount": len(serialized), + "omittedFields": omitted_fields, + "omittedSources": omitted_sources, + } + + +def build_snapshot_manifest( + *, + as_of: int, + snapshot_version: str, + sections: Mapping[str, Any], + omitted_fields: Sequence[Any] | None = None, + omitted_sources: Sequence[Any] | None = None, + final_report: Mapping[str, Any] | None = None, +) -> dict[str, Any]: + """构造稳定 manifest;最终报告只允许摘要、定位、评分和哈希。""" + + normalized_as_of = normalize_chapter(as_of) + if normalized_as_of is None: + raise SnapshotError("as_of 必须是正整数章号") + if not snapshot_version.strip(): + raise SnapshotError("snapshot_version 不能为空") + if final_report is not None: + _validate_final_report(final_report) + + manifest: dict[str, Any] = { + "snapshotVersion": snapshot_version, + "asOfChapter": normalized_as_of, + "sections": { + str(section): _source_record(str(section), value) + for section, value in sorted(sections.items(), key=lambda pair: str(pair[0])) + }, + "omittedFields": list(omitted_fields or []), + "omittedSources": list(omitted_sources or []), + "finalReport": copy.deepcopy(final_report or {}), + } + manifest["manifestSha256"] = sha256_value(manifest) + return manifest + + +def build_snapshot(data: Mapping[str, Any], as_of: int, snapshot_version: str) -> dict[str, Any]: + """从最小 JSON 输入生成冻结后的结构化快照和 manifest。""" + + safe = copy.deepcopy(dict(data)) + all_omitted: list[dict[str, Any]] = [] + + if isinstance(safe.get("milestones"), list): + milestones, omitted = filter_milestones(safe["milestones"], as_of) + safe["milestones"] = milestones + all_omitted.extend({"source": "milestones", **item} for item in omitted) + + if isinstance(safe.get("outlineWindows"), list): + windows, omitted = filter_outline_windows(safe["outlineWindows"], as_of) + safe["outlineWindows"] = windows + all_omitted.extend({"source": "outlineWindows", **item} for item in omitted) + + if isinstance(safe.get("cards"), list): + frozen_cards: list[dict[str, Any]] = [] + for card_index, card in enumerate(safe["cards"]): + if not isinstance(card, Mapping): + all_omitted.append(_omitted(card_index, "card_not_an_object", card)) + continue + projected = copy.deepcopy(dict(card)) + history_key = next( + (key for key in ("milestones", "演变历程") if isinstance(projected.get(key), list)), + None, + ) + if history_key is not None: + history, omitted = filter_milestones(projected[history_key], as_of) + projected[history_key] = history + all_omitted.extend( + {"source": f"cards[{card_index}].{history_key}", **item} + for item in omitted + ) + frozen_cards.append(projected) + safe["cards"] = frozen_cards + + safe = remove_terminal_fields(safe) + manifest = build_snapshot_manifest( + as_of=as_of, + snapshot_version=snapshot_version, + sections={"snapshot": {"sourceId": "frozen-snapshot", "sourceVersion": snapshot_version, "payload": safe}}, + omitted_sources=all_omitted, + ) + return {"snapshot": safe, "manifest": manifest} + + +def _parse_args() -> argparse.Namespace: + parser = argparse.ArgumentParser(description="构造 as_of 章号冻结快照") + parser.add_argument("--input", type=Path, required=True, help="结构化 JSON 输入") + parser.add_argument("--output", type=Path, required=True, help="输出 JSON 路径") + parser.add_argument("--as-of", type=int, required=True, dest="as_of") + parser.add_argument("--snapshot-version", default="next_fine_outline_replay_v0") + return parser.parse_args() + + +def main() -> int: + args = _parse_args() + data = json.loads(args.input.read_text(encoding="utf-8")) + result = build_snapshot(data, args.as_of, args.snapshot_version) + args.output.write_text(_safe_json(result) + "\n", encoding="utf-8") + return 0 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/.claude/skills/replay-eval/scripts/check_snapshot.py b/.claude/skills/replay-eval/scripts/check_snapshot.py new file mode 100644 index 0000000..fd3b0e7 --- /dev/null +++ b/.claude/skills/replay-eval/scripts/check_snapshot.py @@ -0,0 +1,315 @@ +#!/usr/bin/env python3 +"""回放快照的 fail-closed 校验。 + +输入是冻结脚本产出的结构化 manifest 和运行时登记信息;本模块不访问数据库、不调用模型。 +""" + +from __future__ import annotations + +import argparse +import json +from pathlib import Path +from typing import Any, Mapping, Sequence + +from build_snapshot import normalize_chapter, normalize_chapter_range + + +STATUS_READY = "ready" +STATUS_BLOCKED_AUTHORIZATION = "blocked_authorization" +STATUS_INVALID_SNAPSHOT = "invalid_snapshot" +STATUS_INVALID_ARM_DIFF = "invalid_arm_diff" +STATUS_TARGET_SOURCE_FORBIDDEN = "target_source_forbidden" +STATUS_SCHEMA_INVALID = "schema_invalid" + +FORBIDDEN_SOURCE_STATUSES = frozenset( + {"revoked", "delisted", "recalled", "blocked", "owner_missing", "unauthorized"} +) +ALLOWED_SOURCE_STATUSES = frozenset({"active", "approved", "authorized", "licensed"}) +FORBIDDEN_COPYRIGHT_STATUSES = frozenset( + {"unauthorized", "unlicensed", "revoked", "expired", "blocked"} +) +CARD_KEYS = frozenset( + { + "arm", + "card", + "cards", + "cardInjection", + "cardSection", + "cardSections", + "cardManifest", + "cardIds", + "cardSourceIds", + "cardStrategy", + "l2-card", + "l2-placebo", + "card_injection", + "card_section", + "card_sections", + "card_manifest", + "card_ids", + "card_source_ids", + "card_strategy", + "l2_card", + "l2_placebo", + } +) +NORMALIZED_CARD_KEYS = frozenset(key.lower() for key in CARD_KEYS) +REQUIRED_CANDIDATE_FIELDS = ( + "targetChapter", + "chapterGoal", + "keyEvents", + "entities", + "foreshadowing", + "stateChanges", + "hook", + "unknowns", + "assumptions", +) +FORBIDDEN_CANDIDATE_FIELDS = frozenset( + {"body", "raw", "rawText", "content", "正文", "原文", "正文全文", "完整目标细纲"} +) + + +def _result(status: str, errors: Sequence[str] = (), warnings: Sequence[str] = ()) -> dict[str, Any]: + return {"status": status, "ok": status == STATUS_READY, "errors": list(errors), "warnings": list(warnings)} + + +def _field(value: Mapping[str, Any], *keys: str) -> Any: + for key in keys: + if key in value: + return value[key] + return None + + +def check_authorization(authorization: Mapping[str, Any] | None) -> dict[str, Any]: + """授权信息缺失、用途不符或来源进入危险状态时关闭评测。""" + + if not isinstance(authorization, Mapping): + return _result(STATUS_BLOCKED_AUTHORIZATION, ["缺少不可变授权快照"]) + snapshot = _field(authorization, "authorizationSnapshot", "authorization_snapshot") + if not isinstance(snapshot, Mapping) or not snapshot: + return _result(STATUS_BLOCKED_AUTHORIZATION, ["授权快照为空"]) + + source_status = str(_field(authorization, "sourceStatus", "source_status") or "").lower() + if not source_status: + return _result(STATUS_BLOCKED_AUTHORIZATION, ["缺少 sourceStatus"]) + if source_status in FORBIDDEN_SOURCE_STATUSES: + return _result(STATUS_BLOCKED_AUTHORIZATION, [f"来源状态禁止评测: {source_status}"]) + if source_status not in ALLOWED_SOURCE_STATUSES: + return _result(STATUS_BLOCKED_AUTHORIZATION, [f"来源状态未登记,拒绝评测: {source_status}"]) + + copyright_status = str( + _field(authorization, "copyrightStatus", "copyright_status") or "" + ).lower() + if copyright_status in FORBIDDEN_COPYRIGHT_STATUSES: + return _result(STATUS_BLOCKED_AUTHORIZATION, [f"版权状态禁止评测: {copyright_status}"]) + + allowed = _field(authorization, "allowedPurpose", "allowed_purpose") + if isinstance(allowed, str): + allowed = [allowed] + if not isinstance(allowed, Sequence) or isinstance(allowed, (str, bytes)): + return _result(STATUS_BLOCKED_AUTHORIZATION, ["allowedPurpose 不是用途列表"]) + if "offline_evaluation" not in allowed: + return _result(STATUS_BLOCKED_AUTHORIZATION, ["allowedPurpose 不包含 offline_evaluation"]) + return _result(STATUS_READY) + + +def check_target_sources(target_chapter: int, sources: Sequence[Mapping[str, Any]]) -> dict[str, Any]: + """规划器侧来源不得包含目标章或更晚章号。""" + + target = normalize_chapter(target_chapter) + if target is None: + return _result(STATUS_INVALID_SNAPSHOT, ["目标章号无效"]) + errors: list[str] = [] + for index, source in enumerate(sources): + if not isinstance(source, Mapping): + errors.append(f"source[{index}] 不是对象") + continue + value = _field(source, "chapter", "chapterNo", "chapter_no", "章", "章号") + bounds = normalize_chapter_range(value) + if bounds is None: + value = _field( + source, + "chapterRange", + "chapter_range", + "range", + "章节范围", + ) + bounds = normalize_chapter_range(value) + if bounds is None and "from_order" in source: + value = f"{source.get('from_order')}-{source.get('to_order')}" + bounds = normalize_chapter_range(value) + if bounds is not None and bounds[1] >= target: + errors.append(f"source[{index}] 包含目标章或未来章: {bounds[0]}-{bounds[1]}") + return _result(STATUS_TARGET_SOURCE_FORBIDDEN if errors else STATUS_READY, errors) + + +def _without_card_fields(value: Any) -> Any: + """去掉三臂允许变化的 arm/card 分区,保留所有公共输入用于字节级比较。""" + + if isinstance(value, list): + return [_without_card_fields(item) for item in value] + if not isinstance(value, Mapping): + return value + result: dict[str, Any] = {} + for key, item in value.items(): + key_text = str(key) + if key_text.lower() in NORMALIZED_CARD_KEYS or key_text in {"卡", "知识卡", "卡片"}: + continue + result[key_text] = _without_card_fields(item) + return result + + +def check_arm_manifests( + manifests: Mapping[str, Mapping[str, Any]], + required_arms: Sequence[str] | None = None, +) -> dict[str, Any]: + """确保三臂除卡注入区外完全一致。""" + + required = set(required_arms or ("outline_only", "outline_plus_cards", "outline_plus_placebo_cards")) + actual = set(manifests) + missing = sorted(required - actual) + if missing: + return _result(STATUS_INVALID_ARM_DIFF, [f"缺少评测臂: {','.join(missing)}"]) + unexpected = sorted(actual - required) + if unexpected: + return _result(STATUS_INVALID_ARM_DIFF, [f"存在未登记评测臂: {','.join(unexpected)}"]) + if any(not isinstance(manifest, Mapping) for manifest in manifests.values()): + return _result(STATUS_INVALID_ARM_DIFF, ["评测臂 manifest 必须是对象"]) + + names = sorted(required) + baseline = _without_card_fields(manifests[names[0]]) + differences = [name for name in names[1:] if _without_card_fields(manifests[name]) != baseline] + if differences: + return _result(STATUS_INVALID_ARM_DIFF, [f"公共输入区不一致: {','.join(differences)}"]) + return _result(STATUS_READY) + + +def check_candidate_output( + candidate: Mapping[str, Any], + target_chapter: int, + source_catalog: Sequence[Mapping[str, Any]] = (), +) -> dict[str, Any]: + """校验细纲候选的结构边界,不判断内容是否命中标准答案。""" + + if not isinstance(candidate, Mapping): + return _result(STATUS_SCHEMA_INVALID, ["候选不是对象"]) + expected_target = normalize_chapter(target_chapter) + if expected_target is None: + return _result(STATUS_INVALID_SNAPSHOT, ["目标章号无效"]) + errors = [f"缺少字段: {field}" for field in REQUIRED_CANDIDATE_FIELDS if field not in candidate] + forbidden = sorted(set(candidate) & FORBIDDEN_CANDIDATE_FIELDS) + if forbidden: + errors.append(f"候选包含正文/原文字段: {','.join(forbidden)}") + for field in ("keyEvents", "entities", "foreshadowing", "stateChanges", "unknowns", "assumptions"): + if field in candidate and not isinstance(candidate[field], list): + errors.append(f"字段必须是数组: {field}") + for field in ("chapterGoal", "hook"): + if field in candidate and not isinstance(candidate[field], str): + errors.append(f"字段必须是字符串: {field}") + events = candidate.get("keyEvents") + if isinstance(events, list): + event_ids = [item.get("id") for item in events if isinstance(item, Mapping) and item.get("id")] + if len(event_ids) != len(set(event_ids)): + errors.append("keyEvents 包含重复事件 id") + if any(not isinstance(item, Mapping) for item in events): + errors.append("keyEvents 每项必须是对象") + if any(isinstance(item, Mapping) and not item.get("id") for item in events): + errors.append("keyEvents 每项必须有 id") + source_errors: list[str] = [] + source_refs = candidate.get("sourceRefs") + if source_refs is not None: + if not isinstance(source_refs, list): + errors.append("sourceRefs 必须是数组") + else: + catalog = { + str(source.get("sourceId")): source + for source in source_catalog + if isinstance(source, Mapping) and source.get("sourceId") + } + for index, reference in enumerate(source_refs): + source = reference if isinstance(reference, Mapping) else catalog.get(str(reference)) + if source is None: + if source_catalog: + errors.append(f"sourceRefs[{index}] 未登记来源") + continue + source_value = _field( + source, + "chapter", + "chapterNo", + "chapter_no", + "章", + "章号", + "chapterRange", + "chapter_range", + "range", + ) + bounds = normalize_chapter_range(source_value) + if bounds is None and "from_order" in source: + bounds = normalize_chapter_range(f"{source.get('from_order')}-{source.get('to_order')}") + if bounds is not None and bounds[1] >= expected_target: + source_errors.append( + f"sourceRefs[{index}] 包含目标章或未来章: {bounds[0]}-{bounds[1]}" + ) + actual_target = normalize_chapter(candidate.get("targetChapter")) + if actual_target != expected_target: + errors.append(f"候选目标章错误: expected={expected_target}, actual={actual_target}") + if source_errors: + return _result(STATUS_TARGET_SOURCE_FORBIDDEN, source_errors) + if errors: + return _result(STATUS_SCHEMA_INVALID, errors) + return _result(STATUS_READY) + + +def check_replay( + *, + authorization: Mapping[str, Any] | None, + target_chapter: int, + planner_sources: Sequence[Mapping[str, Any]], + arm_manifests: Mapping[str, Mapping[str, Any]], + required_arms: Sequence[str] | None = None, +) -> dict[str, Any]: + """执行回放前置门,任何一项失败都不允许进入模型调用。""" + + checks = [ + check_authorization(authorization), + check_target_sources(target_chapter, planner_sources), + check_arm_manifests(arm_manifests, required_arms), + ] + failures = [check for check in checks if not check["ok"]] + if failures: + return _result( + failures[0]["status"], + [error for check in failures for error in check["errors"]], + [warning for check in failures for warning in check["warnings"]], + ) + return _result(STATUS_READY) + + +def _parse_args() -> argparse.Namespace: + parser = argparse.ArgumentParser(description="校验回放快照和三臂输入") + parser.add_argument("--authorization", type=Path, required=True) + parser.add_argument("--sources", type=Path, required=True) + parser.add_argument("--arms", type=Path, required=True) + parser.add_argument("--target-chapter", type=int, required=True) + parser.add_argument("--output", type=Path, required=True) + return parser.parse_args() + + +def main() -> int: + args = _parse_args() + authorization = json.loads(args.authorization.read_text(encoding="utf-8")) + sources = json.loads(args.sources.read_text(encoding="utf-8")) + arms = json.loads(args.arms.read_text(encoding="utf-8")) + result = check_replay( + authorization=authorization, + target_chapter=args.target_chapter, + planner_sources=sources, + arm_manifests=arms, + ) + args.output.write_text(json.dumps(result, ensure_ascii=False, sort_keys=True) + "\n", encoding="utf-8") + return 0 if result["ok"] else 2 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/.claude/skills/replay-eval/scripts/fine_outline_rubric.py b/.claude/skills/replay-eval/scripts/fine_outline_rubric.py new file mode 100644 index 0000000..d79a9ca --- /dev/null +++ b/.claude/skills/replay-eval/scripts/fine_outline_rubric.py @@ -0,0 +1,77 @@ +#!/usr/bin/env python3 +"""细纲回放 rubric 的确定性校验器。 + +它不替代 judge 打分,只检查评分报告是否使用正确维度、分数范围和证据字段。 +""" + +from __future__ import annotations + +from typing import Any, Mapping + + +RUBRIC_PROFILE = "fine_outline_replay" +DIMENSIONS = ( + "structure_completeness", + "direction_causality", + "order_pacing", + "entity_state", + "foreshadowing_action", + "handoff_hook", +) +PROSE_DIMENSIONS = frozenset( + {"style_fit", "readability", "文风一致性", "文笔", "pacing_tension", "information_density"} +) + + +def validate_scores(scores: Mapping[str, Any]) -> list[str]: + """返回报告问题;每个维度必须有 1-5 分和非空证据。""" + + errors: list[str] = [] + if not isinstance(scores, Mapping): + return ["scores 必须是对象"] + missing = [dimension for dimension in DIMENSIONS if dimension not in scores] + if missing: + errors.append(f"缺少 rubric 维度: {','.join(missing)}") + unexpected = sorted(set(scores) - set(DIMENSIONS)) + if unexpected: + errors.append(f"存在未登记 rubric 维度: {','.join(unexpected)}") + prose = sorted(set(unexpected) & PROSE_DIMENSIONS) + if prose: + errors.append(f"细纲 rubric 禁止正文质量维度: {','.join(prose)}") + for dimension in DIMENSIONS: + value = scores.get(dimension) + if not isinstance(value, Mapping): + errors.append(f"维度必须包含 score/evidence 对象: {dimension}") + continue + score = value.get("score") + if isinstance(score, bool) or not isinstance(score, (int, float)) or not 1 <= score <= 5: + errors.append(f"分数必须在 1-5: {dimension}") + evidence = value.get("evidence") + if not isinstance(evidence, str) or not evidence.strip(): + errors.append(f"分数缺少证据: {dimension}") + return errors + + +def stability_warning( + first: Mapping[str, float], + second: Mapping[str, float], + threshold: float = 0.5, +) -> dict[str, Any]: + """比较两次评审,返回差异和是否需要人工复核。""" + + gaps = { + dimension: abs(float(first[dimension]) - float(second[dimension])) + for dimension in DIMENSIONS + if dimension in first and dimension in second + } + return {"stable": all(gap <= threshold for gap in gaps.values()), "gaps": gaps} + + +def validate_report(report: Mapping[str, Any]) -> list[str]: + """校验一个评委报告的 profile 和评分结构。""" + + errors: list[str] = [] + if report.get("profile") != RUBRIC_PROFILE: + errors.append(f"profile 必须是 {RUBRIC_PROFILE}") + errors.extend(validate_scores(report.get("scores", {}))) + return errors diff --git a/.claude/skills/replay-eval/scripts/test_build_snapshot.py b/.claude/skills/replay-eval/scripts/test_build_snapshot.py new file mode 100644 index 0000000..03e855e --- /dev/null +++ b/.claude/skills/replay-eval/scripts/test_build_snapshot.py @@ -0,0 +1,114 @@ +#!/usr/bin/env python3 +"""build_snapshot.py 的无网络离线测试。""" + +import pathlib +import sys +import unittest + +sys.path.insert(0, str(pathlib.Path(__file__).resolve().parent)) +from build_snapshot import ( # noqa: E402 + SnapshotError, + build_snapshot, + build_snapshot_manifest, + filter_milestones, + filter_outline_windows, + normalize_chapter, + remove_terminal_fields, +) + + +class BuildSnapshotTest(unittest.TestCase): + def test_normalize_chapter_rejects_guessing(self): + self.assertEqual(normalize_chapter(488), 488) + self.assertEqual(normalize_chapter(" 488 "), 488) + self.assertIsNone(normalize_chapter(True)) + self.assertIsNone(normalize_chapter(488.0)) + self.assertIsNone(normalize_chapter("第488章")) + + def test_milestones_are_frozen_by_absolute_chapter(self): + kept, omitted = filter_milestones( + [ + {"章": 487, "台阶": "已知"}, + {"章": "488", "台阶": "已知字符串"}, + {"章": "489", "台阶": "未来"}, + {"章": "487-490", "台阶": "跨冻结点"}, + {"台阶": "没有章号"}, + ], + 488, + ) + self.assertEqual([item["台阶"] for item in kept], ["已知", "已知字符串"]) + self.assertEqual( + [item["reason"] for item in omitted], + ["future_or_crosses_as_of", "future_or_crosses_as_of", "missing_or_unbounded_chapter"], + ) + self.assertTrue(all("台阶" not in item for item in omitted)) + + def test_outline_windows_require_complete_to_order(self): + kept, omitted = filter_outline_windows( + [ + {"window_no": 2, "from_order": 401, "to_order": 430}, + {"window_no": 1, "from_order": 431, "to_order": 490}, + {"window_no": 3, "from_order": "bad", "to_order": 500}, + ], + 488, + ) + self.assertEqual([item["from_order"] for item in kept], [401]) + self.assertEqual([item["reason"] for item in omitted], ["future_or_crosses_as_of", "missing_or_invalid_window_bounds"]) + + def test_terminal_fields_are_removed_recursively(self): + value = { + "name": "苏铭", + "current_state": "终局状态", + "nested": {"future_arc": "未来计划", "safe": "保留"}, + "items": [{"终态摘要": "不应出现"}, {"safe": "保留"}], + } + self.assertEqual(remove_terminal_fields(value), {"name": "苏铭", "nested": {"safe": "保留"}, "items": [{}, {"safe": "保留"}]}) + + def test_nested_card_history_is_frozen(self): + result = build_snapshot( + { + "cards": [ + { + "name": "生物机甲", + "milestones": [ + {"chapter": 488, "step": "可见"}, + {"chapter": 489, "step": "未来"}, + ], + "current_state": "终态字段不得透传", + } + ] + }, + 488, + "v0", + ) + card = result["snapshot"]["cards"][0] + self.assertEqual(card["milestones"], [{"chapter": 488, "step": "可见"}]) + self.assertNotIn("current_state", card) + self.assertEqual(result["manifest"]["omittedSources"][0]["reason"], "future_or_crosses_as_of") + + def test_manifest_is_stable_and_does_not_store_payload(self): + kwargs = { + "as_of": 488, + "snapshot_version": "v0", + "sections": {"l0": {"sourceId": "task", "sourceVersion": "1", "payload": {"target": 489}}}, + "final_report": {"status": "ready", "sourceHash": "abc"}, + } + first = build_snapshot_manifest(**kwargs) + second = build_snapshot_manifest(**kwargs) + self.assertEqual(first, second) + self.assertNotIn("payload", first["sections"]["l0"]) + self.assertEqual(len(first["manifestSha256"]), 64) + self.assertNotIn("原书", str(first["finalReport"])) + + def test_final_report_rejects_original_text_fields(self): + with self.assertRaises(SnapshotError): + build_snapshot_manifest( + as_of=488, + snapshot_version="v0", + sections={}, + final_report={"raw_text": "原书全文"}, + ) + + +if __name__ == "__main__": + unittest.main() diff --git a/.claude/skills/replay-eval/scripts/test_check_snapshot.py b/.claude/skills/replay-eval/scripts/test_check_snapshot.py new file mode 100644 index 0000000..1fcff72 --- /dev/null +++ b/.claude/skills/replay-eval/scripts/test_check_snapshot.py @@ -0,0 +1,119 @@ +#!/usr/bin/env python3 +"""check_snapshot.py 的无网络离线测试。""" + +import pathlib +import sys +import unittest + +sys.path.insert(0, str(pathlib.Path(__file__).resolve().parent)) +from check_snapshot import ( # noqa: E402 + STATUS_BLOCKED_AUTHORIZATION, + STATUS_INVALID_ARM_DIFF, + STATUS_READY, + STATUS_SCHEMA_INVALID, + STATUS_TARGET_SOURCE_FORBIDDEN, + check_arm_manifests, + check_authorization, + check_candidate_output, + check_replay, + check_target_sources, +) + + +AUTH = { + "sourceStatus": "active", + "allowedPurpose": ["offline_evaluation"], + "authorizationSnapshot": {"id": "auth-1", "version": "v1"}, +} + + +def arm(common, cards): + return {**common, "cardInjection": cards} + + +class CheckSnapshotTest(unittest.TestCase): + def test_authorization_is_fail_closed(self): + self.assertEqual(check_authorization(None)["status"], STATUS_BLOCKED_AUTHORIZATION) + self.assertTrue(check_authorization(AUTH)["ok"]) + denied = {**AUTH, "allowedPurpose": ["read"]} + self.assertEqual(check_authorization(denied)["status"], STATUS_BLOCKED_AUTHORIZATION) + unknown_status = {**AUTH, "sourceStatus": "temporary"} + self.assertEqual(check_authorization(unknown_status)["status"], STATUS_BLOCKED_AUTHORIZATION) + unlicensed = {**AUTH, "copyrightStatus": "unlicensed"} + self.assertEqual(check_authorization(unlicensed)["status"], STATUS_BLOCKED_AUTHORIZATION) + + def test_target_source_is_blocked(self): + allowed = check_target_sources(489, [{"chapter": 488}, {"from_order": 450, "to_order": 482}]) + self.assertEqual(allowed["status"], STATUS_READY) + blocked = check_target_sources(489, [{"chapter": 489}]) + self.assertEqual(blocked["status"], STATUS_TARGET_SOURCE_FORBIDDEN) + + def test_arm_common_input_must_match(self): + common = {"snapshotVersion": "v0", "asOfChapter": 488, "l0": {"target": 489}} + manifests = { + "outline_only": arm(common, []), + "outline_plus_cards": arm(common, [{"id": 1}]), + "outline_plus_placebo_cards": arm(common, [{"id": 2}]), + } + self.assertTrue(check_arm_manifests(manifests)["ok"]) + extra = {**manifests, "unregistered_arm": arm(common, [])} + self.assertEqual(check_arm_manifests(extra)["status"], STATUS_INVALID_ARM_DIFF) + changed = dict(manifests) + changed["outline_plus_cards"] = arm({**common, "l0": {"target": 490}}, [{"id": 1}]) + self.assertEqual(check_arm_manifests(changed)["status"], STATUS_INVALID_ARM_DIFF) + + def test_two_arm_smoke_can_be_explicit(self): + common = {"snapshotVersion": "v0", "asOfChapter": 488} + manifests = {"outline_only": arm(common, []), "outline_plus_cards": arm(common, [{"id": 1}])} + self.assertTrue(check_arm_manifests(manifests, ["outline_only", "outline_plus_cards"])["ok"]) + + def test_candidate_contract_is_structural(self): + candidate = { + "targetChapter": 489, + "chapterGoal": "突破", + "keyEvents": [], + "entities": [], + "foreshadowing": [], + "stateChanges": [], + "hook": "悬念", + "unknowns": [], + "assumptions": [], + } + self.assertTrue(check_candidate_output(candidate, 489)["ok"]) + bad = {**candidate, "正文全文": "原文"} + self.assertEqual(check_candidate_output(bad, 489)["status"], STATUS_SCHEMA_INVALID) + bad_type = {**candidate, "keyEvents": "事件"} + self.assertEqual(check_candidate_output(bad_type, 489)["status"], STATUS_SCHEMA_INVALID) + duplicate = {**candidate, "keyEvents": [{"id": "event-1"}, {"id": "event-1"}]} + self.assertEqual(check_candidate_output(duplicate, 489)["status"], STATUS_SCHEMA_INVALID) + missing_id = {**candidate, "keyEvents": [{"event": "没有 id"}]} + self.assertEqual(check_candidate_output(missing_id, 489)["status"], STATUS_SCHEMA_INVALID) + future_ref = {**candidate, "sourceRefs": ["target-scaffold"]} + self.assertEqual( + check_candidate_output( + future_ref, + 489, + [{"sourceId": "target-scaffold", "chapter": 489}], + )["status"], + STATUS_TARGET_SOURCE_FORBIDDEN, + ) + + def test_replay_fails_closed_before_model(self): + common = {"snapshotVersion": "v0", "asOfChapter": 488} + manifests = { + "outline_only": arm(common, []), + "outline_plus_cards": arm(common, [{"id": 1}]), + "outline_plus_placebo_cards": arm(common, [{"id": 2}]), + } + result = check_replay( + authorization=AUTH, + target_chapter=489, + planner_sources=[{"chapter": 489}], + arm_manifests=manifests, + ) + self.assertFalse(result["ok"]) + self.assertEqual(result["status"], STATUS_TARGET_SOURCE_FORBIDDEN) + + +if __name__ == "__main__": + unittest.main() diff --git a/.claude/skills/replay-eval/scripts/test_rubric.py b/.claude/skills/replay-eval/scripts/test_rubric.py new file mode 100644 index 0000000..2ebe557 --- /dev/null +++ b/.claude/skills/replay-eval/scripts/test_rubric.py @@ -0,0 +1,54 @@ +#!/usr/bin/env python3 +"""细纲 rubric 的离线回归测试。""" + +import pathlib +import sys +import unittest + +sys.path.insert(0, str(pathlib.Path(__file__).resolve().parent)) +from fine_outline_rubric import ( # noqa: E402 + DIMENSIONS, + RUBRIC_PROFILE, + stability_warning, + validate_report, + validate_scores, +) + + +def valid_scores(): + return { + dimension: {"score": 4, "evidence": f"证据-{dimension}"} + for dimension in DIMENSIONS + } + + +class FineOutlineRubricTest(unittest.TestCase): + def test_every_dimension_requires_evidence(self): + self.assertEqual(validate_scores(valid_scores()), []) + missing_evidence = valid_scores() + missing_evidence[DIMENSIONS[0]] = {"score": 4} + self.assertTrue(any("缺少证据" in error for error in validate_scores(missing_evidence))) + + def test_prose_dimensions_are_rejected(self): + scores = valid_scores() + scores["style_fit"] = {"score": 5, "evidence": "不应出现"} + self.assertTrue(any("禁止正文质量维度" in error for error in validate_scores(scores))) + + def test_profile_and_score_range_are_checked(self): + report = {"profile": RUBRIC_PROFILE, "scores": valid_scores()} + self.assertEqual(validate_report(report), []) + bad = {"profile": "quality_gate", "scores": valid_scores()} + bad["scores"][DIMENSIONS[1]] = {"score": 6, "evidence": "超范围"} + self.assertEqual(len(validate_report(bad)), 2) + + def test_large_reviewer_gap_warns(self): + first = {dimension: 4 for dimension in DIMENSIONS} + second = {dimension: 4 for dimension in DIMENSIONS} + second[DIMENSIONS[2]] = 5 + result = stability_warning(first, second) + self.assertFalse(result["stable"]) + self.assertEqual(result["gaps"][DIMENSIONS[2]], 1.0) + + +if __name__ == "__main__": + unittest.main() diff --git a/meta/chains/README.md b/meta/chains/README.md index ad3a8d6..2296635 100644 --- a/meta/chains/README.md +++ b/meta/chains/README.md @@ -21,6 +21,7 @@ agent 的提示词按**变化轴**拆三段,不做"一个 agent 一个大 prom | expansion 扩写 | `expansion` | 写作→writer | generation | 同 continuation | 已建 | | polish 润色 | `polish` | 写作→writer | generation | 同 continuation(只动表达层) | 已建 | | planning 规划 | `planning` | 规划→planner | planning | read-context(planning 视图)→槽位→用户确认(confirm) | 已建 | +| fine_outline 细纲规划 | `fine-outline` | 规划→planner | planning | read-context(fine_outline 冻结视图)→槽位→detect→quality-gate | 评测版首建(2026-07-19) | | full_parse 拆书 | `parse-book` | 分析→extractor | extraction | import 分章→逐章槽位→db 落 draft→管理员确认(G3 门) | 已建;B2 PG 版首验 | | extraction 章后抽取 | `extract-knowledge` | 分析→extractor | extraction | 采纳后触发→槽位→草稿/冲突队列→confirm | 已建;C5 首验 | | validation / consistency_check 检测 | `detect` | 检测→detector | detection | read-context(detection 视图)→槽位→报告落评审/ | 已建 |