治理: Skill 测试治理第一阶段——harness 控制平面 + 实现测试迁出运行时目录

范围(不含 design-story-foundation、docs/、humanization/README.md 等进行中改动):

1. 新增 harness/ 控制平面
   - skill_harness.py 静态审计:32 个运行时 Skill 的 frontmatter/manifest/文档污染,当前 0 问题
   - run_selected.py 选择性执行器:manifest 与磁盘一一对账、依赖阻断、
     空跑与 skip-only 失败关闭、AST 测试形状门
   - manifests/skills.json:32 个 Skill 的合同责任方与协作领域登记
   - manifests/test-inventory.json:81 个测试资产登记
   - specs/skill-testing.md 与 README.md:测试分层、证据边界与 harness 职责

2. 实现测试从 .claude/skills/*/scripts/ 迁至 tests/skills/<skill>/
   - 71 个测试文件迁移并修复项目根与临时目录运行导入
   - 数据库触发器测试宽泛异常收窄为 psycopg.errors.RaiseException
   - 抽取离线大测试拆出真实 PG smoke(默认阻断,不计入离线通过)
   - 抽取 presence 去重边界拆出独立测试:493 + 78 = 571 项检查不变

3. 运行时文档清理
   - 13 个 SKILL.md 移除自测/离线验证段落、测试命令与测试文件事实源表述,
     只保留运行时合同;业务运行合同、额度、授权与离线模式均保留

4. SoT 同步
   - AGENTS.md:新增 Skill 领域索引(7 个合同责任方分组,覆盖 32 个运行时 Skill)
   - 领域 07:测试入口改由 harness/manifests/ 登记,SKILL.md 不承载测试命令
   - humanization 覆盖矩阵:活动测试路径同步迁移

验证证据: harness 自测 15 项 + runner 自测 13 项通过;静态审计 32 Skill / 0 问题;
73 个非数据库测试通过;8 个集成条目中 6 个 PostgreSQL 项被依赖门明确阻断;
py_compile 与 git diff --check 通过。未连接 PostgreSQL、网络、真实模型或额度。

已知边界: 真正 skill_behavior_eval 仍为 0,尚未验证任何 Skill 自然语言行为;
evaluate-frozen-replay 的 raw 存储边界冲突留待单独治理。
This commit is contained in:
zizi 2026-08-19 01:50:20 +08:00
parent 476bb71da4
commit c9f69d9d6d
97 changed files with 5247 additions and 821 deletions

View File

@ -31,7 +31,7 @@ Agent 与 Skill 领域拥有角色职责、可调用能力合同、确定性工
6. 输入产出落库:哪些输入和产出必须进库可见,见索引 §3。 6. 输入产出落库:哪些输入和产出必须进库可见,见索引 §3。
7. 稳定错误码和失败恢复。 7. 稳定错误码和失败恢复。
8. raw、授权、预算和审计边界。 8. raw、授权、预算和审计边界。
9. 离线测试与验收命令。 9. 开发验证与行为评测由 `harness/manifests/` 登记;`SKILL.md` 不承载测试命令、测试文件路径或测试结论。Skill 运行时需要执行的业务 dry-run/评测操作仍属于运行合同。
数据库是权威:Skill 对自己读写哪些表负责,声明失败时如何关闭,并确保经手的输入和产出都落库。没落库的输入产出,在系统视角里等于不存在。只读看板只查库渲染,不替 Skill 写任何数据。 数据库是权威:Skill 对自己读写哪些表负责,声明失败时如何关闭,并确保经手的输入和产出都落库。没落库的输入产出,在系统视角里等于不存在。只读看板只查库渲染,不替 Skill 写任何数据。
@ -49,7 +49,7 @@ Agent 与 Skill 领域拥有角色职责、可调用能力合同、确定性工
- 默认从仓库任意工作目录调用,必须自行解析项目根和输入绝对路径,不能依赖调用者先 `cd` 到特定目录。 - 默认从仓库任意工作目录调用,必须自行解析项目根和输入绝对路径,不能依赖调用者先 `cd` 到特定目录。
- 机械事实必须结构化输出稳定状态和错误码;人读日志是补充,不是唯一接口。 - 机械事实必须结构化输出稳定状态和错误码;人读日志是补充,不是唯一接口。
- Tool 不调用模型,除非所属 Skill 明确声明该步骤本质需要模型。 - Tool 不调用模型,除非所属 Skill 明确声明该步骤本质需要模型。
- Tool 变更必须有离线测试、`py_compile` 和 `git diff --check` 证据。 - Tool 变更必须有登记在 `harness/manifests/` 的相关实现测试、`py_compile` 和 `git diff --check` 证据;测试结果只证明机械合同,不升级为 Skill 行为或内容质量结论。
## 5. ReAct 工作方式 ## 5. ReAct 工作方式

View File

@ -24,9 +24,6 @@ description: 通过唯一受控入口查询或修改 muse-example PostgreSQL,
# 大对象(raw 全文)经 stdin 传 JSON 数组(避开 shell 转义 / ARG_MAX): # 大对象(raw 全文)经 stdin 传 JSON 数组(避开 shell 转义 / ARG_MAX):
.venv/bin/python .claude/skills/access-database/scripts/db.py execparams "INSERT ... VALUES (%s,%s)" --stdin < params.json # params.json = ["<sha>", "<完整全文>"] .venv/bin/python .claude/skills/access-database/scripts/db.py execparams "INSERT ... VALUES (%s,%s)" --stdin < params.json # params.json = ["<sha>", "<完整全文>"]
# execparams 参数装载逻辑离线自测(不连库)
.venv/bin/python .claude/skills/access-database/scripts/test_db_params.py
# SQL 文件应用:整文件一个事务,失败全回滚 # SQL 文件应用:整文件一个事务,失败全回滚
.venv/bin/python .claude/skills/access-database/scripts/db.py apply db/ddl/91-example实验私货.sql .venv/bin/python .claude/skills/access-database/scripts/db.py apply db/ddl/91-example实验私货.sql

View File

@ -7,7 +7,7 @@ description: 通过 New-API 的统一治理入口调用内容模型,执行额
创始人拍板(2026-07-13):清洗与拆书的内容生产 LLM **全部走 New-API 的 MiniMax-M3**;主会话(Fable5)只固化 agent/提示词/skill 与发起调用。本 skill 是唯一出口。 创始人拍板(2026-07-13):清洗与拆书的内容生产 LLM **全部走 New-API 的 MiniMax-M3**;主会话(Fable5)只固化 agent/提示词/skill 与发起调用。本 skill 是唯一出口。
管线内容生产调用的**标准入口是 `chat_governed`**(受 5 小时额度窗 + 全局降级链治理);`chat`/`chat` CLI 是不受治理的直连,仅供调试。治理政策的机械事实源是 `.claude/skills/call-content-model/scripts/llm.py` + `test_quota.py`(AGENTS.md §6 点名),模型链切换必须由该 skill 治理并留下日志。 管线内容生产调用的**标准入口是 `chat_governed`**(受 5 小时额度窗 + 全局降级链治理);`chat`/`chat` CLI 是不受治理的直连,仅供调试。治理政策的机械事实源是运行配置、共享额度账本 `example_llm_quota` 和模型运行适配器(`llm.py`),模型链切换必须由该 skill 治理并留下日志。
## 用法 ## 用法

View File

@ -121,13 +121,3 @@ disable-model-invocation: true
`dashboard/server.py` 的 `/ai-flavor` 默认查这三张表,数据库不可用时才明确标注离线回退;页面不会因打开而重新读取原文。 `dashboard/server.py` 的 `/ai-flavor` 默认查这三张表,数据库不可用时才明确标注离线回退;页面不会因打开而重新读取原文。
状态语义:案例卡 `shadow` 只供复核,`canonical` 仅表示获授权且完成评审,`rejected`/`archived` 不进入生成上下文;重验证 `verified` 才能确认、投影样例或消费规则,`stale`(全文哈希变化)、`unavailable`(来源不可得)和 `card_mismatch`(锚点变化)都使当前卡在这些动作上失效,但历史回执保留。 状态语义:案例卡 `shadow` 只供复核,`canonical` 仅表示获授权且完成评审,`rejected`/`archived` 不进入生成上下文;重验证 `verified` 才能确认、投影样例或消费规则,`stale`(全文哈希变化)、`unavailable`(来源不可得)和 `card_mismatch`(锚点变化)都使当前卡在这些动作上失效,但历史回执保留。
## 机械验收
```bash
.venv/bin/python .claude/skills/capture-ai-flavor-cases/scripts/test_capture_cases.py
.venv/bin/python .claude/skills/access-database/scripts/test_skill_catalog.py
git diff --check
```
测试必须覆盖:来源 hash、重验证 verified/stale/unavailable/card_mismatch、未授权 hash-only、重复 ID、live feedback 来源绑定、shadow 不能投影样例、缺重验证回执不能确认、跨作品/反例门,以及候选评测不足时不能 active。

View File

@ -31,10 +31,3 @@ disable-model-invocation: true
- 诊断不得改正文、不得产 patch。 - 诊断不得改正文、不得产 patch。
- `decision_proposal` 只是机械层初步建议;候选/语义命中必须经过功能仲裁,不能当执行指令。 - `decision_proposal` 只是机械层初步建议;候选/语义命中必须经过功能仲裁,不能当执行指令。
- 命中不等于修改命令:发现清单是给仲裁的输入,不是执行指令。 - 命中不等于修改命令:发现清单是给仲裁的输入,不是执行指令。
## 自测
```bash
cd agent-example
.venv/bin/python .claude/skills/diagnose-ai-flavor/scripts/test_diagnose_ai_flavor.py
```

View File

@ -32,16 +32,6 @@ description: 使用固定 Qwen3 嵌入模型将知识草稿或实体批量写入
- **落库字段**:`example_knowledge_embedding(draft_id, content_hash, embed_text, model, dimensions=1024, embedding)`;draft 确认落 entity 后由 confirm 流程把 owner 迁到 `entity_id` 并清空 `draft_id`,关系草稿确认后关闭无 canonical owner 的临时向量。 - **落库字段**:`example_knowledge_embedding(draft_id, content_hash, embed_text, model, dimensions=1024, embedding)`;draft 确认落 entity 后由 confirm 流程把 owner 迁到 `entity_id` 并清空 `draft_id`,关系草稿确认后关闭无 canonical owner 的临时向量。
- 汇报:新嵌 N、跳过 M、失败 K;幂等、失活和冲突原因均输出可追踪明细。 - 汇报:新嵌 N、跳过 M、失败 K;幂等、失活和冲突原因均输出可追踪明细。
## 离线验证
```bash
.venv/bin/python .claude/skills/embed-knowledge/scripts/test_embed_drafts_offline.py
.venv/bin/python -m py_compile .claude/skills/embed-knowledge/scripts/embed_drafts.py \
.claude/skills/embed-knowledge/scripts/test_embed_drafts_offline.py
```
离线测试只使用 fake connection 检查并发顺序、SQL 条件和 owner 反例,不连接真实数据库,不调用 embedding 或 reset。
## 红线 ## 红线
- 调用必须 `trust_env=False`(系统代理会假 502);令牌用 `MUSE_AI_NEW_API_TOKEN`(勿用管理令牌,打 /v1 报无效)。 - 调用必须 `trust_env=False`(系统代理会假 502);令牌用 `MUSE_AI_NEW_API_TOKEN`(勿用管理令牌,打 /v1 报无效)。

View File

@ -36,11 +36,3 @@ disable-model-invocation: true
- 只读已确认正文与作者样张;不得从草稿或未确认候选归纳基线。 - 只读已确认正文与作者样张;不得从草稿或未确认候选归纳基线。
- 基线必须人工确认(`--reviewer` 必填,数据库 CHECK 兜底)。 - 基线必须人工确认(`--reviewer` 必填,数据库 CHECK 兜底)。
- 本技能不产一个字正文。 - 本技能不产一个字正文。
## 自测
```bash
cd agent-example
.venv/bin/python .claude/skills/establish-voice-baseline/scripts/test_establish_voice_baseline.py
.venv/bin/python humanization/tests/test_humanization_v2.py
```

View File

@ -48,14 +48,14 @@ disable-model-invocation: true
## 编排入口 ## 编排入口
- `scripts/run_replay.py --mode dry_run`:只执行授权、来源、冻结和三臂 manifest 预检,不调用模型;这是首个机制 smoke 入口。 - `scripts/run_replay.py --mode dry_run`:只执行授权、来源、冻结和三臂 manifest 预检,不调用模型;用于确认评测执行的前置条件。
- `freeze-context/scripts/load_reference_work.py`:从 PostgreSQL 只读事务组装仓库外临时配置;读取 `example_reference_authorization_snapshot` 当前原文件版本的最新快照,组装 authorization 外层与 snapshot。缺授权记录仍生成可审计配置,但送入 `run_replay` 后必须保持 `blocked_authorization`。 - `freeze-context/scripts/load_reference_work.py`:从 PostgreSQL 只读事务组装仓库外临时配置;读取 `example_reference_authorization_snapshot` 当前原文件版本的最新快照,组装 authorization 外层与 snapshot。缺授权记录仍生成可审计配置,但送入 `run_replay` 后必须保持 `blocked_authorization`。
- `scripts/run_replay.py --mode execute`:在全部前置门通过后,依次执行三臂 planner、整组 schema、逐臂盲 detector、两个独立盲 judge、rubric 校验、稳定性门和去盲汇总;`--output-dir` 必须位于仓库外的临时目录。 - `scripts/run_replay.py --mode execute`:在全部前置门通过后,依次执行三臂 planner、整组 schema、逐臂盲 detector、两个独立盲 judge、rubric 校验、稳定性门和去盲汇总;`--output-dir` 必须位于仓库外的临时目录。
- detector 输入输出均为 JSON;输入只有匿名候选 ID、候选和公共冻结到 `as_of` 的规划上下文,不含 arm 名、`cardInjection`、`cardManifest`、任何臂特有卡内容、目标章 proxy 或其他评委结果。卡注入合法性只由确定性预检负责。报告不合约属于系统失败,须在 judge 前失败关闭;合同合法的 `failed` / `needs_evidence` 属于候选质量信号,三臂均须保留候选与报告并继续盲评和 Gate,Gate A 只由 C 臂高严重度残留与硬约束覆盖率裁决质量失败。 - detector 输入输出均为 JSON;输入只有匿名候选 ID、候选和公共冻结到 `as_of` 的规划上下文,不含 arm 名、`cardInjection`、`cardManifest`、任何臂特有卡内容、目标章 proxy 或其他评委结果。卡注入合法性只由确定性预检负责。报告不合约属于系统失败,须在 judge 前失败关闭;合同合法的 `failed` / `needs_evidence` 属于候选质量信号,三臂均须保留候选与报告并继续盲评和 Gate,Gate A 只由 C 臂高严重度残留与硬约束覆盖率裁决质量失败。
- 合法的非 `passed` 报告在安全 manifest 与 CAS 中只保存按臂分组的 `semantic-diagnostic-v1`:固定结果枚举、固定原因/区段枚举、调用次数和有界阻断计数。完整报告只留受控 raw 和编排器内存中的 Gate builder 输入;安全输出不得保存错误消息、正文、引文、事实文本、补证查询/原因、任何检测项 ID/字段路径、纠错草稿或模型原始字段。检测器不合约时同样只保存固定闭集诊断并失败关闭。 - 合法的非 `passed` 报告在安全 manifest 与 CAS 中只保存按臂分组的 `semantic-diagnostic-v1`:固定结果枚举、固定原因/区段枚举、调用次数和有界阻断计数。完整报告只留受控 raw 和编排器内存中的 Gate builder 输入;安全输出不得保存错误消息、正文、引文、事实文本、补证查询/原因、任何检测项 ID/字段路径、纠错草稿或模型原始字段。检测器不合约时同样只保存固定闭集诊断并失败关闭。
- 任一样本发生系统失败、检测器不合约、盲评无效/不稳定、预算或 raw/CAS 失败后,整轮已无法构建完整 Gate 输入,必须立即停止后续样本并进入统一迁移;不得继续调用模型消耗预算,也不得因失败删除本轮诊断 raw。 - 任一样本发生系统失败、检测器不合约、盲评无效/不稳定、预算或 raw/CAS 失败后,整轮已无法构建完整 Gate 输入,必须立即停止后续样本并进入统一迁移;不得继续调用模型消耗预算,也不得因失败删除本轮诊断 raw。
- 两个 judge 使用不同 `judgeId` 和独立无会话进程。第二个 judge 的匿名候选顺序必须与第一个完全相反;任何 rubric 不合约标记 `judge_invalid`,任一同维差值大于 `0.5` 标记 `judge_unstable`,两者都不得标记 `completed`。 - 两个 judge 使用不同 `judgeId` 和独立无会话进程。第二个 judge 的匿名候选顺序必须与第一个完全相反;任何 rubric 不合约标记 `judge_invalid`,任一同维差值大于 `0.5` 标记 `judge_unstable`,两者都不得标记 `completed`。
- 只有三臂 schema 合法、detector 均产出合同合法报告、双 judge rubric 与稳定性门通过,才去盲生成逐维 `B-A` / `C-A` 差值矩阵并标记 `completed`;detector 报告合同合法不等于其质量终态必须为 `passed`。`--detector-bin`、`--judge-primary-bin`、`--judge-secondary-bin` 可分别指定本地 runner;未指定时复用 `--planner-bin`,测试只能使用 fake binary。 - 只有三臂 schema 合法、detector 均产出合同合法报告、双 judge rubric 与稳定性门通过,才去盲生成逐维 `B-A` / `C-A` 差值矩阵并标记 `completed`;detector 报告合同合法不等于其质量终态必须为 `passed`。`--detector-bin`、`--judge-primary-bin`、`--judge-secondary-bin` 可分别指定本地 runner;未指定时复用 `--planner-bin`;离线评测可使用 fake adapter,真实评测执行使用配置绑定的 runner。
- planner、detector、judge 子进程统一受 `--timeout-seconds` 限制,默认 300 秒;任一超时分别落盘 `planner_timeout`、`detector_timeout`、`judge_timeout`,不得继续进入后续阶段或标记 `completed`。 - planner、detector、judge 子进程统一受 `--timeout-seconds` 限制,默认 300 秒;任一超时分别落盘 `planner_timeout`、`detector_timeout`、`judge_timeout`,不得继续进入后续阶段或标记 `completed`。
- `scripts/write_report.py`:从 `run_result.json` 生成独立严格 schema 的安全摘要,只接受受限标识符、枚举、数字、短安全摘要和 SHA-256;不会读取候选正文,也不会把候选路径以外的原始响应写入报告。 - `scripts/write_report.py`:从 `run_result.json` 生成独立严格 schema 的安全摘要,只接受受限标识符、枚举、数字、短安全摘要和 SHA-256;不会读取候选正文,也不会把候选路径以外的原始响应写入报告。
@ -67,7 +67,7 @@ disable-model-invocation: true
真实五章装配还必须在 `commonControls` 预注册选择器版本、选择器规范 JSON 的原始字节 SHA-256、`inputProvenance=oracle_reference_scaffold` 和统一 `maxContextChars=140000`。五个样本的 A/C 补充原文预算统一预注册为 2000 Unicode code point;loader 在数据库读取前机械核对这些值,不得根据实际文本降低或抬高预算。连续四章基线装不下、A/C 任一臂补充原文不足 2000,或两臂 WriterCreativeInput hash 相同/允许差异为空时均失败关闭。 真实五章装配还必须在 `commonControls` 预注册选择器版本、选择器规范 JSON 的原始字节 SHA-256、`inputProvenance=oracle_reference_scaffold` 和统一 `maxContextChars=140000`。五个样本的 A/C 补充原文预算统一预注册为 2000 Unicode code point;loader 在数据库读取前机械核对这些值,不得根据实际文本降低或抬高预算。连续四章基线装不下、A/C 任一臂补充原文不足 2000,或两臂 WriterCreativeInput hash 相同/允许差异为空时均失败关闭。
每个样本必须提供 `writerContextInput`。仓内基础配置只允许 `contentMode=sanitized_contract_fixture`,并为 A/C 各提供恰好 2000 code point 的明确脱敏合成补充原文,用于机械证明 WriterContext 合同、章号冻结和创作输入差异;这些文本不得来自原书。loader 装配真实临时配置后必须改为 `canonical_frozen_prose`。生产 `--execute` 在创建 vault 或 runner 前拒绝脱敏夹具;脱敏夹具只可由离线测试适配器执行。任何模式都不得把原书全文、完整目标细纲或标准答案写入 Git。 每个样本必须提供 `writerContextInput`。仓内基础配置只允许 `contentMode=sanitized_contract_fixture`,并为 A/C 各提供恰好 2000 code point 的明确脱敏合成补充原文,用于机械证明 WriterContext 合同、章号冻结和创作输入差异;这些文本不得来自原书。loader 装配真实临时配置后必须改为 `canonical_frozen_prose`。生产 `--execute` 在创建 vault 或 runner 前拒绝脱敏夹具;脱敏夹具只可由离线评测适配器执行。任何模式都不得把原书全文、完整目标细纲或标准答案写入 Git。
三臂都必须构造并校验完整 `WriterContext v1`,且固定 `mode=diagnostic_only`、`purpose=evaluation`、`acceptanceEligible=false`。adapter 随后投影 `WriterCreativeInput v2`;writer 模型只看到创作投影,不看到运行身份、manifest、hash、实验臂或验收状态: 三臂都必须构造并校验完整 `WriterContext v1`,且固定 `mode=diagnostic_only`、`purpose=evaluation`、`acceptanceEligible=false`。adapter 随后投影 `WriterCreativeInput v2`;writer 模型只看到创作投影,不看到运行身份、manifest、hash、实验臂或验收状态:
@ -99,8 +99,6 @@ dry-run 只输出计划、manifest 和上下文摘要,不调用 writer、seman
正式 execute 还必须显式传 `--raw-archive-dir <仓外绝对目录>`;缺失、相对路径、位于本轮输出目录内或与临时 vault 跨文件系统时,均在建立 vault 和调用模型前失败关闭。运行结束以 `rawDisposition.status=migrated` 和 `raw-vault-migration-receipt-v1` 证明产物已迁移,不再以删除后的 `closed` 作为完成条件。 正式 execute 还必须显式传 `--raw-archive-dir <仓外绝对目录>`;缺失、相对路径、位于本轮输出目录内或与临时 vault 跨文件系统时,均在建立 vault 和调用模型前失败关闭。运行结束以 `rawDisposition.status=migrated` 和 `raw-vault-migration-receipt-v1` 证明产物已迁移,不再以删除后的 `closed` 作为完成条件。
离线完整链冒烟由 `scripts/test_run_writer_replay.py` 的单样本三臂 production fake 用例承担:它真实经过 writer 投影、机械门、semantic detector、双评委、raw vault、CAS 和 GateInputBuilder,但所有模型均为本地确定性 fake,不产生费用、不得作为 Gate 样本结果。真实一次调用能力只由 `refresh_runtime_probe.py` 的合成 writer 探针验证;它不代表 detector/judge 或五样本 Gate 已通过。
预算合同分为不可变的 `plannedCalls` 和独立安全上限 `maxCalls`。Gate A 当前预注册 writer 60 次(15 次基础写作 + 最多 45 次篇幅修订)、semantic detector 24 次(15 次基础检测 + 9 次纠错/API 重试余量)、blind judge 45 次(最多三位评委,每位基础调用后最多两个格式纠错/API 重试槽位);各 `maxCalls=150`、单次 cap `$5`、总预算 `$2250`。启动预留只按 `plannedCalls * maxBudgetUsdPerCall`,即 `$645`;`$2250` 是三角色各 150 次安全容量对应的有限执行上限,不是预计消费。本预算授权不等于正式 `--execute` 授权。每次调用前账本同时检查角色计划槽位、`maxCalls`、累计实际成本和未执行计划的最坏预留;调用后只能使用可信 `ExecutionReceipt.totalCostUsd` 结算。缺回执成本、重复/错序结算、单次 cap 或总预算越界均 fail closed,未触发的修订、纠错与第三评计划必须保留在 `remainingPlannedCalls`。 预算合同分为不可变的 `plannedCalls` 和独立安全上限 `maxCalls`。Gate A 当前预注册 writer 60 次(15 次基础写作 + 最多 45 次篇幅修订)、semantic detector 24 次(15 次基础检测 + 9 次纠错/API 重试余量)、blind judge 45 次(最多三位评委,每位基础调用后最多两个格式纠错/API 重试槽位);各 `maxCalls=150`、单次 cap `$5`、总预算 `$2250`。启动预留只按 `plannedCalls * maxBudgetUsdPerCall`,即 `$645`;`$2250` 是三角色各 150 次安全容量对应的有限执行上限,不是预计消费。本预算授权不等于正式 `--execute` 授权。每次调用前账本同时检查角色计划槽位、`maxCalls`、累计实际成本和未执行计划的最坏预留;调用后只能使用可信 `ExecutionReceipt.totalCostUsd` 结算。缺回执成本、重复/错序结算、单次 cap 或总预算越界均 fail closed,未触发的修订、纠错与第三评计划必须保留在 `remainingPlannedCalls`。
## 稳定运行纪律 ## 稳定运行纪律
@ -124,12 +122,12 @@ Gate 裁决器 `writer_gate.py`(归属 `score-content-quality`,本 Skill 只
`scripts/refresh_runtime_probe.py` 是刷新 `executionAuthorization.runtimeProbe` 的唯一通道。正式执行门要求探针的 `executionProfileSha256` 必须等于当前 writer profile 的身份哈希;writer 合同(schema/prompt/模型/CLI/预算)升级后旧探针必然失配,门以 `EXECUTE_PROBE_CONTRACT_MISMATCH` 失败关闭,此时必须用当前合同重测探针,不得手动改探针值绕过。 `scripts/refresh_runtime_probe.py` 是刷新 `executionAuthorization.runtimeProbe` 的唯一通道。正式执行门要求探针的 `executionProfileSha256` 必须等于当前 writer profile 的身份哈希;writer 合同(schema/prompt/模型/CLI/预算)升级后旧探针必然失配,门以 `EXECUTE_PROBE_CONTRACT_MISMATCH` 失败关闭,此时必须用当前合同重测探针,不得手动改探针值绕过。
- 输入合同:只读 `--config` 指定的 base 配置,只用 `executionProfiles.writer`(探针只验证 writer 合同),不读 semantic/judge,不读 `oracleTruthPacks`/`samples`/raw 物料。探针输入是固定、极小、全合成的能力探针 prompt(一段自包含的合成微任务),绝不使用原书全文、真实正文或 raw。 - 输入合同:只读 `--config` 指定的 base 配置,只用 `executionProfiles.writer`(探针只验证 writer 合同),不读 semantic/judge,不读 `oracleTruthPacks`/`samples`/raw 物料。探针输入是固定、极小、全合成的能力探针 prompt(一段自包含的合成微任务),绝不使用原书全文、真实正文或 raw。
- 调用合同:Claude 调用经可注入 invoker 发起,默认 invoker = `claude_runtime.run_claude`(真实调用层是薄薄一层);`--dry-run` 注入固定假产出,只走通「构建 profile + 重签 + 写文件」链路供冒烟,不发起真实调用。 - 调用合同:Claude 调用经可注入 invoker 发起,默认 invoker = `claude_runtime.run_claude`(真实调用层是薄薄一层);`--dry-run` 注入固定假产出,只走通「构建 profile + 重签 + 写文件」链路供离线预览,不发起真实调用。
- 输出合同:把绑定新合同的探针替换 `executionAuthorization.runtimeProbe`,并按与门校验逐字节同源的规范自哈希算法重签该探针的 `receiptSha256`(`budget`/`rawRetention` 原样保留、各自自哈希不变),把刷新后的完整配置写到 `--output` 指定的【新文件】,不就地覆盖原配置,便于 review diff 后再替换。新探针字段集与现有 runtimeProbe 完全一致,不缺不多。 - 输出合同:把绑定新合同的探针替换 `executionAuthorization.runtimeProbe`,并按与门校验逐字节同源的规范自哈希算法重签该探针的 `receiptSha256`(`budget`/`rawRetention` 原样保留、各自自哈希不变),把刷新后的完整配置写到 `--output` 指定的【新文件】,不就地覆盖原配置,便于 review diff 后再替换。新探针字段集与现有 runtimeProbe 完全一致,不缺不多。
- 失败关闭:调用失败 / 结构化输出 schema 不过 / 单次成本超冻结 cap / 超 deadline / 回执不可信(模型不匹配、标记错误、非零退出、异常终止、API 错误、身份哈希指向旧合同、输出哈希与产出不一致)→ 不写 successful 探针、不产出刷新配置、退出非零,原配置保持不动。 - 失败关闭:调用失败 / 结构化输出 schema 不过 / 单次成本超冻结 cap / 超 deadline / 回执不可信(模型不匹配、标记错误、非零退出、异常终止、API 错误、身份哈希指向旧合同、输出哈希与产出不一致)→ 不写 successful 探针、不产出刷新配置、退出非零,原配置保持不动。
- 审计合同:标准输出只记录做了什么(探针身份哈希、结构化输出哈希、回执哈希、成本、是否成功、输出文件),绝不打印完整 prompt/response、raw 路径或供应商原始响应。 - 审计合同:标准输出只记录做了什么(探针身份哈希、结构化输出哈希、回执哈希、成本、是否成功、输出文件),绝不打印完整 prompt/response、raw 路径或供应商原始响应。
离线用法(冒烟,不调模型): 离线用法(不调用模型):
```bash ```bash
.venv/bin/python .claude/skills/evaluate-frozen-replay/scripts/refresh_runtime_probe.py \ .venv/bin/python .claude/skills/evaluate-frozen-replay/scripts/refresh_runtime_probe.py \

View File

@ -25,9 +25,3 @@ disable-model-invocation: true
- 不从全局 Claude 会话继承业务状态,不静默换模型或供应商。 - 不从全局 Claude 会话继承业务状态,不静默换模型或供应商。
- 不自行决定候选是否通过,不写 Candidate、Quality 或 Canonical 数据。 - 不自行决定候选是否通过,不写 Candidate、Quality 或 Canonical 数据。
- 不拥有 CAS、raw vault、运行登记或回执修复实现。 - 不拥有 CAS、raw vault、运行登记或回执修复实现。
## 离线验证
```bash
.venv/bin/python .claude/skills/execute-claude-task/scripts/test_claude_runtime.py
```

View File

@ -43,10 +43,3 @@ disable-model-invocation: true
- 写入 `muse_knowledge_draft`、`example_upgrade_window`、`example_upgrade_alias`、`example_upgrade_presence`、`example_upgrade_card_state`、`example_upgrade_audit` 和 `example_knowledge_embedding`。 - 写入 `muse_knowledge_draft`、`example_upgrade_window`、`example_upgrade_alias`、`example_upgrade_presence`、`example_upgrade_card_state`、`example_upgrade_audit` 和 `example_knowledge_embedding`。
- `SOURCE_TYPE="upgrade_book"`、`upgrade-reset` 等是数据库历史兼容值,不随 Skill 改名。 - `SOURCE_TYPE="upgrade_book"`、`upgrade-reset` 等是数据库历史兼容值,不随 Skill 改名。
- 不提供修复、重置、备份、恢复或存量迁移入口;这些动作统一由 `maintain-work-extraction` 承担。 - 不提供修复、重置、备份、恢复或存量迁移入口;这些动作统一由 `maintain-work-extraction` 承担。
## 离线验证
```bash
.venv/bin/python .claude/skills/extract-work-knowledge/scripts/test_upgrade_work_lock_offline.py
.venv/bin/python .claude/skills/extract-work-knowledge/scripts/test_parse_upgrade_offline.py
```

View File

@ -69,10 +69,3 @@ reset 软删目标 `upgrade_book` 草稿和无 Canonical owner 的活向量,
- 备份与 preview 默认只读;execute 只修改目标租户、目标作品、`upgrade_book` 边界内的派生状态。 - 备份与 preview 默认只读;execute 只修改目标租户、目标作品、`upgrade_book` 边界内的派生状态。
- 不生成新抽取内容,不提供正常 `run/windows/status`。 - 不生成新抽取内容,不提供正常 `run/windows/status`。
- `migrate_upgrade_windows.py` 当前只输出迁移计划,不落库;真实写入仍需另行 gate 与授权。 - `migrate_upgrade_windows.py` 当前只输出迁移计划,不落库;真实写入仍需另行 gate 与授权。
## 离线验证
```bash
.venv/bin/python .claude/skills/maintain-work-extraction/scripts/test_backup_upgrade_work_offline.py
.venv/bin/python .claude/skills/maintain-work-extraction/scripts/test_reset_upgrade_work_offline.py
```

View File

@ -31,11 +31,3 @@ disable-model-invocation: true
- 读:`example_voice_baseline` 当前 canonical 版本;规则/样例读 `humanization/` Git 资产。 - 读:`example_voice_baseline` 当前 canonical 版本;规则/样例读 `humanization/` Git 资产。
- 写:`example_run`(幂等 upsert,记录规则指纹、声音账 hash 和投影条数)。 - 写:`example_run`(幂等 upsert,记录规则指纹、声音账 hash 和投影条数)。
- 失败:作品不匹配、声音账非 canonical、规则装载门失败或 guidance 超合同直接失败关闭。 - 失败:作品不匹配、声音账非 canonical、规则装载门失败或 guidance 超合同直接失败关闭。
## 自测
```bash
cd agent-example
.venv/bin/python .claude/skills/prevent-ai-flavor/scripts/test_prevent_ai_flavor.py
.venv/bin/python .claude/skills/assemble-context/scripts/test_assemble_writer_context.py
```

View File

@ -32,11 +32,3 @@ disable-model-invocation: true
- 读取运行、调用、raw 和回执当前状态,只为校验绑定与幂等。 - 读取运行、调用、raw 和回执当前状态,只为校验绑定与幂等。
- 写入 `example_run`、`example_llm_call`、`example_raw_lease`、`example_raw_content`、`example_run_receipt` 和相应质量失败记录。 - 写入 `example_run`、`example_llm_call`、`example_raw_lease`、`example_raw_content`、`example_run_receipt` 和相应质量失败记录。
- 不调用 Claude/New-API,不生成候选,不判断内容是否通过。 - 不调用 Claude/New-API,不生成候选,不判断内容是否通过。
## 离线验证
```bash
.venv/bin/python .claude/skills/record-run-evidence/scripts/test_file_cas.py
.venv/bin/python .claude/skills/record-run-evidence/scripts/test_raw_vault.py
.venv/bin/python .claude/skills/record-run-evidence/scripts/test_run_registry.py
```

View File

@ -30,10 +30,3 @@ disable-model-invocation: true
- 没有诊断产物、作者授权或功能仲裁不得修订。 - 没有诊断产物、作者授权或功能仲裁不得修订。
- 硬门失败不得以风格分、盲评结果抵消;旧诊断不能跨规则库/正文 hash 复用。 - 硬门失败不得以风格分、盲评结果抵消;旧诊断不能跨规则库/正文 hash 复用。
- 语义级不变量(因果、POV、伏笔状态)机械未覆盖的,如实列 `unresolved_risks`,不假装验证完成。 - 语义级不变量(因果、POV、伏笔状态)机械未覆盖的,如实列 `unresolved_risks`,不假装验证完成。
## 自测
```bash
cd agent-example
.venv/bin/python .claude/skills/revise-ai-flavor/scripts/test_revise_ai_flavor.py
```

View File

@ -23,7 +23,8 @@ SoT 按主题分域,不做跨主题的全局排序。可执行脚本与书面
| [`.agent/docs/architecture/domains/`](.agent/docs/architecture/domains/_index.md) | 本仓领域边界、数据权威、落库合同和领域协作的 SoT。 | | [`.agent/docs/architecture/domains/`](.agent/docs/architecture/domains/_index.md) | 本仓领域边界、数据权威、落库合同和领域协作的 SoT。 |
| [`meta/schemas/`](meta/schemas/) | 23 型结构本体的字段合同;库内 payload 结构以该合同为准。 | | [`meta/schemas/`](meta/schemas/) | 23 型结构本体的字段合同;库内 payload 结构以该合同为准。 |
| [`meta/chains/`](meta/chains/) | scenario、purpose、功能 skill、角色槽位和保护节点的链路登记。 | | [`meta/chains/`](meta/chains/) | scenario、purpose、功能 skill、角色槽位和保护节点的链路登记。 |
| [`.claude/skills/`](.claude/skills/) | 每个 `SKILL.md` 定义能力的输入、输出、红线、数据库读写合同、输入产出落库和自测入口;同目录 `scripts/` 是确定性实现与机械门。 | | [`.claude/skills/`](.claude/skills/) | 每个 `SKILL.md` 定义项目运行时能力的输入、输出、红线、数据库读写合同和输入产出落库;同目录 `scripts/` 是运行时确定性实现与机械门,开发验证和行为评测入口由 `harness/manifests/` 登记。 |
| [`tests/skills/`](tests/skills/) | 项目运行时 Skill 的实现测试、集成测试和 fake pipeline 测试;它们提供回归证据,不拥有运行时合同,也不等同于 Skill 行为评测。 |
| [`.agent/`](.agent/_index.md) | 跨任务长期知识;领域设计位于 `docs/architecture/domains/`,新增、删除或重命名必须同步各级 `_index.md`。 | | [`.agent/`](.agent/_index.md) | 跨任务长期知识;领域设计位于 `docs/architecture/domains/`,新增、删除或重命名必须同步各级 `_index.md`。 |
| [`docs/`](docs/) | 单次任务探索、计划、评测资料、样张和历史执行证据;任务完成后把稳定结论蒸馏到 `.agent/`,不得长期拥有领域定义。 | | [`docs/`](docs/) | 单次任务探索、计划、评测资料、样张和历史执行证据;任务完成后把稳定结论蒸馏到 `.agent/`,不得长期拥有领域定义。 |
| [`README.md`](README.md) | 项目背景和历史路线概览。目录、阶段、技能数量和存储方式等描述可能陈旧,不得覆盖本文件、`meta/`、skill 或磁盘事实。 | | [`README.md`](README.md) | 项目背景和历史路线概览。目录、阶段、技能数量和存储方式等描述可能陈旧,不得覆盖本文件、`meta/`、skill 或磁盘事实。 |
@ -46,7 +47,9 @@ agent-example/
│ ├── ddl/ # 可审计 DDL / 迁移文件 │ ├── ddl/ # 可审计 DDL / 迁移文件
│ ├── 表映射.md │ ├── 表映射.md
│ └── 连接信息.md │ └── 连接信息.md
├── humanization/ # 去 AI 味与人感资产层(合同/规则/样例/执行骨架;验证期自治,升华进 Muse 时同步回父仓;20 项研究覆盖矩阵见 humanization/research/) ├── humanization/ # 去 AI 味 Skill 能力域(规则/样例/执行骨架;运行时资产目标为数据库)
├── harness/ # 项目验证与外部评测索引、清单和调度支架
├── tests/ # 运行时 Skill 实现测试(按 Skill 归档)
├── knowledge/ # 仓内参考资产;未经绑定、授权不得进入上下文 ├── knowledge/ # 仓内参考资产;未经绑定、授权不得进入上下文
├── docs/ # 设计、评测、样张与历史执行记录 ├── docs/ # 设计、评测、样张与历史执行记录
├── .venv/ # 本地 Python 运行环境 ├── .venv/ # 本地 Python 运行环境
@ -57,15 +60,23 @@ agent-example/
5 个角色:`writer`、`planner`、`extractor`、`detector`、`judge`。角色身份在 `.claude/agents/*.md`,具体功能合同不复制进角色文件。 5 个角色:`writer`、`planner`、`extractor`、`detector`、`judge`。角色身份在 `.claude/agents/*.md`,具体功能合同不复制进角色文件。
技能按能力分组如下,实际清单以 `.claude/skills/*/SKILL.md` 为准,不手填数量镜像: ### Skill 领域索引(项目运行时 Skill)
- 数据与检索:`access-database`、`import-book`、`embed-knowledge`、`search-knowledge`、`call-content-model`。 实际清单以 `.claude/skills/*/SKILL.md` 为准;本表维护合同责任方、协作领域和领域 SoT,不复制各 Skill 的完整合同。每个运行时 Skill 必须登记一个合同责任方(业务领域或平台领域),但可以同时消费或影响多个协作领域;跨域调用、场景关系和保护节点在 `meta/chains/` 登记。合同责任方表示谁维护该 Skill 的稳定能力合同,不表示 Skill 只能属于一个业务领域。
- 执行、证据与上下文:`execute-claude-task`、`record-run-evidence`、`freeze-context`、`assemble-context`。
- 流程与主权治理:`decide-candidate`。 | 合同责任方 / 能力域 | 领域 SoT | 运行时 Skill |
- 清洗、拆解与知识:`clean-book-text`、`deconstruct-book`、`extract-chapter-knowledge`、`review-knowledge-cards`、`extract-work-knowledge`、`maintain-work-extraction`。 |---|---|---|
- 去 AI 味与人感(父仓专题-09 五技能先行验证):`establish-voice-baseline`(定基线)、`prevent-ai-flavor`(前置预防)、`diagnose-ai-flavor`(诊断)、`revise-ai-flavor`(修订)、`capture-ai-flavor-cases`(挖掘/案例采集);共享资产层在 `humanization/`,接力铁律见 `meta/chains/`。 | 平台运行与证据 | 07-Agent 与 Skill、08-数据权威与可视化 | `access-database`、`call-content-model`、`execute-claude-task`、`record-run-evidence` |
- 规划与写作:`design-story-foundation`、`plan-story`、`plan-chapter`、`write-next-chapter`、`rewrite-selection`、`expand-scene`、`polish-prose`。 | 上下文与知识检索 | 02-实体、04-上下文、08-数据权威与可视化 | `assemble-context`、`embed-knowledge`、`freeze-context`、`search-knowledge` |
- 检测与质量评测:`check-content-consistency`、`score-content-quality`、`optimize-content-quality`、`evaluate-frozen-replay`。 | 导入、清洗与抽取 | 02-实体、05-创作流程 | `clean-book-text`、`deconstruct-book`、`extract-chapter-knowledge`、`extract-work-knowledge`、`import-book`、`maintain-work-extraction`、`review-knowledge-cards` |
| 规划与作品基础 | 05-创作流程 | `design-story-foundation`、`plan-chapter`、`plan-story` |
| 写作与候选主权 | 01-作品、05-创作流程 | `decide-candidate`、`expand-scene`、`polish-prose`、`rewrite-selection`、`write-next-chapter` |
| 质量与回放评测 | 06-质量与复利、05-创作流程 | `check-content-consistency`、`evaluate-frozen-replay`、`optimize-content-quality`、`score-content-quality` |
| 去 AI 味与人感 | 06-质量与复利、父仓专题-09 | `capture-ai-flavor-cases`、`diagnose-ai-flavor`、`establish-voice-baseline`、`prevent-ai-flavor`、`revise-ai-flavor` |
`humanization/` 是“去 AI 味与人感”运行时 Skill 家族的能力域:`src/deai/` 是共享实现,规则、样例、案例和声音资产是该域的数据依赖,运行时以数据库为权威;仓内 YAML/JSON 在数据库化完成前只作为迁移种子、离线夹具或结构合同。规则记录不各自注册为 Skill,Skill 负责动作和消费边界。
Skill 领域列表的新增、删除、改名或主领域调整,必须同时检查 `.claude/skills/`、`meta/chains/README.md` 和相关领域 `_index.md`;不得只改本表造成索引漂移。
上下文目标合同以 [上下文领域 SoT](.agent/docs/architecture/domains/04-上下文领域.md) 为准:数据库读取器是核心实现,库内检索加速(向量)只做候选召回;任何命中都要回读库行并校验 hash。Skill 对自己读写哪些表负责,并把经手的输入和产出落库;没落库的输入产出在系统视角里等于不存在。 上下文目标合同以 [上下文领域 SoT](.agent/docs/architecture/domains/04-上下文领域.md) 为准:数据库读取器是核心实现,库内检索加速(向量)只做候选召回;任何命中都要回读库行并校验 hash。Skill 对自己读写哪些表负责,并把经手的输入和产出落库;没落库的输入产出在系统视角里等于不存在。
@ -95,7 +106,7 @@ agent-example/
## 6. 模型边界 ## 6. 模型边界
- 清洗、抽卡、范式拆取及其模型调用统一走 `call-content-model` Skill,不裸调 New-API。治理政策固定为 5 小时额度窗:MiniMax 模型累计花费上限 `$24`,全模型成功调用上限 `6000`;机械事实源是 `.claude/skills/call-content-model/scripts/llm.py` 及 `test_quota.py`,模型链切换必须由该 Skill 治理并留下日志。 - 清洗、抽卡、范式拆取及其模型调用统一走 `call-content-model` Skill,不裸调 New-API。治理政策固定为 5 小时额度窗:MiniMax 模型累计花费上限 `$24`,全模型成功调用上限 `6000`;运行适配器、正式配置和账本是额度合同的事实源,`.claude/skills/call-content-model/scripts/llm.py` 是实现,`test_quota.py` 只提供回归证据;模型链切换必须由该 Skill 治理并留下日志。
- 角色模型归属:`planner`/`writer`/`judge` 固定 `opus`;`extractor`/`detector` 可用其它模型(非必须降级)。拆书/导入侧抽取经 `call-content-model`/`deconstruct-book` Skill 走 MiniMax-M3,不走角色 model 派发;创作期章后抽取作为角色派发,可用 `opus`。 - 角色模型归属:`planner`/`writer`/`judge` 固定 `opus`;`extractor`/`detector` 可用其它模型(非必须降级)。拆书/导入侧抽取经 `call-content-model`/`deconstruct-book` Skill 走 MiniMax-M3,不走角色 model 派发;创作期章后抽取作为角色派发,可用 `opus`。
- 确定性脚本、合同校验、快照冻结、泄漏审计和报告生成不调用模型;除非对应 `SKILL.md` 明确声明模型步骤,不得把机械任务升级为模型任务。 - 确定性脚本、合同校验、快照冻结、泄漏审计和报告生成不调用模型;除非对应 `SKILL.md` 明确声明模型步骤,不得把机械任务升级为模型任务。
- Claude 生成或评测只在对应任务 SoT、显式预算、固定执行配置和原文用途授权全部满足后运行;任一前置门失败都应关闭执行。 - Claude 生成或评测只在对应任务 SoT、显式预算、固定执行配置和原文用途授权全部满足后运行;任一前置门失败都应关闭执行。
@ -130,12 +141,12 @@ agent-example/
- 明确区分**已验证事实**、**推断**和**假设**。报告必须给出证据来源;没有机械输出或运行证据时,不声称完成、修复或通过。 - 明确区分**已验证事实**、**推断**和**假设**。报告必须给出证据来源;没有机械输出或运行证据时,不声称完成、修复或通过。
- 禁止从 `n=1` 样本推出普适结论。至少分析假阴、假阳、样本偏差和混淆因素;需要判断卡或模型效果时使用同任务、同模型、同预算、同公共上下文的对照,并把不稳定样本排除在方向结论之外。 - 禁止从 `n=1` 样本推出普适结论。至少分析假阴、假阳、样本偏差和混淆因素;需要判断卡或模型效果时使用同任务、同模型、同预算、同公共上下文的对照,并把不稳定样本排除在方向结论之外。
- Python 一律使用仓内解释器 `.venv/bin/python`。先读目标 skill 的 `SKILL.md`,再运行其已存在的相关自测;例如: - Python 一律使用仓内解释器 `.venv/bin/python`。先读目标 Skill 的 `SKILL.md`,再按 `harness/manifests/` 登记的类别和依赖选择相关验证;下面命令仅是现有局部验证入口示例:
```bash ```bash
.venv/bin/python .claude/skills/plan-chapter/scripts/test_contract.py .venv/bin/python tests/skills/plan-chapter/test_contract.py
.venv/bin/python .claude/skills/call-content-model/scripts/test_quota.py .venv/bin/python tests/skills/call-content-model/test_quota.py
.venv/bin/python .claude/skills/assemble-context/scripts/test_writer_contract.py .venv/bin/python tests/skills/assemble-context/test_writer_contract.py
git diff --check git diff --check
``` ```
@ -163,7 +174,7 @@ git diff --check
1. **给谁用**:主会话编排?某角色 agent?还是别的 skill?消费者必须明确、单一。 1. **给谁用**:主会话编排?某角色 agent?还是别的 skill?消费者必须明确、单一。
2. **目的单一**:一个 skill 只实现一个能力。功能并列、又当编排又当执行的,拆。 2. **目的单一**:一个 skill 只实现一个能力。功能并列、又当编排又当执行的,拆。
3. **可靠性与稳定性**:失败是否明确失败关闭(不静默降级、不返回假成功)?错误是否带稳定码、不泄漏原文/密钥?有无离线自测覆盖关键路径与失败路径?确定性步骤是否真不调模型? 3. **可靠性与稳定性**:失败是否明确失败关闭(不静默降级、不返回假成功)?错误是否带稳定码、不泄漏原文/密钥?有无离线自测覆盖关键路径与失败路径?确定性步骤是否真不调模型?
4. **符合 skill 规范**:`SKILL.md` 是否清楚定义输入、输出、红线、用法?`scripts/` 是否其确定性实现与自测入口?是否走 `.venv`、不裸调 PG/New-API/模型? 4. **符合 skill 规范**:`SKILL.md` 是否清楚定义输入、输出、红线、用法?`scripts/` 是否只包含运行时确定性实现与机械门?相关实现测试和行为评测是否已在 `harness/manifests/` 登记?是否走 `.venv`、不裸调 PG/New-API/模型?
5. **边界一致**:引用的路径、合同、字段是否与现行 SoT(本文件、`.agent/`、`meta/` 和库内正式内容)一致?数据库是正式内容权威,Skill 必须声明读写哪些表、失败如何关闭、输入产出如何落库;引用已失效合同的不得执行。 5. **边界一致**:引用的路径、合同、字段是否与现行 SoT(本文件、`.agent/`、`meta/` 和库内正式内容)一致?数据库是正式内容权威,Skill 必须声明读写哪些表、失败如何关闭、输入产出如何落库;引用已失效合同的不得执行。
命名采用小写 `动作-对象`:目录名必须等于 frontmatter `name`,名称表达可调用能力,不复用 `scenario`、内部模块名或含混阶段词。`scenario`、`source_type`、updater 和备份逻辑键是独立稳定标识,不随 Skill 改名。 命名采用小写 `动作-对象`:目录名必须等于 frontmatter `name`,名称表达可调用能力,不复用 `scenario`、内部模块名或含混阶段词。`scenario`、`source_type`、updater 和备份逻辑键是独立稳定标识,不随 Skill 改名。

68
harness/README.md Normal file
View File

@ -0,0 +1,68 @@
# agent-example Harness
> 项目开发、验证和行为评测支架的入口与索引。
> 这里的文档和工具服务开发者与评测编排,不会作为任何 Skill 的运行时提示词自动注入 Agent。
## 1. 作用
`agent-example` 同时包含三种不同性质的验证工作:
- 运行时文档的静态卫生检查;
- 确定性工具、数据库边界和 fake pipeline 的机械测试;
- 将 Skill 交给 Agent 后的外部行为评测。
Harness 负责把这三类工作分开编排并保留证据,不能用其中一类的通过结果冒充另一类结论。
## 2. 目录索引
```text
harness/
├── README.md # 本入口:项目 harness 介绍、边界和索引
├── specs/
│ └── skill-testing.md # Skill 测试与评测规范(唯一规范正文)
├── skill_harness.py # 运行时 Skill 文档静态审计入口
├── run_selected.py # 按 manifest 选择性执行测试
├── test_skill_harness.py # 静态审计器离线测试
├── test_run_selected.py # 选择性执行器离线测试
├── manifests/
│ ├── skills.json # 运行时 Skill 责任方与协作领域清单
│ └── test-inventory.json # 测试分类、依赖、副作用与证据等级
└── evals/<skill>/ # Skill 行为评测夹具与适配器(尚未建立)
项目级实现测试位于 `../tests/skills/<skill>/`,不进入运行时 Skill 目录。
```
现有专项支架:
| 路径 | 责任 | 证据边界 |
|---|---|---|
| `humanization/eval/run_eval.py` | 去 AI 味规则合同和回归陷阱试跑 | 规则合同回放,不证明文学效果 |
| `.claude/skills/evaluate-frozen-replay/` | 正文/细纲冻结回放、盲评和质量门 | 评测编排与质量证据,不是通用 Skill 单测入口 |
| `harness/` | 跨 Skill 的开发验证与行为评测治理 | 只编排和裁决证据,不拥有业务合同 |
## 3. 权威关系
- `.claude/skills/*/SKILL.md`:Skill 的运行时行为合同;不放开发测试说明。
- `.claude/skills/*/scripts/`:Skill 使用的运行时确定性实现与机械门。
- `tests/skills/<skill>/`:对应 Skill 的实现测试、集成测试和 fake pipeline 测试;测试性质由 harness 清单标注。
- `harness/specs/`:测试和评测规范;不被业务 Skill 当作运行时指令读取。
- `harness/manifests/`:测试/评测登记与执行配置;不复制业务字段合同。
- `harness/evals/`:外部行为评测案例、适配器和报告生成逻辑。
- `.agent/docs/architecture/domains/`:领域与数据权威 SoT;harness 只引用,不复制领域定义。
- `docs/`:单次设计、运行证据和历史资料;不替代 harness 规范。
## 4. 当前状态
- `skill_harness.py` 已实现运行时 Skill 文档与 `skills.json` 的只读静态审计。
- `test-inventory.json` 已登记实现测试、集成测试、fake pipeline、领域评测和 harness 自测的依赖与证据等级。
- `run_selected.py` 已实现显式选择、磁盘/manifest 对账、危险依赖阻断、超时和执行证据检查;它不提供默认全仓一键通过结论。
- 项目运行时 Skill 的实现测试以 `tests/skills/<skill>/` 为目标位置,物理现状以 `test-inventory.json` 为准。
- `skill_behavior_eval` 当前登记数量为 0;没有外部 Agent/模型行为证据时,不声称 Skill 内容有效。
- PostgreSQL、网络和真实模型证据未由离线结果替代,是否执行仍受授权和预算约束。
## 5. 运行边界
- 默认只允许静态检查和明确标注的离线测试;
- PostgreSQL、网络、真实模型和额度探针必须显式选择并满足授权、预算和环境前置条件;
- 任何测试失败、依赖缺失或未发现测试都必须显式返回状态,不能静默跳过;
- 报告必须区分已验证事实、推断和未验证假设。

View File

@ -0,0 +1,308 @@
{
"schema_version": 1,
"skills": [
{
"name": "access-database",
"contract_owner": "平台运行与证据",
"collaborates_with": [
"上下文与知识检索",
"导入、清洗与抽取",
"规划与作品基础",
"写作与候选主权",
"质量与回放评测",
"去 AI 味与人感"
],
"skill_path": ".claude/skills/access-database/SKILL.md"
},
{
"name": "assemble-context",
"contract_owner": "上下文与知识检索",
"collaborates_with": [
"规划与作品基础",
"写作与候选主权",
"质量与回放评测",
"去 AI 味与人感"
],
"skill_path": ".claude/skills/assemble-context/SKILL.md"
},
{
"name": "call-content-model",
"contract_owner": "平台运行与证据",
"collaborates_with": [
"导入、清洗与抽取",
"质量与回放评测"
],
"skill_path": ".claude/skills/call-content-model/SKILL.md"
},
{
"name": "capture-ai-flavor-cases",
"contract_owner": "去 AI 味与人感",
"collaborates_with": [
"质量与回放评测",
"写作与候选主权"
],
"skill_path": ".claude/skills/capture-ai-flavor-cases/SKILL.md"
},
{
"name": "check-content-consistency",
"contract_owner": "质量与回放评测",
"collaborates_with": [
"上下文与知识检索",
"规划与作品基础",
"写作与候选主权"
],
"skill_path": ".claude/skills/check-content-consistency/SKILL.md"
},
{
"name": "clean-book-text",
"contract_owner": "导入、清洗与抽取",
"collaborates_with": [
"平台运行与证据"
],
"skill_path": ".claude/skills/clean-book-text/SKILL.md"
},
{
"name": "decide-candidate",
"contract_owner": "写作与候选主权",
"collaborates_with": [
"平台运行与证据",
"质量与回放评测",
"导入、清洗与抽取",
"规划与作品基础"
],
"skill_path": ".claude/skills/decide-candidate/SKILL.md"
},
{
"name": "deconstruct-book",
"contract_owner": "导入、清洗与抽取",
"collaborates_with": [
"平台运行与证据",
"上下文与知识检索",
"质量与回放评测"
],
"skill_path": ".claude/skills/deconstruct-book/SKILL.md"
},
{
"name": "design-story-foundation",
"contract_owner": "规划与作品基础",
"collaborates_with": [],
"skill_path": ".claude/skills/design-story-foundation/SKILL.md"
},
{
"name": "diagnose-ai-flavor",
"contract_owner": "去 AI 味与人感",
"collaborates_with": [
"质量与回放评测",
"写作与候选主权"
],
"skill_path": ".claude/skills/diagnose-ai-flavor/SKILL.md"
},
{
"name": "embed-knowledge",
"contract_owner": "上下文与知识检索",
"collaborates_with": [
"导入、清洗与抽取"
],
"skill_path": ".claude/skills/embed-knowledge/SKILL.md"
},
{
"name": "establish-voice-baseline",
"contract_owner": "去 AI 味与人感",
"collaborates_with": [
"上下文与知识检索",
"写作与候选主权"
],
"skill_path": ".claude/skills/establish-voice-baseline/SKILL.md"
},
{
"name": "evaluate-frozen-replay",
"contract_owner": "质量与回放评测",
"collaborates_with": [
"平台运行与证据",
"上下文与知识检索",
"写作与候选主权"
],
"skill_path": ".claude/skills/evaluate-frozen-replay/SKILL.md"
},
{
"name": "execute-claude-task",
"contract_owner": "平台运行与证据",
"collaborates_with": [
"质量与回放评测"
],
"skill_path": ".claude/skills/execute-claude-task/SKILL.md"
},
{
"name": "expand-scene",
"contract_owner": "写作与候选主权",
"collaborates_with": [
"上下文与知识检索",
"质量与回放评测"
],
"skill_path": ".claude/skills/expand-scene/SKILL.md"
},
{
"name": "extract-chapter-knowledge",
"contract_owner": "导入、清洗与抽取",
"collaborates_with": [
"平台运行与证据",
"上下文与知识检索",
"写作与候选主权"
],
"skill_path": ".claude/skills/extract-chapter-knowledge/SKILL.md"
},
{
"name": "extract-work-knowledge",
"contract_owner": "导入、清洗与抽取",
"collaborates_with": [
"平台运行与证据",
"上下文与知识检索"
],
"skill_path": ".claude/skills/extract-work-knowledge/SKILL.md"
},
{
"name": "freeze-context",
"contract_owner": "上下文与知识检索",
"collaborates_with": [
"平台运行与证据",
"质量与回放评测"
],
"skill_path": ".claude/skills/freeze-context/SKILL.md"
},
{
"name": "import-book",
"contract_owner": "导入、清洗与抽取",
"collaborates_with": [
"平台运行与证据"
],
"skill_path": ".claude/skills/import-book/SKILL.md"
},
{
"name": "maintain-work-extraction",
"contract_owner": "导入、清洗与抽取",
"collaborates_with": [
"平台运行与证据"
],
"skill_path": ".claude/skills/maintain-work-extraction/SKILL.md"
},
{
"name": "optimize-content-quality",
"contract_owner": "质量与回放评测",
"collaborates_with": [
"上下文与知识检索",
"写作与候选主权"
],
"skill_path": ".claude/skills/optimize-content-quality/SKILL.md"
},
{
"name": "plan-chapter",
"contract_owner": "规划与作品基础",
"collaborates_with": [
"上下文与知识检索",
"质量与回放评测"
],
"skill_path": ".claude/skills/plan-chapter/SKILL.md"
},
{
"name": "plan-story",
"contract_owner": "规划与作品基础",
"collaborates_with": [
"平台运行与证据",
"上下文与知识检索",
"写作与候选主权"
],
"skill_path": ".claude/skills/plan-story/SKILL.md"
},
{
"name": "polish-prose",
"contract_owner": "写作与候选主权",
"collaborates_with": [
"上下文与知识检索",
"质量与回放评测",
"去 AI 味与人感"
],
"skill_path": ".claude/skills/polish-prose/SKILL.md"
},
{
"name": "prevent-ai-flavor",
"contract_owner": "去 AI 味与人感",
"collaborates_with": [
"上下文与知识检索",
"写作与候选主权"
],
"skill_path": ".claude/skills/prevent-ai-flavor/SKILL.md"
},
{
"name": "record-run-evidence",
"contract_owner": "平台运行与证据",
"collaborates_with": [
"上下文与知识检索",
"导入、清洗与抽取",
"规划与作品基础",
"写作与候选主权",
"质量与回放评测",
"去 AI 味与人感"
],
"skill_path": ".claude/skills/record-run-evidence/SKILL.md"
},
{
"name": "review-knowledge-cards",
"contract_owner": "导入、清洗与抽取",
"collaborates_with": [
"平台运行与证据",
"质量与回放评测"
],
"skill_path": ".claude/skills/review-knowledge-cards/SKILL.md"
},
{
"name": "revise-ai-flavor",
"contract_owner": "去 AI 味与人感",
"collaborates_with": [
"质量与回放评测",
"写作与候选主权"
],
"skill_path": ".claude/skills/revise-ai-flavor/SKILL.md"
},
{
"name": "rewrite-selection",
"contract_owner": "写作与候选主权",
"collaborates_with": [
"上下文与知识检索",
"质量与回放评测",
"去 AI 味与人感"
],
"skill_path": ".claude/skills/rewrite-selection/SKILL.md"
},
{
"name": "score-content-quality",
"contract_owner": "质量与回放评测",
"collaborates_with": [
"平台运行与证据",
"上下文与知识检索",
"写作与候选主权"
],
"skill_path": ".claude/skills/score-content-quality/SKILL.md"
},
{
"name": "search-knowledge",
"contract_owner": "上下文与知识检索",
"collaborates_with": [
"规划与作品基础",
"写作与候选主权",
"质量与回放评测"
],
"skill_path": ".claude/skills/search-knowledge/SKILL.md"
},
{
"name": "write-next-chapter",
"contract_owner": "写作与候选主权",
"collaborates_with": [
"平台运行与证据",
"上下文与知识检索",
"质量与回放评测",
"去 AI 味与人感"
],
"skill_path": ".claude/skills/write-next-chapter/SKILL.md"
}
]
}

File diff suppressed because it is too large Load Diff

1027
harness/run_selected.py Normal file

File diff suppressed because it is too large Load Diff

870
harness/skill_harness.py Normal file
View File

@ -0,0 +1,870 @@
#!/usr/bin/env python3
"""对项目运行时 Skill 文档执行只读静态审计。
本模块只读取 `.claude/skills/*/SKILL.md`,不连接数据库、网络或模型,也不修改
工作树。机器调用使用 :func:`audit_skills`,命令行入口同时支持人读摘要和 JSON。
"""
from __future__ import annotations
import argparse
import json
import re
import sys
from collections import Counter
from pathlib import Path
from typing import Any, Iterable, Optional, Sequence
SCHEMA_VERSION = 1
SKILLS_ROOT = Path(".claude") / "skills"
DEFAULT_MANIFEST_PATH = Path("harness") / "manifests" / "skills.json"
_FRONTMATTER_NAME = re.compile(r"^\s*name\s*:\s*(.*?)\s*$")
_COVERAGE_TOOL_PATTERN = re.compile(
r"(?:\bcoverage\.py\b|\bpytest-cov\b|"
r"(?<![A-Za-z0-9_])--cov(?:[-A-Za-z0-9_]+)?(?:\s|=|:|$)|"
r"\bcoverage\s+report\b)",
re.IGNORECASE,
)
_EXPLICIT_TEST_COVERAGE_PATTERN = re.compile(
r"(?:"
r"(?:测试(?:用例|代码)?|单元(?:测试)?|集成(?:测试)?|回归测试|开发测试|开发验证|"
r"代码|分支|语句|行|函数|方法)\s*覆盖率"
r"|(?:unit(?:\s+test)?|integration(?:\s+test)?|regression\s+test|"
r"test|code|branch|statement|line|function|method)\s+coverage"
r")",
re.IGNORECASE,
)
_COVERAGE_PERCENT_PATTERN = re.compile(
r"覆盖率\s*(?:达到|为|至少|不低于|不少于|应达到|>=|>|:|:)?\s*"
r"\d+(?:\.\d+)?\s*[%%]",
re.IGNORECASE,
)
_BUSINESS_COVERAGE_CONTEXT_PATTERN = re.compile(
r"(?:细纲|事件|硬约束|软约束|伏笔|实体|设定|情节|场景|知识|事实|正文|"
r"作品|业务|候选|规划|大纲|章)[^。\n]{0,24}覆盖率"
r"|覆盖率[^。\n]{0,24}(?:细纲|事件|硬约束|软约束|伏笔|实体|设定|情节|"
r"场景|知识|事实|正文|作品|业务|候选|规划|大纲|章)",
re.IGNORECASE,
)
# 这些规则只针对运行时 Skill 文档中的明显开发验证残留;业务正文中的普通“检查”
# 等词不在扫描范围内。规则保持确定性,方便报告中的 code 作为机械门禁输入。
_DEVELOPMENT_HEADING_PATTERN = re.compile(r"^\s*(#{1,6})\s*(.*?)\s*$")
_DEVELOPMENT_SECTION_TITLE_PATTERN = re.compile(
r"(?:自测|测试|开发验证|离线验证|开发测试|离线测试)",
re.IGNORECASE,
)
_DEVELOPMENT_COMMAND_PATTERN = re.compile(
r"(?<![A-Za-z0-9_])check_contract\.py(?![A-Za-z0-9_])",
re.IGNORECASE,
)
_POLLUTION_RULES = (
(
"pytest_reference",
re.compile(r"(?<![A-Za-z0-9_])pytest(?![A-Za-z0-9_])", re.IGNORECASE),
"发现 pytest 测试入口",
),
(
"unittest_reference",
re.compile(r"(?<![A-Za-z0-9_])unittest(?![A-Za-z0-9_])", re.IGNORECASE),
"发现 unittest 测试入口",
),
(
"test_file_reference",
re.compile(
r"(?:\btest_[A-Za-z0-9][A-Za-z0-9_.-]*\.py\b"
r"|\b[A-Za-z0-9][A-Za-z0-9_.-]*_test\.py\b"
r"|test_\*\.py|\*_test\.py)",
re.IGNORECASE,
),
"发现测试文件路径或文件名模式",
),
(
"test_pass_declaration",
re.compile(
r"(?:"
r"(?:所有|全部)?\s*测试\s*(?:已|已经|均|都|全部)?\s*(?:通过|成功)"
r"|(?:通过|成功)\s*测试"
r"|(?:all|the)\s+tests?\s+(?:all\s+)?(?:pass(?:ed|es|ing)?|succeed(?:ed|s|ing)?|successful)"
r"|tests?\s+(?:all\s+)?(?:pass(?:ed|es|ing)?|succeed(?:ed|s|ing)?|successful)"
r")",
re.IGNORECASE,
),
"发现测试通过声明",
),
)
def _resolve_root(root: str | Path) -> Path:
"""把审计根目录解析为绝对路径,但不要求它已经存在。"""
path = Path(root).expanduser()
if not path.is_absolute():
path = Path.cwd() / path
return path.resolve()
def _relative_path(path: Path, root: Path) -> str:
"""返回报告中的稳定 POSIX 相对路径。"""
try:
return path.resolve().relative_to(root).as_posix()
except ValueError:
return path.resolve().as_posix()
def _issue(
code: str,
message: str,
*,
path: Optional[str] = None,
line: Optional[int] = None,
**details: Any,
) -> dict[str, Any]:
result: dict[str, Any] = {
"code": code,
"severity": "error",
"message": message,
}
if path is not None:
result["path"] = path
if line is not None:
result["line"] = line
result.update(details)
return result
def _unquote_scalar(value: str) -> str:
"""解析 Skill frontmatter 中足够支持 name 的简单标量。"""
value = value.strip()
if len(value) >= 2 and value[0] == value[-1] == '"':
try:
parsed = json.loads(value)
except json.JSONDecodeError:
return value[1:-1]
return parsed if isinstance(parsed, str) else value
if len(value) >= 2 and value[0] == value[-1] == "'":
return value[1:-1].replace("''", "'")
# 允许常见的 YAML 行尾注释,不引入 YAML 依赖。
return re.sub(r"\s+#.*$", "", value).strip()
def _parse_frontmatter(text: str) -> tuple[Optional[dict[str, str]], Optional[int], Optional[str]]:
"""读取 frontmatter 的简单键值视图。
返回值为 ``(fields, closing_line, error_code)``。该解析器只需要识别 name,
不试图替代完整 YAML 解析器;不合法或缺失的边界会明确进入失败报告。
"""
lines = text.splitlines()
if not lines or lines[0].lstrip("\ufeff").strip() != "---":
return None, None, "frontmatter_missing"
closing_line: Optional[int] = None
for index in range(1, len(lines)):
if lines[index].strip() in {"---", "..."}:
closing_line = index + 1
break
if closing_line is None:
return None, None, "frontmatter_unclosed"
fields: dict[str, str] = {}
for line in lines[1 : closing_line - 1]:
match = _FRONTMATTER_NAME.match(line)
if match:
fields["name"] = _unquote_scalar(match.group(1))
return fields, closing_line, None
def _snippet(line: str) -> str:
compact = line.strip()
return compact if len(compact) <= 200 else compact[:197] + "..."
def _is_development_coverage_reference(line: str) -> bool:
"""只把明确落在开发测试语境中的 coverage 文字判为污染。"""
if _COVERAGE_TOOL_PATTERN.search(line) or _EXPLICIT_TEST_COVERAGE_PATTERN.search(line):
return True
if not _COVERAGE_PERCENT_PATTERN.search(line):
return False
return not _BUSINESS_COVERAGE_CONTEXT_PATTERN.search(line)
def _inspect_skill_file(path: Path, root: Path) -> tuple[dict[str, Any], list[dict[str, Any]]]:
relative = _relative_path(path, root)
entry: dict[str, Any] = {
"directory": path.parent.name,
"skill_path": relative,
"name": None,
"frontmatter_present": False,
}
issues: list[dict[str, Any]] = []
try:
text = path.read_text(encoding="utf-8")
except (OSError, UnicodeError) as exc:
issues.append(
_issue(
"skill_file_unreadable",
"无法读取 SKILL.md",
path=relative,
error=str(exc),
)
)
return entry, issues
if not text.strip():
issues.append(
_issue(
"empty_skill_file",
"SKILL.md 为空",
path=relative,
line=1,
)
)
return entry, issues
fields, closing_line, frontmatter_error = _parse_frontmatter(text)
if frontmatter_error is not None:
line = 1 if frontmatter_error == "frontmatter_missing" else max(1, len(text.splitlines()))
messages = {
"frontmatter_missing": "缺少以 --- 开始的 frontmatter",
"frontmatter_unclosed": "frontmatter 未闭合",
}
issues.append(
_issue(
frontmatter_error,
messages[frontmatter_error],
path=relative,
line=line,
)
)
else:
entry["frontmatter_present"] = True
name = fields.get("name") if fields is not None else None
entry["name"] = name or None
if not name:
issues.append(
_issue(
"frontmatter_name_missing",
"frontmatter 缺少非空 name",
path=relative,
line=1,
)
)
elif name != path.parent.name:
issues.append(
_issue(
"name_directory_mismatch",
"frontmatter name 与 Skill 目录名不一致",
path=relative,
line=2,
expected=path.parent.name,
actual=name,
)
)
development_section_level: Optional[int] = None
for line_number, line in enumerate(text.splitlines(), start=1):
heading_match = _DEVELOPMENT_HEADING_PATTERN.match(line)
if heading_match:
heading_level = len(heading_match.group(1))
heading_title = re.sub(r"\s+#+\s*$", "", heading_match.group(2)).strip()
if (
development_section_level is not None
and heading_level <= development_section_level
):
development_section_level = None
if (
2 <= heading_level <= 6
and _DEVELOPMENT_SECTION_TITLE_PATTERN.search(heading_title)
):
development_section_level = heading_level
issues.append(
_issue(
"development_test_heading",
"发现开发测试章节标题",
path=relative,
line=line_number,
snippet=_snippet(line),
)
)
for code, pattern, message in _POLLUTION_RULES:
if pattern.search(line):
issues.append(
_issue(
code,
message,
path=relative,
line=line_number,
snippet=_snippet(line),
)
)
if (
development_section_level is not None
and _DEVELOPMENT_COMMAND_PATTERN.search(line)
):
issues.append(
_issue(
"development_test_command_reference",
"发现开发测试章节中的合同测试命令",
path=relative,
line=line_number,
snippet=_snippet(line),
)
)
if _is_development_coverage_reference(line):
issues.append(
_issue(
"coverage_reference",
"发现开发测试语境中的覆盖率声明或工具名",
path=relative,
line=line_number,
snippet=_snippet(line),
)
)
return entry, issues
def _scan_skill_directories(root: Path) -> tuple[list[dict[str, Any]], list[dict[str, Any]]]:
skills_root = root / SKILLS_ROOT
relative_skills_root = _relative_path(skills_root, root)
entries: list[dict[str, Any]] = []
issues: list[dict[str, Any]] = []
if not skills_root.exists():
issues.append(
_issue(
"skills_directory_missing",
"运行时 Skill 目录不存在",
path=relative_skills_root,
)
)
return entries, issues
if not skills_root.is_dir():
issues.append(
_issue(
"skills_directory_not_directory",
"运行时 Skill 路径不是目录",
path=relative_skills_root,
)
)
return entries, issues
try:
children = sorted(skills_root.iterdir(), key=lambda item: item.name)
except OSError as exc:
issues.append(
_issue(
"skills_directory_unreadable",
"无法列出运行时 Skill 目录",
path=relative_skills_root,
error=str(exc),
)
)
return entries, issues
directories = [child for child in children if child.is_dir()]
if not directories:
issues.append(
_issue(
"no_skills_found",
"运行时 Skill 目录中未发现 Skill 子目录",
path=relative_skills_root,
)
)
return entries, issues
for directory in directories:
skill_file = directory / "SKILL.md"
relative_skill_file = _relative_path(skill_file, root)
disk_entry: dict[str, Any] = {
"directory": directory.name,
"skill_path": relative_skill_file,
"name": None,
"frontmatter_present": False,
}
if not skill_file.exists():
entries.append(disk_entry)
issues.append(
_issue(
"missing_skill_file",
"Skill 目录缺少 SKILL.md",
path=relative_skill_file,
)
)
continue
if not skill_file.is_file():
entries.append(disk_entry)
issues.append(
_issue(
"skill_file_not_file",
"SKILL.md 路径不是普通文件",
path=relative_skill_file,
)
)
continue
entry, file_issues = _inspect_skill_file(skill_file, root)
entries.append(entry)
issues.extend(file_issues)
return entries, issues
def _resolve_manifest_path(root: Path, manifest: str | Path | None) -> Path:
path = DEFAULT_MANIFEST_PATH if manifest is None else Path(manifest).expanduser()
if not path.is_absolute():
path = root / path
return path.resolve()
def _normalise_manifest_skill_path(
root: Path, value: str
) -> tuple[Optional[str], Optional[Path]]:
declared = Path(value)
if declared.is_absolute():
return None, None
candidate = (root / declared).resolve()
try:
relative = candidate.relative_to(root).as_posix()
except ValueError:
return None, None
return relative, candidate
def _validate_manifest(
root: Path,
entries: Sequence[dict[str, Any]],
manifest_path: Path,
) -> tuple[dict[str, Any], list[dict[str, Any]]]:
"""验证 Skill 登记表,并要求它与磁盘目录一一对应。"""
relative_manifest = _relative_path(manifest_path, root)
metadata: dict[str, Any] = {
"path": relative_manifest,
"loaded": False,
"skills_declared": 0,
}
issues: list[dict[str, Any]] = []
if not manifest_path.exists():
issues.append(
_issue(
"manifest_missing",
"Skill manifest 不存在",
path=relative_manifest,
)
)
return metadata, issues
if not manifest_path.is_file():
issues.append(
_issue(
"manifest_not_file",
"Skill manifest 路径不是普通文件",
path=relative_manifest,
)
)
return metadata, issues
try:
manifest_text = manifest_path.read_text(encoding="utf-8")
manifest = json.loads(manifest_text)
except UnicodeError as exc:
issues.append(
_issue(
"manifest_unreadable",
"无法按 UTF-8 读取 Skill manifest",
path=relative_manifest,
error=str(exc),
)
)
return metadata, issues
except OSError as exc:
issues.append(
_issue(
"manifest_unreadable",
"无法读取 Skill manifest",
path=relative_manifest,
error=str(exc),
)
)
return metadata, issues
except json.JSONDecodeError as exc:
issues.append(
_issue(
"manifest_invalid_json",
"Skill manifest 不是合法 JSON",
path=relative_manifest,
line=exc.lineno,
column=exc.colno,
error=exc.msg,
)
)
return metadata, issues
metadata["loaded"] = True
if not isinstance(manifest, dict):
issues.append(
_issue(
"manifest_invalid_structure",
"Skill manifest 顶层必须是对象",
path=relative_manifest,
)
)
return metadata, issues
raw_skills = manifest.get("skills")
if not isinstance(raw_skills, list):
issues.append(
_issue(
"manifest_skills_invalid",
"Skill manifest.skills 必须是数组",
path=relative_manifest,
)
)
return metadata, issues
metadata["skills_declared"] = len(raw_skills)
disk_paths = {
str(entry["skill_path"])
for entry in entries
if isinstance(entry.get("skill_path"), str)
}
declared_paths: set[str] = set()
path_occurrences: dict[str, list[int]] = {}
name_occurrences: dict[str, list[int]] = {}
skills_root = (root / SKILLS_ROOT).resolve()
for index, item in enumerate(raw_skills):
if not isinstance(item, dict):
issues.append(
_issue(
"manifest_entry_invalid",
"Skill manifest 条目必须是对象",
path=relative_manifest,
entry_index=index,
)
)
continue
name = item.get("name")
if (
not isinstance(name, str)
or not name.strip()
or "/" in name
or "\\" in name
or name in {".", ".."}
):
issues.append(
_issue(
"manifest_name_invalid",
"manifest 条目的 name 必须是非空目录名",
path=relative_manifest,
entry_index=index,
actual=name,
)
)
else:
name_occurrences.setdefault(name, []).append(index)
contract_owner = item.get("contract_owner")
if not isinstance(contract_owner, str) or not contract_owner.strip():
issues.append(
_issue(
"manifest_contract_owner_invalid",
"manifest 条目的 contract_owner 必须非空",
path=relative_manifest,
entry_index=index,
)
)
collaborations = item.get("collaborates_with")
if (
not isinstance(collaborations, list)
or any(not isinstance(value, str) or not value.strip() for value in collaborations)
):
issues.append(
_issue(
"manifest_collaborates_with_invalid",
"manifest 条目的 collaborates_with 必须是非空字符串数组",
path=relative_manifest,
entry_index=index,
)
)
declared_skill_path = item.get("skill_path")
if not isinstance(declared_skill_path, str) or not declared_skill_path.strip():
issues.append(
_issue(
"manifest_skill_path_invalid",
"manifest 条目的 skill_path 必须是非空相对路径",
path=relative_manifest,
entry_index=index,
actual=declared_skill_path,
)
)
continue
normalised_path, candidate = _normalise_manifest_skill_path(
root, declared_skill_path
)
if normalised_path is None or candidate is None:
issues.append(
_issue(
"manifest_skill_path_invalid",
"manifest 条目的 skill_path 必须位于项目根目录内",
path=relative_manifest,
entry_index=index,
actual=declared_skill_path,
)
)
continue
declared_paths.add(normalised_path)
path_occurrences.setdefault(normalised_path, []).append(index)
if candidate != candidate.parent / "SKILL.md" or not candidate.is_relative_to(skills_root):
issues.append(
_issue(
"manifest_skill_path_invalid",
"manifest 条目的 skill_path 必须指向 .claude/skills/*/SKILL.md",
path=normalised_path,
entry_index=index,
)
)
if not candidate.exists():
issues.append(
_issue(
"manifest_skill_path_missing",
"manifest 登记的 Skill 路径不存在",
path=normalised_path,
entry_index=index,
)
)
elif not candidate.is_file():
issues.append(
_issue(
"manifest_skill_path_not_file",
"manifest 登记的 Skill 路径不是普通文件",
path=normalised_path,
entry_index=index,
)
)
directory_name = candidate.parent.name
if isinstance(name, str) and name.strip() and name != directory_name:
issues.append(
_issue(
"manifest_name_directory_mismatch",
"manifest name 与 skill_path 的目录名不一致",
path=normalised_path,
entry_index=index,
expected=directory_name,
actual=name,
)
)
if isinstance(name, str) and name.strip():
expected_path = (root / SKILLS_ROOT / name / "SKILL.md").resolve()
expected_relative = _relative_path(expected_path, root)
if normalised_path != expected_relative:
issues.append(
_issue(
"manifest_skill_path_mismatch",
"manifest skill_path 与 name 对应的 Skill 路径不一致",
path=normalised_path,
entry_index=index,
expected=expected_relative,
actual=normalised_path,
)
)
for path, indexes in sorted(path_occurrences.items()):
if len(indexes) > 1:
issues.append(
_issue(
"manifest_duplicate_skill_path",
"manifest 重复登记同一个 Skill 路径",
path=path,
entry_indexes=indexes,
)
)
for name, indexes in sorted(name_occurrences.items()):
if len(indexes) > 1:
issues.append(
_issue(
"manifest_duplicate_name",
"manifest 重复登记同一个 Skill name",
path=relative_manifest,
name=name,
entry_indexes=indexes,
)
)
for path in sorted(disk_paths - declared_paths):
issues.append(
_issue(
"manifest_skill_missing",
"磁盘 Skill 未在 manifest 中登记",
path=path,
)
)
for path in sorted(declared_paths - disk_paths):
issues.append(
_issue(
"manifest_skill_extra",
"manifest 登记了磁盘中不存在的额外 Skill",
path=path,
)
)
return metadata, issues
def _duplicate_name_issues(entries: Iterable[dict[str, Any]]) -> list[dict[str, Any]]:
by_name: dict[str, list[str]] = {}
for entry in entries:
name = entry.get("name")
if isinstance(name, str) and name:
by_name.setdefault(name, []).append(str(entry["skill_path"]))
issues: list[dict[str, Any]] = []
for name in sorted(by_name):
paths = sorted(by_name[name])
if len(paths) > 1:
issues.append(
_issue(
"duplicate_name",
"多个 Skill 使用相同的 frontmatter name",
path=paths[0],
name=name,
paths=paths,
)
)
return issues
def _summarize(issues: Sequence[dict[str, Any]], skill_count: int) -> dict[str, Any]:
counts = Counter(str(issue["code"]) for issue in issues)
return {
"status": "passed" if not issues else "failed",
"skills_scanned": skill_count,
"issue_count": len(issues),
"issue_codes": {code: counts[code] for code in sorted(counts)},
}
def audit_skills(
root: str | Path = ".", manifest: str | Path | None = None
) -> dict[str, Any]:
"""审计 Skill 文档与 manifest,并返回 JSON 可序列化报告。"""
root_path = _resolve_root(root)
manifest_path = _resolve_manifest_path(root_path, manifest)
report: dict[str, Any] = {
"schema_version": SCHEMA_VERSION,
"root": str(root_path),
"skills_root": SKILLS_ROOT.as_posix(),
"manifest": {
"path": _relative_path(manifest_path, root_path),
"loaded": False,
},
"skills_scanned": 0,
"skills": [],
"issues": [],
}
if not root_path.exists():
report["issues"] = [
_issue("root_missing", "审计根目录不存在", path=".")
]
elif not root_path.is_dir():
report["issues"] = [
_issue("root_not_directory", "审计根路径不是目录", path=".")
]
else:
entries, issues = _scan_skill_directories(root_path)
issues.extend(_duplicate_name_issues(entries))
manifest_info, manifest_issues = _validate_manifest(
root_path, entries, manifest_path
)
issues.extend(manifest_issues)
report["manifest"] = manifest_info
report["skills"] = entries
report["skills_scanned"] = len(entries)
report["issues"] = issues
report["summary"] = _summarize(report["issues"], report["skills_scanned"])
report["issue_count"] = len(report["issues"])
report["ok"] = not report["issues"]
report["status"] = "passed" if report["ok"] else "failed"
return report
def human_summary(report: dict[str, Any]) -> str:
"""把结构化报告渲染为简洁的人读摘要。"""
status = "通过" if report["ok"] else "失败"
lines = [
f"Skill 静态审计:{status}",
f"扫描 Skill:{report['skills_scanned']} 个;问题:{report['issue_count']} 个",
]
issues = report["issues"]
if not issues:
lines.append("未发现 frontmatter、目录命名、空文件或开发测试污染问题。")
return "\n".join(lines)
lines.append("问题明细:")
for issue in issues:
location = str(issue.get("path", "<root>"))
if issue.get("line") is not None:
location += f":{issue['line']}"
lines.append(f"- {location} [{issue['code']}] {issue['message']}")
return "\n".join(lines)
def _build_parser() -> argparse.ArgumentParser:
parser = argparse.ArgumentParser(
description="只读扫描项目运行时 .claude/skills/*/SKILL.md 的静态卫生审计器。"
)
parser.add_argument(
"--root",
default=".",
help="项目根目录,默认为当前目录",
)
parser.add_argument(
"--manifest",
default=None,
help="Skill manifest 路径;相对路径按项目根目录解析,默认 harness/manifests/skills.json",
)
parser.add_argument(
"--json",
action="store_true",
help="输出结构化 JSON 报告",
)
parser.add_argument(
"--quiet",
action="store_true",
help="不输出人读摘要;与 --json 同用时仍输出 JSON",
)
return parser
def main(argv: Optional[Sequence[str]] = None) -> int:
args = _build_parser().parse_args(argv)
report = audit_skills(args.root, args.manifest)
if args.json:
print(json.dumps(report, ensure_ascii=False, indent=2, sort_keys=True))
elif not args.quiet:
print(human_summary(report))
return 0 if report["ok"] else 1
if __name__ == "__main__":
sys.exit(main())

View File

@ -0,0 +1,115 @@
# Skill 测试与评测规范
> 状态:生效规范
> 适用范围:`agent-example/.claude/skills/*/SKILL.md`、其确定性工具和 Skill 行为评测。
> Owner:`agent-example/harness/`
> 本文件是开发与评测支架规范,不属于任何 Skill 的运行时提示词。
## 1. 三层边界
### 1.1 Skill 运行时指令
`SKILL.md` 只描述 Agent 执行该 Skill 时必须知道的稳定合同:唯一目的、消费者、输入、输出、schema/version、允许读取、允许副作用、禁止动作、数据库读写、授权、预算、raw 和失败关闭规则,以及实际执行所需的业务工具与流程。
`SKILL.md` 不承载开发者测试说明、测试文件路径、测试框架命令、夹具细节或测试结论。
禁止在运行时 Skill 文本中出现:
- `## 自测`、`## 测试`、`## 开发验证` 等开发测试章节;
- `test_*.py`、`*_test.py`、pytest、unittest 等实现测试入口;
- “本测试通过”“测试覆盖 N 项”等开发证据;
- 把测试文件、测试常量或测试输出说成事实源或业务合同。
评测 Skill 的生产/评测执行命令可以保留,但必须是用户请求该评测时实际执行的业务步骤,而不是验证实现代码的开发测试。
### 1.2 确定性工具验证
工具验证回答“Python、数据库或状态机实现是否守住机械合同”,不回答模型是否写得好,也不证明 Skill 的自然语言指令有效。
- 单元测试、离线合同测试和数据库集成测试属于这一层;
- 项目运行时 Skill 的实现测试统一放在 `tests/skills/<skill>/`,领域包保留自己的 `tests/`;两者都必须在 harness 的开发验证清单中登记;
- 测试通过只能证明列出的机械行为,不得升级为模型质量或产品可用性结论;
- 真实数据库、网络、模型和额度探针必须显式标记并单独授权。
### 1.3 Skill 行为评测
行为评测回答“把 Skill 内容交给 Agent 后,Agent 是否按合同行动”。评测由 harness 从外部驱动,不能由 `SKILL.md` 自己宣布通过。
行为评测至少区分:
- 正向触发:应该使用该 Skill 的任务;
- 负向触发:不应该使用该 Skill 的任务;
- 输入缺失与越界;
- 输出合同与失败关闭;
- 关键禁止动作;
- 多样例稳定性和已知混淆项。
行为评测结果必须保留输入、Skill 版本指纹、模型/运行配置、输出摘要、判定证据和失败原因。小样本合同回放不等于文学质量、通用效果或生产完成。
## 2. Harness 职责
Harness 只负责从外部发现、运行、收集和裁决测试/评测证据:
1. 扫描运行时 Skill 文档中的开发测试污染;
2. 登记和分类确定性工具测试、集成测试、fake pipeline 测试和真实探针;
3. 驱动 Skill 行为评测案例;
4. 生成带证据类型和失败原因的结构化报告;
5. 防止空跑、静默跳过、宽泛异常吞错和证据级别越权。
Harness 不负责:
- 修改 `SKILL.md` 或业务代码;
- 判断文学质量;
- 用字符串出现证明自然语言合同有效;
- 把 fake、离线回放或小样本结果升级成生产结论。
## 3. 证据分级
从低到高仅表示证据类型,不允许自动跨级:
1. 静态结构检查:frontmatter、路径、禁用污染模式;
2. 确定性离线测试:纯函数、schema、状态机和失败分支;
3. 真实依赖集成测试:PostgreSQL、文件系统或外部服务;
4. Skill 行为评测:外部 Agent/模型运行与结构化裁决;
5. 人工内容评审:文学质量、声音和语义效果。
任一层通过都不能代替更高层证据。没有行为评测,不得声称 Skill 内容有效;没有真实依赖证据,不得声称生产链路可用。
## 4. 改动门禁
### 只改运行时合同
- 通过 harness 的运行时文档污染扫描;
- 更新受影响的行为评测案例,或记录只改措辞、不改变行为的理由;
- 不新增把测试细节塞回 `SKILL.md` 的说明。
### 只改确定性工具
- 更新针对变更不变量的离线测试;
- 对真实数据库、网络和模型测试单独标记;
- 保留失败关闭和副作用边界证据。
### 改变 Skill 意图或边界
- 更新行为评测清单和正/负向案例;
- 检查消费者、Agent 槽位和 `meta/chains` 映射;
- 重新执行相关层级的验证,不得只跑 Python 单测。
## 5. 明确禁止
- 用 `assertIn` 检查几个词出现,就宣称 Skill 内容正确;
- 用测试文件或测试输出作为运行时事实源;
- 用 `except Exception: pass` 把错误依赖、连接失败或实现错误当成预期拒绝;
- 用任意总入口的“零测试”结果当成通过;
- 把离线 fake、模型探针或小样本合同回放写成真实质量结论;
- 为了让 harness 变绿而修改业务合同、降低断言或静默跳过测试。
## 6. 完成定义
本治理任务只有同时满足以下条件,才可称为完成:
- 运行时 `SKILL.md` 不再携带开发测试说明;
- harness 能机械发现并阻断明显的测试污染;
- 实现测试、集成测试和 Skill 行为评测的证据类型可区分;
- 改动范围内的回归测试真实执行,失败不会被吞掉;
- 报告明确区分已验证事实、推断和未验证的模型质量假设。

View File

@ -0,0 +1,438 @@
#!/usr/bin/env python3
"""run_selected 的标准库离线回归测试。
测试项目、manifest 和 .venv/bin/python 均在临时目录中生成,不连接数据库、网络或模型。
"""
from __future__ import annotations
import io
import json
import stat
import sys
import tempfile
import unittest
from contextlib import redirect_stdout
from pathlib import Path
from typing import Any, Iterable
try:
from .run_selected import main
except ImportError: # 允许直接执行 `.venv/bin/python harness/test_run_selected.py`
from run_selected import main
class RunSelectedTests(unittest.TestCase):
def make_project(
self,
entries: Iterable[dict[str, Any]],
*,
generated_scope: str | None = None,
) -> tuple[tempfile.TemporaryDirectory[str], Path]:
temporary = tempfile.TemporaryDirectory()
root = Path(temporary.name)
(root / ".venv" / "bin").mkdir(parents=True)
fake_python = root / ".venv" / "bin" / "python"
fake_python.write_text(
f"#!{sys.executable}\n"
"import pathlib\n"
"import sys\n"
"import time\n"
"script = pathlib.Path(sys.argv[1])\n"
"(pathlib.Path.cwd() / 'invocations.log').open('a', encoding='utf-8').write(script.name + '\\n')\n"
"if script.name == 'empty.py':\n"
" raise SystemExit(0)\n"
"if script.name == 'skip.py':\n"
" print('1 skipped')\n"
" raise SystemExit(0)\n"
"if script.name == 'fail.py':\n"
" print('failure stdout')\n"
" print('failure stderr', file=sys.stderr)\n"
" raise SystemExit(7)\n"
"if script.name == 'sleep.py':\n"
" time.sleep(2)\n"
"if script.name == 'noisy.py':\n"
" print('x' * 5000)\n"
"else:\n"
" print('ran ' + script.name)\n",
encoding="utf-8",
)
fake_python.chmod(fake_python.stat().st_mode | stat.S_IXUSR | stat.S_IXGRP | stat.S_IXOTH)
manifest_entries = []
for entry in entries:
entry_copy = dict(entry)
path = root / str(entry_copy["path"])
if entry_copy.pop("create", True):
path.parent.mkdir(parents=True, exist_ok=True)
path.write_text("# temporary test script\n", encoding="utf-8")
manifest_entries.append(entry_copy)
manifest_path = root / "manifest.json"
manifest_payload: dict[str, Any] = {
"schema_version": 1,
"entries": manifest_entries,
}
if generated_scope is not None:
manifest_payload["generated_scope"] = generated_scope
manifest_path.write_text(
json.dumps(manifest_payload, indent=2) + "\n",
encoding="utf-8",
)
self.addCleanup(temporary.cleanup)
return temporary, manifest_path
def invoke(self, root: Path, manifest: Path, *arguments: str) -> tuple[int, dict[str, Any]]:
output = io.StringIO()
with redirect_stdout(output):
return_code = main(
[
"--root",
str(root),
"--manifest",
str(manifest),
*arguments,
"--json",
]
)
return return_code, json.loads(output.getvalue())
@staticmethod
def entry(
path: str,
*,
owner: str = "alpha",
kind: str = "tool_unit",
scope: str = "runtime_skill",
requires: list[str] | None = None,
create: bool = True,
) -> dict[str, Any]:
return {
"path": path,
"owner_skill_or_domain": owner,
"kind": kind,
"scope": scope,
"requires": ["offline"] if requires is None else requires,
"create": create,
}
def test_selector_alias_and_intersection_are_reported_as_json(self) -> None:
_, manifest = self.make_project(
[
self.entry("pass.py", owner="alpha", kind="tool_unit"),
self.entry("other.py", owner="beta", kind="tool_contract"),
]
)
root = manifest.parent
return_code, report = self.invoke(
root,
manifest,
"--skill",
"alpha",
"--kind",
"tool_unit",
"--path",
"pass.py",
)
self.assertEqual(return_code, 0)
self.assertEqual(report["status"], "passed")
self.assertEqual(report["selected_count"], 1)
self.assertEqual(report["entries"][0]["path"], "pass.py")
self.assertEqual(report["entries"][0]["returncode"], 0)
self.assertIn("ran pass.py", report["entries"][0]["stdout"])
def test_offline_dependency_is_blocked_without_running_child(self) -> None:
_, manifest = self.make_project(
[self.entry("db.py", requires=["offline", "postgresql"])]
)
root = manifest.parent
return_code, report = self.invoke(root, manifest, "--path", "db.py")
result = report["entries"][0]
self.assertEqual(return_code, 1)
self.assertEqual(report["status"], "failed")
self.assertEqual(result["status"], "blocked_dependency")
self.assertEqual(result["requires"], ["offline", "postgresql"])
self.assertEqual(result["blocked_requires"], ["postgresql"])
self.assertFalse((root / "invocations.log").exists())
def test_allow_requires_runs_and_still_reports_dependency(self) -> None:
_, manifest = self.make_project(
[self.entry("db.py", requires=["offline", "postgresql"])]
)
root = manifest.parent
return_code, report = self.invoke(
root,
manifest,
"--path",
"db.py",
"--allow-requires",
"postgresql",
)
result = report["entries"][0]
self.assertEqual(return_code, 0)
self.assertEqual(result["status"], "passed")
self.assertEqual(result["requires"], ["offline", "postgresql"])
self.assertEqual((root / "invocations.log").read_text(encoding="utf-8"), "db.py\n")
def test_nonzero_and_timeout_have_distinct_structured_statuses(self) -> None:
_, manifest = self.make_project(
[self.entry("fail.py"), self.entry("sleep.py")]
)
root = manifest.parent
return_code, report = self.invoke(
root,
manifest,
"--path",
"fail.py",
"--path",
"sleep.py",
"--timeout-seconds",
"0.5",
)
self.assertEqual(return_code, 1)
results = {entry["path"]: entry for entry in report["entries"]}
self.assertEqual(results["fail.py"]["status"], "failed")
self.assertEqual(results["fail.py"]["returncode"], 7)
self.assertIn("failure stdout", results["fail.py"]["stdout"])
self.assertIn("failure stderr", results["fail.py"]["stderr"])
self.assertEqual(results["sleep.py"]["status"], "timeout")
self.assertIsNone(results["sleep.py"]["returncode"])
self.assertEqual(results["sleep.py"]["error"]["code"], "timeout")
def test_no_matches_and_missing_path_are_not_silent(self) -> None:
_, manifest = self.make_project(
[
self.entry("missing.py", create=False),
self.entry("present.py"),
]
)
root = manifest.parent
return_code, no_match = self.invoke(root, manifest, "--kind", "does_not_exist")
self.assertEqual(return_code, 1)
self.assertEqual(no_match["status"], "no_matches")
self.assertEqual(no_match["entries"], [])
self.assertEqual(no_match["issues"][0]["code"], "no_matches")
return_code, missing = self.invoke(root, manifest, "--path", "missing.py")
self.assertEqual(return_code, 1)
self.assertEqual(missing["entries"][0]["status"], "not_found")
self.assertEqual(missing["entries"][0]["error"]["code"], "path_not_found")
def test_invalid_manifest_and_selector_requirement_are_structured(self) -> None:
temporary = tempfile.TemporaryDirectory()
self.addCleanup(temporary.cleanup)
root = Path(temporary.name)
manifest = root / "manifest.json"
manifest.write_text("{broken", encoding="utf-8")
return_code, invalid_manifest = self.invoke(root, manifest, "--path", "x.py")
self.assertEqual(return_code, 1)
self.assertEqual(invalid_manifest["status"], "manifest_invalid")
self.assertEqual(invalid_manifest["issues"][0]["code"], "manifest_invalid_json")
valid_root, valid_manifest = self.make_project([self.entry("pass.py")])
return_code, invalid_selector = self.invoke(valid_root, valid_manifest)
self.assertEqual(return_code, 1)
self.assertEqual(invalid_selector["status"], "invalid_selector")
self.assertEqual(invalid_selector["issues"][0]["code"], "selector_required")
def test_all_offline_is_explicit_opt_in_and_output_is_summary_only(self) -> None:
_, manifest = self.make_project(
[
self.entry("pass.py"),
self.entry("db.py", requires=["postgresql"]),
self.entry("noisy.py"),
]
)
root = manifest.parent
return_code, report = self.invoke(root, manifest, "--all-offline")
self.assertEqual(return_code, 1)
self.assertEqual(report["selected_count"], 3)
statuses = {entry["path"]: entry["status"] for entry in report["entries"]}
self.assertEqual(statuses["pass.py"], "passed")
self.assertEqual(statuses["db.py"], "blocked_dependency")
self.assertEqual(statuses["noisy.py"], "passed")
noisy = next(entry for entry in report["entries"] if entry["path"] == "noisy.py")
self.assertLess(len(noisy["stdout"]), 2001)
self.assertIn("output summary truncated", noisy["stdout"])
def test_generated_scope_rejects_unregistered_disk_asset(self) -> None:
_, manifest = self.make_project(
[self.entry("tests/skills/registered/test_registered.py")],
generated_scope="temporary test asset inventory",
)
root = manifest.parent
unregistered = root / "tests" / "skills" / "new" / "test_unregistered.py"
unregistered.parent.mkdir(parents=True)
unregistered.write_text("# unregistered test asset\n", encoding="utf-8")
return_code, report = self.invoke(root, manifest, "--all-offline")
self.assertEqual(return_code, 1)
self.assertEqual(report["status"], "manifest_invalid")
self.assertEqual(report["entries"], [])
self.assertEqual(
[
issue["path"]
for issue in report["issues"]
if issue["code"] == "manifest_test_asset_missing"
],
["tests/skills/new/test_unregistered.py"],
)
self.assertFalse((root / "invocations.log").exists())
def test_generated_scope_rejects_manifest_extra_asset(self) -> None:
_, manifest = self.make_project(
[
self.entry(
"tests/skills/removed/test_removed.py",
create=False,
)
],
generated_scope="temporary test asset inventory",
)
root = manifest.parent
return_code, report = self.invoke(root, manifest, "--all-offline")
self.assertEqual(return_code, 1)
self.assertEqual(report["status"], "manifest_invalid")
self.assertEqual(report["entries"], [])
self.assertEqual(
[
issue["path"]
for issue in report["issues"]
if issue["code"] == "manifest_test_asset_extra"
],
["tests/skills/removed/test_removed.py"],
)
self.assertFalse((root / "invocations.log").exists())
def test_generated_scope_ignores_non_test_helpers_under_test_roots(self) -> None:
_, manifest = self.make_project(
[self.entry("tests/skills/registered/test_registered.py")],
generated_scope="temporary test asset inventory",
)
root = manifest.parent
registered_test = root / "tests" / "skills" / "registered" / "test_registered.py"
registered_test.write_text("def test_registered():\n pass\n", encoding="utf-8")
(root / "tests" / "skills" / "registered" / "helper.py").write_text(
"VALUE = 1\n", encoding="utf-8"
)
(root / "humanization" / "tests" / "helper.py").parent.mkdir(
parents=True, exist_ok=True
)
(root / "humanization" / "tests" / "helper.py").write_text(
"VALUE = 2\n", encoding="utf-8"
)
return_code, report = self.invoke(root, manifest, "--all-offline")
self.assertEqual(return_code, 0)
self.assertEqual(report["status"], "passed")
self.assertEqual(report["manifest"]["test_assets_scanned"], 1)
self.assertEqual(report["entries"][0]["path"], "tests/skills/registered/test_registered.py")
def test_generated_scope_blocks_empty_and_print_only_scripts_before_child(self) -> None:
entries = [
self.entry("tests/skills/empty/test_empty.py"),
self.entry("tests/skills/print_only/test_print_only.py"),
]
_, manifest = self.make_project(
entries,
generated_scope="temporary test asset inventory",
)
root = manifest.parent
(root / "tests" / "skills" / "empty" / "test_empty.py").write_text(
"", encoding="utf-8"
)
(root / "tests" / "skills" / "print_only" / "test_print_only.py").write_text(
"print('not a test')\n", encoding="utf-8"
)
return_code, report = self.invoke(root, manifest, "--all-offline")
self.assertEqual(return_code, 1)
self.assertEqual(report["status"], "failed")
results = {entry["path"]: entry for entry in report["entries"]}
for path in (
"tests/skills/empty/test_empty.py",
"tests/skills/print_only/test_print_only.py",
):
self.assertEqual(results[path]["status"], "failed")
self.assertIsNone(results[path]["returncode"])
self.assertEqual(results[path]["error"]["code"], "test_shape_missing")
self.assertFalse((root / "invocations.log").exists())
def test_generated_scope_accepts_supported_python_shapes(self) -> None:
entries = [
self.entry("tests/skills/function/test_function.py"),
self.entry("tests/skills/class/test_class.py"),
self.entry("harness/test_main_entry.py"),
]
_, manifest = self.make_project(
entries,
generated_scope="temporary test asset inventory",
)
root = manifest.parent
(root / "tests" / "skills" / "function" / "test_function.py").write_text(
"def test_function():\n pass\n", encoding="utf-8"
)
(root / "tests" / "skills" / "class" / "test_class.py").write_text(
"import unittest\n\nclass Fixture(unittest.TestCase):\n pass\n",
encoding="utf-8",
)
(root / "harness" / "test_main_entry.py").write_text(
"if __name__ == \"__main__\":\n print(\"script\")\n",
encoding="utf-8",
)
return_code, report = self.invoke(root, manifest, "--all-offline")
self.assertEqual(return_code, 0)
self.assertEqual(report["status"], "passed")
self.assertEqual(
[entry["status"] for entry in report["entries"]],
["passed", "passed", "passed"],
)
self.assertEqual(
(root / "invocations.log").read_text(encoding="utf-8").splitlines(),
["test_function.py", "test_class.py", "test_main_entry.py"],
)
def test_zero_exit_without_execution_evidence_fails_closed(self) -> None:
_, manifest = self.make_project(
[self.entry("empty.py"), self.entry("skip.py")]
)
root = manifest.parent
return_code, report = self.invoke(
root,
manifest,
"--path",
"empty.py",
"--path",
"skip.py",
)
self.assertEqual(return_code, 1)
results = {entry["path"]: entry for entry in report["entries"]}
for path in ("empty.py", "skip.py"):
self.assertEqual(results[path]["status"], "failed")
self.assertEqual(results[path]["returncode"], 0)
self.assertEqual(results[path]["error"]["code"], "no_execution_evidence")
if __name__ == "__main__":
unittest.main(verbosity=2)

View File

@ -0,0 +1,420 @@
#!/usr/bin/env python3
"""skill_harness 的纯标准库离线回归测试。
所有夹具都在临时目录中构造,不读取当前仓库的 Skill,也不依赖数据库、网络或模型。
"""
from __future__ import annotations
import io
import json
import tempfile
import unittest
from contextlib import redirect_stdout
from pathlib import Path
from typing import Optional
try:
from .skill_harness import audit_skills, main
except ImportError: # 允许直接执行 `.venv/bin/python harness/test_skill_harness.py`
from skill_harness import audit_skills, main
class SkillHarnessTests(unittest.TestCase):
def make_skill(
self,
root: Path,
directory: str,
*,
name: Optional[str] = None,
body: str = "# 合同\n\n只描述运行时行为。\n",
) -> Path:
skill_dir = root / ".claude" / "skills" / directory
skill_dir.mkdir(parents=True, exist_ok=True)
skill_path = skill_dir / "SKILL.md"
frontmatter_name = directory if name is None else name
skill_path.write_text(
f"---\nname: {frontmatter_name}\ndescription: 离线夹具\n---\n{body}",
encoding="utf-8",
)
self.write_manifest(root)
return skill_path
def write_manifest(
self,
root: Path,
entries: Optional[list[dict[str, object]]] = None,
*,
path: Optional[Path] = None,
) -> Path:
if entries is None:
entries = []
skills_root = root / ".claude" / "skills"
if skills_root.exists():
for directory in sorted(skills_root.iterdir(), key=lambda item: item.name):
if directory.is_dir():
entries.append(
{
"name": directory.name,
"contract_owner": "测试夹具",
"collaborates_with": [],
"skill_path": (
Path(".claude")
/ "skills"
/ directory.name
/ "SKILL.md"
).as_posix(),
}
)
manifest_path = path or root / "harness" / "manifests" / "skills.json"
manifest_path.parent.mkdir(parents=True, exist_ok=True)
manifest_path.write_text(
json.dumps(
{"schema_version": 1, "skills": entries},
ensure_ascii=False,
indent=2,
)
+ "\n",
encoding="utf-8",
)
return manifest_path
def test_passing_project_is_independent_of_current_repository(self) -> None:
with tempfile.TemporaryDirectory() as temporary:
root = Path(temporary)
self.make_skill(root, "alpha")
report = audit_skills(root)
self.assertTrue(report["ok"])
self.assertEqual(report["status"], "passed")
self.assertEqual(report["skills_scanned"], 1)
self.assertEqual(report["issues"], [])
def test_pollution_is_reported_for_each_obvious_category(self) -> None:
with tempfile.TemporaryDirectory() as temporary:
root = Path(temporary)
self.make_skill(
root,
"polluted",
body=(
"## 自测\n"
"pytest -q\n"
"import unittest\n"
"run test_sample.py and sample_test.py\n"
"覆盖率达到 100%。\n"
"测试通过。\n"
),
)
report = audit_skills(root)
codes = {issue["code"] for issue in report["issues"]}
self.assertFalse(report["ok"])
self.assertTrue(
{
"development_test_heading",
"pytest_reference",
"unittest_reference",
"test_file_reference",
"coverage_reference",
"test_pass_declaration",
}.issubset(codes)
)
def test_development_headings_support_multiple_levels(self) -> None:
with tempfile.TemporaryDirectory() as temporary:
root = Path(temporary)
self.make_skill(
root,
"nested-development-sections",
body=(
"### 自测\n"
"这里不能登记测试命令。\n"
"## 离线验证\n"
"这里也不能登记测试命令。\n"
),
)
report = audit_skills(root)
headings = [
issue
for issue in report["issues"]
if issue["code"] == "development_test_heading"
]
self.assertFalse(report["ok"])
self.assertEqual(len(headings), 2)
self.assertEqual([issue["line"] for issue in headings], [5, 7])
def test_check_contract_command_requires_development_section_context(self) -> None:
with tempfile.TemporaryDirectory() as temporary:
root = Path(temporary)
self.make_skill(
root,
"contextual-contract-command",
body=(
"业务合同字段名可以写 check_contract.py。\n"
"## 离线验证\n"
"python check_contract.py\n"
"## 业务说明\n"
"普通业务 check_contract.py 不是测试入口。\n"
),
)
report = audit_skills(root)
command_issues = [
issue
for issue in report["issues"]
if issue["code"] == "development_test_command_reference"
]
self.assertFalse(report["ok"])
self.assertEqual(len(command_issues), 1)
self.assertEqual(command_issues[0]["line"], 7)
def test_business_coverage_terms_are_not_development_pollution(self) -> None:
with tempfile.TemporaryDirectory() as temporary:
root = Path(temporary)
self.make_skill(
root,
"business-coverage",
body=(
"细纲覆盖率达到 100%。\n"
"事件覆盖率为 80%。\n"
"硬约束覆盖率 100%。\n"
),
)
report = audit_skills(root)
self.assertTrue(report["ok"])
self.assertNotIn(
"coverage_reference",
{issue["code"] for issue in report["issues"]},
)
def test_development_coverage_terms_are_reported(self) -> None:
with tempfile.TemporaryDirectory() as temporary:
root = Path(temporary)
self.make_skill(
root,
"development-coverage",
body=(
"coverage.py\n"
"pytest-cov\n"
"pytest --cov=harness\n"
"coverage report\n"
"测试覆盖率达到 90%。\n"
"覆盖率达到 80%。\n"
),
)
report = audit_skills(root)
self.assertFalse(report["ok"])
self.assertIn(
"coverage_reference",
{issue["code"] for issue in report["issues"]},
)
def test_frontmatter_name_must_match_skill_directory(self) -> None:
with tempfile.TemporaryDirectory() as temporary:
root = Path(temporary)
self.make_skill(root, "alpha", name="beta")
report = audit_skills(root)
self.assertFalse(report["ok"])
self.assertIn(
"name_directory_mismatch",
{issue["code"] for issue in report["issues"]},
)
def test_duplicate_frontmatter_names_are_reported(self) -> None:
with tempfile.TemporaryDirectory() as temporary:
root = Path(temporary)
self.make_skill(root, "alpha", name="shared")
self.make_skill(root, "beta", name="shared")
report = audit_skills(root)
duplicates = [
issue for issue in report["issues"] if issue["code"] == "duplicate_name"
]
self.assertFalse(report["ok"])
self.assertEqual(len(duplicates), 1)
self.assertEqual(duplicates[0]["name"], "shared")
self.assertEqual(len(duplicates[0]["paths"]), 2)
def test_clean_manifest_can_be_selected_explicitly(self) -> None:
with tempfile.TemporaryDirectory() as temporary:
root = Path(temporary)
self.make_skill(root, "alpha")
manifest_path = self.write_manifest(
root,
path=root / "clean-manifest.json",
)
output = io.StringIO()
with redirect_stdout(output):
return_code = main(
[
"--root",
str(root),
"--manifest",
str(manifest_path),
"--json",
"--quiet",
]
)
payload = json.loads(output.getvalue())
self.assertEqual(return_code, 0)
self.assertTrue(payload["ok"])
self.assertTrue(payload["manifest"]["loaded"])
self.assertEqual(payload["manifest"]["path"], "clean-manifest.json")
def test_error_manifest_fails_with_structured_issues(self) -> None:
with tempfile.TemporaryDirectory() as temporary:
root = Path(temporary)
self.make_skill(root, "alpha")
self.make_skill(root, "beta")
broken_manifest = self.write_manifest(
root,
entries=[
{
"name": "alpha",
"contract_owner": "",
"collaborates_with": ["上下文与知识检索", 3],
"skill_path": ".claude/skills/beta/SKILL.md",
},
{
"name": "alpha",
"contract_owner": "测试夹具",
"collaborates_with": [],
"skill_path": ".claude/skills/beta/SKILL.md",
},
{
"name": "extra",
"contract_owner": "测试夹具",
"collaborates_with": [],
"skill_path": ".claude/skills/extra/SKILL.md",
},
],
path=root / "broken-manifest.json",
)
output = io.StringIO()
with redirect_stdout(output):
return_code = main(
[
"--root",
str(root),
"--manifest",
str(broken_manifest),
"--json",
]
)
payload = json.loads(output.getvalue())
codes = {issue["code"] for issue in payload["issues"]}
self.assertEqual(return_code, 1)
self.assertFalse(payload["ok"])
self.assertTrue(all(isinstance(issue, dict) for issue in payload["issues"]))
self.assertTrue(
{
"manifest_contract_owner_invalid",
"manifest_collaborates_with_invalid",
"manifest_duplicate_skill_path",
"manifest_duplicate_name",
"manifest_skill_missing",
"manifest_skill_extra",
"manifest_skill_path_mismatch",
"manifest_name_directory_mismatch",
"manifest_skill_path_missing",
}.issubset(codes)
)
def test_missing_manifest_fails_closed(self) -> None:
with tempfile.TemporaryDirectory() as temporary:
root = Path(temporary)
self.make_skill(root, "alpha")
(root / "harness" / "manifests" / "skills.json").unlink()
report = audit_skills(root)
self.assertFalse(report["ok"])
self.assertIn(
"manifest_missing",
{issue["code"] for issue in report["issues"]},
)
def test_missing_skills_directory_fails_closed(self) -> None:
with tempfile.TemporaryDirectory() as temporary:
root = Path(temporary)
report = audit_skills(root)
self.assertFalse(report["ok"])
self.assertIn(
"skills_directory_missing",
{issue["code"] for issue in report["issues"]},
)
def test_skill_directory_without_skill_file_fails(self) -> None:
with tempfile.TemporaryDirectory() as temporary:
root = Path(temporary)
(root / ".claude" / "skills" / "missing").mkdir(parents=True)
report = audit_skills(root)
self.assertFalse(report["ok"])
self.assertIn(
"missing_skill_file",
{issue["code"] for issue in report["issues"]},
)
def test_empty_skill_file_is_reported(self) -> None:
with tempfile.TemporaryDirectory() as temporary:
root = Path(temporary)
skill_path = self.make_skill(root, "empty")
skill_path.write_text("", encoding="utf-8")
report = audit_skills(root)
self.assertFalse(report["ok"])
self.assertIn(
"empty_skill_file",
{issue["code"] for issue in report["issues"]},
)
def test_json_cli_and_exit_code(self) -> None:
with tempfile.TemporaryDirectory() as temporary:
root = Path(temporary)
self.make_skill(root, "alpha")
output = io.StringIO()
with redirect_stdout(output):
return_code = main(["--root", str(root), "--json", "--quiet"])
payload = json.loads(output.getvalue())
self.assertEqual(return_code, 0)
self.assertTrue(payload["ok"])
self.assertEqual(payload["skills_scanned"], 1)
(root / ".claude" / "skills" / "alpha" / "SKILL.md").write_text(
"---\nname: alpha\n---\n## 测试\n",
encoding="utf-8",
)
output = io.StringIO()
with redirect_stdout(output):
return_code = main(["--root", str(root), "--json"])
failed_payload = json.loads(output.getvalue())
self.assertEqual(return_code, 1)
self.assertFalse(failed_payload["ok"])
if __name__ == "__main__":
unittest.main(verbosity=2)

View File

@ -73,7 +73,7 @@ capabilities:
evidence_sources: [no-ai-slop, neuro-book, humanizer] evidence_sources: [no-ai-slop, neuro-book, humanizer]
owner_skill: prevent-ai-flavor owner_skill: prevent-ai-flavor
implementation: .claude/skills/prevent-ai-flavor/scripts/prevent_ai_flavor.py implementation: .claude/skills/prevent-ai-flavor/scripts/prevent_ai_flavor.py
test: .claude/skills/assemble-context/scripts/test_assemble_writer_context.py test: tests/skills/assemble-context/test_assemble_writer_context.py
status: implemented status: implemented
- id: sf_snf_boundary_regression - id: sf_snf_boundary_regression
evidence_sources: [speak-human-tw, shuorenhua, neuro-book] evidence_sources: [speak-human-tw, shuorenhua, neuro-book]
@ -98,14 +98,14 @@ capabilities:
evidence_sources: [inkos, oh-story-claudecode, neuro-book] evidence_sources: [inkos, oh-story-claudecode, neuro-book]
owner_skill: revise-ai-flavor owner_skill: revise-ai-flavor
implementation: humanization/src/deai/pipeline.py implementation: humanization/src/deai/pipeline.py
test: .claude/skills/revise-ai-flavor/scripts/test_revise_ai_flavor.py test: tests/skills/revise-ai-flavor/test_revise_ai_flavor.py
status: implemented status: implemented
note: 当前以最多3轮合同、复扫和pairwise no_gain实现;跨轮最佳快照编排仍由上层负责 note: 当前以最多3轮合同、复扫和pairwise no_gain实现;跨轮最佳快照编排仍由上层负责
- id: source_revalidation - id: source_revalidation
evidence_sources: [shuorenhua, neuro-book] evidence_sources: [shuorenhua, neuro-book]
owner_skill: capture-ai-flavor-cases owner_skill: capture-ai-flavor-cases
implementation: .claude/skills/capture-ai-flavor-cases/scripts/capture_cases.py implementation: .claude/skills/capture-ai-flavor-cases/scripts/capture_cases.py
test: .claude/skills/capture-ai-flavor-cases/scripts/test_capture_cases.py test: tests/skills/capture-ai-flavor-cases/test_capture_cases.py
status: implemented status: implemented
- id: case_to_rule_lifecycle - id: case_to_rule_lifecycle
evidence_sources: [unslop, shuorenhua, neuro-book] evidence_sources: [unslop, shuorenhua, neuro-book]

View File

@ -6,7 +6,8 @@ import re
import unittest import unittest
DDL_PATH = pathlib.Path(__file__).resolve().parents[4] / "db" / "ddl" / "96-example参考作品授权快照.sql" PROJECT_ROOT = pathlib.Path(__file__).resolve().parents[3]
DDL_PATH = PROJECT_ROOT / "db" / "ddl" / "96-example参考作品授权快照.sql"
class AuthorizationSnapshotDdlTest(unittest.TestCase): class AuthorizationSnapshotDdlTest(unittest.TestCase):

View File

@ -1,14 +1,16 @@
#!/usr/bin/env python3 #!/usr/bin/env python3
"""db execparams 参数装载逻辑离线自测(不连库)。 """db execparams 参数装载逻辑离线自测(不连库)。
跑法(仓库根目录):.venv/bin/python .claude/skills/access-database/scripts/test_db_params.py 跑法(仓库根目录):.venv/bin/python tests/skills/access-database/test_db_params.py
""" """
import io import io
import json import json
import os
import sys import sys
from pathlib import Path
sys.path.insert(0, os.path.dirname(os.path.abspath(__file__))) PROJECT_ROOT = Path(__file__).resolve().parents[3]
SCRIPTS_DIR = PROJECT_ROOT / ".claude" / "skills" / "access-database" / "scripts"
sys.path.insert(0, str(SCRIPTS_DIR))
import click # noqa: E402 import click # noqa: E402

View File

@ -2,9 +2,14 @@
"""Skill 目录命名与 frontmatter 一致性的离线测试。""" """Skill 目录命名与 frontmatter 一致性的离线测试。"""
import pathlib import pathlib
import sys
import tempfile import tempfile
import unittest import unittest
PROJECT_ROOT = pathlib.Path(__file__).resolve().parents[3]
SCRIPTS_DIR = PROJECT_ROOT / ".claude" / "skills" / "access-database" / "scripts"
sys.path.insert(0, str(SCRIPTS_DIR))
from sync_agent_registry import validate_skill_catalog from sync_agent_registry import validate_skill_catalog

View File

@ -9,7 +9,9 @@ import pathlib
import sys import sys
import unittest import unittest
sys.path.insert(0, str(pathlib.Path(__file__).resolve().parent)) PROJECT_ROOT = pathlib.Path(__file__).resolve().parents[3]
SCRIPT_DIR = PROJECT_ROOT / ".claude" / "skills" / "assemble-context" / "scripts"
sys.path.insert(0, str(SCRIPT_DIR))
from assemble_writer_context import ( # noqa: E402 from assemble_writer_context import ( # noqa: E402
AssemblyError, AssemblyError,

View File

@ -5,7 +5,9 @@ import pathlib
import sys import sys
import unittest import unittest
sys.path.insert(0, str(pathlib.Path(__file__).resolve().parent)) PROJECT_ROOT = pathlib.Path(__file__).resolve().parents[3]
SCRIPT_DIR = PROJECT_ROOT / ".claude" / "skills" / "assemble-context" / "scripts"
sys.path.insert(0, str(SCRIPT_DIR))
from retrieve_writer_sources import RetrievalError, load_confirmed_fine_outline # noqa: E402 from retrieve_writer_sources import RetrievalError, load_confirmed_fine_outline # noqa: E402

View File

@ -14,7 +14,9 @@ import pathlib
import sys import sys
import unittest import unittest
sys.path.insert(0, str(pathlib.Path(__file__).resolve().parent)) PROJECT_ROOT = pathlib.Path(__file__).resolve().parents[3]
SCRIPT_DIR = PROJECT_ROOT / ".claude" / "skills" / "assemble-context" / "scripts"
sys.path.insert(0, str(SCRIPT_DIR))
from assemble_writer_context import AssemblyError, _outline_contract # noqa: E402 from assemble_writer_context import AssemblyError, _outline_contract # noqa: E402

View File

@ -9,7 +9,9 @@ import pathlib
import sys import sys
import unittest import unittest
sys.path.insert(0, str(pathlib.Path(__file__).resolve().parent)) PROJECT_ROOT = pathlib.Path(__file__).resolve().parents[3]
SCRIPT_DIR = PROJECT_ROOT / ".claude" / "skills" / "assemble-context" / "scripts"
sys.path.insert(0, str(SCRIPT_DIR))
from retrieve_writer_sources import load_confirmed_pattern_bindings # noqa: E402 from retrieve_writer_sources import load_confirmed_pattern_bindings # noqa: E402

View File

@ -8,10 +8,18 @@ import pathlib
import sys import sys
import unittest import unittest
SCRIPT_DIR = pathlib.Path(__file__).resolve().parent PROJECT_ROOT = pathlib.Path(__file__).resolve().parents[3]
sys.path.insert(0, str(SCRIPT_DIR)) ASSEMBLE_CONTEXT_SCRIPTS = PROJECT_ROOT / ".claude" / "skills" / "assemble-context" / "scripts"
sys.path.insert(0, str(SCRIPT_DIR.parents[1] / "freeze-context" / "scripts")) FREEZE_CONTEXT_SCRIPTS = PROJECT_ROOT / ".claude" / "skills" / "freeze-context" / "scripts"
sys.path.insert(0, str(SCRIPT_DIR.parents[1] / "search-knowledge" / "scripts")) SEARCH_KNOWLEDGE_SCRIPTS = PROJECT_ROOT / ".claude" / "skills" / "search-knowledge" / "scripts"
EMBED_KNOWLEDGE_SCRIPTS = PROJECT_ROOT / ".claude" / "skills" / "embed-knowledge" / "scripts"
for script_dir in (
ASSEMBLE_CONTEXT_SCRIPTS,
FREEZE_CONTEXT_SCRIPTS,
SEARCH_KNOWLEDGE_SCRIPTS,
EMBED_KNOWLEDGE_SCRIPTS,
):
sys.path.insert(0, str(script_dir))
from load_reference_work import begin_read_snapshot # noqa: E402 from load_reference_work import begin_read_snapshot # noqa: E402
from search import search_cards # noqa: E402 from search import search_cards # noqa: E402

View File

@ -9,7 +9,9 @@ import pathlib
import sys import sys
import unittest import unittest
sys.path.insert(0, str(pathlib.Path(__file__).resolve().parent)) PROJECT_ROOT = pathlib.Path(__file__).resolve().parents[3]
SCRIPT_DIR = PROJECT_ROOT / ".claude" / "skills" / "assemble-context" / "scripts"
sys.path.insert(0, str(SCRIPT_DIR))
from retrieve_writer_sources import derive_style_constraints, load_confirmed_style # noqa: E402 from retrieve_writer_sources import derive_style_constraints, load_confirmed_style # noqa: E402

View File

@ -9,7 +9,9 @@ import pathlib
import sys import sys
import unittest import unittest
sys.path.insert(0, str(pathlib.Path(__file__).resolve().parent)) PROJECT_ROOT = pathlib.Path(__file__).resolve().parents[3]
SCRIPT_DIR = PROJECT_ROOT / ".claude" / "skills" / "assemble-context" / "scripts"
sys.path.insert(0, str(SCRIPT_DIR))
from writer_contract import ( # noqa: E402 from writer_contract import ( # noqa: E402
ContractError, ContractError,

View File

@ -9,7 +9,9 @@ import sys
import types import types
sys.path.insert(0, str(pathlib.Path(__file__).resolve().parent)) PROJECT_ROOT = pathlib.Path(__file__).resolve().parents[3]
SCRIPT_DIR = PROJECT_ROOT / ".claude" / "skills" / "call-content-model" / "scripts"
sys.path.insert(0, str(SCRIPT_DIR))
import llm # noqa: E402 import llm # noqa: E402

View File

@ -10,7 +10,9 @@ import sys
import types import types
from datetime import datetime, timedelta from datetime import datetime, timedelta
sys.path.insert(0, str(pathlib.Path(__file__).resolve().parent)) PROJECT_ROOT = pathlib.Path(__file__).resolve().parents[3]
SCRIPT_DIR = PROJECT_ROOT / ".claude" / "skills" / "call-content-model" / "scripts"
sys.path.insert(0, str(SCRIPT_DIR))
import llm # noqa: E402 import llm # noqa: E402
# 费率缓存预置为兜底表:cost_usd/chat_governed 记账时 get_pricing() 直接命中缓存,绝不触网 # 费率缓存预置为兜底表:cost_usd/chat_governed 记账时 get_pricing() 直接命中缓存,绝不触网

View File

@ -3,10 +3,16 @@
from __future__ import annotations from __future__ import annotations
import sys
import tempfile import tempfile
import unittest import unittest
from pathlib import Path from pathlib import Path
from unittest.mock import patch from unittest.mock import patch
PROJECT_ROOT = Path(__file__).resolve().parents[3]
SCRIPT_DIR = PROJECT_ROOT / ".claude" / "skills" / "capture-ai-flavor-cases" / "scripts"
sys.path.insert(0, str(SCRIPT_DIR))
import yaml import yaml
from capture_cases import ( from capture_cases import (
@ -233,20 +239,20 @@ class CaptureCasesTest(unittest.TestCase):
self.assertEqual("narration", sample["carrier"]) self.assertEqual("narration", sample["carrier"])
def test_shipped_fixtures_pass_the_same_validator(self): def test_shipped_fixtures_pass_the_same_validator(self):
root = Path(__file__).resolve().parents[1] / "references" / "fixtures" root = SCRIPT_DIR.parent / "references" / "fixtures"
for name in ("backfill-hash-only.yaml", "canonical-samples.yaml"): for name in ("backfill-hash-only.yaml", "canonical-samples.yaml"):
data = yaml.safe_load((root / name).read_text(encoding="utf-8")) data = yaml.safe_load((root / name).read_text(encoding="utf-8"))
for card in data["cards"]: for card in data["cards"]:
validate_card(card) validate_card(card)
def test_shipped_rule_seed_is_candidate_only(self): def test_shipped_rule_seed_is_candidate_only(self):
root = Path(__file__).resolve().parents[1] / "references" / "fixtures" root = SCRIPT_DIR.parent / "references" / "fixtures"
data = yaml.safe_load((root / "rule-candidates.yaml").read_text(encoding="utf-8")) data = yaml.safe_load((root / "rule-candidates.yaml").read_text(encoding="utf-8"))
self.assertTrue(data["rules"]) self.assertTrue(data["rules"])
self.assertTrue(all(rule["status"] == "candidate" for rule in data["rules"])) self.assertTrue(all(rule["status"] == "candidate" for rule in data["rules"]))
def test_shipped_revalidation_report_is_structurally_usable(self): def test_shipped_revalidation_report_is_structurally_usable(self):
root = Path(__file__).resolve().parents[1] / "references" / "fixtures" root = SCRIPT_DIR.parent / "references" / "fixtures"
data = yaml.safe_load((root / "revalidation-2026-08-14.json").read_text(encoding="utf-8")) data = yaml.safe_load((root / "revalidation-2026-08-14.json").read_text(encoding="utf-8"))
self.assertEqual("ai-flavor-revalidation-v1", data["schema_version"]) self.assertEqual("ai-flavor-revalidation-v1", data["schema_version"])
self.assertTrue(data["usable"]) self.assertTrue(data["usable"])

View File

@ -4,7 +4,7 @@
验证:WriterContext + 候选能被投影成通过闭集校验的 semantic-detector-input-v3; 验证:WriterContext + 候选能被投影成通过闭集校验的 semantic-detector-input-v3;
sourceRef 多余字段被清洗;身份字段严格绑定;哈希自洽。 sourceRef 多余字段被清洗;身份字段严格绑定;哈希自洽。
跑法:.venv/bin/python .claude/skills/check-content-consistency/scripts/test_build_semantic_input.py 跑法:.venv/bin/python tests/skills/check-content-consistency/test_build_semantic_input.py
""" """
from __future__ import annotations from __future__ import annotations
@ -14,12 +14,13 @@ import pathlib
import sys import sys
import unittest import unittest
SCRIPT_DIR = pathlib.Path(__file__).resolve().parent PROJECT_ROOT = pathlib.Path(__file__).resolve().parents[3]
SKILLS_DIR = SCRIPT_DIR.parents[1] SKILLS_DIR = PROJECT_ROOT / ".claude" / "skills"
for path in (SCRIPT_DIR, TEST_DIR = PROJECT_ROOT / "tests" / "skills" / "check-content-consistency"
SKILLS_DIR / "check-content-consistency" / "scripts", SCRIPT_DIR = SKILLS_DIR / "check-content-consistency" / "scripts"
SKILLS_DIR / "write-next-chapter" / "scripts", CONTINUATION_DIR = SKILLS_DIR / "write-next-chapter" / "scripts"
SKILLS_DIR / "assemble-context" / "scripts"): READ_CONTEXT_DIR = SKILLS_DIR / "assemble-context" / "scripts"
for path in (TEST_DIR, SCRIPT_DIR, CONTINUATION_DIR, READ_CONTEXT_DIR):
if str(path) not in sys.path: if str(path) not in sys.path:
sys.path.insert(0, str(path)) sys.path.insert(0, str(path))

View File

@ -9,12 +9,15 @@ import pathlib
import sys import sys
import unittest import unittest
SCRIPT_DIR = pathlib.Path(__file__).resolve().parent PROJECT_ROOT = pathlib.Path(__file__).resolve().parents[3]
SKILLS_DIR = SCRIPT_DIR.parents[1] SKILLS_DIR = PROJECT_ROOT / ".claude" / "skills"
SCRIPT_DIR = SKILLS_DIR / "check-content-consistency" / "scripts"
CONTINUATION_DIR = SKILLS_DIR / "write-next-chapter" / "scripts" CONTINUATION_DIR = SKILLS_DIR / "write-next-chapter" / "scripts"
READ_CONTEXT_DIR = SKILLS_DIR / "assemble-context" / "scripts" READ_CONTEXT_DIR = SKILLS_DIR / "assemble-context" / "scripts"
for path in (SCRIPT_DIR, CONTINUATION_DIR, READ_CONTEXT_DIR): WRITER_TEST_DIR = PROJECT_ROOT / "tests" / "skills" / "write-next-chapter"
sys.path.insert(0, str(path)) for path in (SCRIPT_DIR, CONTINUATION_DIR, READ_CONTEXT_DIR, WRITER_TEST_DIR):
if str(path) not in sys.path:
sys.path.insert(0, str(path))
from test_run_writer import _bound_context # noqa: E402 from test_run_writer import _bound_context # noqa: E402
from check_writer_candidate import check_writer_candidate # noqa: E402 from check_writer_candidate import check_writer_candidate # noqa: E402

View File

@ -10,7 +10,10 @@ import sys
import unittest import unittest
from typing import Any, Mapping, Sequence from typing import Any, Mapping, Sequence
sys.path.insert(0, str(pathlib.Path(__file__).resolve().parent)) PROJECT_ROOT = pathlib.Path(__file__).resolve().parents[3]
SCRIPT_DIR = PROJECT_ROOT / ".claude" / "skills" / "check-content-consistency" / "scripts"
if str(SCRIPT_DIR) not in sys.path:
sys.path.insert(0, str(SCRIPT_DIR))
from run_writer_semantic_detector import ( # noqa: E402 from run_writer_semantic_detector import ( # noqa: E402
SEMANTIC_DETECTOR_REPORT_JSON_SCHEMA, SEMANTIC_DETECTOR_REPORT_JSON_SCHEMA,

View File

@ -16,7 +16,9 @@ from unittest.mock import patch
from click.testing import CliRunner from click.testing import CliRunner
sys.path.insert(0, str(pathlib.Path(__file__).resolve().parent)) PROJECT_ROOT = pathlib.Path(__file__).resolve().parents[3]
SCRIPT_DIR = PROJECT_ROOT / ".claude" / "skills" / "clean-book-text" / "scripts"
sys.path.insert(0, str(SCRIPT_DIR))
import clean_detect # noqa: E402 import clean_detect # noqa: E402

View File

@ -5,7 +5,8 @@ import sys
import unittest import unittest
SCRIPT_DIR = pathlib.Path(__file__).resolve().parent PROJECT_ROOT = pathlib.Path(__file__).resolve().parents[3]
SCRIPT_DIR = PROJECT_ROOT / ".claude" / "skills" / "decide-candidate" / "scripts"
sys.path.insert(0, str(SCRIPT_DIR)) sys.path.insert(0, str(SCRIPT_DIR))
import confirm_knowledge as confirm # noqa: E402 import confirm_knowledge as confirm # noqa: E402

View File

@ -4,7 +4,7 @@
模型只能提类型化增量,证据引文必须真实出现在候选正文中(不得编造证据); 模型只能提类型化增量,证据引文必须真实出现在候选正文中(不得编造证据);
字段闭集、类型闭集、payload 合同、重复 ID 一律失败关闭。 字段闭集、类型闭集、payload 合同、重复 ID 一律失败关闭。
跑法:.venv/bin/python .claude/skills/decide-candidate/scripts/test_fact_delta.py 跑法:.venv/bin/python tests/skills/decide-candidate/test_fact_delta.py
""" """
from __future__ import annotations from __future__ import annotations
@ -14,8 +14,9 @@ import pathlib
import sys import sys
import unittest import unittest
SCRIPT_DIR = pathlib.Path(__file__).resolve().parent PROJECT_ROOT = pathlib.Path(__file__).resolve().parents[3]
SKILLS_DIR = SCRIPT_DIR.parents[1] SKILLS_DIR = PROJECT_ROOT / ".claude" / "skills"
SCRIPT_DIR = SKILLS_DIR / "decide-candidate" / "scripts"
for path in (SCRIPT_DIR, SKILLS_DIR / "assemble-context" / "scripts"): for path in (SCRIPT_DIR, SKILLS_DIR / "assemble-context" / "scripts"):
if str(path) not in sys.path: if str(path) not in sys.path:
sys.path.insert(0, str(path)) sys.path.insert(0, str(path))

View File

@ -11,7 +11,7 @@
测试数据 unittest-delta- 前缀隔离;清理时短暂禁用账本防删触发器(try/finally 恢复)。 测试数据 unittest-delta- 前缀隔离;清理时短暂禁用账本防删触发器(try/finally 恢复)。
跑法(需 Tailscale 内网可达 muse-example): 跑法(需 Tailscale 内网可达 muse-example):
.venv/bin/python .claude/skills/decide-candidate/scripts/test_fact_delta_db.py .venv/bin/python tests/skills/decide-candidate/test_fact_delta_db.py
""" """
from __future__ import annotations from __future__ import annotations
@ -21,8 +21,11 @@ import pathlib
import sys import sys
import uuid import uuid
SCRIPT_DIR = pathlib.Path(__file__).resolve().parent from psycopg.errors import RaiseException
SKILLS_DIR = SCRIPT_DIR.parents[1]
PROJECT_ROOT = pathlib.Path(__file__).resolve().parents[3]
SKILLS_DIR = PROJECT_ROOT / ".claude" / "skills"
SCRIPT_DIR = SKILLS_DIR / "decide-candidate" / "scripts"
for path in (SCRIPT_DIR, SKILLS_DIR / "access-database" / "scripts"): for path in (SCRIPT_DIR, SKILLS_DIR / "access-database" / "scripts"):
if str(path) not in sys.path: if str(path) not in sys.path:
sys.path.insert(0, str(path)) sys.path.insert(0, str(path))
@ -266,7 +269,7 @@ def test_ledger_append_only(work_id: int) -> None:
raise AssertionError("账本必须 append-only") raise AssertionError("账本必须 append-only")
except AssertionError: except AssertionError:
raise raise
except Exception: except RaiseException:
pass pass
try: try:
with connect() as conn: with connect() as conn:
@ -275,7 +278,7 @@ def test_ledger_append_only(work_id: int) -> None:
raise AssertionError("账本必须 append-only") raise AssertionError("账本必须 append-only")
except AssertionError: except AssertionError:
raise raise
except Exception: except RaiseException:
pass pass

View File

@ -10,7 +10,7 @@
测试数据 unittest-proj- 前缀隔离,结束物理清理(本表可变,直接 DELETE)。 测试数据 unittest-proj- 前缀隔离,结束物理清理(本表可变,直接 DELETE)。
跑法(需 Tailscale 内网可达 muse-example): 跑法(需 Tailscale 内网可达 muse-example):
.venv/bin/python .claude/skills/decide-candidate/scripts/test_projection_db.py .venv/bin/python tests/skills/decide-candidate/test_projection_db.py
""" """
from __future__ import annotations from __future__ import annotations
@ -20,8 +20,11 @@ import pathlib
import sys import sys
import uuid import uuid
SCRIPT_DIR = pathlib.Path(__file__).resolve().parent from psycopg.errors import RaiseException
SKILLS_DIR = SCRIPT_DIR.parents[1]
PROJECT_ROOT = pathlib.Path(__file__).resolve().parents[3]
SKILLS_DIR = PROJECT_ROOT / ".claude" / "skills"
SCRIPT_DIR = SKILLS_DIR / "decide-candidate" / "scripts"
for path in (SCRIPT_DIR, SKILLS_DIR / "access-database" / "scripts"): for path in (SCRIPT_DIR, SKILLS_DIR / "access-database" / "scripts"):
if str(path) not in sys.path: if str(path) not in sys.path:
sys.path.insert(0, str(path)) sys.path.insert(0, str(path))
@ -175,7 +178,7 @@ def test_failed_never_masquerades_completed(work_id: int) -> None:
raise AssertionError("触发器必须拒绝 failed→completed") raise AssertionError("触发器必须拒绝 failed→completed")
except AssertionError: except AssertionError:
raise raise
except Exception: except RaiseException:
pass pass
# 显式 retry 才能回 pending,且 attempt+1 # 显式 retry 才能回 pending,且 attempt+1
retried = retry_projection(pending[0]) retried = retry_projection(pending[0])
@ -202,7 +205,7 @@ def test_stale_projections_cannot_report_outcomes(work_id: int) -> None:
raise AssertionError("stale→completed 必须被拒绝") raise AssertionError("stale→completed 必须被拒绝")
except AssertionError: except AssertionError:
raise raise
except Exception: except RaiseException:
pass pass
try: try:
with connect() as conn: with connect() as conn:
@ -212,7 +215,7 @@ def test_stale_projections_cannot_report_outcomes(work_id: int) -> None:
raise AssertionError("stale→failed 必须被拒绝") raise AssertionError("stale→failed 必须被拒绝")
except AssertionError: except AssertionError:
raise raise
except Exception: except RaiseException:
pass pass
# stale 走 retry 恢复:attempt+1 回 pending,随后可以正常完成 # stale 走 retry 恢复:attempt+1 回 pending,随后可以正常完成
retried = retry_projection(stale_row[0]) retried = retry_projection(stale_row[0])

View File

@ -12,7 +12,7 @@
清理时短暂禁用其防删触发器(try/finally 保证恢复)。 清理时短暂禁用其防删触发器(try/finally 保证恢复)。
跑法(需 Tailscale 内网可达 muse-example): 跑法(需 Tailscale 内网可达 muse-example):
.venv/bin/python .claude/skills/decide-candidate/scripts/test_write_canonical_db.py .venv/bin/python tests/skills/decide-candidate/test_write_canonical_db.py
""" """
from __future__ import annotations from __future__ import annotations
@ -24,8 +24,9 @@ import uuid
from typing import Any from typing import Any
from unittest import mock from unittest import mock
SCRIPT_DIR = pathlib.Path(__file__).resolve().parent PROJECT_ROOT = pathlib.Path(__file__).resolve().parents[3]
SKILLS_DIR = SCRIPT_DIR.parents[1] SKILLS_DIR = PROJECT_ROOT / ".claude" / "skills"
SCRIPT_DIR = SKILLS_DIR / "decide-candidate" / "scripts"
for path in (SCRIPT_DIR, SKILLS_DIR / "access-database" / "scripts"): for path in (SCRIPT_DIR, SKILLS_DIR / "access-database" / "scripts"):
if str(path) not in sys.path: if str(path) not in sys.path:
sys.path.insert(0, str(path)) sys.path.insert(0, str(path))

View File

@ -14,8 +14,9 @@ import pathlib
import sys import sys
import unittest import unittest
SCRIPT_DIR = pathlib.Path(__file__).resolve().parent PROJECT_ROOT = pathlib.Path(__file__).resolve().parents[3]
SKILLS_DIR = SCRIPT_DIR.parents[1] SKILLS_DIR = PROJECT_ROOT / ".claude" / "skills"
SCRIPT_DIR = SKILLS_DIR / "decide-candidate" / "scripts"
CONTINUATION_DIR = SKILLS_DIR / "write-next-chapter" / "scripts" CONTINUATION_DIR = SKILLS_DIR / "write-next-chapter" / "scripts"
READ_CONTEXT_DIR = SKILLS_DIR / "assemble-context" / "scripts" READ_CONTEXT_DIR = SKILLS_DIR / "assemble-context" / "scripts"
for path in (SCRIPT_DIR, CONTINUATION_DIR, READ_CONTEXT_DIR): for path in (SCRIPT_DIR, CONTINUATION_DIR, READ_CONTEXT_DIR):

View File

@ -11,7 +11,9 @@ import unittest
from unittest.mock import Mock, patch from unittest.mock import Mock, patch
sys.path.insert(0, str(pathlib.Path(__file__).resolve().parent)) PROJECT_ROOT = pathlib.Path(__file__).resolve().parents[3]
SCRIPT_DIR = PROJECT_ROOT / ".claude" / "skills" / "deconstruct-book" / "scripts"
sys.path.insert(0, str(SCRIPT_DIR))
import parse_llm as pll # noqa: E402 import parse_llm as pll # noqa: E402

View File

@ -6,8 +6,9 @@ import sys
import click import click
HERE = pathlib.Path(__file__).resolve().parent PROJECT_ROOT = pathlib.Path(__file__).resolve().parents[3]
sys.path.insert(0, str(HERE)) SCRIPT_DIR = PROJECT_ROOT / ".claude" / "skills" / "deconstruct-book" / "scripts"
sys.path.insert(0, str(SCRIPT_DIR))
import parse_outline as po # noqa: E402 import parse_outline as po # noqa: E402

View File

@ -7,7 +7,9 @@ import tempfile
import unittest import unittest
from unittest.mock import patch from unittest.mock import patch
sys.path.insert(0, str(pathlib.Path(__file__).resolve().parent)) PROJECT_ROOT = pathlib.Path(__file__).resolve().parents[3]
SCRIPT_DIR = PROJECT_ROOT / ".claude" / "skills" / "diagnose-ai-flavor" / "scripts"
sys.path.insert(0, str(SCRIPT_DIR))
import diagnose_ai_flavor as diag # noqa: E402 import diagnose_ai_flavor as diag # noqa: E402

View File

@ -11,7 +11,8 @@ from unittest.mock import MagicMock, Mock, patch
from click.testing import CliRunner from click.testing import CliRunner
SCRIPT_DIR = pathlib.Path(__file__).resolve().parent PROJECT_ROOT = pathlib.Path(__file__).resolve().parents[3]
SCRIPT_DIR = PROJECT_ROOT / ".claude" / "skills" / "embed-knowledge" / "scripts"
sys.path.insert(0, str(SCRIPT_DIR)) sys.path.insert(0, str(SCRIPT_DIR))
import embed_drafts as embed # noqa: E402 import embed_drafts as embed # noqa: E402

View File

@ -6,7 +6,9 @@ import tempfile
import unittest import unittest
from unittest.mock import patch from unittest.mock import patch
sys.path.insert(0, str(pathlib.Path(__file__).resolve().parent)) PROJECT_ROOT = pathlib.Path(__file__).resolve().parents[3]
SCRIPT_DIR = PROJECT_ROOT / ".claude" / "skills" / "establish-voice-baseline" / "scripts"
sys.path.insert(0, str(SCRIPT_DIR))
import establish_voice_baseline as base # noqa: E402 import establish_voice_baseline as base # noqa: E402

View File

@ -7,7 +7,11 @@ import pathlib
import sys import sys
import unittest import unittest
sys.path.insert(0, str(pathlib.Path(__file__).resolve().parent)) PROJECT_ROOT = pathlib.Path(__file__).resolve().parents[3]
SCRIPT_DIR = PROJECT_ROOT / ".claude" / "skills" / "evaluate-frozen-replay" / "scripts"
if str(SCRIPT_DIR) not in sys.path:
sys.path.insert(0, str(SCRIPT_DIR))
from fine_outline_detector import validate_detector_report # noqa: E402 from fine_outline_detector import validate_detector_report # noqa: E402
@ -76,7 +80,9 @@ class FineOutlineDetectorTest(unittest.TestCase):
self.assertTrue(validate_detector_report(report, "blind-1")["ok"]) self.assertTrue(validate_detector_report(report, "blind-1")["ok"])
skill_path = ( skill_path = (
pathlib.Path(__file__).resolve().parents[2] PROJECT_ROOT
/ ".claude"
/ "skills"
/ "check-content-consistency" / "check-content-consistency"
/ "SKILL.md" / "SKILL.md"
) )

View File

@ -8,9 +8,12 @@ import pathlib
import sys import sys
import unittest import unittest
SCRIPT_DIR = pathlib.Path(__file__).resolve().parent PROJECT_ROOT = pathlib.Path(__file__).resolve().parents[3]
QUALITY_GATE_DIR = SCRIPT_DIR.parents[1] / "score-content-quality" / "scripts" SKILLS_DIR = PROJECT_ROOT / ".claude" / "skills"
for _import_dir in (SCRIPT_DIR, QUALITY_GATE_DIR): SCRIPT_DIR = SKILLS_DIR / "evaluate-frozen-replay" / "scripts"
TEST_DIR = PROJECT_ROOT / "tests" / "skills" / "evaluate-frozen-replay"
QUALITY_GATE_DIR = SKILLS_DIR / "score-content-quality" / "scripts"
for _import_dir in (SCRIPT_DIR, TEST_DIR, QUALITY_GATE_DIR):
if str(_import_dir) not in sys.path: if str(_import_dir) not in sys.path:
sys.path.insert(0, str(_import_dir)) sys.path.insert(0, str(_import_dir))

View File

@ -14,12 +14,24 @@ import unittest
from datetime import datetime, timedelta, timezone from datetime import datetime, timedelta, timezone
from unittest.mock import patch from unittest.mock import patch
SCRIPT_DIR = pathlib.Path(__file__).resolve().parent PROJECT_ROOT = pathlib.Path(__file__).resolve().parents[3]
READ_CONTEXT_SCRIPTS = SCRIPT_DIR.parents[1] / "assemble-context" / "scripts" SKILLS_DIR = PROJECT_ROOT / ".claude" / "skills"
DETECT_SCRIPTS = SCRIPT_DIR.parents[1] / "check-content-consistency" / "scripts" SCRIPT_DIR = SKILLS_DIR / "evaluate-frozen-replay" / "scripts"
CONTINUATION_SCRIPTS = SCRIPT_DIR.parents[1] / "write-next-chapter" / "scripts" TEST_DIR = PROJECT_ROOT / "tests" / "skills" / "evaluate-frozen-replay"
for _path in (SCRIPT_DIR, READ_CONTEXT_SCRIPTS, DETECT_SCRIPTS, CONTINUATION_SCRIPTS): READ_CONTEXT_SCRIPTS = SKILLS_DIR / "assemble-context" / "scripts"
sys.path.insert(0, str(_path)) DETECT_SCRIPTS = SKILLS_DIR / "check-content-consistency" / "scripts"
CONTINUATION_SCRIPTS = SKILLS_DIR / "write-next-chapter" / "scripts"
CONTINUATION_TESTS = PROJECT_ROOT / "tests" / "skills" / "write-next-chapter"
for _path in (
SCRIPT_DIR,
TEST_DIR,
READ_CONTEXT_SCRIPTS,
DETECT_SCRIPTS,
CONTINUATION_SCRIPTS,
CONTINUATION_TESTS,
):
if str(_path) not in sys.path:
sys.path.insert(0, str(_path))
import load_writer_reference_work as loader # noqa: E402 import load_writer_reference_work as loader # noqa: E402
import run_writer_replay as replay_module # noqa: E402 import run_writer_replay as replay_module # noqa: E402

View File

@ -13,8 +13,11 @@ import pathlib
import sys import sys
import unittest import unittest
SCRIPT_DIR = pathlib.Path(__file__).resolve().parent PROJECT_ROOT = pathlib.Path(__file__).resolve().parents[3]
SKILLS_DIR = SCRIPT_DIR.parents[1] SKILLS_DIR = PROJECT_ROOT / ".claude" / "skills"
SCRIPT_DIR = SKILLS_DIR / "evaluate-frozen-replay" / "scripts"
TEST_DIR = PROJECT_ROOT / "tests" / "skills" / "evaluate-frozen-replay"
ASSEMBLE_CONTEXT_TESTS = PROJECT_ROOT / "tests" / "skills" / "assemble-context"
READ_CONTEXT_SCRIPTS = SKILLS_DIR / "assemble-context" / "scripts" READ_CONTEXT_SCRIPTS = SKILLS_DIR / "assemble-context" / "scripts"
# 端到端链路测试要导入回放包(run_writer_replay.sample),其依赖执行、证据与 # 端到端链路测试要导入回放包(run_writer_replay.sample),其依赖执行、证据与
# 评分三个 Skill 的 scripts 目录,路径口径与 test_run_writer_replay 保持一致。 # 评分三个 Skill 的 scripts 目录,路径口径与 test_run_writer_replay 保持一致。
@ -23,12 +26,15 @@ EVIDENCE_SCRIPTS = SKILLS_DIR / "record-run-evidence" / "scripts"
QUALITY_GATE_SCRIPTS = SKILLS_DIR / "score-content-quality" / "scripts" QUALITY_GATE_SCRIPTS = SKILLS_DIR / "score-content-quality" / "scripts"
for _path in ( for _path in (
SCRIPT_DIR, SCRIPT_DIR,
TEST_DIR,
ASSEMBLE_CONTEXT_TESTS,
READ_CONTEXT_SCRIPTS, READ_CONTEXT_SCRIPTS,
EXECUTION_SCRIPTS, EXECUTION_SCRIPTS,
EVIDENCE_SCRIPTS, EVIDENCE_SCRIPTS,
QUALITY_GATE_SCRIPTS, QUALITY_GATE_SCRIPTS,
): ):
sys.path.insert(0, str(_path)) if str(_path) not in sys.path:
sys.path.insert(0, str(_path))
import load_writer_reference_work as loader # noqa: E402 import load_writer_reference_work as loader # noqa: E402
from load_writer_reference_work import ( # noqa: E402 from load_writer_reference_work import ( # noqa: E402

View File

@ -21,15 +21,14 @@ import unittest
from typing import Any, Mapping from typing import Any, Mapping
from unittest import mock from unittest import mock
SCRIPT_DIR = pathlib.Path(__file__).resolve().parent PROJECT_ROOT = pathlib.Path(__file__).resolve().parents[3]
EXECUTION_DIR = SCRIPT_DIR.parents[1] / "execute-claude-task" / "scripts" SKILLS_DIR = PROJECT_ROOT / ".claude" / "skills"
QUALITY_GATE_DIR = SCRIPT_DIR.parents[1] / "score-content-quality" / "scripts" SCRIPT_DIR = SKILLS_DIR / "evaluate-frozen-replay" / "scripts"
if str(SCRIPT_DIR) not in sys.path: EXECUTION_DIR = SKILLS_DIR / "execute-claude-task" / "scripts"
sys.path.insert(0, str(SCRIPT_DIR)) QUALITY_GATE_DIR = SKILLS_DIR / "score-content-quality" / "scripts"
if str(EXECUTION_DIR) not in sys.path: for _import_dir in (SCRIPT_DIR, EXECUTION_DIR, QUALITY_GATE_DIR):
sys.path.insert(0, str(EXECUTION_DIR)) if str(_import_dir) not in sys.path:
if str(QUALITY_GATE_DIR) not in sys.path: sys.path.insert(0, str(_import_dir))
sys.path.insert(0, str(QUALITY_GATE_DIR))
import refresh_runtime_probe as refresh_module # noqa: E402 import refresh_runtime_probe as refresh_module # noqa: E402
from refresh_runtime_probe import ( # noqa: E402 from refresh_runtime_probe import ( # noqa: E402

View File

@ -11,10 +11,14 @@ import tempfile
import unittest import unittest
from unittest.mock import patch from unittest.mock import patch
SCRIPT_DIR = pathlib.Path(__file__).resolve().parent PROJECT_ROOT = pathlib.Path(__file__).resolve().parents[3]
QUALITY_GATE_DIR = SCRIPT_DIR.parents[1] / "score-content-quality" / "scripts" SKILLS_DIR = PROJECT_ROOT / ".claude" / "skills"
sys.path.insert(0, str(SCRIPT_DIR)) SCRIPT_DIR = SKILLS_DIR / "evaluate-frozen-replay" / "scripts"
sys.path.insert(0, str(QUALITY_GATE_DIR)) TEST_DIR = PROJECT_ROOT / "tests" / "skills" / "evaluate-frozen-replay"
QUALITY_GATE_DIR = SKILLS_DIR / "score-content-quality" / "scripts"
for _import_dir in (SCRIPT_DIR, TEST_DIR, QUALITY_GATE_DIR):
if str(_import_dir) not in sys.path:
sys.path.insert(0, str(_import_dir))
from run_replay import _parse_args, _planner_prompt, run_replay as _run_replay # noqa: E402 from run_replay import _parse_args, _planner_prompt, run_replay as _run_replay # noqa: E402
from test_writer_gate import build_input, source_bundle # noqa: E402 from test_writer_gate import build_input, source_bundle # noqa: E402
from writer_gate import issue_gate_report_and_receipt # noqa: E402 from writer_gate import issue_gate_report_and_receipt # noqa: E402

View File

@ -19,8 +19,10 @@ from dataclasses import replace
from decimal import Decimal from decimal import Decimal
from datetime import datetime, timedelta, timezone from datetime import datetime, timedelta, timezone
SCRIPT_DIR = pathlib.Path(__file__).resolve().parent TEST_DIR = pathlib.Path(__file__).resolve().parent
SKILLS_DIR = SCRIPT_DIR.parents[1] PROJECT_ROOT = pathlib.Path(__file__).resolve().parents[3]
SKILLS_DIR = PROJECT_ROOT / ".claude" / "skills"
SCRIPT_DIR = SKILLS_DIR / "evaluate-frozen-replay" / "scripts"
READ_CONTEXT_DIR = SKILLS_DIR / "assemble-context" / "scripts" READ_CONTEXT_DIR = SKILLS_DIR / "assemble-context" / "scripts"
EXECUTION_DIR = SKILLS_DIR / "execute-claude-task" / "scripts" EXECUTION_DIR = SKILLS_DIR / "execute-claude-task" / "scripts"
EVIDENCE_DIR = SKILLS_DIR / "record-run-evidence" / "scripts" EVIDENCE_DIR = SKILLS_DIR / "record-run-evidence" / "scripts"
@ -32,7 +34,8 @@ for import_path in (
EVIDENCE_DIR, EVIDENCE_DIR,
QUALITY_GATE_DIR, QUALITY_GATE_DIR,
): ):
sys.path.insert(0, str(import_path)) if str(import_path) not in sys.path:
sys.path.insert(0, str(import_path))
import run_writer_replay as replay_module # noqa: E402 import run_writer_replay as replay_module # noqa: E402
import run_writer_replay.execute as execute_module # noqa: E402 import run_writer_replay.execute as execute_module # noqa: E402

View File

@ -8,7 +8,10 @@ import pathlib
import sys import sys
import unittest import unittest
sys.path.insert(0, str(pathlib.Path(__file__).resolve().parent)) PROJECT_ROOT = pathlib.Path(__file__).resolve().parents[3]
SCRIPT_DIR = PROJECT_ROOT / ".claude" / "skills" / "evaluate-frozen-replay" / "scripts"
if str(SCRIPT_DIR) not in sys.path:
sys.path.insert(0, str(SCRIPT_DIR))
from writer_eval_preregister import ( # noqa: E402 from writer_eval_preregister import ( # noqa: E402
PreregistrationError, PreregistrationError,

View File

@ -10,10 +10,13 @@ import sys
import tempfile import tempfile
import unittest import unittest
SCRIPT_DIR = pathlib.Path(__file__).resolve().parent PROJECT_ROOT = pathlib.Path(__file__).resolve().parents[3]
QUALITY_GATE_DIR = SCRIPT_DIR.parents[1] / "score-content-quality" / "scripts" SKILLS_DIR = PROJECT_ROOT / ".claude" / "skills"
EVIDENCE_DIR = SCRIPT_DIR.parents[1] / "record-run-evidence" / "scripts" SCRIPT_DIR = SKILLS_DIR / "evaluate-frozen-replay" / "scripts"
for _import_dir in (SCRIPT_DIR, QUALITY_GATE_DIR, EVIDENCE_DIR): TEST_DIR = PROJECT_ROOT / "tests" / "skills" / "evaluate-frozen-replay"
QUALITY_GATE_DIR = SKILLS_DIR / "score-content-quality" / "scripts"
EVIDENCE_DIR = SKILLS_DIR / "record-run-evidence" / "scripts"
for _import_dir in (SCRIPT_DIR, TEST_DIR, QUALITY_GATE_DIR, EVIDENCE_DIR):
if str(_import_dir) not in sys.path: if str(_import_dir) not in sys.path:
sys.path.insert(0, str(_import_dir)) sys.path.insert(0, str(_import_dir))

View File

@ -17,7 +17,8 @@ import unittest
from unittest import mock from unittest import mock
from decimal import Decimal from decimal import Decimal
SCRIPT_DIR = pathlib.Path(__file__).resolve().parent PROJECT_ROOT = pathlib.Path(__file__).resolve().parents[3]
SCRIPT_DIR = PROJECT_ROOT / ".claude" / "skills" / "execute-claude-task" / "scripts"
sys.path.insert(0, str(SCRIPT_DIR)) sys.path.insert(0, str(SCRIPT_DIR))
from claude_runtime import ( # noqa: E402 from claude_runtime import ( # noqa: E402

View File

@ -3,7 +3,9 @@
import pathlib import pathlib
import sys import sys
sys.path.insert(0, str(pathlib.Path(__file__).resolve().parent)) PROJECT_ROOT = pathlib.Path(__file__).resolve().parents[3]
SCRIPT_DIR = PROJECT_ROOT / ".claude" / "skills" / "extract-chapter-knowledge" / "scripts"
sys.path.insert(0, str(SCRIPT_DIR))
from extract_knowledge import ExtractionContractError, normalize_extraction, salvage_extraction # noqa: E402 from extract_knowledge import ExtractionContractError, normalize_extraction, salvage_extraction # noqa: E402

View File

@ -13,13 +13,12 @@
本测试**不触碰真实服务**——质量修复命令只通过 fake DB 验证 preview、CAS 和原子写入, 本测试**不触碰真实服务**——质量修复命令只通过 fake DB 验证 preview、CAS 和原子写入,
不在离线自测中执行真实 work/window。 不在离线自测中执行真实 work/window。
跑法:仓根 `.venv/bin/python .claude/skills/extract-work-knowledge/scripts/test_parse_upgrade_offline.py` 跑法:仓库根目录 `.venv/bin/python tests/skills/extract-work-knowledge/test_parse_upgrade_offline.py`
""" """
import json import json
import hashlib import hashlib
import inspect import inspect
from io import StringIO from io import StringIO
import os
import pathlib import pathlib
import sys import sys
from contextlib import nullcontext from contextlib import nullcontext
@ -29,8 +28,10 @@ from unittest.mock import patch
from click.testing import CliRunner from click.testing import CliRunner
# 与 upgrade 同目录:直接 import 触发其 sys.path 装配(含 parse-book/embed/llm scripts),随后可导 embed_drafts # 从生产 scripts 导入 upgrade;upgrade 自己装配 parse-book/embed/llm 的生产模块路径。
sys.path.insert(0, str(pathlib.Path(__file__).resolve().parent)) PROJECT_ROOT = pathlib.Path(__file__).resolve().parents[3]
SCRIPT_DIR = PROJECT_ROOT / ".claude" / "skills" / "extract-work-knowledge" / "scripts"
sys.path.insert(0, str(SCRIPT_DIR))
import upgrade as pu # noqa: E402 import upgrade as pu # noqa: E402
import embed_drafts # noqa: E402 (型修正验证在 embed skill 本体) import embed_drafts # noqa: E402 (型修正验证在 embed skill 本体)
@ -4724,17 +4725,6 @@ def test_recovery_and_redo_rejection_contracts():
check("redo-run首行机械拒绝", False, detail="_run 接受了 redo") check("redo-run首行机械拒绝", False, detail="_run 接受了 redo")
def test_real_pg_rollback_smoke_entry():
"""真实 PG 入口默认不运行;显式开启后验证 public 窗状态写入可回滚。"""
if os.getenv("MUSE_REAL_PG_ROLLBACK_SMOKE") != "1":
check("real-pg-rollback-smoke默认跳过", True)
return
work_id = int(os.getenv("MUSE_REAL_PG_ROLLBACK_SMOKE_WORK_ID", "8"))
pu.real_pg_rollback_smoke(work_id)
check("real-pg-rollback-smoke真实public窗回滚", True)
class _LegacyRecoveryConn: class _LegacyRecoveryConn:
"""恢复命令离线夹具:用安全墓碑及异常软删验证 SQL 过滤。""" """恢复命令离线夹具:用安全墓碑及异常软删验证 SQL 过滤。"""
@ -5847,538 +5837,7 @@ def test_quality_repair_rejects_alias_and_presence_duplicates():
) )
# ── presence 冗余收口(work8 十组双行:保留 MIN(id) 软删 MAX(id))──
# 离线 fixture 用字符串时间戳代替数据库驱动的 datetime:窄合同只比较同组两行是否相等,
# _sha256_json 计算摘要时统一 default=str 转写,二者行为一致。
_PRESENCE_DEDUPE_BASE_TIME = "2026-07-01 12:00:00"
def _presence_dedupe_rows():
"""构造满足窄合同的 work8 presence 行 fixture。
十组双行:每组 observation 不同、create_time 相同、恰好一个待删 ID 命中 DELETE_IDS、
保留 ID 是组内较小者;另有四个单行,证明快照会跳过不构成冗余的 key。
"""
rows = []
for index, delete_id in enumerate(sorted(pu.PRESENCE_DEDUPE_DELETE_IDS)):
keep_id = 11501 + index
entity_type = "location" if index % 2 == 0 else "item"
name = f"重复实体{index}"
for row_id, observation in ((keep_id, f"观察A{index}"),
(delete_id, f"观察B{index}")):
rows.append((row_id, 8, 57 + index, 403 + index, entity_type, name,
observation, "upgrade", _PRESENCE_DEDUPE_BASE_TIME,
False, pu.TENANT))
for index in range(4):
rows.append((11400 + index, 8, 10 + index, 20 + index, "character",
f"单体实体{index}", f"单体观察{index}",
"upgrade", _PRESENCE_DEDUPE_BASE_TIME, False, pu.TENANT))
return sorted(rows, key=lambda row: row[0])
def _presence_dedupe_state():
return {
"rows": _presence_dedupe_rows(),
"writes": [],
"queries": [],
"commits": 0,
}
def _mutate_presence_row(state, row_id, column, value):
"""替换指定行的单列;元组不可变,整体重建后写回。"""
state["rows"] = [
row[:column] + (value,) + row[column + 1:] if row[0] == row_id else row
for row in state["rows"]
]
class _PresenceDedupeConn:
"""presence 冗余收口 fake DB:只实现本命令的读快照与软删 SQL。
其他任何域(draft/window/alias/card_state/audit/embedding)的 SQL 会落到末尾
AssertionError,因此「零副作用」无需逐条枚举禁写语句即可离线断言。fake 只做
「未软删」粗过滤;列数/work/tenant/deleted 类型等窄合同行像校验正是离线断言对象,
fake 不能代为过滤。
"""
def __init__(self, state):
self.state = state
def __enter__(self):
return self
def __exit__(self, exc_type, exc, tb):
return False
def execute(self, query, params=()):
normalized = " ".join(query.split())
self.state.setdefault("queries", []).append(normalized)
if normalized.startswith("SET TRANSACTION") or normalized.startswith("LOCK TABLE"):
return _CardResult()
if normalized.startswith(
"SELECT id, work_id, window_no, chapter_no, entity_type, name, observation,"):
rows = [
row for row in sorted(self.state["rows"], key=lambda item: item[0])
if len(row) > 9 and not row[9]
]
return _CardResult(rows=rows)
if normalized.startswith("UPDATE example_upgrade_presence SET deleted=TRUE"):
tenant_id, work_id, delete_ids = params
assert tenant_id == pu.TENANT
delete_set = set(delete_ids)
hit = []
new_rows = []
for row in self.state["rows"]:
if row[1] == work_id and row[0] in delete_set and not row[9]:
row = row[:9] + (True,) + row[10:]
hit.append((row[0],))
new_rows.append(row)
self.state["rows"] = new_rows
self.state["writes"].append("UPDATE example_upgrade_presence")
return _CardResult(rows=hit)
raise AssertionError(f"presence 冗余收口 fake DB 未覆盖 SQL:{normalized}")
def commit(self):
self.state["commits"] += 1
def _presence_dedupe_cli(state, args):
"""离线执行 repair-presence-duplicates:注入 fake DB 与空放同书锁。"""
conn = _PresenceDedupeConn(state)
with patch.object(pu.psycopg, "connect", return_value=conn), \
patch.object(pu, "upgrade_work_lock", return_value=nullcontext()):
result = CliRunner().invoke(pu.maintenance_cli, ["repair-presence-duplicates"] + args)
return result, conn
def _presence_dedupe_preview_sha(state):
"""离线取 preview 的 confirmation_sha,作为 execute 的合法输入。"""
result, _ = _presence_dedupe_cli(state, ["--work-id", "8", "--preview"])
assert result.exit_code == 0, result.output
return json.loads(result.output)["confirmation_sha"]
def _presence_snapshot_raises(label, state, work_id=pu.PRESENCE_DEDUPE_WORK_ID):
"""断言快照以 CompensationFenceConflict 拒绝,且没有发出任何写入。"""
conn = _PresenceDedupeConn(state)
try:
pu._presence_dedupe_capture_snapshot(conn, work_id)
except pu.CompensationFenceConflict as exc:
check(f"presence-dedupe-{label}失败关闭", True, detail=str(exc))
else:
check(f"presence-dedupe-{label}失败关闭", False,
detail="未抛 CompensationFenceConflict")
check(
f"presence-dedupe-{label}零副作用",
state["writes"] == [] and state["commits"] == 0,
)
def test_presence_dedupe_actions_planner_guard():
"""planner 级收口守卫:同观察幂等留首条、不同观察失败关闭、非 presence 原样透传。"""
presence_a = ("presence", None, 403, "location", "重复星体", "观察一")
presence_a_dup = ("presence", None, 403, "location", "重复星体", "观察一")
presence_b = ("presence", None, 404, "item", "镜面护盾", "观察二")
alias_action = ("alias", 7, "规范名", "别名", "ai")
new_action = ("new", -1, {"名称": "新实体"}, set(), [])
actions = [presence_a, alias_action, presence_a_dup, new_action, presence_b, ()]
original = list(actions)
deduped = pu._dedupe_presence_actions(actions)
check(
"presence-dedupe-planner同观察去重保留首条",
deduped == [presence_a, alias_action, new_action, presence_b, ()]
and deduped[0] is presence_a,
detail=str(deduped),
)
check("presence-dedupe-重复应用幂等",
pu._dedupe_presence_actions(deduped) == deduped)
check("presence-dedupe-不修改输入列表", actions == original)
conflict = [presence_a,
("presence", None, 403, "location", "重复星体", "另一个观察")]
try:
pu._dedupe_presence_actions(conflict)
except RuntimeError as exc:
check("presence-dedupe-不同观察失败关闭",
"observation 冲突" in str(exc), detail=str(exc))
else:
check("presence-dedupe-不同观察失败关闭", False, detail="未抛 RuntimeError")
for label, bad, keyword in (
("结构缺一元", ("presence", None, 403, "location", "重复星体"),
"结构非法"),
("结构多一元",
("presence", None, 403, "location", "重复星体", "观察一", "extra"),
"结构非法"),
("observation非字符串", ("presence", None, 403, "location", "重复星体", None),
"必须是字符串"),
):
try:
pu._dedupe_presence_actions([bad])
except RuntimeError as exc:
check(f"presence-dedupe-{label}抛错", keyword in str(exc), detail=str(exc))
else:
check(f"presence-dedupe-{label}抛错", False, detail="未抛 RuntimeError")
def test_presence_dedupe_snapshot_success():
"""快照成功路径:恰好 10 组/20 行,待删集合精确等于 DELETE_IDS,保留为组内较小者。"""
state = _presence_dedupe_state()
snapshot = pu._presence_dedupe_capture_snapshot(
_PresenceDedupeConn(state), pu.PRESENCE_DEDUPE_WORK_ID,
)
check(
"presence-dedupe-快照组数行数精确",
len(snapshot["groups"]) == pu.PRESENCE_DEDUPE_GROUP_COUNT
and len(snapshot["rows"]) == pu.PRESENCE_DEDUPE_ROW_COUNT
and snapshot["contract"] == pu.PRESENCE_DEDUPE_CONTRACT
and snapshot["work_id"] == pu.PRESENCE_DEDUPE_WORK_ID,
)
check(
"presence-dedupe-待删集合精确等于DELETE_IDS",
snapshot["delete_ids"] == sorted(pu.PRESENCE_DEDUPE_DELETE_IDS),
detail=str(snapshot["delete_ids"]),
)
check(
"presence-dedupe-保留为组内较小ID",
snapshot["keep_ids"] == list(range(11501, 11511))
and all(group["keep_id"] < group["delete_id"] for group in snapshot["groups"])
and all(len(group["rows"]) == 2 for group in snapshot["groups"]),
detail=str(snapshot["keep_ids"]),
)
check(
"presence-dedupe-行按id排序且排除单体行",
[row[0] for row in snapshot["rows"]]
== sorted(row[0] for row in snapshot["rows"])
and not any(row[0] in range(11400, 11404) for row in snapshot["rows"]),
)
def test_presence_dedupe_snapshot_fail_closed():
"""各类窄合同漂移都必须在写入前以 CompensationFenceConflict 逐条拒绝。"""
# 额外冗余组:组数 11 ≠ 10(新增组复用合同内待删 ID,确保失败点落在组数校验)。
state = _presence_dedupe_state()
state["rows"].extend((
(11581, 8, 90, 500, "event", "额外重复组", "观察C",
"upgrade", _PRESENCE_DEDUPE_BASE_TIME, False, pu.TENANT),
(11582, 8, 90, 500, "event", "额外重复组", "观察D",
"upgrade", _PRESENCE_DEDUPE_BASE_TIME, False, pu.TENANT),
))
_presence_snapshot_raises("额外冗余组", state)
# 三行组不是「恰好两行」。
state = _presence_dedupe_state()
state["rows"].append(
(11601, 8, 57, 403, "location", "重复实体0", "观察E",
"upgrade", _PRESENCE_DEDUPE_BASE_TIME, False, pu.TENANT)
)
_presence_snapshot_raises("三行冗余组", state)
# 缺行:一组只剩单行被跳过 → 组数 9 ≠ 10。
state = _presence_dedupe_state()
state["rows"] = [row for row in state["rows"]
if row[0] != sorted(pu.PRESENCE_DEDUPE_DELETE_IDS)[0]]
_presence_snapshot_raises("缺行", state)
# 同组两行 create_time 漂移。
state = _presence_dedupe_state()
_mutate_presence_row(state, 11501, 8, "2026-07-02 09:30:00")
_presence_snapshot_raises("create_time漂移", state)
# 待删 ID 命中 0:组内两行都不在 DELETE_IDS。
state = _presence_dedupe_state()
_mutate_presence_row(state, 11501, 0, 12001)
_mutate_presence_row(state, 11582, 0, 12002)
_presence_snapshot_raises("待删ID无命中", state)
# 待删 ID 命中 2:组内两行都在 DELETE_IDS。
state = _presence_dedupe_state()
_mutate_presence_row(state, 11501, 0, 11693)
_presence_snapshot_raises("待删ID双命中", state)
# 保留不是组内较小者:待删 ID 反而是组内较小者。
state = _presence_dedupe_state()
_mutate_presence_row(state, 11501, 0, 11900)
_presence_snapshot_raises("保留非组内较小", state)
# 同组两行 observation 相同(byte-exact 重复违反窄合同)。
state = _presence_dedupe_state()
_mutate_presence_row(state, 11582, 6, "观察A0")
_presence_snapshot_raises("observation相同", state)
# work_id ≠ 8 拒绝。
_presence_snapshot_raises("work不符", _presence_dedupe_state(), work_id=9)
# 窄合同行像漂移:列数不足。
state = _presence_dedupe_state()
state["rows"] = [row[:10] if row[0] == 11501 else row for row in state["rows"]]
_presence_snapshot_raises("列数不符", state)
# 窄合同行像漂移:work 列与查询目标不符。
state = _presence_dedupe_state()
_mutate_presence_row(state, 11501, 1, 9)
_presence_snapshot_raises("work列不符", state)
# 窄合同行像漂移:deleted 不是严格 False(整数 0 也必须拒绝)。
state = _presence_dedupe_state()
_mutate_presence_row(state, 11501, 9, 0)
_presence_snapshot_raises("deleted类型不符", state)
# 窄合同行像漂移:tenant 列不符。
state = _presence_dedupe_state()
_mutate_presence_row(state, 11501, 10, "other-tenant")
_presence_snapshot_raises("tenant不符", state)
def test_presence_dedupe_confirmation_sha_binds_rows_and_ids():
"""confirmation_sha 必须确定且绑定二十行完整内容与保留/删除 ID,任一来源漂移即变化。"""
base_snapshot = pu._presence_dedupe_capture_snapshot(
_PresenceDedupeConn(_presence_dedupe_state()), pu.PRESENCE_DEDUPE_WORK_ID,
)
base_sha = pu._presence_dedupe_confirmation_sha(base_snapshot)
again_snapshot = pu._presence_dedupe_capture_snapshot(
_PresenceDedupeConn(_presence_dedupe_state()), pu.PRESENCE_DEDUPE_WORK_ID,
)
check(
"presence-dedupe-confirmation_sha确定性",
base_sha == pu._presence_dedupe_confirmation_sha(again_snapshot)
and len(base_sha) == 64,
detail=base_sha,
)
for label, row_id, column, value in (
("observation", 11582, 6, "被篡改的观察"),
("保留行ID", 11501, 0, 11001),
("creator", 11501, 7, "篡改者"),
):
state = _presence_dedupe_state()
_mutate_presence_row(state, row_id, column, value)
drifted = pu._presence_dedupe_capture_snapshot(
_PresenceDedupeConn(state), pu.PRESENCE_DEDUPE_WORK_ID,
)
check(
f"presence-dedupe-sha感知{label}漂移",
pu._presence_dedupe_confirmation_sha(drifted) != base_sha,
)
def test_presence_dedupe_cli_preview_read_only():
"""preview:RR READ ONLY 输出十组摘要与 SHA,不写任何表,不触嵌入/模型。"""
state = _presence_dedupe_state()
result, _ = _presence_dedupe_cli(state, ["--work-id", "8", "--preview"])
check("presence-dedupe-preview成功", result.exit_code == 0, detail=result.output)
data = json.loads(result.output)
check(
"presence-dedupe-preview摘要完整",
data["mode"] == "preview"
and data["work_id"] == pu.PRESENCE_DEDUPE_WORK_ID
and data["group_count"] == pu.PRESENCE_DEDUPE_GROUP_COUNT
and data["row_count"] == pu.PRESENCE_DEDUPE_ROW_COUNT
and data["delete_ids"] == sorted(pu.PRESENCE_DEDUPE_DELETE_IDS)
and data["keep_ids"] == list(range(11501, 11511))
and len(data["confirmation_sha"]) == 64
and data["deleted_ids"] is None
and data["audit_rows_written"] == 0
and len(data["groups"]) == pu.PRESENCE_DEDUPE_GROUP_COUNT,
detail=result.output[:500],
)
check(
"presence-dedupe-preview不输出observation原文",
"观察A0" not in result.output and "观察B0" not in result.output,
)
check(
"presence-dedupe-preview零副作用",
state["writes"] == [] and state["commits"] == 0
and all(not row[9] for row in state["rows"]),
)
check(
"presence-dedupe-preview首条SQL为RR只读",
state["queries"][0] == "SET TRANSACTION ISOLATION LEVEL REPEATABLE READ, READ ONLY",
detail=str(state["queries"][:2]),
)
check(
"presence-dedupe-preview读快照不加行锁",
not any(query.endswith("FOR UPDATE") for query in state["queries"]),
detail=str(state["queries"]),
)
def test_presence_dedupe_cli_execute_soft_deletes_exactly_ten():
"""execute:锁内重算精确匹配后同事务只软删十个精确 ID;其他域零写入。"""
state = _presence_dedupe_state()
sha = _presence_dedupe_preview_sha(_presence_dedupe_state())
result, _ = _presence_dedupe_cli(
state,
["--work-id", "8", "--execute", "--confirmation-sha", sha,
"--confirm-no-live-process"],
)
check("presence-dedupe-execute成功", result.exit_code == 0, detail=result.output)
data = json.loads(result.output)
check(
"presence-dedupe-execute只软删十个精确ID",
data["mode"] == "execute"
and data["deleted_ids"] == sorted(pu.PRESENCE_DEDUPE_DELETE_IDS)
and sorted(row[0] for row in state["rows"] if row[9] is True)
== sorted(pu.PRESENCE_DEDUPE_DELETE_IDS),
detail=result.output[:500],
)
check(
"presence-dedupe-execute保留行与单体行全部留存",
sorted(row[0] for row in state["rows"] if not row[9])
== sorted(list(range(11501, 11511)) + list(range(11400, 11404))),
detail=str(sorted(row[0] for row in state["rows"] if not row[9])),
)
check(
"presence-dedupe-execute同事务单次提交",
state["writes"] == ["UPDATE example_upgrade_presence"] and state["commits"] == 1,
detail=str(state["writes"]),
)
mutating = [
query for query in state["queries"]
if query.startswith(("UPDATE ", "INSERT ", "DELETE "))
]
check(
"presence-dedupe-execute写SQL仅presence软删",
len(mutating) == 1
and mutating[0].startswith("UPDATE example_upgrade_presence SET deleted=TRUE")
and "RETURNING id" in mutating[0],
detail=str(mutating),
)
check(
"presence-dedupe-execute锁内带行锁重算快照",
any(
query.startswith("SELECT id, work_id, window_no, chapter_no")
and query.endswith("FOR UPDATE")
for query in state["queries"]
),
detail=str(state["queries"]),
)
def test_presence_dedupe_execute_second_run_fails_closed():
"""execute 成功后二次运行:软删后每组只剩单行不成十组,必须失败关闭且无新写入。"""
state = _presence_dedupe_state()
sha = _presence_dedupe_preview_sha(_presence_dedupe_state())
args = ["--work-id", "8", "--execute", "--confirmation-sha", sha,
"--confirm-no-live-process"]
first, _ = _presence_dedupe_cli(state, args)
check("presence-dedupe-首次execute成功", first.exit_code == 0,
detail=first.output)
second, _ = _presence_dedupe_cli(state, args)
check(
"presence-dedupe-二次execute失败关闭",
second.exit_code != 0 and "冗余组数量不精确" in second.output,
detail=second.output,
)
check(
"presence-dedupe-二次execute无新写入",
state["writes"] == ["UPDATE example_upgrade_presence"] and state["commits"] == 1,
detail=str(state["writes"]),
)
def test_presence_dedupe_execute_rejects_stale_or_drifted_snapshot():
"""旧 SHA 与 preview/execute 间快照漂移都必须在写入前失败关闭。"""
# 旧/伪 confirmation-sha:锁内重算后精确匹配拒绝。
state = _presence_dedupe_state()
stale, _ = _presence_dedupe_cli(
state,
["--work-id", "8", "--execute", "--confirmation-sha", "0" * 64,
"--confirm-no-live-process"],
)
check("presence-dedupe-旧SHA失败关闭", stale.exit_code != 0,
detail=stale.output)
check(
"presence-dedupe-旧SHA零写入",
state["writes"] == [] and state["commits"] == 0
and all(not row[9] for row in state["rows"]),
)
# preview 与 execute 之间漂移(observation 被改):锁内重算 SHA 不匹配。
sha = _presence_dedupe_preview_sha(_presence_dedupe_state())
drifted = _presence_dedupe_state()
_mutate_presence_row(drifted, 11582, 6, "被篡改的观察")
result, _ = _presence_dedupe_cli(
drifted,
["--work-id", "8", "--execute", "--confirmation-sha", sha,
"--confirm-no-live-process"],
)
check(
"presence-dedupe-锁内漂移失败关闭",
result.exit_code != 0 and "不匹配" in result.output,
detail=result.output,
)
check(
"presence-dedupe-锁内漂移零写入",
drifted["writes"] == [] and drifted["commits"] == 0
and all(not row[9] for row in drifted["rows"]),
)
def test_presence_dedupe_cli_argument_contracts():
"""preview/execute 互斥;execute 必须同时提供 SHA 与无活进程确认。"""
cases = (
("preview与execute同给", ["--work-id", "8", "--preview", "--execute"]),
("两模式都不给", ["--work-id", "8"]),
("execute缺SHA", ["--work-id", "8", "--execute", "--confirm-no-live-process"]),
("execute缺进程确认",
["--work-id", "8", "--execute", "--confirmation-sha", "ab" * 32]),
("work不符", ["--work-id", "9", "--preview"]),
)
for label, args in cases:
state = _presence_dedupe_state()
result, _ = _presence_dedupe_cli(state, args)
check(f"presence-dedupe-{label}拒绝", result.exit_code != 0,
detail=result.output)
check(
f"presence-dedupe-{label}零副作用",
state["writes"] == [] and state["commits"] == 0,
)
def test_presence_dedupe_cli_preview_fail_closed_on_drift():
"""CLI 层各类快照漂移都在 preview 被拒,且零副作用。"""
first_delete_id = sorted(pu.PRESENCE_DEDUPE_DELETE_IDS)[0]
cases = (
("额外冗余组", lambda state: state["rows"].extend((
(11581, 8, 90, 500, "event", "额外重复组", "观察C",
"upgrade", _PRESENCE_DEDUPE_BASE_TIME, False, pu.TENANT),
(11582, 8, 90, 500, "event", "额外重复组", "观察D",
"upgrade", _PRESENCE_DEDUPE_BASE_TIME, False, pu.TENANT),
))),
("observation相同",
lambda state: _mutate_presence_row(state, first_delete_id, 6, "观察A0")),
("缺行", lambda state: state["rows"].remove(
next(row for row in state["rows"] if row[0] == first_delete_id))),
("tenant不符",
lambda state: _mutate_presence_row(state, 11501, 10, "other-tenant")),
)
for label, mutate in cases:
state = _presence_dedupe_state()
mutate(state)
result, _ = _presence_dedupe_cli(state, ["--work-id", "8", "--preview"])
check(f"presence-dedupe-preview{label}拒绝", result.exit_code != 0,
detail=result.output)
check(
f"presence-dedupe-preview{label}零副作用",
state["writes"] == [] and state["commits"] == 0,
)
if __name__ == "__main__": if __name__ == "__main__":
@ -6443,7 +5902,7 @@ if __name__ == "__main__":
test_prompts_disciplines, test_window_fence_hashes_and_markers, test_prompts_disciplines, test_window_fence_hashes_and_markers,
test_capture_window_input_rejects_incomplete_or_empty_content, test_capture_window_input_rejects_incomplete_or_empty_content,
test_two_phase_external_calls_and_atomic_embedding_contract, test_two_phase_external_calls_and_atomic_embedding_contract,
test_recovery_and_redo_rejection_contracts, test_real_pg_rollback_smoke_entry, test_recovery_and_redo_rejection_contracts,
test_compensation_attempt_baseline_exact_allows_old_non_allowlist_tombstone, test_compensation_attempt_baseline_exact_allows_old_non_allowlist_tombstone,
test_compensation_attempt_baseline_input_drift_blocks, test_compensation_attempt_baseline_input_drift_blocks,
test_compensation_attempt_baseline_state_drift_blocks, test_compensation_attempt_baseline_state_drift_blocks,
@ -6458,21 +5917,8 @@ if __name__ == "__main__":
test_quality_repair_rejects_final_snapshot_drift_and_bad_embedding, test_quality_repair_rejects_final_snapshot_drift_and_bad_embedding,
test_quality_repair_rejects_invalid_vector_dimension_and_values, test_quality_repair_rejects_invalid_vector_dimension_and_values,
test_quality_repair_rejects_alias_and_presence_duplicates, test_quality_repair_rejects_alias_and_presence_duplicates,
test_presence_dedupe_actions_planner_guard,
test_presence_dedupe_snapshot_success,
test_presence_dedupe_snapshot_fail_closed,
test_presence_dedupe_confirmation_sha_binds_rows_and_ids,
test_presence_dedupe_cli_preview_read_only,
test_presence_dedupe_cli_execute_soft_deletes_exactly_ten,
test_presence_dedupe_execute_second_run_fails_closed,
test_presence_dedupe_execute_rejects_stale_or_drifted_snapshot,
test_presence_dedupe_cli_argument_contracts,
test_presence_dedupe_cli_preview_fail_closed_on_drift,
test_capture_window_state_binds_current_window_and_preserves_other_status, test_capture_window_state_binds_current_window_and_preserves_other_status,
test_load_window_material_orders_blocks_like_capture_input, test_load_window_material_orders_blocks_like_capture_input,
test_run_defaults_enforce_safety_contract): test_run_defaults_enforce_safety_contract):
fn() fn()
if os.getenv("MUSE_REAL_PG_ROLLBACK_SMOKE") == "1": print(f"\n全部离线自测通过:{_passed} 项(未发网络/嵌入/LLM 调用)")
print(f"\n全部自测通过:{_passed} 项(含真实 PG public 窗 rollback smoke;未发模型/嵌入调用)")
else:
print(f"\n全部离线自测通过:{_passed} 项(真实 PG smoke 未启用;未发网络/嵌入/LLM 调用)")

View File

@ -0,0 +1,28 @@
#!/usr/bin/env python3
"""Opt-in real PostgreSQL rollback smoke entry point."""
import os
import pathlib
import sys
PROJECT_ROOT = pathlib.Path(__file__).resolve().parents[3]
SCRIPT_DIR = PROJECT_ROOT / ".claude" / "skills" / "extract-work-knowledge" / "scripts"
sys.path.insert(0, str(SCRIPT_DIR))
import upgrade as pu # noqa: E402
def main():
if os.getenv("MUSE_REAL_PG_ROLLBACK_SMOKE") != "1":
print(
"BLOCKED: set MUSE_REAL_PG_ROLLBACK_SMOKE=1 to run the real PostgreSQL rollback smoke"
)
return 2
work_id = int(os.getenv("MUSE_REAL_PG_ROLLBACK_SMOKE_WORK_ID", "8"))
pu.real_pg_rollback_smoke(work_id)
print(f"PASS: real PostgreSQL rollback smoke completed for work_id={work_id}")
return 0
if __name__ == "__main__":
raise SystemExit(main())

View File

@ -0,0 +1,599 @@
#!/usr/bin/env python3
"""presence 冗余收口(work8 十组双行)的纯逻辑离线自测。
红线:**不连真实库、不发任何网络/嵌入/LLM 调用**——只使用内存 fake DB 验证
presence 去重的 planner、快照、确认摘要和 preview/execute 事务合同。
跑法:仓库根目录 `.venv/bin/python tests/skills/extract-work-knowledge/test_presence_dedupe.py`;
也支持从其他工作目录通过该文件的绝对路径运行。
"""
import json
import pathlib
import sys
from contextlib import nullcontext
from unittest.mock import patch
from click.testing import CliRunner
# 从原生产 scripts 导入 upgrade;测试替身只隔离数据库连接,不替代生产逻辑。
PROJECT_ROOT = pathlib.Path(__file__).resolve().parents[3]
SCRIPT_DIR = PROJECT_ROOT / ".claude" / "skills" / "extract-work-knowledge" / "scripts"
sys.path.insert(0, str(SCRIPT_DIR))
import upgrade as pu # noqa: E402
_passed = 0
def check(name, cond, detail=""):
"""单项断言:通过打 [PASS],失败抛 AssertionError(带上下文,令 CI/人工一眼定位)。"""
global _passed
assert cond, f"[FAIL] {name} :: {detail}"
_passed += 1
print(f"[PASS] {name}")
class _CardResult:
"""为 presence 去重离线测试提供最小查询结果对象。"""
def __init__(self, row=None, rows=None):
self.row = row
self.rows = rows or []
def fetchone(self):
"""返回预置的单行结果。"""
return self.row
def fetchall(self):
"""返回预置的多行结果。"""
return self.rows
# ── presence 冗余收口(work8 十组双行:保留 MIN(id) 软删 MAX(id))──
# 离线 fixture 用字符串时间戳代替数据库驱动的 datetime:窄合同只比较同组两行是否相等,
# _sha256_json 计算摘要时统一 default=str 转写,二者行为一致。
_PRESENCE_DEDUPE_BASE_TIME = "2026-07-01 12:00:00"
def _presence_dedupe_rows():
"""构造满足窄合同的 work8 presence 行 fixture。
十组双行:每组 observation 不同、create_time 相同、恰好一个待删 ID 命中 DELETE_IDS、
保留 ID 是组内较小者;另有四个单行,证明快照会跳过不构成冗余的 key。
"""
rows = []
for index, delete_id in enumerate(sorted(pu.PRESENCE_DEDUPE_DELETE_IDS)):
keep_id = 11501 + index
entity_type = "location" if index % 2 == 0 else "item"
name = f"重复实体{index}"
for row_id, observation in ((keep_id, f"观察A{index}"),
(delete_id, f"观察B{index}")):
rows.append((row_id, 8, 57 + index, 403 + index, entity_type, name,
observation, "upgrade", _PRESENCE_DEDUPE_BASE_TIME,
False, pu.TENANT))
for index in range(4):
rows.append((11400 + index, 8, 10 + index, 20 + index, "character",
f"单体实体{index}", f"单体观察{index}",
"upgrade", _PRESENCE_DEDUPE_BASE_TIME, False, pu.TENANT))
return sorted(rows, key=lambda row: row[0])
def _presence_dedupe_state():
return {
"rows": _presence_dedupe_rows(),
"writes": [],
"queries": [],
"commits": 0,
}
def _mutate_presence_row(state, row_id, column, value):
"""替换指定行的单列;元组不可变,整体重建后写回。"""
state["rows"] = [
row[:column] + (value,) + row[column + 1:] if row[0] == row_id else row
for row in state["rows"]
]
class _PresenceDedupeConn:
"""presence 冗余收口 fake DB:只实现本命令的读快照与软删 SQL。
其他任何域(draft/window/alias/card_state/audit/embedding)的 SQL 会落到末尾
AssertionError,因此「零副作用」无需逐条枚举禁写语句即可离线断言。fake 只做
「未软删」粗过滤;列数/work/tenant/deleted 类型等窄合同行像校验正是离线断言对象,
fake 不能代为过滤。
"""
def __init__(self, state):
self.state = state
def __enter__(self):
return self
def __exit__(self, exc_type, exc, tb):
return False
def execute(self, query, params=()):
normalized = " ".join(query.split())
self.state.setdefault("queries", []).append(normalized)
if normalized.startswith("SET TRANSACTION") or normalized.startswith("LOCK TABLE"):
return _CardResult()
if normalized.startswith(
"SELECT id, work_id, window_no, chapter_no, entity_type, name, observation,"):
rows = [
row for row in sorted(self.state["rows"], key=lambda item: item[0])
if len(row) > 9 and not row[9]
]
return _CardResult(rows=rows)
if normalized.startswith("UPDATE example_upgrade_presence SET deleted=TRUE"):
tenant_id, work_id, delete_ids = params
assert tenant_id == pu.TENANT
delete_set = set(delete_ids)
hit = []
new_rows = []
for row in self.state["rows"]:
if row[1] == work_id and row[0] in delete_set and not row[9]:
row = row[:9] + (True,) + row[10:]
hit.append((row[0],))
new_rows.append(row)
self.state["rows"] = new_rows
self.state["writes"].append("UPDATE example_upgrade_presence")
return _CardResult(rows=hit)
raise AssertionError(f"presence 冗余收口 fake DB 未覆盖 SQL:{normalized}")
def commit(self):
self.state["commits"] += 1
def _presence_dedupe_cli(state, args):
"""离线执行 repair-presence-duplicates:注入 fake DB 与空放同书锁。"""
conn = _PresenceDedupeConn(state)
with patch.object(pu.psycopg, "connect", return_value=conn), \
patch.object(pu, "upgrade_work_lock", return_value=nullcontext()):
result = CliRunner().invoke(pu.maintenance_cli, ["repair-presence-duplicates"] + args)
return result, conn
def _presence_dedupe_preview_sha(state):
"""离线取 preview 的 confirmation_sha,作为 execute 的合法输入。"""
result, _ = _presence_dedupe_cli(state, ["--work-id", "8", "--preview"])
assert result.exit_code == 0, result.output
return json.loads(result.output)["confirmation_sha"]
def _presence_snapshot_raises(label, state, work_id=pu.PRESENCE_DEDUPE_WORK_ID):
"""断言快照以 CompensationFenceConflict 拒绝,且没有发出任何写入。"""
conn = _PresenceDedupeConn(state)
try:
pu._presence_dedupe_capture_snapshot(conn, work_id)
except pu.CompensationFenceConflict as exc:
check(f"presence-dedupe-{label}失败关闭", True, detail=str(exc))
else:
check(f"presence-dedupe-{label}失败关闭", False,
detail="未抛 CompensationFenceConflict")
check(
f"presence-dedupe-{label}零副作用",
state["writes"] == [] and state["commits"] == 0,
)
def test_presence_dedupe_actions_planner_guard():
"""planner 级收口守卫:同观察幂等留首条、不同观察失败关闭、非 presence 原样透传。"""
presence_a = ("presence", None, 403, "location", "重复星体", "观察一")
presence_a_dup = ("presence", None, 403, "location", "重复星体", "观察一")
presence_b = ("presence", None, 404, "item", "镜面护盾", "观察二")
alias_action = ("alias", 7, "规范名", "别名", "ai")
new_action = ("new", -1, {"名称": "新实体"}, set(), [])
actions = [presence_a, alias_action, presence_a_dup, new_action, presence_b, ()]
original = list(actions)
deduped = pu._dedupe_presence_actions(actions)
check(
"presence-dedupe-planner同观察去重保留首条",
deduped == [presence_a, alias_action, new_action, presence_b, ()]
and deduped[0] is presence_a,
detail=str(deduped),
)
check("presence-dedupe-重复应用幂等",
pu._dedupe_presence_actions(deduped) == deduped)
check("presence-dedupe-不修改输入列表", actions == original)
conflict = [presence_a,
("presence", None, 403, "location", "重复星体", "另一个观察")]
try:
pu._dedupe_presence_actions(conflict)
except RuntimeError as exc:
check("presence-dedupe-不同观察失败关闭",
"observation 冲突" in str(exc), detail=str(exc))
else:
check("presence-dedupe-不同观察失败关闭", False, detail="未抛 RuntimeError")
for label, bad, keyword in (
("结构缺一元", ("presence", None, 403, "location", "重复星体"),
"结构非法"),
("结构多一元",
("presence", None, 403, "location", "重复星体", "观察一", "extra"),
"结构非法"),
("observation非字符串", ("presence", None, 403, "location", "重复星体", None),
"必须是字符串"),
):
try:
pu._dedupe_presence_actions([bad])
except RuntimeError as exc:
check(f"presence-dedupe-{label}抛错", keyword in str(exc), detail=str(exc))
else:
check(f"presence-dedupe-{label}抛错", False, detail="未抛 RuntimeError")
def test_presence_dedupe_snapshot_success():
"""快照成功路径:恰好 10 组/20 行,待删集合精确等于 DELETE_IDS,保留为组内较小者。"""
state = _presence_dedupe_state()
snapshot = pu._presence_dedupe_capture_snapshot(
_PresenceDedupeConn(state), pu.PRESENCE_DEDUPE_WORK_ID,
)
check(
"presence-dedupe-快照组数行数精确",
len(snapshot["groups"]) == pu.PRESENCE_DEDUPE_GROUP_COUNT
and len(snapshot["rows"]) == pu.PRESENCE_DEDUPE_ROW_COUNT
and snapshot["contract"] == pu.PRESENCE_DEDUPE_CONTRACT
and snapshot["work_id"] == pu.PRESENCE_DEDUPE_WORK_ID,
)
check(
"presence-dedupe-待删集合精确等于DELETE_IDS",
snapshot["delete_ids"] == sorted(pu.PRESENCE_DEDUPE_DELETE_IDS),
detail=str(snapshot["delete_ids"]),
)
check(
"presence-dedupe-保留为组内较小ID",
snapshot["keep_ids"] == list(range(11501, 11511))
and all(group["keep_id"] < group["delete_id"] for group in snapshot["groups"])
and all(len(group["rows"]) == 2 for group in snapshot["groups"]),
detail=str(snapshot["keep_ids"]),
)
check(
"presence-dedupe-行按id排序且排除单体行",
[row[0] for row in snapshot["rows"]]
== sorted(row[0] for row in snapshot["rows"])
and not any(row[0] in range(11400, 11404) for row in snapshot["rows"]),
)
def test_presence_dedupe_snapshot_fail_closed():
"""各类窄合同漂移都必须在写入前以 CompensationFenceConflict 逐条拒绝。"""
# 额外冗余组:组数 11 ≠ 10(新增组复用合同内待删 ID,确保失败点落在组数校验)。
state = _presence_dedupe_state()
state["rows"].extend((
(11581, 8, 90, 500, "event", "额外重复组", "观察C",
"upgrade", _PRESENCE_DEDUPE_BASE_TIME, False, pu.TENANT),
(11582, 8, 90, 500, "event", "额外重复组", "观察D",
"upgrade", _PRESENCE_DEDUPE_BASE_TIME, False, pu.TENANT),
))
_presence_snapshot_raises("额外冗余组", state)
# 三行组不是「恰好两行」。
state = _presence_dedupe_state()
state["rows"].append(
(11601, 8, 57, 403, "location", "重复实体0", "观察E",
"upgrade", _PRESENCE_DEDUPE_BASE_TIME, False, pu.TENANT)
)
_presence_snapshot_raises("三行冗余组", state)
# 缺行:一组只剩单行被跳过 → 组数 9 ≠ 10。
state = _presence_dedupe_state()
state["rows"] = [row for row in state["rows"]
if row[0] != sorted(pu.PRESENCE_DEDUPE_DELETE_IDS)[0]]
_presence_snapshot_raises("缺行", state)
# 同组两行 create_time 漂移。
state = _presence_dedupe_state()
_mutate_presence_row(state, 11501, 8, "2026-07-02 09:30:00")
_presence_snapshot_raises("create_time漂移", state)
# 待删 ID 命中 0:组内两行都不在 DELETE_IDS。
state = _presence_dedupe_state()
_mutate_presence_row(state, 11501, 0, 12001)
_mutate_presence_row(state, 11582, 0, 12002)
_presence_snapshot_raises("待删ID无命中", state)
# 待删 ID 命中 2:组内两行都在 DELETE_IDS。
state = _presence_dedupe_state()
_mutate_presence_row(state, 11501, 0, 11693)
_presence_snapshot_raises("待删ID双命中", state)
# 保留不是组内较小者:待删 ID 反而是组内较小者。
state = _presence_dedupe_state()
_mutate_presence_row(state, 11501, 0, 11900)
_presence_snapshot_raises("保留非组内较小", state)
# 同组两行 observation 相同(byte-exact 重复违反窄合同)。
state = _presence_dedupe_state()
_mutate_presence_row(state, 11582, 6, "观察A0")
_presence_snapshot_raises("observation相同", state)
# work_id ≠ 8 拒绝。
_presence_snapshot_raises("work不符", _presence_dedupe_state(), work_id=9)
# 窄合同行像漂移:列数不足。
state = _presence_dedupe_state()
state["rows"] = [row[:10] if row[0] == 11501 else row for row in state["rows"]]
_presence_snapshot_raises("列数不符", state)
# 窄合同行像漂移:work 列与查询目标不符。
state = _presence_dedupe_state()
_mutate_presence_row(state, 11501, 1, 9)
_presence_snapshot_raises("work列不符", state)
# 窄合同行像漂移:deleted 不是严格 False(整数 0 也必须拒绝)。
state = _presence_dedupe_state()
_mutate_presence_row(state, 11501, 9, 0)
_presence_snapshot_raises("deleted类型不符", state)
# 窄合同行像漂移:tenant 列不符。
state = _presence_dedupe_state()
_mutate_presence_row(state, 11501, 10, "other-tenant")
_presence_snapshot_raises("tenant不符", state)
def test_presence_dedupe_confirmation_sha_binds_rows_and_ids():
"""confirmation_sha 必须确定且绑定二十行完整内容与保留/删除 ID,任一来源漂移即变化。"""
base_snapshot = pu._presence_dedupe_capture_snapshot(
_PresenceDedupeConn(_presence_dedupe_state()), pu.PRESENCE_DEDUPE_WORK_ID,
)
base_sha = pu._presence_dedupe_confirmation_sha(base_snapshot)
again_snapshot = pu._presence_dedupe_capture_snapshot(
_PresenceDedupeConn(_presence_dedupe_state()), pu.PRESENCE_DEDUPE_WORK_ID,
)
check(
"presence-dedupe-confirmation_sha确定性",
base_sha == pu._presence_dedupe_confirmation_sha(again_snapshot)
and len(base_sha) == 64,
detail=base_sha,
)
for label, row_id, column, value in (
("observation", 11582, 6, "被篡改的观察"),
("保留行ID", 11501, 0, 11001),
("creator", 11501, 7, "篡改者"),
):
state = _presence_dedupe_state()
_mutate_presence_row(state, row_id, column, value)
drifted = pu._presence_dedupe_capture_snapshot(
_PresenceDedupeConn(state), pu.PRESENCE_DEDUPE_WORK_ID,
)
check(
f"presence-dedupe-sha感知{label}漂移",
pu._presence_dedupe_confirmation_sha(drifted) != base_sha,
)
def test_presence_dedupe_cli_preview_read_only():
"""preview:RR READ ONLY 输出十组摘要与 SHA,不写任何表,不触嵌入/模型。"""
state = _presence_dedupe_state()
result, _ = _presence_dedupe_cli(state, ["--work-id", "8", "--preview"])
check("presence-dedupe-preview成功", result.exit_code == 0, detail=result.output)
data = json.loads(result.output)
check(
"presence-dedupe-preview摘要完整",
data["mode"] == "preview"
and data["work_id"] == pu.PRESENCE_DEDUPE_WORK_ID
and data["group_count"] == pu.PRESENCE_DEDUPE_GROUP_COUNT
and data["row_count"] == pu.PRESENCE_DEDUPE_ROW_COUNT
and data["delete_ids"] == sorted(pu.PRESENCE_DEDUPE_DELETE_IDS)
and data["keep_ids"] == list(range(11501, 11511))
and len(data["confirmation_sha"]) == 64
and data["deleted_ids"] is None
and data["audit_rows_written"] == 0
and len(data["groups"]) == pu.PRESENCE_DEDUPE_GROUP_COUNT,
detail=result.output[:500],
)
check(
"presence-dedupe-preview不输出observation原文",
"观察A0" not in result.output and "观察B0" not in result.output,
)
check(
"presence-dedupe-preview零副作用",
state["writes"] == [] and state["commits"] == 0
and all(not row[9] for row in state["rows"]),
)
check(
"presence-dedupe-preview首条SQL为RR只读",
state["queries"][0] == "SET TRANSACTION ISOLATION LEVEL REPEATABLE READ, READ ONLY",
detail=str(state["queries"][:2]),
)
check(
"presence-dedupe-preview读快照不加行锁",
not any(query.endswith("FOR UPDATE") for query in state["queries"]),
detail=str(state["queries"]),
)
def test_presence_dedupe_cli_execute_soft_deletes_exactly_ten():
"""execute:锁内重算精确匹配后同事务只软删十个精确 ID;其他域零写入。"""
state = _presence_dedupe_state()
sha = _presence_dedupe_preview_sha(_presence_dedupe_state())
result, _ = _presence_dedupe_cli(
state,
["--work-id", "8", "--execute", "--confirmation-sha", sha,
"--confirm-no-live-process"],
)
check("presence-dedupe-execute成功", result.exit_code == 0, detail=result.output)
data = json.loads(result.output)
check(
"presence-dedupe-execute只软删十个精确ID",
data["mode"] == "execute"
and data["deleted_ids"] == sorted(pu.PRESENCE_DEDUPE_DELETE_IDS)
and sorted(row[0] for row in state["rows"] if row[9] is True)
== sorted(pu.PRESENCE_DEDUPE_DELETE_IDS),
detail=result.output[:500],
)
check(
"presence-dedupe-execute保留行与单体行全部留存",
sorted(row[0] for row in state["rows"] if not row[9])
== sorted(list(range(11501, 11511)) + list(range(11400, 11404))),
detail=str(sorted(row[0] for row in state["rows"] if not row[9])),
)
check(
"presence-dedupe-execute同事务单次提交",
state["writes"] == ["UPDATE example_upgrade_presence"] and state["commits"] == 1,
detail=str(state["writes"]),
)
mutating = [
query for query in state["queries"]
if query.startswith(("UPDATE ", "INSERT ", "DELETE "))
]
check(
"presence-dedupe-execute写SQL仅presence软删",
len(mutating) == 1
and mutating[0].startswith("UPDATE example_upgrade_presence SET deleted=TRUE")
and "RETURNING id" in mutating[0],
detail=str(mutating),
)
check(
"presence-dedupe-execute锁内带行锁重算快照",
any(
query.startswith("SELECT id, work_id, window_no, chapter_no")
and query.endswith("FOR UPDATE")
for query in state["queries"]
),
detail=str(state["queries"]),
)
def test_presence_dedupe_execute_second_run_fails_closed():
"""execute 成功后二次运行:软删后每组只剩单行不成十组,必须失败关闭且无新写入。"""
state = _presence_dedupe_state()
sha = _presence_dedupe_preview_sha(_presence_dedupe_state())
args = ["--work-id", "8", "--execute", "--confirmation-sha", sha,
"--confirm-no-live-process"]
first, _ = _presence_dedupe_cli(state, args)
check("presence-dedupe-首次execute成功", first.exit_code == 0,
detail=first.output)
second, _ = _presence_dedupe_cli(state, args)
check(
"presence-dedupe-二次execute失败关闭",
second.exit_code != 0 and "冗余组数量不精确" in second.output,
detail=second.output,
)
check(
"presence-dedupe-二次execute无新写入",
state["writes"] == ["UPDATE example_upgrade_presence"] and state["commits"] == 1,
detail=str(state["writes"]),
)
def test_presence_dedupe_execute_rejects_stale_or_drifted_snapshot():
"""旧 SHA 与 preview/execute 间快照漂移都必须在写入前失败关闭。"""
# 旧/伪 confirmation-sha:锁内重算后精确匹配拒绝。
state = _presence_dedupe_state()
stale, _ = _presence_dedupe_cli(
state,
["--work-id", "8", "--execute", "--confirmation-sha", "0" * 64,
"--confirm-no-live-process"],
)
check("presence-dedupe-旧SHA失败关闭", stale.exit_code != 0,
detail=stale.output)
check(
"presence-dedupe-旧SHA零写入",
state["writes"] == [] and state["commits"] == 0
and all(not row[9] for row in state["rows"]),
)
# preview 与 execute 之间漂移(observation 被改):锁内重算 SHA 不匹配。
sha = _presence_dedupe_preview_sha(_presence_dedupe_state())
drifted = _presence_dedupe_state()
_mutate_presence_row(drifted, 11582, 6, "被篡改的观察")
result, _ = _presence_dedupe_cli(
drifted,
["--work-id", "8", "--execute", "--confirmation-sha", sha,
"--confirm-no-live-process"],
)
check(
"presence-dedupe-锁内漂移失败关闭",
result.exit_code != 0 and "不匹配" in result.output,
detail=result.output,
)
check(
"presence-dedupe-锁内漂移零写入",
drifted["writes"] == [] and drifted["commits"] == 0
and all(not row[9] for row in drifted["rows"]),
)
def test_presence_dedupe_cli_argument_contracts():
"""preview/execute 互斥;execute 必须同时提供 SHA 与无活进程确认。"""
cases = (
("preview与execute同给", ["--work-id", "8", "--preview", "--execute"]),
("两模式都不给", ["--work-id", "8"]),
("execute缺SHA", ["--work-id", "8", "--execute", "--confirm-no-live-process"]),
("execute缺进程确认",
["--work-id", "8", "--execute", "--confirmation-sha", "ab" * 32]),
("work不符", ["--work-id", "9", "--preview"]),
)
for label, args in cases:
state = _presence_dedupe_state()
result, _ = _presence_dedupe_cli(state, args)
check(f"presence-dedupe-{label}拒绝", result.exit_code != 0,
detail=result.output)
check(
f"presence-dedupe-{label}零副作用",
state["writes"] == [] and state["commits"] == 0,
)
def test_presence_dedupe_cli_preview_fail_closed_on_drift():
"""CLI 层各类快照漂移都在 preview 被拒,且零副作用。"""
first_delete_id = sorted(pu.PRESENCE_DEDUPE_DELETE_IDS)[0]
cases = (
("额外冗余组", lambda state: state["rows"].extend((
(11581, 8, 90, 500, "event", "额外重复组", "观察C",
"upgrade", _PRESENCE_DEDUPE_BASE_TIME, False, pu.TENANT),
(11582, 8, 90, 500, "event", "额外重复组", "观察D",
"upgrade", _PRESENCE_DEDUPE_BASE_TIME, False, pu.TENANT),
))),
("observation相同",
lambda state: _mutate_presence_row(state, first_delete_id, 6, "观察A0")),
("缺行", lambda state: state["rows"].remove(
next(row for row in state["rows"] if row[0] == first_delete_id))),
("tenant不符",
lambda state: _mutate_presence_row(state, 11501, 10, "other-tenant")),
)
for label, mutate in cases:
state = _presence_dedupe_state()
mutate(state)
result, _ = _presence_dedupe_cli(state, ["--work-id", "8", "--preview"])
check(f"presence-dedupe-preview{label}拒绝", result.exit_code != 0,
detail=result.output)
check(
f"presence-dedupe-preview{label}零副作用",
state["writes"] == [] and state["commits"] == 0,
)
if __name__ == "__main__":
for fn in (test_presence_dedupe_actions_planner_guard,
test_presence_dedupe_snapshot_success,
test_presence_dedupe_snapshot_fail_closed,
test_presence_dedupe_confirmation_sha_binds_rows_and_ids,
test_presence_dedupe_cli_preview_read_only,
test_presence_dedupe_cli_execute_soft_deletes_exactly_ten,
test_presence_dedupe_execute_second_run_fails_closed,
test_presence_dedupe_execute_rejects_stale_or_drifted_snapshot,
test_presence_dedupe_cli_argument_contracts,
test_presence_dedupe_cli_preview_fail_closed_on_drift):
fn()
print(f"\n全部离线自测通过:{_passed} 项(未发网络/嵌入/LLM 调用)")

View File

@ -6,7 +6,8 @@ import sys
import unittest import unittest
SCRIPT_DIR = pathlib.Path(__file__).resolve().parent PROJECT_ROOT = pathlib.Path(__file__).resolve().parents[3]
SCRIPT_DIR = PROJECT_ROOT / ".claude" / "skills" / "extract-work-knowledge" / "scripts"
sys.path.insert(0, str(SCRIPT_DIR)) sys.path.insert(0, str(SCRIPT_DIR))
from upgrade_work_lock import ( # noqa: E402 from upgrade_work_lock import ( # noqa: E402

View File

@ -8,7 +8,9 @@ import pathlib
import sys import sys
import unittest import unittest
sys.path.insert(0, str(pathlib.Path(__file__).resolve().parent)) PROJECT_ROOT = pathlib.Path(__file__).resolve().parents[3]
SCRIPT_DIR = PROJECT_ROOT / ".claude" / "skills" / "freeze-context" / "scripts"
sys.path.insert(0, str(SCRIPT_DIR))
from audit_leakage import audit_snapshot # noqa: E402 from audit_leakage import audit_snapshot # noqa: E402

View File

@ -5,7 +5,9 @@ import pathlib
import sys import sys
import unittest import unittest
sys.path.insert(0, str(pathlib.Path(__file__).resolve().parent)) PROJECT_ROOT = pathlib.Path(__file__).resolve().parents[3]
SCRIPT_DIR = PROJECT_ROOT / ".claude" / "skills" / "freeze-context" / "scripts"
sys.path.insert(0, str(SCRIPT_DIR))
from build_snapshot import ( # noqa: E402 from build_snapshot import ( # noqa: E402
SnapshotError, SnapshotError,
build_snapshot, build_snapshot,

View File

@ -5,7 +5,9 @@ import pathlib
import sys import sys
import unittest import unittest
sys.path.insert(0, str(pathlib.Path(__file__).resolve().parent)) PROJECT_ROOT = pathlib.Path(__file__).resolve().parents[3]
SCRIPT_DIR = PROJECT_ROOT / ".claude" / "skills" / "freeze-context" / "scripts"
sys.path.insert(0, str(SCRIPT_DIR))
from check_snapshot import ( # noqa: E402 from check_snapshot import ( # noqa: E402
STATUS_BLOCKED_AUTHORIZATION, STATUS_BLOCKED_AUTHORIZATION,
STATUS_INVALID_ARM_DIFF, STATUS_INVALID_ARM_DIFF,

View File

@ -10,7 +10,9 @@ import sys
import unittest import unittest
from unittest.mock import patch from unittest.mock import patch
sys.path.insert(0, str(pathlib.Path(__file__).resolve().parent)) PROJECT_ROOT = pathlib.Path(__file__).resolve().parents[3]
SCRIPT_DIR = PROJECT_ROOT / ".claude" / "skills" / "freeze-context" / "scripts"
sys.path.insert(0, str(SCRIPT_DIR))
import load_reference_work as loader # noqa: E402 import load_reference_work as loader # noqa: E402
from load_reference_work import ( # noqa: E402 from load_reference_work import ( # noqa: E402
AdapterError, AdapterError,

View File

@ -12,8 +12,9 @@ from decimal import Decimal
from unittest import mock from unittest import mock
HERE = pathlib.Path(__file__).resolve().parent PROJECT_ROOT = pathlib.Path(__file__).resolve().parents[3]
sys.path.insert(0, str(HERE)) SCRIPT_DIR = PROJECT_ROOT / ".claude" / "skills" / "maintain-work-extraction" / "scripts"
sys.path.insert(0, str(SCRIPT_DIR))
import backup_upgrade_work as backup # noqa: E402 import backup_upgrade_work as backup # noqa: E402

View File

@ -11,9 +11,11 @@ from unittest.mock import Mock, patch
from click.testing import CliRunner from click.testing import CliRunner
SCRIPT_DIR = pathlib.Path(__file__).resolve().parent PROJECT_ROOT = pathlib.Path(__file__).resolve().parents[3]
SCRIPT_DIR = PROJECT_ROOT / ".claude" / "skills" / "maintain-work-extraction" / "scripts"
EXTRACTION_SCRIPTS = PROJECT_ROOT / ".claude" / "skills" / "extract-work-knowledge" / "scripts"
sys.path.insert(0, str(SCRIPT_DIR)) sys.path.insert(0, str(SCRIPT_DIR))
sys.path.insert(0, str(SCRIPT_DIR.parents[1] / "extract-work-knowledge" / "scripts")) sys.path.insert(0, str(EXTRACTION_SCRIPTS))
import upgrade as parse # noqa: E402 import upgrade as parse # noqa: E402
import backup_upgrade_work as backup # noqa: E402 import backup_upgrade_work as backup # noqa: E402

View File

@ -9,11 +9,15 @@ from pathlib import Path
import unittest import unittest
ROOT = Path(__file__).resolve().parents[4] PROJECT_ROOT = Path(__file__).resolve().parents[3]
SKILL = (ROOT / ".claude/skills/plan-chapter/SKILL.md").read_text(encoding="utf-8") SKILL = (
PLANNER = (ROOT / ".claude/agents/planner.md").read_text(encoding="utf-8") PROJECT_ROOT / ".claude" / "skills" / "plan-chapter" / "SKILL.md"
CHAINS = (ROOT / "meta/chains/README.md").read_text(encoding="utf-8") ).read_text(encoding="utf-8")
SCHEMA = (ROOT / "meta/schemas/fine_outline.yaml").read_text(encoding="utf-8") PLANNER = (PROJECT_ROOT / ".claude" / "agents" / "planner.md").read_text(encoding="utf-8")
CHAINS = (PROJECT_ROOT / "meta" / "chains" / "README.md").read_text(encoding="utf-8")
SCHEMA = (
PROJECT_ROOT / "meta" / "schemas" / "fine_outline.yaml"
).read_text(encoding="utf-8")
# 与 meta/schemas/fine_outline.yaml 的必填/推荐集保持一致。 # 与 meta/schemas/fine_outline.yaml 的必填/推荐集保持一致。
REQUIRED_FIELDS = ( REQUIRED_FIELDS = (

View File

@ -10,8 +10,10 @@ import pathlib
import sys import sys
import unittest import unittest
sys.path.insert(0, str(pathlib.Path(__file__).resolve().parent))
sys.path.insert(0, str(pathlib.Path(__file__).resolve().parents[2] / "access-database" / "scripts")) PROJECT_ROOT = pathlib.Path(__file__).resolve().parents[3]
SCRIPT_DIR = PROJECT_ROOT / ".claude" / "skills" / "plan-story" / "scripts"
sys.path.insert(0, str(SCRIPT_DIR))
from persist_planning import check_field_coverage, write_section, _load_schema # noqa: E402 from persist_planning import check_field_coverage, write_section, _load_schema # noqa: E402

View File

@ -1,11 +1,13 @@
#!/usr/bin/env python3 #!/usr/bin/env python3
"""规划执行留痕脚本的离线合同测试。""" """规划执行留痕脚本的离线合同测试。"""
import pathlib import pathlib
import sys import sys
import unittest import unittest
SCRIPT_DIR = pathlib.Path(__file__).resolve().parent PROJECT_ROOT = pathlib.Path(__file__).resolve().parents[3]
SCRIPT_DIR = PROJECT_ROOT / ".claude" / "skills" / "plan-story" / "scripts"
sys.path.insert(0, str(SCRIPT_DIR)) sys.path.insert(0, str(SCRIPT_DIR))
import record_planning_execution as recorder # noqa: E402 import record_planning_execution as recorder # noqa: E402

View File

@ -1,11 +1,13 @@
#!/usr/bin/env python3 #!/usr/bin/env python3
"""确定性规划回执修正的离线合同测试。""" """确定性规划回执修正的离线合同测试。"""
import pathlib import pathlib
import sys import sys
import unittest import unittest
SCRIPT_DIR = pathlib.Path(__file__).resolve().parent PROJECT_ROOT = pathlib.Path(__file__).resolve().parents[3]
SCRIPT_DIR = PROJECT_ROOT / ".claude" / "skills" / "plan-story" / "scripts"
sys.path.insert(0, str(SCRIPT_DIR)) sys.path.insert(0, str(SCRIPT_DIR))
import repair_deterministic_receipt as repair # noqa: E402 import repair_deterministic_receipt as repair # noqa: E402

View File

@ -1,12 +1,14 @@
#!/usr/bin/env python3 #!/usr/bin/env python3
"""范式选择投影的离线测试,不调用向量服务。""" """范式选择投影的离线测试,不调用向量服务。"""
import pathlib import pathlib
import sys import sys
import unittest import unittest
from unittest.mock import patch from unittest.mock import patch
SCRIPT_DIR = pathlib.Path(__file__).resolve().parent PROJECT_ROOT = pathlib.Path(__file__).resolve().parents[3]
SCRIPT_DIR = PROJECT_ROOT / ".claude" / "skills" / "plan-story" / "scripts"
sys.path.insert(0, str(SCRIPT_DIR)) sys.path.insert(0, str(SCRIPT_DIR))
import select_patterns as selector # noqa: E402 import select_patterns as selector # noqa: E402

View File

@ -5,7 +5,9 @@ import sys
import unittest import unittest
from unittest.mock import patch from unittest.mock import patch
sys.path.insert(0, str(pathlib.Path(__file__).resolve().parent)) PROJECT_ROOT = pathlib.Path(__file__).resolve().parents[3]
SCRIPT_DIR = PROJECT_ROOT / ".claude" / "skills" / "prevent-ai-flavor" / "scripts"
sys.path.insert(0, str(SCRIPT_DIR))
import prevent_ai_flavor as prev # noqa: E402 import prevent_ai_flavor as prev # noqa: E402

View File

@ -12,7 +12,8 @@ import threading
import unittest import unittest
from unittest import mock from unittest import mock
SCRIPT_DIR = pathlib.Path(__file__).resolve().parent PROJECT_ROOT = pathlib.Path(__file__).resolve().parents[3]
SCRIPT_DIR = PROJECT_ROOT / ".claude" / "skills" / "record-run-evidence" / "scripts"
sys.path.insert(0, str(SCRIPT_DIR)) sys.path.insert(0, str(SCRIPT_DIR))
from file_cas import CasConflictError, CasRecoveryError, FileCasStore # noqa: E402 from file_cas import CasConflictError, CasRecoveryError, FileCasStore # noqa: E402

View File

@ -3,7 +3,9 @@
import pathlib import pathlib
import sys import sys
sys.path.insert(0, str(pathlib.Path(__file__).resolve().parent)) PROJECT_ROOT = pathlib.Path(__file__).resolve().parents[3]
SCRIPT_DIR = PROJECT_ROOT / ".claude" / "skills" / "record-run-evidence" / "scripts"
sys.path.insert(0, str(SCRIPT_DIR))
from persist_raw import _check_no_secrets # noqa: E402 from persist_raw import _check_no_secrets # noqa: E402

View File

@ -13,7 +13,8 @@ import unittest
from datetime import datetime, timedelta, timezone from datetime import datetime, timedelta, timezone
from unittest import mock from unittest import mock
SCRIPT_DIR = pathlib.Path(__file__).resolve().parent PROJECT_ROOT = pathlib.Path(__file__).resolve().parents[3]
SCRIPT_DIR = PROJECT_ROOT / ".claude" / "skills" / "record-run-evidence" / "scripts"
sys.path.insert(0, str(SCRIPT_DIR)) sys.path.insert(0, str(SCRIPT_DIR))
import raw_vault # noqa: E402 import raw_vault # noqa: E402

View File

@ -5,7 +5,8 @@ import sys
import unittest import unittest
SCRIPT_DIR = pathlib.Path(__file__).resolve().parent PROJECT_ROOT = pathlib.Path(__file__).resolve().parents[3]
SCRIPT_DIR = PROJECT_ROOT / ".claude" / "skills" / "record-run-evidence" / "scripts"
sys.path.insert(0, str(SCRIPT_DIR)) sys.path.insert(0, str(SCRIPT_DIR))
import record_failed_run as failure # noqa: E402 import record_failed_run as failure # noqa: E402

View File

@ -5,7 +5,8 @@ import sys
import unittest import unittest
SCRIPT_DIR = pathlib.Path(__file__).resolve().parent PROJECT_ROOT = pathlib.Path(__file__).resolve().parents[3]
SCRIPT_DIR = PROJECT_ROOT / ".claude" / "skills" / "record-run-evidence" / "scripts"
sys.path.insert(0, str(SCRIPT_DIR)) sys.path.insert(0, str(SCRIPT_DIR))
import repair_receipt_evidence as repair # noqa: E402 import repair_receipt_evidence as repair # noqa: E402

View File

@ -3,7 +3,9 @@
import pathlib import pathlib
import sys import sys
sys.path.insert(0, str(pathlib.Path(__file__).resolve().parent)) PROJECT_ROOT = pathlib.Path(__file__).resolve().parents[3]
SCRIPT_DIR = PROJECT_ROOT / ".claude" / "skills" / "record-run-evidence" / "scripts"
sys.path.insert(0, str(SCRIPT_DIR))
import run_registry # noqa: E402 import run_registry # noqa: E402

View File

@ -7,8 +7,11 @@ import tempfile
import unittest import unittest
from unittest.mock import patch from unittest.mock import patch
sys.path.insert(0, str(pathlib.Path(__file__).resolve().parent)) PROJECT_ROOT = pathlib.Path(__file__).resolve().parents[3]
sys.path.insert(0, str(pathlib.Path(__file__).resolve().parents[3] / "skills" / "diagnose-ai-flavor" / "scripts")) SCRIPT_DIR = PROJECT_ROOT / ".claude" / "skills" / "revise-ai-flavor" / "scripts"
DIAGNOSE_SCRIPT_DIR = PROJECT_ROOT / ".claude" / "skills" / "diagnose-ai-flavor" / "scripts"
sys.path.insert(0, str(SCRIPT_DIR))
sys.path.insert(0, str(DIAGNOSE_SCRIPT_DIR))
import revise_ai_flavor as rev # noqa: E402 import revise_ai_flavor as rev # noqa: E402
import diagnose_ai_flavor as diag # noqa: E402 import diagnose_ai_flavor as diag # noqa: E402

View File

@ -10,7 +10,7 @@
测试行用 unittest-lesson- 前缀的 run_id 隔离,结束物理清理(本表可变)。 测试行用 unittest-lesson- 前缀的 run_id 隔离,结束物理清理(本表可变)。
跑法(需 Tailscale 内网可达 muse-example): 跑法(需 Tailscale 内网可达 muse-example):
.venv/bin/python .claude/skills/score-content-quality/scripts/test_lesson_registry_db.py .venv/bin/python tests/skills/score-content-quality/test_lesson_registry_db.py
""" """
from __future__ import annotations from __future__ import annotations
@ -19,9 +19,13 @@ import pathlib
import sys import sys
import uuid import uuid
SCRIPT_DIR = pathlib.Path(__file__).resolve().parent from psycopg.errors import RaiseException
SKILLS_DIR = SCRIPT_DIR.parents[1]
for path in (SCRIPT_DIR, SKILLS_DIR / "access-database" / "scripts"): PROJECT_ROOT = pathlib.Path(__file__).resolve().parents[3]
SKILLS_DIR = PROJECT_ROOT / ".claude" / "skills"
SCRIPT_DIR = SKILLS_DIR / "score-content-quality" / "scripts"
DB_SCRIPT_DIR = SKILLS_DIR / "access-database" / "scripts"
for path in (SCRIPT_DIR, DB_SCRIPT_DIR):
if str(path) not in sys.path: if str(path) not in sys.path:
sys.path.insert(0, str(path)) sys.path.insert(0, str(path))
@ -96,7 +100,7 @@ def test_no_auto_promotion() -> None:
raise AssertionError("触发器必须拒绝自动升格") raise AssertionError("触发器必须拒绝自动升格")
except AssertionError: except AssertionError:
raise raise
except Exception: except RaiseException:
pass pass
rows = list_lessons(status="proposed") rows = list_lessons(status="proposed")
assert any(row["id"] == out["lesson_id"] for row in rows) assert any(row["id"] == out["lesson_id"] for row in rows)

View File

@ -6,7 +6,10 @@ import inspect
import sys import sys
import unittest import unittest
sys.path.insert(0, str(pathlib.Path(__file__).resolve().parent)) PROJECT_ROOT = pathlib.Path(__file__).resolve().parents[3]
SCRIPT_DIR = PROJECT_ROOT / ".claude" / "skills" / "score-content-quality" / "scripts"
if str(SCRIPT_DIR) not in sys.path:
sys.path.insert(0, str(SCRIPT_DIR))
from fine_outline_rubric import ( # noqa: E402 from fine_outline_rubric import ( # noqa: E402
DIMENSIONS, DIMENSIONS,
RUBRIC_PROFILE, RUBRIC_PROFILE,

View File

@ -11,7 +11,8 @@ import sys
import unittest import unittest
from typing import Any, Mapping from typing import Any, Mapping
SCRIPT_DIR = pathlib.Path(__file__).resolve().parent PROJECT_ROOT = pathlib.Path(__file__).resolve().parents[3]
SCRIPT_DIR = PROJECT_ROOT / ".claude" / "skills" / "score-content-quality" / "scripts"
if str(SCRIPT_DIR) not in sys.path: if str(SCRIPT_DIR) not in sys.path:
sys.path.insert(0, str(SCRIPT_DIR)) sys.path.insert(0, str(SCRIPT_DIR))

View File

@ -7,7 +7,10 @@ import pathlib
import sys import sys
import unittest import unittest
sys.path.insert(0, str(pathlib.Path(__file__).resolve().parent)) PROJECT_ROOT = pathlib.Path(__file__).resolve().parents[3]
SCRIPT_DIR = PROJECT_ROOT / ".claude" / "skills" / "score-content-quality" / "scripts"
if str(SCRIPT_DIR) not in sys.path:
sys.path.insert(0, str(SCRIPT_DIR))
from writer_rubric import ( # noqa: E402 from writer_rubric import ( # noqa: E402
COMMON_DIMENSIONS, COMMON_DIMENSIONS,

View File

@ -3,8 +3,16 @@
from __future__ import annotations from __future__ import annotations
import pathlib
import sys
import unittest import unittest
PROJECT_ROOT = pathlib.Path(__file__).resolve().parents[3]
SCRIPT_DIR = PROJECT_ROOT / ".claude" / "skills" / "search-knowledge" / "scripts"
EMBED_SCRIPT_DIR = PROJECT_ROOT / ".claude" / "skills" / "embed-knowledge" / "scripts"
sys.path.insert(0, str(SCRIPT_DIR))
sys.path.insert(0, str(EMBED_SCRIPT_DIR))
from search import search_cards from search import search_cards

View File

@ -6,7 +6,7 @@
- transition/start_next 的守卫与 token 递增; - transition/start_next 的守卫与 token 递增;
- 该存储可直接替换 InMemoryCasStateStore 驱动 run_writer_pipeline 全链。 - 该存储可直接替换 InMemoryCasStateStore 驱动 run_writer_pipeline 全链。
跑法(仓库根目录):.venv/bin/python .claude/skills/write-next-chapter/scripts/test_candidate_cas.py 跑法(仓库根目录):.venv/bin/python tests/skills/write-next-chapter/test_candidate_cas.py
""" """
from __future__ import annotations from __future__ import annotations
@ -16,11 +16,14 @@ import pathlib
import sys import sys
import unittest import unittest
SCRIPT_DIR = pathlib.Path(__file__).resolve().parent PROJECT_ROOT = pathlib.Path(__file__).resolve().parents[3]
SKILLS_DIR = SCRIPT_DIR.parents[1] SKILLS_DIR = PROJECT_ROOT / ".claude" / "skills"
TEST_DIR = PROJECT_ROOT / "tests" / "skills" / "write-next-chapter"
CHECK_TEST_DIR = PROJECT_ROOT / "tests" / "skills" / "check-content-consistency"
SCRIPT_DIR = SKILLS_DIR / "write-next-chapter" / "scripts"
DETECT_DIR = SKILLS_DIR / "check-content-consistency" / "scripts" DETECT_DIR = SKILLS_DIR / "check-content-consistency" / "scripts"
READ_CONTEXT_DIR = SKILLS_DIR / "assemble-context" / "scripts" READ_CONTEXT_DIR = SKILLS_DIR / "assemble-context" / "scripts"
for path in (SCRIPT_DIR, DETECT_DIR, READ_CONTEXT_DIR): for path in (READ_CONTEXT_DIR, DETECT_DIR, SCRIPT_DIR, TEST_DIR, CHECK_TEST_DIR):
if str(path) not in sys.path: if str(path) not in sys.path:
sys.path.insert(0, str(path)) sys.path.insert(0, str(path))

View File

@ -6,7 +6,7 @@ PASSED 终态、身份不可变,以及条件 UPDATE 在真实并发语义下
测试行用 unittest-cas- 前缀的 run_id,结束前物理清理(本表是可变注册表,无 append-only 约束)。 测试行用 unittest-cas- 前缀的 run_id,结束前物理清理(本表是可变注册表,无 append-only 约束)。
跑法(需 Tailscale 内网可达 muse-example): 跑法(需 Tailscale 内网可达 muse-example):
.venv/bin/python .claude/skills/write-next-chapter/scripts/test_candidate_cas_db.py .venv/bin/python tests/skills/write-next-chapter/test_candidate_cas_db.py
""" """
from __future__ import annotations from __future__ import annotations
@ -15,8 +15,11 @@ import pathlib
import sys import sys
import uuid import uuid
SCRIPT_DIR = pathlib.Path(__file__).resolve().parent from psycopg.errors import RaiseException
SKILLS_DIR = SCRIPT_DIR.parents[1]
PROJECT_ROOT = pathlib.Path(__file__).resolve().parents[3]
SKILLS_DIR = PROJECT_ROOT / ".claude" / "skills"
SCRIPT_DIR = SKILLS_DIR / "write-next-chapter" / "scripts"
for path in (SCRIPT_DIR, SKILLS_DIR / "access-database" / "scripts"): for path in (SCRIPT_DIR, SKILLS_DIR / "access-database" / "scripts"):
if str(path) not in sys.path: if str(path) not in sys.path:
sys.path.insert(0, str(path)) sys.path.insert(0, str(path))
@ -107,7 +110,7 @@ def test_db_trigger_rejects_malformed_direct_updates() -> None:
raise AssertionError(f"触发器必须拒绝: {sql}") raise AssertionError(f"触发器必须拒绝: {sql}")
except AssertionError: except AssertionError:
raise raise
except Exception: except RaiseException:
pass pass
latest = store.latest(run_id) latest = store.latest(run_id)
assert (latest.state, latest.revision) == ("DRAFT", 1), "非法 UPDATE 不得改变链" assert (latest.state, latest.revision) == ("DRAFT", 1), "非法 UPDATE 不得改变链"
@ -130,7 +133,7 @@ def test_start_next_monotonic_at_db_level() -> None:
raise AssertionError("REJECTED 开新轮必须递增 attempt/candidate_version") raise AssertionError("REJECTED 开新轮必须递增 attempt/candidate_version")
except AssertionError: except AssertionError:
raise raise
except Exception: except RaiseException:
pass pass
assert store.latest(run_id).state == "REJECTED", "非法开新轮不得改变链" assert store.latest(run_id).state == "REJECTED", "非法开新轮不得改变链"
@ -146,7 +149,7 @@ def test_insert_must_start_draft_revision_one() -> None:
raise AssertionError("初始行必须是 DRAFT") raise AssertionError("初始行必须是 DRAFT")
except AssertionError: except AssertionError:
raise raise
except Exception: except RaiseException:
pass pass
try: try:
with connect() as conn: with connect() as conn:
@ -157,7 +160,7 @@ def test_insert_must_start_draft_revision_one() -> None:
raise AssertionError("初始 revision 必须是 1") raise AssertionError("初始 revision 必须是 1")
except AssertionError: except AssertionError:
raise raise
except Exception: except RaiseException:
pass pass

View File

@ -5,7 +5,8 @@ import sys
import unittest import unittest
SCRIPT_DIR = pathlib.Path(__file__).resolve().parent PROJECT_ROOT = pathlib.Path(__file__).resolve().parents[3]
SCRIPT_DIR = PROJECT_ROOT / ".claude" / "skills" / "write-next-chapter" / "scripts"
sys.path.insert(0, str(SCRIPT_DIR)) sys.path.insert(0, str(SCRIPT_DIR))
import persist_writer_run as writer_persist # noqa: E402 import persist_writer_run as writer_persist # noqa: E402

View File

@ -12,10 +12,13 @@ import sys
import unittest import unittest
from decimal import Decimal from decimal import Decimal
SCRIPT_DIR = pathlib.Path(__file__).resolve().parent PROJECT_ROOT = pathlib.Path(__file__).resolve().parents[3]
READ_CONTEXT_DIR = SCRIPT_DIR.parents[1] / "assemble-context" / "scripts" SCRIPT_DIR = PROJECT_ROOT / ".claude" / "skills" / "write-next-chapter" / "scripts"
READ_CONTEXT_DIR = PROJECT_ROOT / ".claude" / "skills" / "assemble-context" / "scripts"
TEST_CONTEXT_DIR = PROJECT_ROOT / "tests" / "skills" / "assemble-context"
sys.path.insert(0, str(SCRIPT_DIR)) sys.path.insert(0, str(SCRIPT_DIR))
sys.path.insert(0, str(READ_CONTEXT_DIR)) sys.path.insert(0, str(READ_CONTEXT_DIR))
sys.path.insert(0, str(TEST_CONTEXT_DIR))
from test_writer_contract import valid_context, valid_draft # noqa: E402 from test_writer_contract import valid_context, valid_draft # noqa: E402
from writer_contract import build_writer_creative_input, canonical_json, retrieval_identity # noqa: E402 from writer_contract import build_writer_creative_input, canonical_json, retrieval_identity # noqa: E402

View File

@ -12,11 +12,14 @@ import tempfile
import unittest import unittest
from unittest import mock from unittest import mock
SCRIPT_DIR = pathlib.Path(__file__).resolve().parent PROJECT_ROOT = pathlib.Path(__file__).resolve().parents[3]
SKILLS_DIR = SCRIPT_DIR.parents[1] SKILLS_DIR = PROJECT_ROOT / ".claude" / "skills"
TEST_DIR = PROJECT_ROOT / "tests" / "skills" / "write-next-chapter"
CHECK_TEST_DIR = PROJECT_ROOT / "tests" / "skills" / "check-content-consistency"
SCRIPT_DIR = SKILLS_DIR / "write-next-chapter" / "scripts"
DETECT_DIR = SKILLS_DIR / "check-content-consistency" / "scripts" DETECT_DIR = SKILLS_DIR / "check-content-consistency" / "scripts"
READ_CONTEXT_DIR = SKILLS_DIR / "assemble-context" / "scripts" READ_CONTEXT_DIR = SKILLS_DIR / "assemble-context" / "scripts"
for path in (SCRIPT_DIR, DETECT_DIR, READ_CONTEXT_DIR): for path in (READ_CONTEXT_DIR, DETECT_DIR, SCRIPT_DIR, TEST_DIR, CHECK_TEST_DIR):
sys.path.insert(0, str(path)) sys.path.insert(0, str(path))
from test_check_writer_candidate import _candidate_body, _requirements, _valid_pair # noqa: E402 from test_check_writer_candidate import _candidate_body, _requirements, _valid_pair # noqa: E402

View File

@ -4,7 +4,7 @@
语义状态固化到候选行是接受通道 DB 兜底的依据,绑定/哈希校验必须失败关闭: 语义状态固化到候选行是接受通道 DB 兜底的依据,绑定/哈希校验必须失败关闭:
报告版本、候选哈希、候选版本、上下文哈希、运行 ID 任一不一致都拒绝落库。 报告版本、候选哈希、候选版本、上下文哈希、运行 ID 任一不一致都拒绝落库。
跑法:.venv/bin/python .claude/skills/write-next-chapter/scripts/test_semantic_verdict.py 跑法:.venv/bin/python tests/skills/write-next-chapter/test_semantic_verdict.py
""" """
from __future__ import annotations from __future__ import annotations
@ -14,7 +14,8 @@ import pathlib
import sys import sys
import unittest import unittest
SCRIPT_DIR = pathlib.Path(__file__).resolve().parent PROJECT_ROOT = pathlib.Path(__file__).resolve().parents[3]
SCRIPT_DIR = PROJECT_ROOT / ".claude" / "skills" / "write-next-chapter" / "scripts"
if str(SCRIPT_DIR) not in sys.path: if str(SCRIPT_DIR) not in sys.path:
sys.path.insert(0, str(SCRIPT_DIR)) sys.path.insert(0, str(SCRIPT_DIR))