diff --git a/.agent/docs/architecture/domains/07-Agent与Skill领域.md b/.agent/docs/architecture/domains/07-Agent与Skill领域.md index aded6ab..0a9ddb8 100644 --- a/.agent/docs/architecture/domains/07-Agent与Skill领域.md +++ b/.agent/docs/architecture/domains/07-Agent与Skill领域.md @@ -31,7 +31,7 @@ Agent 与 Skill 领域拥有角色职责、可调用能力合同、确定性工 6. 输入产出落库:哪些输入和产出必须进库可见,见索引 §3。 7. 稳定错误码和失败恢复。 8. raw、授权、预算和审计边界。 -9. 离线测试与验收命令。 +9. 开发验证与行为评测由 `harness/manifests/` 登记;`SKILL.md` 不承载测试命令、测试文件路径或测试结论。Skill 运行时需要执行的业务 dry-run/评测操作仍属于运行合同。 数据库是权威:Skill 对自己读写哪些表负责,声明失败时如何关闭,并确保经手的输入和产出都落库。没落库的输入产出,在系统视角里等于不存在。只读看板只查库渲染,不替 Skill 写任何数据。 @@ -49,7 +49,7 @@ Agent 与 Skill 领域拥有角色职责、可调用能力合同、确定性工 - 默认从仓库任意工作目录调用,必须自行解析项目根和输入绝对路径,不能依赖调用者先 `cd` 到特定目录。 - 机械事实必须结构化输出稳定状态和错误码;人读日志是补充,不是唯一接口。 - Tool 不调用模型,除非所属 Skill 明确声明该步骤本质需要模型。 -- Tool 变更必须有离线测试、`py_compile` 和 `git diff --check` 证据。 +- Tool 变更必须有登记在 `harness/manifests/` 的相关实现测试、`py_compile` 和 `git diff --check` 证据;测试结果只证明机械合同,不升级为 Skill 行为或内容质量结论。 ## 5. ReAct 工作方式 diff --git a/.claude/skills/access-database/SKILL.md b/.claude/skills/access-database/SKILL.md index f024db7..423ba8c 100644 --- a/.claude/skills/access-database/SKILL.md +++ b/.claude/skills/access-database/SKILL.md @@ -24,9 +24,6 @@ description: 通过唯一受控入口查询或修改 muse-example PostgreSQL, # 大对象(raw 全文)经 stdin 传 JSON 数组(避开 shell 转义 / ARG_MAX): .venv/bin/python .claude/skills/access-database/scripts/db.py execparams "INSERT ... VALUES (%s,%s)" --stdin < params.json # params.json = ["", "<完整全文>"] -# execparams 参数装载逻辑离线自测(不连库) -.venv/bin/python .claude/skills/access-database/scripts/test_db_params.py - # SQL 文件应用:整文件一个事务,失败全回滚 .venv/bin/python .claude/skills/access-database/scripts/db.py apply db/ddl/91-example实验私货.sql diff --git a/.claude/skills/call-content-model/SKILL.md b/.claude/skills/call-content-model/SKILL.md index 41686a4..e02f92c 100644 --- a/.claude/skills/call-content-model/SKILL.md +++ b/.claude/skills/call-content-model/SKILL.md @@ -7,7 +7,7 @@ description: 通过 New-API 的统一治理入口调用内容模型,执行额 创始人拍板(2026-07-13):清洗与拆书的内容生产 LLM **全部走 New-API 的 MiniMax-M3**;主会话(Fable5)只固化 agent/提示词/skill 与发起调用。本 skill 是唯一出口。 -管线内容生产调用的**标准入口是 `chat_governed`**(受 5 小时额度窗 + 全局降级链治理);`chat`/`chat` CLI 是不受治理的直连,仅供调试。治理政策的机械事实源是 `.claude/skills/call-content-model/scripts/llm.py` + `test_quota.py`(AGENTS.md §6 点名),模型链切换必须由该 skill 治理并留下日志。 +管线内容生产调用的**标准入口是 `chat_governed`**(受 5 小时额度窗 + 全局降级链治理);`chat`/`chat` CLI 是不受治理的直连,仅供调试。治理政策的机械事实源是运行配置、共享额度账本 `example_llm_quota` 和模型运行适配器(`llm.py`),模型链切换必须由该 skill 治理并留下日志。 ## 用法 diff --git a/.claude/skills/capture-ai-flavor-cases/SKILL.md b/.claude/skills/capture-ai-flavor-cases/SKILL.md index 52ab29d..b483862 100644 --- a/.claude/skills/capture-ai-flavor-cases/SKILL.md +++ b/.claude/skills/capture-ai-flavor-cases/SKILL.md @@ -121,13 +121,3 @@ disable-model-invocation: true `dashboard/server.py` 的 `/ai-flavor` 默认查这三张表,数据库不可用时才明确标注离线回退;页面不会因打开而重新读取原文。 状态语义:案例卡 `shadow` 只供复核,`canonical` 仅表示获授权且完成评审,`rejected`/`archived` 不进入生成上下文;重验证 `verified` 才能确认、投影样例或消费规则,`stale`(全文哈希变化)、`unavailable`(来源不可得)和 `card_mismatch`(锚点变化)都使当前卡在这些动作上失效,但历史回执保留。 - -## 机械验收 - -```bash -.venv/bin/python .claude/skills/capture-ai-flavor-cases/scripts/test_capture_cases.py -.venv/bin/python .claude/skills/access-database/scripts/test_skill_catalog.py -git diff --check -``` - -测试必须覆盖:来源 hash、重验证 verified/stale/unavailable/card_mismatch、未授权 hash-only、重复 ID、live feedback 来源绑定、shadow 不能投影样例、缺重验证回执不能确认、跨作品/反例门,以及候选评测不足时不能 active。 diff --git a/.claude/skills/diagnose-ai-flavor/SKILL.md b/.claude/skills/diagnose-ai-flavor/SKILL.md index 64fde5b..cc02d89 100644 --- a/.claude/skills/diagnose-ai-flavor/SKILL.md +++ b/.claude/skills/diagnose-ai-flavor/SKILL.md @@ -31,10 +31,3 @@ disable-model-invocation: true - 诊断不得改正文、不得产 patch。 - `decision_proposal` 只是机械层初步建议;候选/语义命中必须经过功能仲裁,不能当执行指令。 - 命中不等于修改命令:发现清单是给仲裁的输入,不是执行指令。 - -## 自测 - -```bash -cd agent-example -.venv/bin/python .claude/skills/diagnose-ai-flavor/scripts/test_diagnose_ai_flavor.py -``` diff --git a/.claude/skills/embed-knowledge/SKILL.md b/.claude/skills/embed-knowledge/SKILL.md index c1096ea..2e8dbb8 100644 --- a/.claude/skills/embed-knowledge/SKILL.md +++ b/.claude/skills/embed-knowledge/SKILL.md @@ -32,16 +32,6 @@ description: 使用固定 Qwen3 嵌入模型将知识草稿或实体批量写入 - **落库字段**:`example_knowledge_embedding(draft_id, content_hash, embed_text, model, dimensions=1024, embedding)`;draft 确认落 entity 后由 confirm 流程把 owner 迁到 `entity_id` 并清空 `draft_id`,关系草稿确认后关闭无 canonical owner 的临时向量。 - 汇报:新嵌 N、跳过 M、失败 K;幂等、失活和冲突原因均输出可追踪明细。 -## 离线验证 - -```bash -.venv/bin/python .claude/skills/embed-knowledge/scripts/test_embed_drafts_offline.py -.venv/bin/python -m py_compile .claude/skills/embed-knowledge/scripts/embed_drafts.py \ - .claude/skills/embed-knowledge/scripts/test_embed_drafts_offline.py -``` - -离线测试只使用 fake connection 检查并发顺序、SQL 条件和 owner 反例,不连接真实数据库,不调用 embedding 或 reset。 - ## 红线 - 调用必须 `trust_env=False`(系统代理会假 502);令牌用 `MUSE_AI_NEW_API_TOKEN`(勿用管理令牌,打 /v1 报无效)。 diff --git a/.claude/skills/establish-voice-baseline/SKILL.md b/.claude/skills/establish-voice-baseline/SKILL.md index 733b1e6..23357cb 100644 --- a/.claude/skills/establish-voice-baseline/SKILL.md +++ b/.claude/skills/establish-voice-baseline/SKILL.md @@ -36,11 +36,3 @@ disable-model-invocation: true - 只读已确认正文与作者样张;不得从草稿或未确认候选归纳基线。 - 基线必须人工确认(`--reviewer` 必填,数据库 CHECK 兜底)。 - 本技能不产一个字正文。 - -## 自测 - -```bash -cd agent-example -.venv/bin/python .claude/skills/establish-voice-baseline/scripts/test_establish_voice_baseline.py -.venv/bin/python humanization/tests/test_humanization_v2.py -``` diff --git a/.claude/skills/evaluate-frozen-replay/SKILL.md b/.claude/skills/evaluate-frozen-replay/SKILL.md index 90b1fab..16d8349 100644 --- a/.claude/skills/evaluate-frozen-replay/SKILL.md +++ b/.claude/skills/evaluate-frozen-replay/SKILL.md @@ -48,14 +48,14 @@ disable-model-invocation: true ## 编排入口 -- `scripts/run_replay.py --mode dry_run`:只执行授权、来源、冻结和三臂 manifest 预检,不调用模型;这是首个机制 smoke 入口。 +- `scripts/run_replay.py --mode dry_run`:只执行授权、来源、冻结和三臂 manifest 预检,不调用模型;用于确认评测执行的前置条件。 - `freeze-context/scripts/load_reference_work.py`:从 PostgreSQL 只读事务组装仓库外临时配置;读取 `example_reference_authorization_snapshot` 当前原文件版本的最新快照,组装 authorization 外层与 snapshot。缺授权记录仍生成可审计配置,但送入 `run_replay` 后必须保持 `blocked_authorization`。 - `scripts/run_replay.py --mode execute`:在全部前置门通过后,依次执行三臂 planner、整组 schema、逐臂盲 detector、两个独立盲 judge、rubric 校验、稳定性门和去盲汇总;`--output-dir` 必须位于仓库外的临时目录。 - detector 输入输出均为 JSON;输入只有匿名候选 ID、候选和公共冻结到 `as_of` 的规划上下文,不含 arm 名、`cardInjection`、`cardManifest`、任何臂特有卡内容、目标章 proxy 或其他评委结果。卡注入合法性只由确定性预检负责。报告不合约属于系统失败,须在 judge 前失败关闭;合同合法的 `failed` / `needs_evidence` 属于候选质量信号,三臂均须保留候选与报告并继续盲评和 Gate,Gate A 只由 C 臂高严重度残留与硬约束覆盖率裁决质量失败。 - 合法的非 `passed` 报告在安全 manifest 与 CAS 中只保存按臂分组的 `semantic-diagnostic-v1`:固定结果枚举、固定原因/区段枚举、调用次数和有界阻断计数。完整报告只留受控 raw 和编排器内存中的 Gate builder 输入;安全输出不得保存错误消息、正文、引文、事实文本、补证查询/原因、任何检测项 ID/字段路径、纠错草稿或模型原始字段。检测器不合约时同样只保存固定闭集诊断并失败关闭。 - 任一样本发生系统失败、检测器不合约、盲评无效/不稳定、预算或 raw/CAS 失败后,整轮已无法构建完整 Gate 输入,必须立即停止后续样本并进入统一迁移;不得继续调用模型消耗预算,也不得因失败删除本轮诊断 raw。 - 两个 judge 使用不同 `judgeId` 和独立无会话进程。第二个 judge 的匿名候选顺序必须与第一个完全相反;任何 rubric 不合约标记 `judge_invalid`,任一同维差值大于 `0.5` 标记 `judge_unstable`,两者都不得标记 `completed`。 -- 只有三臂 schema 合法、detector 均产出合同合法报告、双 judge rubric 与稳定性门通过,才去盲生成逐维 `B-A` / `C-A` 差值矩阵并标记 `completed`;detector 报告合同合法不等于其质量终态必须为 `passed`。`--detector-bin`、`--judge-primary-bin`、`--judge-secondary-bin` 可分别指定本地 runner;未指定时复用 `--planner-bin`,测试只能使用 fake binary。 +- 只有三臂 schema 合法、detector 均产出合同合法报告、双 judge rubric 与稳定性门通过,才去盲生成逐维 `B-A` / `C-A` 差值矩阵并标记 `completed`;detector 报告合同合法不等于其质量终态必须为 `passed`。`--detector-bin`、`--judge-primary-bin`、`--judge-secondary-bin` 可分别指定本地 runner;未指定时复用 `--planner-bin`;离线评测可使用 fake adapter,真实评测执行使用配置绑定的 runner。 - planner、detector、judge 子进程统一受 `--timeout-seconds` 限制,默认 300 秒;任一超时分别落盘 `planner_timeout`、`detector_timeout`、`judge_timeout`,不得继续进入后续阶段或标记 `completed`。 - `scripts/write_report.py`:从 `run_result.json` 生成独立严格 schema 的安全摘要,只接受受限标识符、枚举、数字、短安全摘要和 SHA-256;不会读取候选正文,也不会把候选路径以外的原始响应写入报告。 @@ -67,7 +67,7 @@ disable-model-invocation: true 真实五章装配还必须在 `commonControls` 预注册选择器版本、选择器规范 JSON 的原始字节 SHA-256、`inputProvenance=oracle_reference_scaffold` 和统一 `maxContextChars=140000`。五个样本的 A/C 补充原文预算统一预注册为 2000 Unicode code point;loader 在数据库读取前机械核对这些值,不得根据实际文本降低或抬高预算。连续四章基线装不下、A/C 任一臂补充原文不足 2000,或两臂 WriterCreativeInput hash 相同/允许差异为空时均失败关闭。 -每个样本必须提供 `writerContextInput`。仓内基础配置只允许 `contentMode=sanitized_contract_fixture`,并为 A/C 各提供恰好 2000 code point 的明确脱敏合成补充原文,用于机械证明 WriterContext 合同、章号冻结和创作输入差异;这些文本不得来自原书。loader 装配真实临时配置后必须改为 `canonical_frozen_prose`。生产 `--execute` 在创建 vault 或 runner 前拒绝脱敏夹具;脱敏夹具只可由离线测试适配器执行。任何模式都不得把原书全文、完整目标细纲或标准答案写入 Git。 +每个样本必须提供 `writerContextInput`。仓内基础配置只允许 `contentMode=sanitized_contract_fixture`,并为 A/C 各提供恰好 2000 code point 的明确脱敏合成补充原文,用于机械证明 WriterContext 合同、章号冻结和创作输入差异;这些文本不得来自原书。loader 装配真实临时配置后必须改为 `canonical_frozen_prose`。生产 `--execute` 在创建 vault 或 runner 前拒绝脱敏夹具;脱敏夹具只可由离线评测适配器执行。任何模式都不得把原书全文、完整目标细纲或标准答案写入 Git。 三臂都必须构造并校验完整 `WriterContext v1`,且固定 `mode=diagnostic_only`、`purpose=evaluation`、`acceptanceEligible=false`。adapter 随后投影 `WriterCreativeInput v2`;writer 模型只看到创作投影,不看到运行身份、manifest、hash、实验臂或验收状态: @@ -99,8 +99,6 @@ dry-run 只输出计划、manifest 和上下文摘要,不调用 writer、seman 正式 execute 还必须显式传 `--raw-archive-dir <仓外绝对目录>`;缺失、相对路径、位于本轮输出目录内或与临时 vault 跨文件系统时,均在建立 vault 和调用模型前失败关闭。运行结束以 `rawDisposition.status=migrated` 和 `raw-vault-migration-receipt-v1` 证明产物已迁移,不再以删除后的 `closed` 作为完成条件。 -离线完整链冒烟由 `scripts/test_run_writer_replay.py` 的单样本三臂 production fake 用例承担:它真实经过 writer 投影、机械门、semantic detector、双评委、raw vault、CAS 和 GateInputBuilder,但所有模型均为本地确定性 fake,不产生费用、不得作为 Gate 样本结果。真实一次调用能力只由 `refresh_runtime_probe.py` 的合成 writer 探针验证;它不代表 detector/judge 或五样本 Gate 已通过。 - 预算合同分为不可变的 `plannedCalls` 和独立安全上限 `maxCalls`。Gate A 当前预注册 writer 60 次(15 次基础写作 + 最多 45 次篇幅修订)、semantic detector 24 次(15 次基础检测 + 9 次纠错/API 重试余量)、blind judge 45 次(最多三位评委,每位基础调用后最多两个格式纠错/API 重试槽位);各 `maxCalls=150`、单次 cap `$5`、总预算 `$2250`。启动预留只按 `plannedCalls * maxBudgetUsdPerCall`,即 `$645`;`$2250` 是三角色各 150 次安全容量对应的有限执行上限,不是预计消费。本预算授权不等于正式 `--execute` 授权。每次调用前账本同时检查角色计划槽位、`maxCalls`、累计实际成本和未执行计划的最坏预留;调用后只能使用可信 `ExecutionReceipt.totalCostUsd` 结算。缺回执成本、重复/错序结算、单次 cap 或总预算越界均 fail closed,未触发的修订、纠错与第三评计划必须保留在 `remainingPlannedCalls`。 ## 稳定运行纪律 @@ -124,12 +122,12 @@ Gate 裁决器 `writer_gate.py`(归属 `score-content-quality`,本 Skill 只 `scripts/refresh_runtime_probe.py` 是刷新 `executionAuthorization.runtimeProbe` 的唯一通道。正式执行门要求探针的 `executionProfileSha256` 必须等于当前 writer profile 的身份哈希;writer 合同(schema/prompt/模型/CLI/预算)升级后旧探针必然失配,门以 `EXECUTE_PROBE_CONTRACT_MISMATCH` 失败关闭,此时必须用当前合同重测探针,不得手动改探针值绕过。 - 输入合同:只读 `--config` 指定的 base 配置,只用 `executionProfiles.writer`(探针只验证 writer 合同),不读 semantic/judge,不读 `oracleTruthPacks`/`samples`/raw 物料。探针输入是固定、极小、全合成的能力探针 prompt(一段自包含的合成微任务),绝不使用原书全文、真实正文或 raw。 -- 调用合同:Claude 调用经可注入 invoker 发起,默认 invoker = `claude_runtime.run_claude`(真实调用层是薄薄一层);`--dry-run` 注入固定假产出,只走通「构建 profile + 重签 + 写文件」链路供冒烟,不发起真实调用。 +- 调用合同:Claude 调用经可注入 invoker 发起,默认 invoker = `claude_runtime.run_claude`(真实调用层是薄薄一层);`--dry-run` 注入固定假产出,只走通「构建 profile + 重签 + 写文件」链路供离线预览,不发起真实调用。 - 输出合同:把绑定新合同的探针替换 `executionAuthorization.runtimeProbe`,并按与门校验逐字节同源的规范自哈希算法重签该探针的 `receiptSha256`(`budget`/`rawRetention` 原样保留、各自自哈希不变),把刷新后的完整配置写到 `--output` 指定的【新文件】,不就地覆盖原配置,便于 review diff 后再替换。新探针字段集与现有 runtimeProbe 完全一致,不缺不多。 - 失败关闭:调用失败 / 结构化输出 schema 不过 / 单次成本超冻结 cap / 超 deadline / 回执不可信(模型不匹配、标记错误、非零退出、异常终止、API 错误、身份哈希指向旧合同、输出哈希与产出不一致)→ 不写 successful 探针、不产出刷新配置、退出非零,原配置保持不动。 - 审计合同:标准输出只记录做了什么(探针身份哈希、结构化输出哈希、回执哈希、成本、是否成功、输出文件),绝不打印完整 prompt/response、raw 路径或供应商原始响应。 -离线用法(冒烟,不调模型): +离线用法(不调用模型): ```bash .venv/bin/python .claude/skills/evaluate-frozen-replay/scripts/refresh_runtime_probe.py \ diff --git a/.claude/skills/execute-claude-task/SKILL.md b/.claude/skills/execute-claude-task/SKILL.md index 2d65c9b..a181cc6 100644 --- a/.claude/skills/execute-claude-task/SKILL.md +++ b/.claude/skills/execute-claude-task/SKILL.md @@ -25,9 +25,3 @@ disable-model-invocation: true - 不从全局 Claude 会话继承业务状态,不静默换模型或供应商。 - 不自行决定候选是否通过,不写 Candidate、Quality 或 Canonical 数据。 - 不拥有 CAS、raw vault、运行登记或回执修复实现。 - -## 离线验证 - -```bash -.venv/bin/python .claude/skills/execute-claude-task/scripts/test_claude_runtime.py -``` diff --git a/.claude/skills/extract-work-knowledge/SKILL.md b/.claude/skills/extract-work-knowledge/SKILL.md index f4696ce..198093c 100644 --- a/.claude/skills/extract-work-knowledge/SKILL.md +++ b/.claude/skills/extract-work-knowledge/SKILL.md @@ -43,10 +43,3 @@ disable-model-invocation: true - 写入 `muse_knowledge_draft`、`example_upgrade_window`、`example_upgrade_alias`、`example_upgrade_presence`、`example_upgrade_card_state`、`example_upgrade_audit` 和 `example_knowledge_embedding`。 - `SOURCE_TYPE="upgrade_book"`、`upgrade-reset` 等是数据库历史兼容值,不随 Skill 改名。 - 不提供修复、重置、备份、恢复或存量迁移入口;这些动作统一由 `maintain-work-extraction` 承担。 - -## 离线验证 - -```bash -.venv/bin/python .claude/skills/extract-work-knowledge/scripts/test_upgrade_work_lock_offline.py -.venv/bin/python .claude/skills/extract-work-knowledge/scripts/test_parse_upgrade_offline.py -``` diff --git a/.claude/skills/maintain-work-extraction/SKILL.md b/.claude/skills/maintain-work-extraction/SKILL.md index adc455f..6bb8dff 100644 --- a/.claude/skills/maintain-work-extraction/SKILL.md +++ b/.claude/skills/maintain-work-extraction/SKILL.md @@ -69,10 +69,3 @@ reset 软删目标 `upgrade_book` 草稿和无 Canonical owner 的活向量, - 备份与 preview 默认只读;execute 只修改目标租户、目标作品、`upgrade_book` 边界内的派生状态。 - 不生成新抽取内容,不提供正常 `run/windows/status`。 - `migrate_upgrade_windows.py` 当前只输出迁移计划,不落库;真实写入仍需另行 gate 与授权。 - -## 离线验证 - -```bash -.venv/bin/python .claude/skills/maintain-work-extraction/scripts/test_backup_upgrade_work_offline.py -.venv/bin/python .claude/skills/maintain-work-extraction/scripts/test_reset_upgrade_work_offline.py -``` diff --git a/.claude/skills/prevent-ai-flavor/SKILL.md b/.claude/skills/prevent-ai-flavor/SKILL.md index a21eb90..1ceab1a 100644 --- a/.claude/skills/prevent-ai-flavor/SKILL.md +++ b/.claude/skills/prevent-ai-flavor/SKILL.md @@ -31,11 +31,3 @@ disable-model-invocation: true - 读:`example_voice_baseline` 当前 canonical 版本;规则/样例读 `humanization/` Git 资产。 - 写:`example_run`(幂等 upsert,记录规则指纹、声音账 hash 和投影条数)。 - 失败:作品不匹配、声音账非 canonical、规则装载门失败或 guidance 超合同直接失败关闭。 - -## 自测 - -```bash -cd agent-example -.venv/bin/python .claude/skills/prevent-ai-flavor/scripts/test_prevent_ai_flavor.py -.venv/bin/python .claude/skills/assemble-context/scripts/test_assemble_writer_context.py -``` diff --git a/.claude/skills/record-run-evidence/SKILL.md b/.claude/skills/record-run-evidence/SKILL.md index 22ec8bd..3c0e78a 100644 --- a/.claude/skills/record-run-evidence/SKILL.md +++ b/.claude/skills/record-run-evidence/SKILL.md @@ -32,11 +32,3 @@ disable-model-invocation: true - 读取运行、调用、raw 和回执当前状态,只为校验绑定与幂等。 - 写入 `example_run`、`example_llm_call`、`example_raw_lease`、`example_raw_content`、`example_run_receipt` 和相应质量失败记录。 - 不调用 Claude/New-API,不生成候选,不判断内容是否通过。 - -## 离线验证 - -```bash -.venv/bin/python .claude/skills/record-run-evidence/scripts/test_file_cas.py -.venv/bin/python .claude/skills/record-run-evidence/scripts/test_raw_vault.py -.venv/bin/python .claude/skills/record-run-evidence/scripts/test_run_registry.py -``` diff --git a/.claude/skills/revise-ai-flavor/SKILL.md b/.claude/skills/revise-ai-flavor/SKILL.md index 2ec4fef..93b47ef 100644 --- a/.claude/skills/revise-ai-flavor/SKILL.md +++ b/.claude/skills/revise-ai-flavor/SKILL.md @@ -30,10 +30,3 @@ disable-model-invocation: true - 没有诊断产物、作者授权或功能仲裁不得修订。 - 硬门失败不得以风格分、盲评结果抵消;旧诊断不能跨规则库/正文 hash 复用。 - 语义级不变量(因果、POV、伏笔状态)机械未覆盖的,如实列 `unresolved_risks`,不假装验证完成。 - -## 自测 - -```bash -cd agent-example -.venv/bin/python .claude/skills/revise-ai-flavor/scripts/test_revise_ai_flavor.py -``` diff --git a/AGENTS.md b/AGENTS.md index d76d8e6..36fcd62 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -23,7 +23,8 @@ SoT 按主题分域,不做跨主题的全局排序。可执行脚本与书面 | [`.agent/docs/architecture/domains/`](.agent/docs/architecture/domains/_index.md) | 本仓领域边界、数据权威、落库合同和领域协作的 SoT。 | | [`meta/schemas/`](meta/schemas/) | 23 型结构本体的字段合同;库内 payload 结构以该合同为准。 | | [`meta/chains/`](meta/chains/) | scenario、purpose、功能 skill、角色槽位和保护节点的链路登记。 | -| [`.claude/skills/`](.claude/skills/) | 每个 `SKILL.md` 定义能力的输入、输出、红线、数据库读写合同、输入产出落库和自测入口;同目录 `scripts/` 是确定性实现与机械门。 | +| [`.claude/skills/`](.claude/skills/) | 每个 `SKILL.md` 定义项目运行时能力的输入、输出、红线、数据库读写合同和输入产出落库;同目录 `scripts/` 是运行时确定性实现与机械门,开发验证和行为评测入口由 `harness/manifests/` 登记。 | +| [`tests/skills/`](tests/skills/) | 项目运行时 Skill 的实现测试、集成测试和 fake pipeline 测试;它们提供回归证据,不拥有运行时合同,也不等同于 Skill 行为评测。 | | [`.agent/`](.agent/_index.md) | 跨任务长期知识;领域设计位于 `docs/architecture/domains/`,新增、删除或重命名必须同步各级 `_index.md`。 | | [`docs/`](docs/) | 单次任务探索、计划、评测资料、样张和历史执行证据;任务完成后把稳定结论蒸馏到 `.agent/`,不得长期拥有领域定义。 | | [`README.md`](README.md) | 项目背景和历史路线概览。目录、阶段、技能数量和存储方式等描述可能陈旧,不得覆盖本文件、`meta/`、skill 或磁盘事实。 | @@ -46,7 +47,9 @@ agent-example/ │ ├── ddl/ # 可审计 DDL / 迁移文件 │ ├── 表映射.md │ └── 连接信息.md -├── humanization/ # 去 AI 味与人感资产层(合同/规则/样例/执行骨架;验证期自治,升华进 Muse 时同步回父仓;20 项研究覆盖矩阵见 humanization/research/) +├── humanization/ # 去 AI 味 Skill 能力域(规则/样例/执行骨架;运行时资产目标为数据库) +├── harness/ # 项目验证与外部评测索引、清单和调度支架 +├── tests/ # 运行时 Skill 实现测试(按 Skill 归档) ├── knowledge/ # 仓内参考资产;未经绑定、授权不得进入上下文 ├── docs/ # 设计、评测、样张与历史执行记录 ├── .venv/ # 本地 Python 运行环境 @@ -57,15 +60,23 @@ agent-example/ 5 个角色:`writer`、`planner`、`extractor`、`detector`、`judge`。角色身份在 `.claude/agents/*.md`,具体功能合同不复制进角色文件。 -技能按能力分组如下,实际清单以 `.claude/skills/*/SKILL.md` 为准,不手填数量镜像: +### Skill 领域索引(项目运行时 Skill) -- 数据与检索:`access-database`、`import-book`、`embed-knowledge`、`search-knowledge`、`call-content-model`。 -- 执行、证据与上下文:`execute-claude-task`、`record-run-evidence`、`freeze-context`、`assemble-context`。 -- 流程与主权治理:`decide-candidate`。 -- 清洗、拆解与知识:`clean-book-text`、`deconstruct-book`、`extract-chapter-knowledge`、`review-knowledge-cards`、`extract-work-knowledge`、`maintain-work-extraction`。 -- 去 AI 味与人感(父仓专题-09 五技能先行验证):`establish-voice-baseline`(定基线)、`prevent-ai-flavor`(前置预防)、`diagnose-ai-flavor`(诊断)、`revise-ai-flavor`(修订)、`capture-ai-flavor-cases`(挖掘/案例采集);共享资产层在 `humanization/`,接力铁律见 `meta/chains/`。 -- 规划与写作:`design-story-foundation`、`plan-story`、`plan-chapter`、`write-next-chapter`、`rewrite-selection`、`expand-scene`、`polish-prose`。 -- 检测与质量评测:`check-content-consistency`、`score-content-quality`、`optimize-content-quality`、`evaluate-frozen-replay`。 +实际清单以 `.claude/skills/*/SKILL.md` 为准;本表维护合同责任方、协作领域和领域 SoT,不复制各 Skill 的完整合同。每个运行时 Skill 必须登记一个合同责任方(业务领域或平台领域),但可以同时消费或影响多个协作领域;跨域调用、场景关系和保护节点在 `meta/chains/` 登记。合同责任方表示谁维护该 Skill 的稳定能力合同,不表示 Skill 只能属于一个业务领域。 + +| 合同责任方 / 能力域 | 领域 SoT | 运行时 Skill | +|---|---|---| +| 平台运行与证据 | 07-Agent 与 Skill、08-数据权威与可视化 | `access-database`、`call-content-model`、`execute-claude-task`、`record-run-evidence` | +| 上下文与知识检索 | 02-实体、04-上下文、08-数据权威与可视化 | `assemble-context`、`embed-knowledge`、`freeze-context`、`search-knowledge` | +| 导入、清洗与抽取 | 02-实体、05-创作流程 | `clean-book-text`、`deconstruct-book`、`extract-chapter-knowledge`、`extract-work-knowledge`、`import-book`、`maintain-work-extraction`、`review-knowledge-cards` | +| 规划与作品基础 | 05-创作流程 | `design-story-foundation`、`plan-chapter`、`plan-story` | +| 写作与候选主权 | 01-作品、05-创作流程 | `decide-candidate`、`expand-scene`、`polish-prose`、`rewrite-selection`、`write-next-chapter` | +| 质量与回放评测 | 06-质量与复利、05-创作流程 | `check-content-consistency`、`evaluate-frozen-replay`、`optimize-content-quality`、`score-content-quality` | +| 去 AI 味与人感 | 06-质量与复利、父仓专题-09 | `capture-ai-flavor-cases`、`diagnose-ai-flavor`、`establish-voice-baseline`、`prevent-ai-flavor`、`revise-ai-flavor` | + +`humanization/` 是“去 AI 味与人感”运行时 Skill 家族的能力域:`src/deai/` 是共享实现,规则、样例、案例和声音资产是该域的数据依赖,运行时以数据库为权威;仓内 YAML/JSON 在数据库化完成前只作为迁移种子、离线夹具或结构合同。规则记录不各自注册为 Skill,Skill 负责动作和消费边界。 + +Skill 领域列表的新增、删除、改名或主领域调整,必须同时检查 `.claude/skills/`、`meta/chains/README.md` 和相关领域 `_index.md`;不得只改本表造成索引漂移。 上下文目标合同以 [上下文领域 SoT](.agent/docs/architecture/domains/04-上下文领域.md) 为准:数据库读取器是核心实现,库内检索加速(向量)只做候选召回;任何命中都要回读库行并校验 hash。Skill 对自己读写哪些表负责,并把经手的输入和产出落库;没落库的输入产出在系统视角里等于不存在。 @@ -95,7 +106,7 @@ agent-example/ ## 6. 模型边界 -- 清洗、抽卡、范式拆取及其模型调用统一走 `call-content-model` Skill,不裸调 New-API。治理政策固定为 5 小时额度窗:MiniMax 模型累计花费上限 `$24`,全模型成功调用上限 `6000`;机械事实源是 `.claude/skills/call-content-model/scripts/llm.py` 及 `test_quota.py`,模型链切换必须由该 Skill 治理并留下日志。 +- 清洗、抽卡、范式拆取及其模型调用统一走 `call-content-model` Skill,不裸调 New-API。治理政策固定为 5 小时额度窗:MiniMax 模型累计花费上限 `$24`,全模型成功调用上限 `6000`;运行适配器、正式配置和账本是额度合同的事实源,`.claude/skills/call-content-model/scripts/llm.py` 是实现,`test_quota.py` 只提供回归证据;模型链切换必须由该 Skill 治理并留下日志。 - 角色模型归属:`planner`/`writer`/`judge` 固定 `opus`;`extractor`/`detector` 可用其它模型(非必须降级)。拆书/导入侧抽取经 `call-content-model`/`deconstruct-book` Skill 走 MiniMax-M3,不走角色 model 派发;创作期章后抽取作为角色派发,可用 `opus`。 - 确定性脚本、合同校验、快照冻结、泄漏审计和报告生成不调用模型;除非对应 `SKILL.md` 明确声明模型步骤,不得把机械任务升级为模型任务。 - Claude 生成或评测只在对应任务 SoT、显式预算、固定执行配置和原文用途授权全部满足后运行;任一前置门失败都应关闭执行。 @@ -130,12 +141,12 @@ agent-example/ - 明确区分**已验证事实**、**推断**和**假设**。报告必须给出证据来源;没有机械输出或运行证据时,不声称完成、修复或通过。 - 禁止从 `n=1` 样本推出普适结论。至少分析假阴、假阳、样本偏差和混淆因素;需要判断卡或模型效果时使用同任务、同模型、同预算、同公共上下文的对照,并把不稳定样本排除在方向结论之外。 -- Python 一律使用仓内解释器 `.venv/bin/python`。先读目标 skill 的 `SKILL.md`,再运行其已存在的相关自测;例如: +- Python 一律使用仓内解释器 `.venv/bin/python`。先读目标 Skill 的 `SKILL.md`,再按 `harness/manifests/` 登记的类别和依赖选择相关验证;下面命令仅是现有局部验证入口示例: ```bash -.venv/bin/python .claude/skills/plan-chapter/scripts/test_contract.py -.venv/bin/python .claude/skills/call-content-model/scripts/test_quota.py -.venv/bin/python .claude/skills/assemble-context/scripts/test_writer_contract.py +.venv/bin/python tests/skills/plan-chapter/test_contract.py +.venv/bin/python tests/skills/call-content-model/test_quota.py +.venv/bin/python tests/skills/assemble-context/test_writer_contract.py git diff --check ``` @@ -163,7 +174,7 @@ git diff --check 1. **给谁用**:主会话编排?某角色 agent?还是别的 skill?消费者必须明确、单一。 2. **目的单一**:一个 skill 只实现一个能力。功能并列、又当编排又当执行的,拆。 3. **可靠性与稳定性**:失败是否明确失败关闭(不静默降级、不返回假成功)?错误是否带稳定码、不泄漏原文/密钥?有无离线自测覆盖关键路径与失败路径?确定性步骤是否真不调模型? -4. **符合 skill 规范**:`SKILL.md` 是否清楚定义输入、输出、红线、用法?`scripts/` 是否其确定性实现与自测入口?是否走 `.venv`、不裸调 PG/New-API/模型? +4. **符合 skill 规范**:`SKILL.md` 是否清楚定义输入、输出、红线、用法?`scripts/` 是否只包含运行时确定性实现与机械门?相关实现测试和行为评测是否已在 `harness/manifests/` 登记?是否走 `.venv`、不裸调 PG/New-API/模型? 5. **边界一致**:引用的路径、合同、字段是否与现行 SoT(本文件、`.agent/`、`meta/` 和库内正式内容)一致?数据库是正式内容权威,Skill 必须声明读写哪些表、失败如何关闭、输入产出如何落库;引用已失效合同的不得执行。 命名采用小写 `动作-对象`:目录名必须等于 frontmatter `name`,名称表达可调用能力,不复用 `scenario`、内部模块名或含混阶段词。`scenario`、`source_type`、updater 和备份逻辑键是独立稳定标识,不随 Skill 改名。 diff --git a/harness/README.md b/harness/README.md new file mode 100644 index 0000000..ba20451 --- /dev/null +++ b/harness/README.md @@ -0,0 +1,68 @@ +# agent-example Harness + +> 项目开发、验证和行为评测支架的入口与索引。 +> 这里的文档和工具服务开发者与评测编排,不会作为任何 Skill 的运行时提示词自动注入 Agent。 + +## 1. 作用 + +`agent-example` 同时包含三种不同性质的验证工作: + +- 运行时文档的静态卫生检查; +- 确定性工具、数据库边界和 fake pipeline 的机械测试; +- 将 Skill 交给 Agent 后的外部行为评测。 + +Harness 负责把这三类工作分开编排并保留证据,不能用其中一类的通过结果冒充另一类结论。 + +## 2. 目录索引 + +```text +harness/ +├── README.md # 本入口:项目 harness 介绍、边界和索引 +├── specs/ +│ └── skill-testing.md # Skill 测试与评测规范(唯一规范正文) +├── skill_harness.py # 运行时 Skill 文档静态审计入口 +├── run_selected.py # 按 manifest 选择性执行测试 +├── test_skill_harness.py # 静态审计器离线测试 +├── test_run_selected.py # 选择性执行器离线测试 +├── manifests/ +│ ├── skills.json # 运行时 Skill 责任方与协作领域清单 +│ └── test-inventory.json # 测试分类、依赖、副作用与证据等级 +└── evals// # Skill 行为评测夹具与适配器(尚未建立) + +项目级实现测试位于 `../tests/skills//`,不进入运行时 Skill 目录。 +``` + +现有专项支架: + +| 路径 | 责任 | 证据边界 | +|---|---|---| +| `humanization/eval/run_eval.py` | 去 AI 味规则合同和回归陷阱试跑 | 规则合同回放,不证明文学效果 | +| `.claude/skills/evaluate-frozen-replay/` | 正文/细纲冻结回放、盲评和质量门 | 评测编排与质量证据,不是通用 Skill 单测入口 | +| `harness/` | 跨 Skill 的开发验证与行为评测治理 | 只编排和裁决证据,不拥有业务合同 | + +## 3. 权威关系 + +- `.claude/skills/*/SKILL.md`:Skill 的运行时行为合同;不放开发测试说明。 +- `.claude/skills/*/scripts/`:Skill 使用的运行时确定性实现与机械门。 +- `tests/skills//`:对应 Skill 的实现测试、集成测试和 fake pipeline 测试;测试性质由 harness 清单标注。 +- `harness/specs/`:测试和评测规范;不被业务 Skill 当作运行时指令读取。 +- `harness/manifests/`:测试/评测登记与执行配置;不复制业务字段合同。 +- `harness/evals/`:外部行为评测案例、适配器和报告生成逻辑。 +- `.agent/docs/architecture/domains/`:领域与数据权威 SoT;harness 只引用,不复制领域定义。 +- `docs/`:单次设计、运行证据和历史资料;不替代 harness 规范。 + +## 4. 当前状态 + +- `skill_harness.py` 已实现运行时 Skill 文档与 `skills.json` 的只读静态审计。 +- `test-inventory.json` 已登记实现测试、集成测试、fake pipeline、领域评测和 harness 自测的依赖与证据等级。 +- `run_selected.py` 已实现显式选择、磁盘/manifest 对账、危险依赖阻断、超时和执行证据检查;它不提供默认全仓一键通过结论。 +- 项目运行时 Skill 的实现测试以 `tests/skills//` 为目标位置,物理现状以 `test-inventory.json` 为准。 +- `skill_behavior_eval` 当前登记数量为 0;没有外部 Agent/模型行为证据时,不声称 Skill 内容有效。 +- PostgreSQL、网络和真实模型证据未由离线结果替代,是否执行仍受授权和预算约束。 + +## 5. 运行边界 + +- 默认只允许静态检查和明确标注的离线测试; +- PostgreSQL、网络、真实模型和额度探针必须显式选择并满足授权、预算和环境前置条件; +- 任何测试失败、依赖缺失或未发现测试都必须显式返回状态,不能静默跳过; +- 报告必须区分已验证事实、推断和未验证假设。 diff --git a/harness/manifests/skills.json b/harness/manifests/skills.json new file mode 100644 index 0000000..425044c --- /dev/null +++ b/harness/manifests/skills.json @@ -0,0 +1,308 @@ +{ + "schema_version": 1, + "skills": [ + { + "name": "access-database", + "contract_owner": "平台运行与证据", + "collaborates_with": [ + "上下文与知识检索", + "导入、清洗与抽取", + "规划与作品基础", + "写作与候选主权", + "质量与回放评测", + "去 AI 味与人感" + ], + "skill_path": ".claude/skills/access-database/SKILL.md" + }, + { + "name": "assemble-context", + "contract_owner": "上下文与知识检索", + "collaborates_with": [ + "规划与作品基础", + "写作与候选主权", + "质量与回放评测", + "去 AI 味与人感" + ], + "skill_path": ".claude/skills/assemble-context/SKILL.md" + }, + { + "name": "call-content-model", + "contract_owner": "平台运行与证据", + "collaborates_with": [ + "导入、清洗与抽取", + "质量与回放评测" + ], + "skill_path": ".claude/skills/call-content-model/SKILL.md" + }, + { + "name": "capture-ai-flavor-cases", + "contract_owner": "去 AI 味与人感", + "collaborates_with": [ + "质量与回放评测", + "写作与候选主权" + ], + "skill_path": ".claude/skills/capture-ai-flavor-cases/SKILL.md" + }, + { + "name": "check-content-consistency", + "contract_owner": "质量与回放评测", + "collaborates_with": [ + "上下文与知识检索", + "规划与作品基础", + "写作与候选主权" + ], + "skill_path": ".claude/skills/check-content-consistency/SKILL.md" + }, + { + "name": "clean-book-text", + "contract_owner": "导入、清洗与抽取", + "collaborates_with": [ + "平台运行与证据" + ], + "skill_path": ".claude/skills/clean-book-text/SKILL.md" + }, + { + "name": "decide-candidate", + "contract_owner": "写作与候选主权", + "collaborates_with": [ + "平台运行与证据", + "质量与回放评测", + "导入、清洗与抽取", + "规划与作品基础" + ], + "skill_path": ".claude/skills/decide-candidate/SKILL.md" + }, + { + "name": "deconstruct-book", + "contract_owner": "导入、清洗与抽取", + "collaborates_with": [ + "平台运行与证据", + "上下文与知识检索", + "质量与回放评测" + ], + "skill_path": ".claude/skills/deconstruct-book/SKILL.md" + }, + { + "name": "design-story-foundation", + "contract_owner": "规划与作品基础", + "collaborates_with": [], + "skill_path": ".claude/skills/design-story-foundation/SKILL.md" + }, + { + "name": "diagnose-ai-flavor", + "contract_owner": "去 AI 味与人感", + "collaborates_with": [ + "质量与回放评测", + "写作与候选主权" + ], + "skill_path": ".claude/skills/diagnose-ai-flavor/SKILL.md" + }, + { + "name": "embed-knowledge", + "contract_owner": "上下文与知识检索", + "collaborates_with": [ + "导入、清洗与抽取" + ], + "skill_path": ".claude/skills/embed-knowledge/SKILL.md" + }, + { + "name": "establish-voice-baseline", + "contract_owner": "去 AI 味与人感", + "collaborates_with": [ + "上下文与知识检索", + "写作与候选主权" + ], + "skill_path": ".claude/skills/establish-voice-baseline/SKILL.md" + }, + { + "name": "evaluate-frozen-replay", + "contract_owner": "质量与回放评测", + "collaborates_with": [ + "平台运行与证据", + "上下文与知识检索", + "写作与候选主权" + ], + "skill_path": ".claude/skills/evaluate-frozen-replay/SKILL.md" + }, + { + "name": "execute-claude-task", + "contract_owner": "平台运行与证据", + "collaborates_with": [ + "质量与回放评测" + ], + "skill_path": ".claude/skills/execute-claude-task/SKILL.md" + }, + { + "name": "expand-scene", + "contract_owner": "写作与候选主权", + "collaborates_with": [ + "上下文与知识检索", + "质量与回放评测" + ], + "skill_path": ".claude/skills/expand-scene/SKILL.md" + }, + { + "name": "extract-chapter-knowledge", + "contract_owner": "导入、清洗与抽取", + "collaborates_with": [ + "平台运行与证据", + "上下文与知识检索", + "写作与候选主权" + ], + "skill_path": ".claude/skills/extract-chapter-knowledge/SKILL.md" + }, + { + "name": "extract-work-knowledge", + "contract_owner": "导入、清洗与抽取", + "collaborates_with": [ + "平台运行与证据", + "上下文与知识检索" + ], + "skill_path": ".claude/skills/extract-work-knowledge/SKILL.md" + }, + { + "name": "freeze-context", + "contract_owner": "上下文与知识检索", + "collaborates_with": [ + "平台运行与证据", + "质量与回放评测" + ], + "skill_path": ".claude/skills/freeze-context/SKILL.md" + }, + { + "name": "import-book", + "contract_owner": "导入、清洗与抽取", + "collaborates_with": [ + "平台运行与证据" + ], + "skill_path": ".claude/skills/import-book/SKILL.md" + }, + { + "name": "maintain-work-extraction", + "contract_owner": "导入、清洗与抽取", + "collaborates_with": [ + "平台运行与证据" + ], + "skill_path": ".claude/skills/maintain-work-extraction/SKILL.md" + }, + { + "name": "optimize-content-quality", + "contract_owner": "质量与回放评测", + "collaborates_with": [ + "上下文与知识检索", + "写作与候选主权" + ], + "skill_path": ".claude/skills/optimize-content-quality/SKILL.md" + }, + { + "name": "plan-chapter", + "contract_owner": "规划与作品基础", + "collaborates_with": [ + "上下文与知识检索", + "质量与回放评测" + ], + "skill_path": ".claude/skills/plan-chapter/SKILL.md" + }, + { + "name": "plan-story", + "contract_owner": "规划与作品基础", + "collaborates_with": [ + "平台运行与证据", + "上下文与知识检索", + "写作与候选主权" + ], + "skill_path": ".claude/skills/plan-story/SKILL.md" + }, + { + "name": "polish-prose", + "contract_owner": "写作与候选主权", + "collaborates_with": [ + "上下文与知识检索", + "质量与回放评测", + "去 AI 味与人感" + ], + "skill_path": ".claude/skills/polish-prose/SKILL.md" + }, + { + "name": "prevent-ai-flavor", + "contract_owner": "去 AI 味与人感", + "collaborates_with": [ + "上下文与知识检索", + "写作与候选主权" + ], + "skill_path": ".claude/skills/prevent-ai-flavor/SKILL.md" + }, + { + "name": "record-run-evidence", + "contract_owner": "平台运行与证据", + "collaborates_with": [ + "上下文与知识检索", + "导入、清洗与抽取", + "规划与作品基础", + "写作与候选主权", + "质量与回放评测", + "去 AI 味与人感" + ], + "skill_path": ".claude/skills/record-run-evidence/SKILL.md" + }, + { + "name": "review-knowledge-cards", + "contract_owner": "导入、清洗与抽取", + "collaborates_with": [ + "平台运行与证据", + "质量与回放评测" + ], + "skill_path": ".claude/skills/review-knowledge-cards/SKILL.md" + }, + { + "name": "revise-ai-flavor", + "contract_owner": "去 AI 味与人感", + "collaborates_with": [ + "质量与回放评测", + "写作与候选主权" + ], + "skill_path": ".claude/skills/revise-ai-flavor/SKILL.md" + }, + { + "name": "rewrite-selection", + "contract_owner": "写作与候选主权", + "collaborates_with": [ + "上下文与知识检索", + "质量与回放评测", + "去 AI 味与人感" + ], + "skill_path": ".claude/skills/rewrite-selection/SKILL.md" + }, + { + "name": "score-content-quality", + "contract_owner": "质量与回放评测", + "collaborates_with": [ + "平台运行与证据", + "上下文与知识检索", + "写作与候选主权" + ], + "skill_path": ".claude/skills/score-content-quality/SKILL.md" + }, + { + "name": "search-knowledge", + "contract_owner": "上下文与知识检索", + "collaborates_with": [ + "规划与作品基础", + "写作与候选主权", + "质量与回放评测" + ], + "skill_path": ".claude/skills/search-knowledge/SKILL.md" + }, + { + "name": "write-next-chapter", + "contract_owner": "写作与候选主权", + "collaborates_with": [ + "平台运行与证据", + "上下文与知识检索", + "质量与回放评测", + "去 AI 味与人感" + ], + "skill_path": ".claude/skills/write-next-chapter/SKILL.md" + } + ] +} diff --git a/harness/manifests/test-inventory.json b/harness/manifests/test-inventory.json new file mode 100644 index 0000000..bd5c408 --- /dev/null +++ b/harness/manifests/test-inventory.json @@ -0,0 +1,1002 @@ +{ + "schema_version": 1, + "generated_scope": "Current agent-example source test assets: .claude/skills/**/test_*.py and *_test.py, tests/skills/** source files, humanization/tests/** source files, harness/**/test_*.py, dashboard/test_server_display.py, and other obvious source test files; excludes .git, .venv, __pycache__ compiled artifacts, deleted working-tree files, and harness specification documents.", + "entries": [ + { + "path": "tests/skills/access-database/test_authorization_snapshot_ddl.py", + "scope": "runtime_skill", + "owner_skill_or_domain": "access-database", + "kind": "tool_contract", + "evidence_level": "static_structure", + "requires": ["offline", "filesystem"], + "side_effects": ["none"], + "skill_behavior_eval": false, + "classification_confidence": "high", + "classification_basis": "Reads the authorization DDL and applies regex/substring invariants; no service call or Agent/model driver." + }, + { + "path": "tests/skills/access-database/test_db_params.py", + "scope": "runtime_skill", + "owner_skill_or_domain": "access-database", + "kind": "tool_unit", + "evidence_level": "deterministic_offline", + "requires": ["offline"], + "side_effects": ["none"], + "skill_behavior_eval": false, + "classification_confidence": "high", + "classification_basis": "Exercises _read_params with StringIO and Click exceptions; database access is not invoked." + }, + { + "path": "tests/skills/access-database/test_skill_catalog.py", + "scope": "runtime_skill", + "owner_skill_or_domain": "access-database", + "kind": "tool_contract", + "evidence_level": "deterministic_offline", + "requires": ["offline", "filesystem"], + "side_effects": ["filesystem"], + "skill_behavior_eval": false, + "classification_confidence": "high", + "classification_basis": "Validates skill directory/frontmatter rules and writes only temporary fixture files." + }, + { + "path": "tests/skills/assemble-context/test_assemble_writer_context.py", + "scope": "runtime_skill", + "owner_skill_or_domain": "assemble-context", + "kind": "fake_pipeline", + "evidence_level": "deterministic_offline", + "requires": ["offline"], + "side_effects": ["none"], + "skill_behavior_eval": false, + "classification_confidence": "high", + "classification_basis": "Runs A/B/C context assembly with in-memory retrieval repositories and validates projected contracts." + }, + { + "path": "tests/skills/assemble-context/test_fine_outline_reader.py", + "scope": "runtime_skill", + "owner_skill_or_domain": "assemble-context", + "kind": "tool_contract", + "evidence_level": "deterministic_offline", + "requires": ["offline"], + "side_effects": ["none"], + "skill_behavior_eval": false, + "classification_confidence": "high", + "classification_basis": "Uses a fake connection to assert fine-outline SQL filters and fail-closed payload parsing." + }, + { + "path": "tests/skills/assemble-context/test_fine_outline_unification.py", + "scope": "runtime_skill", + "owner_skill_or_domain": "assemble-context", + "kind": "tool_contract", + "evidence_level": "deterministic_offline", + "requires": ["offline"], + "side_effects": ["none"], + "skill_behavior_eval": false, + "classification_confidence": "high", + "classification_basis": "Checks the unified fine-outline field contract and required-field rejection in the assembler." + }, + { + "path": "tests/skills/assemble-context/test_pattern_binding_reader.py", + "scope": "runtime_skill", + "owner_skill_or_domain": "assemble-context", + "kind": "tool_contract", + "evidence_level": "deterministic_offline", + "requires": ["offline"], + "side_effects": ["none"], + "skill_behavior_eval": false, + "classification_confidence": "high", + "classification_basis": "Uses a fake assembly row to verify confirmed pattern-reference projection and empty-selection behavior." + }, + { + "path": "tests/skills/assemble-context/test_retrieve_writer_sources.py", + "scope": "runtime_skill", + "owner_skill_or_domain": "assemble-context", + "kind": "fake_pipeline", + "evidence_level": "deterministic_offline", + "requires": ["offline"], + "side_effects": ["none"], + "skill_behavior_eval": false, + "classification_confidence": "high", + "classification_basis": "Exercises retrieval planning, frozen cards, prose expansion, and replay repositories with fake connections." + }, + { + "path": "tests/skills/assemble-context/test_style_loader.py", + "scope": "runtime_skill", + "owner_skill_or_domain": "assemble-context", + "kind": "tool_contract", + "evidence_level": "deterministic_offline", + "requires": ["offline"], + "side_effects": ["none"], + "skill_behavior_eval": false, + "classification_confidence": "high", + "classification_basis": "Tests style normalization and confirmed-section fallback using an in-memory fake connection." + }, + { + "path": "tests/skills/assemble-context/test_writer_contract.py", + "scope": "runtime_skill", + "owner_skill_or_domain": "assemble-context", + "kind": "tool_contract", + "evidence_level": "deterministic_offline", + "requires": ["offline"], + "side_effects": ["none"], + "skill_behavior_eval": false, + "classification_confidence": "high", + "classification_basis": "Validates WriterContext, creative-input projection, hashes, freeze boundaries, and closed fields in memory." + }, + { + "path": "tests/skills/call-content-model/test_call_persistence.py", + "scope": "runtime_skill", + "owner_skill_or_domain": "call-content-model", + "kind": "tool_contract", + "evidence_level": "deterministic_offline", + "requires": ["offline"], + "side_effects": ["none"], + "skill_behavior_eval": false, + "classification_confidence": "high", + "classification_basis": "Mocks the HTTP session and asserts the persistence event passed to the model adapter." + }, + { + "path": "tests/skills/call-content-model/test_quota.py", + "scope": "runtime_skill", + "owner_skill_or_domain": "call-content-model", + "kind": "tool_unit", + "evidence_level": "deterministic_offline", + "requires": ["offline"], + "side_effects": ["none"], + "skill_behavior_eval": false, + "classification_confidence": "high", + "classification_basis": "Uses fake clocks, quota state, HTTP responses, and chat functions; comments explicitly prohibit real calls." + }, + { + "path": "tests/skills/capture-ai-flavor-cases/test_capture_cases.py", + "scope": "runtime_skill", + "owner_skill_or_domain": "capture-ai-flavor-cases", + "kind": "tool_contract", + "evidence_level": "deterministic_offline", + "requires": ["offline", "filesystem"], + "side_effects": ["filesystem"], + "skill_behavior_eval": false, + "classification_confidence": "high", + "classification_basis": "Covers case-card validation, revalidation, CLI persistence gates, and temporary source/receipt files with persistence mocked." + }, + { + "path": "tests/skills/check-content-consistency/test_build_semantic_input.py", + "scope": "runtime_skill", + "owner_skill_or_domain": "check-content-consistency", + "kind": "tool_contract", + "evidence_level": "deterministic_offline", + "requires": ["offline"], + "side_effects": ["none"], + "skill_behavior_eval": false, + "classification_confidence": "high", + "classification_basis": "Checks deterministic semantic-input projection, source-ref cleaning, identity binding, and hash rejection." + }, + { + "path": "tests/skills/check-content-consistency/test_check_writer_candidate.py", + "scope": "runtime_skill", + "owner_skill_or_domain": "check-content-consistency", + "kind": "tool_contract", + "evidence_level": "deterministic_offline", + "requires": ["offline"], + "side_effects": ["none"], + "skill_behavior_eval": false, + "classification_confidence": "high", + "classification_basis": "Tests the mechanical candidate gate for outline anchors, hashes, length, and forbidden writer fields." + }, + { + "path": "tests/skills/check-content-consistency/test_run_writer_semantic_detector.py", + "scope": "runtime_skill", + "owner_skill_or_domain": "check-content-consistency", + "kind": "fake_pipeline", + "evidence_level": "deterministic_offline", + "requires": ["offline"], + "side_effects": ["none"], + "skill_behavior_eval": false, + "classification_confidence": "high", + "classification_basis": "Drives detector correction and binding paths with SequenceFakeRunner/FakeRunner; no real model is called." + }, + { + "path": "tests/skills/clean-book-text/test_clean_detect_offline.py", + "scope": "runtime_skill", + "owner_skill_or_domain": "clean-book-text", + "kind": "fake_pipeline", + "evidence_level": "deterministic_offline", + "requires": ["offline", "filesystem"], + "side_effects": ["filesystem"], + "skill_behavior_eval": false, + "classification_confidence": "high", + "classification_basis": "Runs the detect CLI against temporary windows while chat_governed and JSON parsing are mocked." + }, + { + "path": "tests/skills/decide-candidate/test_confirm_knowledge_offline.py", + "scope": "runtime_skill", + "owner_skill_or_domain": "decide-candidate", + "kind": "tool_contract", + "evidence_level": "deterministic_offline", + "requires": ["offline"], + "side_effects": ["none"], + "skill_behavior_eval": false, + "classification_confidence": "high", + "classification_basis": "Tests normalization and idempotent confirmation helpers with direct in-memory inputs." + }, + { + "path": "tests/skills/decide-candidate/test_fact_delta.py", + "scope": "runtime_skill", + "owner_skill_or_domain": "decide-candidate", + "kind": "tool_unit", + "evidence_level": "deterministic_offline", + "requires": ["offline"], + "side_effects": ["none"], + "skill_behavior_eval": false, + "classification_confidence": "high", + "classification_basis": "Tests closed delta types, payloads, evidence quotes, and duplicate IDs using pure validation functions." + }, + { + "path": "tests/skills/decide-candidate/test_fact_delta_db.py", + "scope": "runtime_skill", + "owner_skill_or_domain": "decide-candidate", + "kind": "integration", + "evidence_level": "real_dependency_integration", + "requires": ["postgresql"], + "side_effects": ["postgresql"], + "skill_behavior_eval": false, + "classification_confidence": "high", + "classification_basis": "Calls the real db.connect, inserts/accepts/rolls back rows, checks triggers, and cleans test rows." + }, + { + "path": "tests/skills/decide-candidate/test_projection_db.py", + "scope": "runtime_skill", + "owner_skill_or_domain": "decide-candidate", + "kind": "integration", + "evidence_level": "real_dependency_integration", + "requires": ["postgresql"], + "side_effects": ["postgresql"], + "skill_behavior_eval": false, + "classification_confidence": "high", + "classification_basis": "Uses real PostgreSQL connections for projection registration, staleness, retries, trigger checks, and cleanup." + }, + { + "path": "tests/skills/decide-candidate/test_write_canonical_db.py", + "scope": "runtime_skill", + "owner_skill_or_domain": "decide-candidate", + "kind": "integration", + "evidence_level": "real_dependency_integration", + "requires": ["postgresql"], + "side_effects": ["postgresql"], + "skill_behavior_eval": false, + "classification_confidence": "high", + "classification_basis": "Uses real PostgreSQL rows and transactions to test canonical acceptance, CAS, rollback, and database guards." + }, + { + "path": "tests/skills/decide-candidate/test_writer_acceptance.py", + "scope": "runtime_skill", + "owner_skill_or_domain": "decide-candidate", + "kind": "fake_pipeline", + "evidence_level": "deterministic_offline", + "requires": ["offline"], + "side_effects": ["none"], + "skill_behavior_eval": false, + "classification_confidence": "high", + "classification_basis": "Builds self-contained WriterContext/Candidate fixtures and drives Shadow acceptance with an in-memory CAS store." + }, + { + "path": "tests/skills/deconstruct-book/test_parse_llm_offline.py", + "scope": "runtime_skill", + "owner_skill_or_domain": "deconstruct-book", + "kind": "fake_pipeline", + "evidence_level": "deterministic_offline", + "requires": ["offline", "filesystem"], + "side_effects": ["filesystem"], + "skill_behavior_eval": false, + "classification_confidence": "high", + "classification_basis": "Exercises outline repair/cache and chapter selection with mocked M3 calls, fake rows, and temporary cache files." + }, + { + "path": "tests/skills/deconstruct-book/test_parse_outline_offline.py", + "scope": "runtime_skill", + "owner_skill_or_domain": "deconstruct-book", + "kind": "tool_unit", + "evidence_level": "deterministic_offline", + "requires": ["offline"], + "side_effects": ["none"], + "skill_behavior_eval": false, + "classification_confidence": "high", + "classification_basis": "Tests outline-window coverage, bounded retry, sorting, and rendering with a patched chat function." + }, + { + "path": ".claude/skills/design-story-foundation/scripts/test_serial_merge.py", + "scope": "runtime_skill", + "owner_skill_or_domain": "design-story-foundation", + "kind": "tool_unit", + "evidence_level": "deterministic_offline", + "requires": ["offline", "filesystem"], + "side_effects": ["filesystem"], + "skill_behavior_eval": false, + "classification_confidence": "high", + "classification_basis": "Tests packet construction and raw-output parsing with temporary Markdown files; no model or service driver." + }, + { + "path": ".claude/skills/design-story-foundation/scripts/test_validate_candidates.py", + "scope": "runtime_skill", + "owner_skill_or_domain": "design-story-foundation", + "kind": "tool_contract", + "evidence_level": "deterministic_offline", + "requires": ["offline", "filesystem"], + "side_effects": ["filesystem"], + "skill_behavior_eval": false, + "classification_confidence": "high", + "classification_basis": "Validates candidate tree/heading/placeholder/root contracts using temporary candidate files." + }, + { + "path": "tests/skills/diagnose-ai-flavor/test_diagnose_ai_flavor.py", + "scope": "runtime_skill", + "owner_skill_or_domain": "diagnose-ai-flavor", + "kind": "domain_eval", + "evidence_level": "deterministic_offline", + "requires": ["offline", "filesystem"], + "side_effects": ["filesystem"], + "skill_behavior_eval": false, + "classification_confidence": "high", + "classification_basis": "Checks synthetic AI-flavor findings, artifact headers, and CLI persistence/offline behavior; no external judge." + }, + { + "path": "tests/skills/embed-knowledge/test_embed_drafts_offline.py", + "scope": "runtime_skill", + "owner_skill_or_domain": "embed-knowledge", + "kind": "fake_pipeline", + "evidence_level": "deterministic_offline", + "requires": ["offline"], + "side_effects": ["none"], + "skill_behavior_eval": false, + "classification_confidence": "high", + "classification_basis": "Uses fake database connections and an in-memory embedding HTTP session to test owner/lock/bulk flows." + }, + { + "path": "tests/skills/establish-voice-baseline/test_establish_voice_baseline.py", + "scope": "runtime_skill", + "owner_skill_or_domain": "establish-voice-baseline", + "kind": "tool_contract", + "evidence_level": "deterministic_offline", + "requires": ["offline", "filesystem"], + "side_effects": ["filesystem"], + "skill_behavior_eval": false, + "classification_confidence": "high", + "classification_basis": "Validates voice-ledger schema/grounding and CLI file flow with persistence mocked." + }, + { + "path": "tests/skills/evaluate-frozen-replay/test_fine_outline_detector.py", + "scope": "runtime_skill", + "owner_skill_or_domain": "evaluate-frozen-replay", + "kind": "tool_contract", + "evidence_level": "static_structure", + "requires": ["offline", "filesystem"], + "side_effects": ["none"], + "skill_behavior_eval": false, + "classification_confidence": "high", + "classification_basis": "Validates the closed detector report categories and statically reads a related SKILL.md; no Agent/model execution." + }, + { + "path": "tests/skills/evaluate-frozen-replay/test_gate_input_builder.py", + "scope": "runtime_skill", + "owner_skill_or_domain": "evaluate-frozen-replay", + "kind": "fake_pipeline", + "evidence_level": "deterministic_offline", + "requires": ["offline"], + "side_effects": ["none"], + "skill_behavior_eval": false, + "classification_confidence": "high", + "classification_basis": "Builds synthetic receipts/reports and drives GateInputBuilder validation without model or service calls." + }, + { + "path": "tests/skills/evaluate-frozen-replay/test_load_writer_reference_work.py", + "scope": "runtime_skill", + "owner_skill_or_domain": "evaluate-frozen-replay", + "kind": "fake_pipeline", + "evidence_level": "deterministic_offline", + "requires": ["offline", "filesystem"], + "side_effects": ["filesystem"], + "skill_behavior_eval": false, + "classification_confidence": "high", + "classification_basis": "Loads synthetic rows through a fake read-only connection and assembles dry-run Gate A configs with temp files." + }, + { + "path": "tests/skills/evaluate-frozen-replay/test_pattern_reference_injection.py", + "scope": "runtime_skill", + "owner_skill_or_domain": "evaluate-frozen-replay", + "kind": "fake_pipeline", + "evidence_level": "deterministic_offline", + "requires": ["offline", "filesystem"], + "side_effects": ["none"], + "skill_behavior_eval": false, + "classification_confidence": "high", + "classification_basis": "Uses a stub card searcher and dry-run assembly/config round trips to verify A/C pattern projection." + }, + { + "path": "tests/skills/evaluate-frozen-replay/test_refresh_runtime_probe.py", + "scope": "runtime_skill", + "owner_skill_or_domain": "evaluate-frozen-replay", + "kind": "runtime_probe", + "evidence_level": "deterministic_offline", + "requires": ["offline", "filesystem"], + "side_effects": ["filesystem"], + "skill_behavior_eval": false, + "classification_confidence": "high", + "classification_basis": "Exercises runtime-probe refresh and authorization gates with fake invocation results and temp output files." + }, + { + "path": "tests/skills/evaluate-frozen-replay/test_run_replay.py", + "scope": "runtime_skill", + "owner_skill_or_domain": "evaluate-frozen-replay", + "kind": "fake_pipeline", + "evidence_level": "deterministic_offline", + "requires": ["offline", "filesystem", "subprocess"], + "side_effects": ["filesystem"], + "skill_behavior_eval": false, + "classification_confidence": "high", + "classification_basis": "Runs the replay orchestrator with a generated fake-agent executable, temp output, and synthetic planner/detector/judge responses." + }, + { + "path": "tests/skills/evaluate-frozen-replay/test_run_writer_replay.py", + "scope": "runtime_skill", + "owner_skill_or_domain": "evaluate-frozen-replay", + "kind": "fake_pipeline", + "evidence_level": "deterministic_offline", + "requires": ["offline", "filesystem", "raw_vault"], + "side_effects": ["filesystem"], + "skill_behavior_eval": false, + "classification_confidence": "high", + "classification_basis": "Runs the full replay/CAS/raw-vault/Gate path with fake subprocess, semantic, and judge adapters; no real model." + }, + { + "path": "tests/skills/evaluate-frozen-replay/test_writer_eval_preregister.py", + "scope": "runtime_skill", + "owner_skill_or_domain": "evaluate-frozen-replay", + "kind": "tool_contract", + "evidence_level": "deterministic_offline", + "requires": ["offline"], + "side_effects": ["none"], + "skill_behavior_eval": false, + "classification_confidence": "high", + "classification_basis": "Checks deterministic hash sorting, balanced arm assignment, and duplicate rejection for preregistration." + }, + { + "path": "tests/skills/evaluate-frozen-replay/test_writer_gate.py", + "scope": "runtime_skill", + "owner_skill_or_domain": "evaluate-frozen-replay", + "kind": "fake_pipeline", + "evidence_level": "deterministic_offline", + "requires": ["offline", "filesystem"], + "side_effects": ["filesystem"], + "skill_behavior_eval": false, + "classification_confidence": "high", + "classification_basis": "Builds synthetic Gate inputs/reports and tests gate decisions, receipts, tamper detection, and temp CAS output." + }, + { + "path": "tests/skills/execute-claude-task/test_claude_runtime.py", + "scope": "runtime_skill", + "owner_skill_or_domain": "execute-claude-task", + "kind": "runtime_probe", + "evidence_level": "deterministic_offline", + "requires": ["offline", "filesystem"], + "side_effects": ["filesystem"], + "skill_behavior_eval": false, + "classification_confidence": "high", + "classification_basis": "Tests runtime profile/receipt/sandbox/environment handling through a mocked subprocess runner and temp isolation directories." + }, + { + "path": "tests/skills/extract-chapter-knowledge/test_extract_knowledge_offline.py", + "scope": "runtime_skill", + "owner_skill_or_domain": "extract-chapter-knowledge", + "kind": "tool_contract", + "evidence_level": "deterministic_offline", + "requires": ["offline"], + "side_effects": ["none"], + "skill_behavior_eval": false, + "classification_confidence": "high", + "classification_basis": "Checks evidence binding, alias normalization, and salvage drops with pure extraction functions." + }, + { + "path": "tests/skills/extract-work-knowledge/test_parse_upgrade_offline.py", + "scope": "runtime_skill", + "owner_skill_or_domain": "extract-work-knowledge", + "kind": "fake_pipeline", + "evidence_level": "deterministic_offline", + "requires": ["offline"], + "side_effects": ["none"], + "skill_behavior_eval": false, + "classification_confidence": "high", + "classification_basis": "Main path uses fake DB/model/embed adapters and in-memory transaction fixtures; the real PostgreSQL smoke is not part of this offline entry." + }, + { + "path": "tests/skills/extract-work-knowledge/test_presence_dedupe.py", + "scope": "runtime_skill", + "owner_skill_or_domain": "extract-work-knowledge", + "kind": "fake_pipeline", + "evidence_level": "deterministic_offline", + "requires": ["offline"], + "side_effects": ["none"], + "skill_behavior_eval": false, + "classification_confidence": "high", + "classification_basis": "Runs the production presence-dedupe CLI against an in-memory fake database and patched lock/connection boundary; no PostgreSQL, network, model, or embedding call is made." + }, + { + "path": "tests/skills/extract-work-knowledge/test_parse_upgrade_pg_smoke.py", + "scope": "runtime_skill", + "owner_skill_or_domain": "extract-work-knowledge", + "kind": "integration", + "evidence_level": "real_dependency_integration", + "requires": ["postgresql"], + "side_effects": ["postgresql"], + "skill_behavior_eval": false, + "classification_confidence": "high", + "classification_basis": "Explicitly opt-in entry point imports the production upgrade module and calls the real PostgreSQL rollback smoke only when MUSE_REAL_PG_ROLLBACK_SMOKE=1." + }, + { + "path": "tests/skills/extract-work-knowledge/test_upgrade_work_lock_offline.py", + "scope": "runtime_skill", + "owner_skill_or_domain": "extract-work-knowledge", + "kind": "tool_unit", + "evidence_level": "deterministic_offline", + "requires": ["offline"], + "side_effects": ["none"], + "skill_behavior_eval": false, + "classification_confidence": "high", + "classification_basis": "Simulates advisory-lock sessions entirely in memory and asserts lock/release SQL semantics." + }, + { + "path": "tests/skills/freeze-context/test_audit_leakage.py", + "scope": "runtime_skill", + "owner_skill_or_domain": "freeze-context", + "kind": "tool_contract", + "evidence_level": "deterministic_offline", + "requires": ["offline"], + "side_effects": ["none"], + "skill_behavior_eval": false, + "classification_confidence": "high", + "classification_basis": "Tests pure snapshot leakage audit decisions and hash-only findings on synthetic records." + }, + { + "path": "tests/skills/freeze-context/test_build_snapshot.py", + "scope": "runtime_skill", + "owner_skill_or_domain": "freeze-context", + "kind": "tool_contract", + "evidence_level": "deterministic_offline", + "requires": ["offline"], + "side_effects": ["none"], + "skill_behavior_eval": false, + "classification_confidence": "high", + "classification_basis": "Checks chapter/milestone/window freezing, terminal-field removal, manifest closure, and payload omission in memory." + }, + { + "path": "tests/skills/freeze-context/test_check_snapshot.py", + "scope": "runtime_skill", + "owner_skill_or_domain": "freeze-context", + "kind": "tool_contract", + "evidence_level": "deterministic_offline", + "requires": ["offline"], + "side_effects": ["none"], + "skill_behavior_eval": false, + "classification_confidence": "high", + "classification_basis": "Validates authorization, arm manifests, candidate shape, source bounds, and replay preflight before model execution." + }, + { + "path": "tests/skills/freeze-context/test_load_reference_work.py", + "scope": "runtime_skill", + "owner_skill_or_domain": "freeze-context", + "kind": "fake_pipeline", + "evidence_level": "deterministic_offline", + "requires": ["offline"], + "side_effects": ["none"], + "skill_behavior_eval": false, + "classification_confidence": "high", + "classification_basis": "Tests source/auth projection and frozen reference-card loading with a fake read-only connection." + }, + { + "path": "tests/skills/maintain-work-extraction/test_backup_upgrade_work_offline.py", + "scope": "runtime_skill", + "owner_skill_or_domain": "maintain-work-extraction", + "kind": "fake_pipeline", + "evidence_level": "deterministic_offline", + "requires": ["offline", "filesystem"], + "side_effects": ["filesystem"], + "skill_behavior_eval": false, + "classification_confidence": "high", + "classification_basis": "Runs backup/verify/restore paths with fake database rows and temporary backup directories; real DB calls are patched." + }, + { + "path": "tests/skills/maintain-work-extraction/test_reset_upgrade_work_offline.py", + "scope": "runtime_skill", + "owner_skill_or_domain": "maintain-work-extraction", + "kind": "fake_pipeline", + "evidence_level": "deterministic_offline", + "requires": ["offline"], + "side_effects": ["none"], + "skill_behavior_eval": false, + "classification_confidence": "high", + "classification_basis": "Drives reset/backup lock and rollback logic through stateful fake DB connections and patched external boundaries." + }, + { + "path": "tests/skills/plan-chapter/test_contract.py", + "scope": "runtime_skill", + "owner_skill_or_domain": "plan-chapter", + "kind": "tool_contract", + "evidence_level": "static_structure", + "requires": ["offline", "filesystem"], + "side_effects": ["none"], + "skill_behavior_eval": false, + "classification_confidence": "high", + "classification_basis": "Reads SKILL.md, planner prompt, chain registry, and schema to assert documented field/role contracts." + }, + { + "path": "tests/skills/plan-story/test_field_coverage.py", + "scope": "runtime_skill", + "owner_skill_or_domain": "plan-story", + "kind": "tool_contract", + "evidence_level": "deterministic_offline", + "requires": ["offline", "filesystem"], + "side_effects": ["none"], + "skill_behavior_eval": false, + "classification_confidence": "high", + "classification_basis": "Loads the fine-outline schema and checks required/recommended field coverage before any DB write." + }, + { + "path": "tests/skills/plan-story/test_record_planning_execution.py", + "scope": "runtime_skill", + "owner_skill_or_domain": "plan-story", + "kind": "tool_unit", + "evidence_level": "deterministic_offline", + "requires": ["offline"], + "side_effects": ["none"], + "skill_behavior_eval": false, + "classification_confidence": "high", + "classification_basis": "Tests canonical JSON ordering and secret rejection in pure helper functions." + }, + { + "path": "tests/skills/plan-story/test_repair_deterministic_receipt.py", + "scope": "runtime_skill", + "owner_skill_or_domain": "plan-story", + "kind": "tool_unit", + "evidence_level": "deterministic_offline", + "requires": ["offline"], + "side_effects": ["none"], + "skill_behavior_eval": false, + "classification_confidence": "high", + "classification_basis": "Checks deterministic receipt classification and correction projection with in-memory dictionaries." + }, + { + "path": "tests/skills/plan-story/test_select_patterns_offline.py", + "scope": "runtime_skill", + "owner_skill_or_domain": "plan-story", + "kind": "tool_contract", + "evidence_level": "deterministic_offline", + "requires": ["offline"], + "side_effects": ["none"], + "skill_behavior_eval": false, + "classification_confidence": "high", + "classification_basis": "Mocks search/write/record helpers and verifies authorized pattern-reference projection and empty results." + }, + { + "path": "tests/skills/prevent-ai-flavor/test_prevent_ai_flavor.py", + "scope": "runtime_skill", + "owner_skill_or_domain": "prevent-ai-flavor", + "kind": "tool_contract", + "evidence_level": "deterministic_offline", + "requires": ["offline", "filesystem"], + "side_effects": ["filesystem"], + "skill_behavior_eval": false, + "classification_confidence": "high", + "classification_basis": "Checks prevention-contract projection and CLI persistence/offline switches with DB helpers patched." + }, + { + "path": "tests/skills/record-run-evidence/test_file_cas.py", + "scope": "runtime_skill", + "owner_skill_or_domain": "record-run-evidence", + "kind": "integration", + "evidence_level": "real_dependency_integration", + "requires": ["offline", "filesystem"], + "side_effects": ["filesystem"], + "skill_behavior_eval": false, + "classification_confidence": "high", + "classification_basis": "Uses the real FileCasStore against temporary directories to test journal immutability, concurrency, recovery, and permissions." + }, + { + "path": "tests/skills/record-run-evidence/test_persist_raw.py", + "scope": "runtime_skill", + "owner_skill_or_domain": "record-run-evidence", + "kind": "tool_unit", + "evidence_level": "deterministic_offline", + "requires": ["offline"], + "side_effects": ["none"], + "skill_behavior_eval": false, + "classification_confidence": "high", + "classification_basis": "Tests only the raw secret-pattern validator with direct strings." + }, + { + "path": "tests/skills/record-run-evidence/test_raw_vault.py", + "scope": "runtime_skill", + "owner_skill_or_domain": "record-run-evidence", + "kind": "integration", + "evidence_level": "real_dependency_integration", + "requires": ["offline", "filesystem", "raw_vault"], + "side_effects": ["filesystem"], + "skill_behavior_eval": false, + "classification_confidence": "high", + "classification_basis": "Uses the real RawVaultManager and temporary filesystem to test lease ordering, permissions, migration, and recovery." + }, + { + "path": "tests/skills/record-run-evidence/test_record_failed_run.py", + "scope": "runtime_skill", + "owner_skill_or_domain": "record-run-evidence", + "kind": "tool_unit", + "evidence_level": "deterministic_offline", + "requires": ["offline"], + "side_effects": ["none"], + "skill_behavior_eval": false, + "classification_confidence": "high", + "classification_basis": "Checks failure-record helper shape and bounded failure dimensions without persistence." + }, + { + "path": "tests/skills/record-run-evidence/test_repair_receipt_evidence.py", + "scope": "runtime_skill", + "owner_skill_or_domain": "record-run-evidence", + "kind": "tool_unit", + "evidence_level": "deterministic_offline", + "requires": ["offline"], + "side_effects": ["none"], + "skill_behavior_eval": false, + "classification_confidence": "high", + "classification_basis": "Checks receipt eligibility predicates using in-memory values only." + }, + { + "path": "tests/skills/record-run-evidence/test_run_registry.py", + "scope": "runtime_skill", + "owner_skill_or_domain": "record-run-evidence", + "kind": "tool_unit", + "evidence_level": "deterministic_offline", + "requires": ["offline"], + "side_effects": ["none"], + "skill_behavior_eval": false, + "classification_confidence": "high", + "classification_basis": "Tests run ID formatting and terminal-state rejection before database access." + }, + { + "path": "tests/skills/revise-ai-flavor/test_revise_ai_flavor.py", + "scope": "runtime_skill", + "owner_skill_or_domain": "revise-ai-flavor", + "kind": "domain_eval", + "evidence_level": "deterministic_offline", + "requires": ["offline", "filesystem"], + "side_effects": ["filesystem"], + "skill_behavior_eval": false, + "classification_confidence": "high", + "classification_basis": "Runs the synthetic diagnosis/patch/gate lifecycle and CLI report path with persistence and baseline loading mocked." + }, + { + "path": "tests/skills/score-content-quality/test_lesson_registry_db.py", + "scope": "runtime_skill", + "owner_skill_or_domain": "score-content-quality", + "kind": "integration", + "evidence_level": "real_dependency_integration", + "requires": ["postgresql"], + "side_effects": ["postgresql"], + "skill_behavior_eval": false, + "classification_confidence": "high", + "classification_basis": "Uses real PostgreSQL rows and trigger checks for lesson proposal/review/promotion/rejection, then cleans them." + }, + { + "path": "tests/skills/score-content-quality/test_rubric.py", + "scope": "runtime_skill", + "owner_skill_or_domain": "score-content-quality", + "kind": "tool_contract", + "evidence_level": "deterministic_offline", + "requires": ["offline"], + "side_effects": ["none"], + "skill_behavior_eval": false, + "classification_confidence": "high", + "classification_basis": "Validates fine-outline rubric dimensions, evidence requirements, profiles, and stability warnings in memory." + }, + { + "path": "tests/skills/score-content-quality/test_run_writer_blind_judge.py", + "scope": "runtime_skill", + "owner_skill_or_domain": "score-content-quality", + "kind": "fake_pipeline", + "evidence_level": "deterministic_offline", + "requires": ["offline"], + "side_effects": ["none"], + "skill_behavior_eval": false, + "classification_confidence": "high", + "classification_basis": "Drives blind-judge adapter/panel correction with SequenceRunner fake structured outputs; no model endpoint is used." + }, + { + "path": "tests/skills/score-content-quality/test_writer_rubric.py", + "scope": "runtime_skill", + "owner_skill_or_domain": "score-content-quality", + "kind": "tool_contract", + "evidence_level": "deterministic_offline", + "requires": ["offline"], + "side_effects": ["none"], + "skill_behavior_eval": false, + "classification_confidence": "high", + "classification_basis": "Tests five-dimension rubric validation, blind ordering, reviewer adjudication, and structured verdict contracts." + }, + { + "path": "tests/skills/search-knowledge/test_search.py", + "scope": "runtime_skill", + "owner_skill_or_domain": "search-knowledge", + "kind": "tool_contract", + "evidence_level": "deterministic_offline", + "requires": ["offline"], + "side_effects": ["none"], + "skill_behavior_eval": false, + "classification_confidence": "high", + "classification_basis": "Uses a fake connection and embedder to assert public-pattern SQL scope and tenant binding." + }, + { + "path": "tests/skills/write-next-chapter/test_candidate_cas.py", + "scope": "runtime_skill", + "owner_skill_or_domain": "write-next-chapter", + "kind": "fake_pipeline", + "evidence_level": "deterministic_offline", + "requires": ["offline"], + "side_effects": ["none"], + "skill_behavior_eval": false, + "classification_confidence": "high", + "classification_basis": "Uses a fake CAS connection and in-memory writer pipeline to test token transitions and evidence loops." + }, + { + "path": "tests/skills/write-next-chapter/test_candidate_cas_db.py", + "scope": "runtime_skill", + "owner_skill_or_domain": "write-next-chapter", + "kind": "integration", + "evidence_level": "real_dependency_integration", + "requires": ["postgresql"], + "side_effects": ["postgresql"], + "skill_behavior_eval": false, + "classification_confidence": "high", + "classification_basis": "Uses real PostgreSQL CAS rows and direct trigger updates, then removes isolated unittest rows." + }, + { + "path": "tests/skills/write-next-chapter/test_persist_writer_run.py", + "scope": "runtime_skill", + "owner_skill_or_domain": "write-next-chapter", + "kind": "tool_unit", + "evidence_level": "deterministic_offline", + "requires": ["offline"], + "side_effects": ["none"], + "skill_behavior_eval": false, + "classification_confidence": "high", + "classification_basis": "Tests hash normalization and rejection in the writer persistence helper without a database call." + }, + { + "path": "tests/skills/write-next-chapter/test_run_writer.py", + "scope": "runtime_skill", + "owner_skill_or_domain": "write-next-chapter", + "kind": "fake_pipeline", + "evidence_level": "deterministic_offline", + "requires": ["offline", "filesystem"], + "side_effects": ["filesystem"], + "skill_behavior_eval": false, + "classification_confidence": "high", + "classification_basis": "Runs the writer adapter against a FakeRunner/CompletedProcess and temporary profile inputs; no real Claude process or model." + }, + { + "path": "tests/skills/write-next-chapter/test_run_writer_pipeline.py", + "scope": "runtime_skill", + "owner_skill_or_domain": "write-next-chapter", + "kind": "fake_pipeline", + "evidence_level": "deterministic_offline", + "requires": ["offline", "filesystem"], + "side_effects": ["filesystem"], + "skill_behavior_eval": false, + "classification_confidence": "high", + "classification_basis": "Drives the in-memory writer/mechanical/semantic/CAS pipeline and atomic temporary result writes with fake detectors." + }, + { + "path": "tests/skills/write-next-chapter/test_semantic_verdict.py", + "scope": "runtime_skill", + "owner_skill_or_domain": "write-next-chapter", + "kind": "tool_contract", + "evidence_level": "deterministic_offline", + "requires": ["offline"], + "side_effects": ["none"], + "skill_behavior_eval": false, + "classification_confidence": "high", + "classification_basis": "Checks semantic report version, candidate/context/run bindings, and report hash consistency before persistence." + }, + { + "path": "dashboard/test_server_display.py", + "scope": "other", + "owner_skill_or_domain": "dashboard", + "kind": "tool_contract", + "evidence_level": "deterministic_offline", + "requires": ["offline"], + "side_effects": ["none"], + "skill_behavior_eval": false, + "classification_confidence": "high", + "classification_basis": "Tests dashboard display/encoding helpers and synthetic AI-flavor views; the file declares no database connection." + }, + { + "path": "harness/test_run_selected.py", + "scope": "harness", + "owner_skill_or_domain": "harness", + "kind": "harness_self_test", + "evidence_level": "deterministic_offline", + "requires": ["offline", "filesystem", "subprocess"], + "side_effects": ["filesystem", "subprocess"], + "skill_behavior_eval": false, + "classification_confidence": "high", + "classification_basis": "Uses temporary manifests and a fake Python child process to test selector, dependency, timeout, nonzero and output-summary handling." + }, + { + "path": "harness/test_skill_harness.py", + "scope": "harness", + "owner_skill_or_domain": "harness", + "kind": "harness_self_test", + "evidence_level": "deterministic_offline", + "requires": ["offline", "filesystem"], + "side_effects": ["filesystem"], + "skill_behavior_eval": false, + "classification_confidence": "high", + "classification_basis": "Creates temporary SKILL.md/manifest fixtures and tests harness static-audit reports and CLI exit codes." + }, + { + "path": "humanization/tests/test_contracts.py", + "scope": "domain", + "owner_skill_or_domain": "humanization", + "kind": "domain_eval", + "evidence_level": "deterministic_offline", + "requires": ["offline", "filesystem"], + "side_effects": ["none"], + "skill_behavior_eval": false, + "classification_confidence": "high", + "classification_basis": "Checks humanization asset contracts and executes the synthetic U0 patch/review replay; no external Agent/model driver." + }, + { + "path": "humanization/tests/test_framework_coverage.py", + "scope": "domain", + "owner_skill_or_domain": "humanization", + "kind": "tool_contract", + "evidence_level": "static_structure", + "requires": ["offline", "filesystem"], + "side_effects": ["none"], + "skill_behavior_eval": false, + "classification_confidence": "high", + "classification_basis": "Reads the research coverage YAML and checks capability owners, implementation paths, and status values." + }, + { + "path": "humanization/tests/test_humanization_v2.py", + "scope": "domain", + "owner_skill_or_domain": "humanization", + "kind": "domain_eval", + "evidence_level": "deterministic_offline", + "requires": ["offline", "filesystem"], + "side_effects": ["filesystem"], + "skill_behavior_eval": false, + "classification_confidence": "high", + "classification_basis": "Evaluates synthetic voice/rule/carrier gates and lifecycle fixtures with temporary files; no external Agent/model reads SKILL.md." + } + ], + "summary": { + "entry_count": 81, + "by_scope": { + "runtime_skill": 75, + "domain": 3, + "harness": 2, + "other": 1 + }, + "by_kind": { + "tool_unit": 13, + "tool_contract": 30, + "integration": 8, + "fake_pipeline": 22, + "runtime_probe": 2, + "skill_behavior_eval": 0, + "domain_eval": 4, + "harness_self_test": 2 + }, + "by_evidence_level": { + "static_structure": 4, + "deterministic_offline": 69, + "real_dependency_integration": 8 + } + } +} diff --git a/harness/run_selected.py b/harness/run_selected.py new file mode 100644 index 0000000..9db241d --- /dev/null +++ b/harness/run_selected.py @@ -0,0 +1,1027 @@ +#!/usr/bin/env python3 +"""按 manifest 选择并执行离线测试条目的最小 harness。 + +该入口只在明确选择条目或传入 ``--all-offline`` 时运行,不承担全仓默认门禁。 +执行器自身不连接数据库、网络或模型;危险依赖只有在 ``--allow-requires`` 明确 +放行后才会启动对应子进程。 +""" + +from __future__ import annotations + +import argparse +import ast +import json +import os +import re +import subprocess +import sys +import time +from collections import Counter +from pathlib import Path +from typing import Any, Iterable, Optional, Sequence + +SCHEMA_VERSION = 1 +DEFAULT_MANIFEST_PATH = Path("harness") / "manifests" / "test-inventory.json" +BLOCKED_REQUIRES = frozenset({"postgresql", "network", "model", "credentials"}) +ENTRY_STATUSES = frozenset( + {"passed", "failed", "timeout", "blocked_dependency", "not_found"} +) +OUTPUT_SUMMARY_LIMIT = 2000 +DEFAULT_TIMEOUT_SECONDS = 60.0 +_TEST_ASSET_SCAN_ROOTS = ( + Path(".claude") / "skills", + Path("tests"), + Path("humanization") / "tests", + Path("harness"), + Path("dashboard"), +) +_IGNORED_TEST_ASSET_PARTS = frozenset({".git", ".venv", "__pycache__"}) +_NO_TESTS_OUTPUT_PATTERN = re.compile( + r"(?im)^\s*(?:" + r"no\s+tests?(?:\s+(?:were\s+)?(?:found|collected|run|ran|to\s+run))?" + r"|collected\s+0\s+items?" + r"|ran\s+0\s+tests?" + r"|0\s+tests?\s+ran" + r")\b[^\n]*$" +) +_SKIP_ONLY_LINE_PATTERN = re.compile( + r"(?i)^\s*(?:(?:\d+\s+)?skipped\b.*|" + r"skip(?:ped)?\s*(?:[:-]\s*.*)?)\s*$" +) +_UNITTEST_SKIP_SUMMARY_PATTERN = re.compile( + r"(?im)^\s*OK\s*\(\s*skipped\s*=\s*(\d+)\s*\)\s*$" +) +_UNITTEST_RUN_PATTERN = re.compile(r"(?i)\bRan\s+(\d+)\s+tests?\b") +_EXECUTION_RESULT_PATTERN = re.compile( + r"(?i)\b(?:\d+\s+(?:passed|failed|errors?|xfailed|xpassed)|" + r"(?:ran|collected)\s+[1-9]\d*)\b" +) + + +def _resolve_root(root: str | Path) -> Path: + path = Path(root).expanduser() + if not path.is_absolute(): + path = Path.cwd() / path + return path.resolve() + + +def _resolve_manifest_path(root: Path, manifest: str | Path | None) -> Path: + path = DEFAULT_MANIFEST_PATH if manifest is None else Path(manifest).expanduser() + if not path.is_absolute(): + path = root / path + return path.resolve() + + +def _relative_path(path: Path, root: Path) -> str: + try: + return path.resolve().relative_to(root).as_posix() + except ValueError: + return path.resolve().as_posix() + + +def _issue(code: str, message: str, **details: Any) -> dict[str, Any]: + issue: dict[str, Any] = {"code": code, "message": message} + issue.update(details) + return issue + + +def _normalise_manifest_entry_path(root: Path, value: str) -> tuple[Optional[str], Optional[str]]: + """校验并规范 manifest 中的相对路径。""" + + declared = Path(value) + if declared.is_absolute(): + return None, "manifest_path_not_relative" + try: + candidate = (root / declared).resolve() + relative = candidate.relative_to(root).as_posix() + except (OSError, RuntimeError, ValueError): + return None, "manifest_path_outside_root" + if not relative or relative == ".": + return None, "manifest_path_empty" + return relative, None + + +def _normalise_selector_path(root: Path, value: str) -> str: + """把路径选择器转换为与 manifest 相同的 POSIX 相对表示。""" + + declared = Path(value).expanduser() + if declared.is_absolute(): + try: + return declared.resolve().relative_to(root).as_posix() + except (OSError, RuntimeError, ValueError): + return declared.resolve().as_posix() + try: + return (root / declared).resolve().relative_to(root).as_posix() + except (OSError, RuntimeError, ValueError): + return declared.as_posix() + + +def _string_values(values: Iterable[str] | None) -> list[str]: + if values is None: + return [] + return [str(value).strip() for value in values if str(value).strip()] + + +def _is_named_python_test(path: Path) -> bool: + return path.suffix.casefold() == ".py" and ( + path.name.startswith("test_") or path.name.endswith("_test.py") + ) + + +def _is_test_asset_path(relative: str) -> bool: + path = Path(relative) + parts = path.parts + if not parts or any(part in _IGNORED_TEST_ASSET_PARTS for part in parts): + return False + if path.name == "__init__.py" or path.suffix.casefold() == ".pyc": + return False + + if parts[:2] == (".claude", "skills"): + return len(parts) >= 3 and path.suffix.casefold() == ".py" and ( + path.name.startswith("test_") or path.name.endswith("_test.py") + ) + if parts[0] == "tests": + return len(parts) > 1 and _is_named_python_test(path) + if parts[:2] == ("humanization", "tests"): + return len(parts) > 2 and _is_named_python_test(path) + if parts[0] in {"harness", "dashboard"}: + return len(parts) > 1 and path.suffix.casefold() == ".py" and path.name.startswith( + "test_" + ) + return False + + +def _scan_test_assets(root: Path) -> tuple[set[str], list[dict[str, Any]]]: + """扫描 generated_scope 约定的测试资产,失败时保留明确的扫描问题。""" + + assets: set[str] = set() + issues: list[dict[str, Any]] = [] + + def onerror(error: OSError) -> None: + error_path = getattr(error, "filename", None) + path = Path(str(error_path)) if error_path else root + issues.append( + _issue( + "test_asset_scan_failed", + "无法完整扫描测试资产目录", + path=_relative_path(path, root), + error=str(error), + ) + ) + + for relative_root in _TEST_ASSET_SCAN_ROOTS: + scan_root = root / relative_root + try: + if not scan_root.exists(): + continue + if not scan_root.is_dir(): + issues.append( + _issue( + "test_asset_scan_root_invalid", + "测试资产扫描根路径不是目录", + path=_relative_path(scan_root, root), + ) + ) + continue + except OSError as exc: + issues.append( + _issue( + "test_asset_scan_failed", + "无法检查测试资产扫描根目录", + path=_relative_path(scan_root, root), + error=str(exc), + ) + ) + continue + + for current, directory_names, file_names in os.walk( + scan_root, + topdown=True, + followlinks=False, + onerror=onerror, + ): + directory_names[:] = sorted( + name + for name in directory_names + if name not in _IGNORED_TEST_ASSET_PARTS + ) + for file_name in sorted(file_names): + candidate = Path(current) / file_name + if candidate.is_symlink(): + continue + try: + if not candidate.is_file(): + continue + relative = candidate.relative_to(root).as_posix() + except (OSError, ValueError) as exc: + issues.append( + _issue( + "test_asset_scan_failed", + "无法读取测试资产路径", + path=_relative_path(candidate, root), + error=str(exc), + ) + ) + continue + if _is_test_asset_path(relative): + assets.add(relative) + + return assets, issues + + +def _load_manifest( + root: Path, manifest_path: Path +) -> tuple[Optional[list[dict[str, Any]]], dict[str, Any], list[dict[str, Any]]]: + """读取并校验执行所需的 manifest 字段,任何合同错误都关闭执行。""" + + metadata: dict[str, Any] = { + "path": _relative_path(manifest_path, root), + "loaded": False, + "generated_scope_present": False, + } + issues: list[dict[str, Any]] = [] + + if not manifest_path.exists(): + issues.append( + _issue("manifest_missing", "测试 manifest 不存在", path=metadata["path"]) + ) + return None, metadata, issues + if not manifest_path.is_file(): + issues.append( + _issue("manifest_not_file", "测试 manifest 路径不是普通文件", path=metadata["path"]) + ) + return None, metadata, issues + + try: + raw = json.loads(manifest_path.read_text(encoding="utf-8")) + except UnicodeError as exc: + issues.append( + _issue( + "manifest_unreadable", + "测试 manifest 不是合法 UTF-8 文本", + path=metadata["path"], + error=str(exc), + ) + ) + return None, metadata, issues + except OSError as exc: + issues.append( + _issue( + "manifest_unreadable", + "无法读取测试 manifest", + path=metadata["path"], + error=str(exc), + ) + ) + return None, metadata, issues + except json.JSONDecodeError as exc: + issues.append( + _issue( + "manifest_invalid_json", + "测试 manifest 不是合法 JSON", + path=metadata["path"], + line=exc.lineno, + column=exc.colno, + error=exc.msg, + ) + ) + return None, metadata, issues + + metadata["loaded"] = True + if not isinstance(raw, dict): + issues.append( + _issue( + "manifest_invalid_structure", + "测试 manifest 顶层必须是对象", + path=metadata["path"], + ) + ) + return None, metadata, issues + + metadata["generated_scope_present"] = "generated_scope" in raw + if metadata["generated_scope_present"]: + generated_scope = raw["generated_scope"] + if not isinstance(generated_scope, str) or not generated_scope.strip(): + issues.append( + _issue( + "manifest_generated_scope_invalid", + "测试 manifest 的 generated_scope 必须是非空字符串", + path=metadata["path"], + actual=generated_scope, + ) + ) + else: + metadata["generated_scope"] = generated_scope.strip() + + if "schema_version" in raw and raw["schema_version"] != SCHEMA_VERSION: + issues.append( + _issue( + "manifest_schema_version_invalid", + "测试 manifest 的 schema_version 不受支持", + path=metadata["path"], + expected=SCHEMA_VERSION, + actual=raw["schema_version"], + ) + ) + + raw_entries = raw.get("entries") + if not isinstance(raw_entries, list): + issues.append( + _issue( + "manifest_entries_invalid", + "测试 manifest.entries 必须是数组", + path=metadata["path"], + ) + ) + return None, metadata, issues + + entries: list[dict[str, Any]] = [] + seen_paths: dict[str, int] = {} + required_string_fields = ("scope", "owner_skill_or_domain", "kind") + + for index, raw_entry in enumerate(raw_entries): + if not isinstance(raw_entry, dict): + issues.append( + _issue( + "manifest_entry_invalid", + "测试 manifest 条目必须是对象", + path=metadata["path"], + entry_index=index, + ) + ) + continue + + declared_path = raw_entry.get("path") + if not isinstance(declared_path, str) or not declared_path.strip(): + issues.append( + _issue( + "manifest_entry_path_invalid", + "测试 manifest 条目的 path 必须是非空字符串", + path=metadata["path"], + entry_index=index, + actual=declared_path, + ) + ) + continue + normalised_path, path_error = _normalise_manifest_entry_path(root, declared_path) + if path_error is not None or normalised_path is None: + issues.append( + _issue( + path_error or "manifest_entry_path_invalid", + "测试 manifest 条目的 path 必须位于项目根目录内且为相对路径", + path=metadata["path"], + entry_index=index, + actual=declared_path, + ) + ) + continue + + if normalised_path in seen_paths: + issues.append( + _issue( + "manifest_duplicate_path", + "测试 manifest 重复登记同一个 path", + path=normalised_path, + entry_indexes=[seen_paths[normalised_path], index], + ) + ) + else: + seen_paths[normalised_path] = index + + entry: dict[str, Any] = {"path": normalised_path} + for field in required_string_fields: + value = raw_entry.get(field) + if not isinstance(value, str) or not value.strip(): + issues.append( + _issue( + f"manifest_entry_{field}_invalid", + f"测试 manifest 条目的 {field} 必须是非空字符串", + path=metadata["path"], + entry_index=index, + actual=value, + ) + ) + else: + entry[field] = value.strip() + + requires = raw_entry.get("requires") + if not isinstance(requires, list) or any( + not isinstance(requirement, str) or not requirement.strip() + for requirement in requires + ): + issues.append( + _issue( + "manifest_entry_requires_invalid", + "测试 manifest 条目的 requires 必须是字符串数组", + path=metadata["path"], + entry_index=index, + actual=requires, + ) + ) + else: + entry["requires"] = [requirement.strip() for requirement in requires] + + if all(field in entry for field in (*required_string_fields, "requires")): + entries.append(entry) + + if issues: + return None, metadata, issues + return entries, metadata, [] + + +def _reconcile_test_assets( + root: Path, + entries: Sequence[dict[str, Any]], + manifest_metadata: dict[str, Any], +) -> list[dict[str, Any]]: + """对 generated_scope 清单做磁盘路径与登记路径的一一对账。""" + + if not manifest_metadata.get("generated_scope_present"): + return [] + + disk_assets, issues = _scan_test_assets(root) + declared_assets = {entry["path"] for entry in entries} + manifest_metadata["test_assets_scanned"] = len(disk_assets) + + for path in sorted(disk_assets - declared_assets): + issues.append( + _issue( + "manifest_test_asset_missing", + "磁盘测试资产未在 manifest 中登记", + path=path, + ) + ) + for path in sorted(declared_assets - disk_assets): + issues.append( + _issue( + "manifest_test_asset_extra", + "manifest 登记了磁盘中不存在的测试资产", + path=path, + ) + ) + return issues + + +def _matches( + entry: dict[str, Any], + *, + paths: Sequence[str], + owners: Sequence[str], + kinds: Sequence[str], + scopes: Sequence[str], +) -> bool: + return ( + (not paths or entry["path"] in paths) + and (not owners or entry["owner_skill_or_domain"] in owners) + and (not kinds or entry["kind"] in kinds) + and (not scopes or entry["scope"] in scopes) + ) + + +def _blocked_requires(requires: Sequence[str], allow_requires: set[str]) -> list[str]: + return [ + requirement + for requirement in requires + if requirement.casefold() in BLOCKED_REQUIRES + and requirement.casefold() not in allow_requires + ] + + +def _coerce_output(value: Any) -> str: + if value is None: + return "" + if isinstance(value, bytes): + return value.decode("utf-8", errors="replace") + return str(value) + + +def _output_summary(value: Any, limit: int = OUTPUT_SUMMARY_LIMIT) -> str: + """只保留有限摘要,避免把子进程 raw 输出写入报告。""" + + text = _coerce_output(value).strip() + if len(text) <= limit: + return text + marker = "\n...[output summary truncated]...\n" + available = max(0, limit - len(marker)) + head = available // 2 + tail = available - head + return text[:head] + marker + text[-tail:] + + +def _is_no_execution_evidence(stdout: Any, stderr: Any) -> bool: + """判断 0 返回码是否只有空输出或明确的跳过/无测试结果。""" + + text = "\n".join( + value for value in (_coerce_output(stdout), _coerce_output(stderr)) if value + ).strip() + if not text: + return True + + if _NO_TESTS_OUTPUT_PATTERN.search(text): + return _EXECUTION_RESULT_PATTERN.search(text) is None + + summary_match = _UNITTEST_SKIP_SUMMARY_PATTERN.search(text) + if summary_match: + run_match = _UNITTEST_RUN_PATTERN.search(text) + if run_match is None: + return True + return int(run_match.group(1)) == int(summary_match.group(1)) + + lines = [line.strip() for line in text.splitlines() if line.strip()] + return bool(lines) and all(_SKIP_ONLY_LINE_PATTERN.match(line) for line in lines) + + +def _dotted_name(node: ast.AST) -> Optional[str]: + if isinstance(node, ast.Name): + return node.id + if isinstance(node, ast.Attribute): + parent = _dotted_name(node.value) + return f"{parent}.{node.attr}" if parent else None + return None + + +def _test_case_aliases(tree: ast.AST) -> tuple[set[str], set[str]]: + unittest_names = {"unittest"} + test_case_names = {"TestCase"} + for node in ast.walk(tree): + if isinstance(node, ast.Import): + for alias in node.names: + if alias.name == "unittest": + unittest_names.add(alias.asname or alias.name) + elif isinstance(node, ast.ImportFrom) and node.module: + if node.module == "unittest" or node.module.startswith("unittest."): + for alias in node.names: + if alias.name == "TestCase": + test_case_names.add(alias.asname or alias.name) + return unittest_names, test_case_names + + +def _is_test_case_class( + node: ast.ClassDef, + unittest_names: set[str], + test_case_names: set[str], +) -> bool: + for base in node.bases: + if isinstance(base, ast.Name) and base.id in test_case_names: + return True + dotted = _dotted_name(base) + if dotted and dotted.endswith(".TestCase"): + module_name = dotted[: -len(".TestCase")] + if module_name in unittest_names or any( + module_name.startswith(f"{name}.") for name in unittest_names + ): + return True + return False + + +def _is_main_guard(test: ast.AST) -> bool: + if not isinstance(test, ast.Compare): + return False + values = [test.left, *test.comparators] + for left, right in zip(values, values[1:]): + if ( + isinstance(left, ast.Name) + and left.id == "__name__" + and isinstance(right, ast.Constant) + and right.value == "__main__" + ) or ( + isinstance(right, ast.Name) + and right.id == "__name__" + and isinstance(left, ast.Constant) + and left.value == "__main__" + ): + return True + return False + + +def _test_shape_error(candidate: Path, root: Path) -> Optional[dict[str, Any]]: + try: + source = candidate.read_text(encoding="utf-8") + except (OSError, UnicodeError) as exc: + return { + "code": "test_shape_unreadable", + "message": "无法读取 Python 测试文件进行形状检查", + "path": _relative_path(candidate, root), + "error": str(exc), + } + + try: + tree = ast.parse(source, filename=str(candidate)) + except SyntaxError as exc: + return { + "code": "test_shape_invalid", + "message": "Python 测试文件无法通过 AST 解析", + "path": _relative_path(candidate, root), + "line": exc.lineno, + "column": exc.offset, + "error": exc.msg, + } + + has_test_function = any( + isinstance(node, (ast.FunctionDef, ast.AsyncFunctionDef)) + and node.name.startswith("test_") + for node in ast.walk(tree) + ) + unittest_names, test_case_names = _test_case_aliases(tree) + has_test_case = any( + isinstance(node, ast.ClassDef) + and _is_test_case_class(node, unittest_names, test_case_names) + for node in ast.walk(tree) + ) + has_main_entry = any( + isinstance(node, ast.If) and _is_main_guard(node.test) + for node in ast.walk(tree) + ) + if has_test_function or has_test_case or has_main_entry: + return None + + return { + "code": "test_shape_missing", + "message": ( + "Python 测试文件必须包含 test_ 函数、unittest.TestCase 测试类" + "或 __main__ 入口" + ), + "path": _relative_path(candidate, root), + } + + +def _entry_base(entry: dict[str, Any]) -> dict[str, Any]: + return { + "path": entry["path"], + "status": "failed", + "returncode": None, + "duration_ms": 0, + "requires": list(entry["requires"]), + "stdout": "", + "stderr": "", + } + + +def _run_entry( + root: Path, + python_executable: Path, + entry: dict[str, Any], + *, + timeout_seconds: float, + allow_requires: set[str], + enforce_test_shape: bool = False, +) -> dict[str, Any]: + result = _entry_base(entry) + candidate = root / Path(entry["path"]) + + try: + path_exists = candidate.is_file() + except OSError as exc: + path_exists = False + result["error"] = {"code": "path_unreadable", "message": str(exc)} + if not path_exists: + result["status"] = "not_found" + result.setdefault( + "error", + {"code": "path_not_found", "message": "测试路径不存在或不是普通文件"}, + ) + return result + + if enforce_test_shape and candidate.suffix.casefold() == ".py": + shape_error = _test_shape_error(candidate, root) + if shape_error is not None: + result["error"] = shape_error + return result + + blocked = _blocked_requires(entry["requires"], allow_requires) + if blocked: + result["status"] = "blocked_dependency" + result["blocked_requires"] = blocked + return result + + if not python_executable.is_file(): + result["error"] = { + "code": "python_executable_not_found", + "message": "项目 .venv/bin/python 不存在", + } + return result + if not os.access(python_executable, os.X_OK): + result["error"] = { + "code": "python_executable_not_executable", + "message": "项目 .venv/bin/python 不可执行", + } + return result + + started = time.monotonic() + try: + completed = subprocess.run( + [str(python_executable), entry["path"]], + cwd=str(root), + stdout=subprocess.PIPE, + stderr=subprocess.PIPE, + text=True, + encoding="utf-8", + errors="replace", + timeout=timeout_seconds, + check=False, + ) + except subprocess.TimeoutExpired as exc: + result["status"] = "timeout" + result["duration_ms"] = max(0, round((time.monotonic() - started) * 1000)) + result["stdout"] = _output_summary(exc.stdout) + result["stderr"] = _output_summary(exc.stderr) + result["error"] = { + "code": "timeout", + "message": f"子进程超过 {timeout_seconds:g} 秒限制", + } + return result + except OSError as exc: + result["duration_ms"] = max(0, round((time.monotonic() - started) * 1000)) + result["stdout"] = _output_summary(getattr(exc, "stdout", "")) + result["stderr"] = _output_summary(getattr(exc, "stderr", "")) + result["error"] = {"code": "subprocess_start_failed", "message": str(exc)} + return result + + result["returncode"] = completed.returncode + result["duration_ms"] = max(0, round((time.monotonic() - started) * 1000)) + result["stdout"] = _output_summary(completed.stdout) + result["stderr"] = _output_summary(completed.stderr) + if completed.returncode == 0 and _is_no_execution_evidence( + completed.stdout, completed.stderr + ): + result["status"] = "failed" + result["error"] = { + "code": "no_execution_evidence", + "message": "子进程返回 0,但没有发现测试执行证据", + } + else: + result["status"] = "passed" if completed.returncode == 0 else "failed" + if result["status"] == "failed": + result["error"] = { + "code": "nonzero_returncode", + "message": f"子进程返回码为 {completed.returncode}", + } + return result + + +def _selectors_present( + paths: Sequence[str], owners: Sequence[str], kinds: Sequence[str], scopes: Sequence[str] +) -> bool: + return bool(paths or owners or kinds or scopes) + + +def _base_report( + root: Path, + manifest_path: Path, + *, + paths: Sequence[str], + owners: Sequence[str], + kinds: Sequence[str], + scopes: Sequence[str], + all_offline: bool, + offline_only: bool, + allow_requires: Sequence[str], + timeout_seconds: float, +) -> dict[str, Any]: + return { + "schema_version": SCHEMA_VERSION, + "status": "failed", + "root": str(root), + "manifest": {"path": _relative_path(manifest_path, root), "loaded": False}, + "selectors": { + "path": list(paths), + "owner": list(owners), + "kind": list(kinds), + "scope": list(scopes), + "all_offline": all_offline, + }, + "offline_only": offline_only, + "allow_requires": list(allow_requires), + "timeout_seconds": timeout_seconds, + "selected_count": 0, + "entries": [], + "issues": [], + "summary": {"selected": 0, "by_status": {}}, + } + + +def run_selected( + root: str | Path = ".", + manifest: str | Path | None = None, + *, + paths: Sequence[str] | None = None, + owners: Sequence[str] | None = None, + skills: Sequence[str] | None = None, + kinds: Sequence[str] | None = None, + scopes: Sequence[str] | None = None, + all_offline: bool = False, + offline_only: bool = True, + allow_requires: Sequence[str] | None = None, + timeout_seconds: float = DEFAULT_TIMEOUT_SECONDS, +) -> dict[str, Any]: + """执行 manifest 中被选择的条目并返回 JSON 可序列化报告。""" + + root_path = _resolve_root(root) + manifest_path = _resolve_manifest_path(root_path, manifest) + path_values = [_normalise_selector_path(root_path, value) for value in _string_values(paths)] + owner_values = _string_values(owners) + _string_values(skills) + kind_values = _string_values(kinds) + scope_values = _string_values(scopes) + allow_values = _string_values(allow_requires) + allow_set = {value.casefold() for value in allow_values} + report = _base_report( + root_path, + manifest_path, + paths=path_values, + owners=owner_values, + kinds=kind_values, + scopes=scope_values, + all_offline=all_offline, + offline_only=offline_only, + allow_requires=allow_values, + timeout_seconds=timeout_seconds, + ) + + if not _selectors_present(path_values, owner_values, kind_values, scope_values) and not all_offline: + report["status"] = "invalid_selector" + report["issues"] = [ + _issue( + "selector_required", + "至少需要 --path、--owner/--skill、--kind 或 --scope,或显式传入 --all-offline", + ) + ] + return report + + if not root_path.exists(): + report["status"] = "root_not_found" + report["issues"] = [_issue("root_missing", "项目根目录不存在", path=".")] + return report + if not root_path.is_dir(): + report["status"] = "root_invalid" + report["issues"] = [_issue("root_not_directory", "项目根路径不是目录", path=".")] + return report + + entries, manifest_metadata, manifest_issues = _load_manifest(root_path, manifest_path) + report["manifest"] = manifest_metadata + if manifest_issues or entries is None: + report["status"] = "manifest_invalid" + report["issues"] = manifest_issues + return report + + reconciliation_issues = _reconcile_test_assets( + root_path, + entries, + manifest_metadata, + ) + if reconciliation_issues: + report["status"] = "manifest_invalid" + report["issues"] = reconciliation_issues + return report + + has_selectors = _selectors_present(path_values, owner_values, kind_values, scope_values) + if all_offline and not has_selectors: + selected = list(entries) + else: + selected = [ + entry + for entry in entries + if _matches( + entry, + paths=path_values, + owners=owner_values, + kinds=kind_values, + scopes=scope_values, + ) + ] + + if not selected: + report["status"] = "no_matches" + report["issues"] = [ + _issue( + "no_matches", + "没有条目匹配当前选择器", + selectors=report["selectors"], + ) + ] + return report + + python_executable = root_path / ".venv" / "bin" / "python" + results = [ + _run_entry( + root_path, + python_executable, + entry, + timeout_seconds=timeout_seconds, + allow_requires=allow_set, + enforce_test_shape=bool(manifest_metadata.get("generated_scope_present")), + ) + for entry in selected + ] + counts = Counter(result["status"] for result in results) + report["entries"] = results + report["selected_count"] = len(results) + report["summary"] = { + "selected": len(results), + "by_status": {status: counts[status] for status in sorted(counts)}, + } + report["status"] = "passed" if all(result["status"] == "passed" for result in results) else "failed" + return report + + +def human_summary(report: dict[str, Any]) -> str: + """输出不包含完整子进程 raw 的人读摘要。""" + + lines = [ + f"选择性测试:{report['status']}", + f"项目根:{report['root']}", + f"选择条目:{report['selected_count']}", + ] + for entry in report["entries"]: + lines.append( + f"- {entry['path']} [{entry['status']}] " + f"returncode={entry['returncode']} duration_ms={entry['duration_ms']}" + ) + for issue in report["issues"]: + lines.append(f"- issue [{issue['code']}] {issue['message']}") + return "\n".join(lines) + + +def _positive_timeout(value: str) -> float: + try: + parsed = float(value) + except ValueError as exc: + raise argparse.ArgumentTypeError("--timeout-seconds 必须是数字") from exc + if parsed <= 0: + raise argparse.ArgumentTypeError("--timeout-seconds 必须大于 0") + return parsed + + +def _build_parser() -> argparse.ArgumentParser: + parser = argparse.ArgumentParser(description="按 manifest 选择性执行项目测试。") + parser.add_argument("--root", default=".", help="项目根目录,默认为当前目录") + parser.add_argument( + "--manifest", + default=None, + help="manifest 路径;相对路径按项目根目录解析,默认 harness/manifests/test-inventory.json", + ) + parser.add_argument("--path", dest="paths", action="append", default=[], help="按相对路径选择条目,可重复") + parser.add_argument( + "--owner", + "--skill", + dest="owners", + action="append", + default=[], + help="按 owner_skill_or_domain 选择条目,可使用 --skill 作为别名", + ) + parser.add_argument("--kind", dest="kinds", action="append", default=[], help="按 kind 选择条目,可重复") + parser.add_argument("--scope", dest="scopes", action="append", default=[], help="按 scope 选择条目,可重复") + parser.add_argument( + "--all-offline", + action="store_true", + help="没有其它 selector 时选择 manifest 全部条目;危险依赖仍逐条报告为 blocked_dependency", + ) + parser.add_argument( + "--offline-only", + action="store_true", + default=True, + help="启用离线依赖门;默认开启,危险依赖须配合 --allow-requires", + ) + parser.add_argument( + "--allow-requires", + dest="allow_requires", + action="append", + default=[], + help="显式放行依赖名,可重复;例如 --allow-requires postgresql", + ) + parser.add_argument( + "--timeout-seconds", + type=_positive_timeout, + default=DEFAULT_TIMEOUT_SECONDS, + help="每个子进程的超时秒数,默认 60", + ) + parser.add_argument("--json", action="store_true", help="输出结构化 JSON 报告") + return parser + + +def _exit_code(report: dict[str, Any]) -> int: + return 0 if report["status"] == "passed" else 1 + + +def main(argv: Optional[Sequence[str]] = None) -> int: + args = _build_parser().parse_args(argv) + report = run_selected( + args.root, + args.manifest, + paths=args.paths, + owners=args.owners, + kinds=args.kinds, + scopes=args.scopes, + all_offline=args.all_offline, + offline_only=args.offline_only, + allow_requires=args.allow_requires, + timeout_seconds=args.timeout_seconds, + ) + if args.json: + print(json.dumps(report, ensure_ascii=False, indent=2, sort_keys=True)) + else: + print(human_summary(report)) + return _exit_code(report) + + +if __name__ == "__main__": + sys.exit(main()) diff --git a/harness/skill_harness.py b/harness/skill_harness.py new file mode 100644 index 0000000..a706895 --- /dev/null +++ b/harness/skill_harness.py @@ -0,0 +1,870 @@ +#!/usr/bin/env python3 +"""对项目运行时 Skill 文档执行只读静态审计。 + +本模块只读取 `.claude/skills/*/SKILL.md`,不连接数据库、网络或模型,也不修改 +工作树。机器调用使用 :func:`audit_skills`,命令行入口同时支持人读摘要和 JSON。 +""" + +from __future__ import annotations + +import argparse +import json +import re +import sys +from collections import Counter +from pathlib import Path +from typing import Any, Iterable, Optional, Sequence + + +SCHEMA_VERSION = 1 +SKILLS_ROOT = Path(".claude") / "skills" +DEFAULT_MANIFEST_PATH = Path("harness") / "manifests" / "skills.json" + +_FRONTMATTER_NAME = re.compile(r"^\s*name\s*:\s*(.*?)\s*$") +_COVERAGE_TOOL_PATTERN = re.compile( + r"(?:\bcoverage\.py\b|\bpytest-cov\b|" + r"(?=|>|:|:)?\s*" + r"\d+(?:\.\d+)?\s*[%%]", + re.IGNORECASE, +) +_BUSINESS_COVERAGE_CONTEXT_PATTERN = re.compile( + r"(?:细纲|事件|硬约束|软约束|伏笔|实体|设定|情节|场景|知识|事实|正文|" + r"作品|业务|候选|规划|大纲|章)[^。\n]{0,24}覆盖率" + r"|覆盖率[^。\n]{0,24}(?:细纲|事件|硬约束|软约束|伏笔|实体|设定|情节|" + r"场景|知识|事实|正文|作品|业务|候选|规划|大纲|章)", + re.IGNORECASE, +) + +# 这些规则只针对运行时 Skill 文档中的明显开发验证残留;业务正文中的普通“检查” +# 等词不在扫描范围内。规则保持确定性,方便报告中的 code 作为机械门禁输入。 +_DEVELOPMENT_HEADING_PATTERN = re.compile(r"^\s*(#{1,6})\s*(.*?)\s*$") +_DEVELOPMENT_SECTION_TITLE_PATTERN = re.compile( + r"(?:自测|测试|开发验证|离线验证|开发测试|离线测试)", + re.IGNORECASE, +) +_DEVELOPMENT_COMMAND_PATTERN = re.compile( + r"(? Path: + """把审计根目录解析为绝对路径,但不要求它已经存在。""" + + path = Path(root).expanduser() + if not path.is_absolute(): + path = Path.cwd() / path + return path.resolve() + + +def _relative_path(path: Path, root: Path) -> str: + """返回报告中的稳定 POSIX 相对路径。""" + + try: + return path.resolve().relative_to(root).as_posix() + except ValueError: + return path.resolve().as_posix() + + +def _issue( + code: str, + message: str, + *, + path: Optional[str] = None, + line: Optional[int] = None, + **details: Any, +) -> dict[str, Any]: + result: dict[str, Any] = { + "code": code, + "severity": "error", + "message": message, + } + if path is not None: + result["path"] = path + if line is not None: + result["line"] = line + result.update(details) + return result + + +def _unquote_scalar(value: str) -> str: + """解析 Skill frontmatter 中足够支持 name 的简单标量。""" + + value = value.strip() + if len(value) >= 2 and value[0] == value[-1] == '"': + try: + parsed = json.loads(value) + except json.JSONDecodeError: + return value[1:-1] + return parsed if isinstance(parsed, str) else value + if len(value) >= 2 and value[0] == value[-1] == "'": + return value[1:-1].replace("''", "'") + # 允许常见的 YAML 行尾注释,不引入 YAML 依赖。 + return re.sub(r"\s+#.*$", "", value).strip() + + +def _parse_frontmatter(text: str) -> tuple[Optional[dict[str, str]], Optional[int], Optional[str]]: + """读取 frontmatter 的简单键值视图。 + + 返回值为 ``(fields, closing_line, error_code)``。该解析器只需要识别 name, + 不试图替代完整 YAML 解析器;不合法或缺失的边界会明确进入失败报告。 + """ + + lines = text.splitlines() + if not lines or lines[0].lstrip("\ufeff").strip() != "---": + return None, None, "frontmatter_missing" + + closing_line: Optional[int] = None + for index in range(1, len(lines)): + if lines[index].strip() in {"---", "..."}: + closing_line = index + 1 + break + if closing_line is None: + return None, None, "frontmatter_unclosed" + + fields: dict[str, str] = {} + for line in lines[1 : closing_line - 1]: + match = _FRONTMATTER_NAME.match(line) + if match: + fields["name"] = _unquote_scalar(match.group(1)) + return fields, closing_line, None + + +def _snippet(line: str) -> str: + compact = line.strip() + return compact if len(compact) <= 200 else compact[:197] + "..." + + +def _is_development_coverage_reference(line: str) -> bool: + """只把明确落在开发测试语境中的 coverage 文字判为污染。""" + + if _COVERAGE_TOOL_PATTERN.search(line) or _EXPLICIT_TEST_COVERAGE_PATTERN.search(line): + return True + if not _COVERAGE_PERCENT_PATTERN.search(line): + return False + return not _BUSINESS_COVERAGE_CONTEXT_PATTERN.search(line) + + +def _inspect_skill_file(path: Path, root: Path) -> tuple[dict[str, Any], list[dict[str, Any]]]: + relative = _relative_path(path, root) + entry: dict[str, Any] = { + "directory": path.parent.name, + "skill_path": relative, + "name": None, + "frontmatter_present": False, + } + issues: list[dict[str, Any]] = [] + + try: + text = path.read_text(encoding="utf-8") + except (OSError, UnicodeError) as exc: + issues.append( + _issue( + "skill_file_unreadable", + "无法读取 SKILL.md", + path=relative, + error=str(exc), + ) + ) + return entry, issues + + if not text.strip(): + issues.append( + _issue( + "empty_skill_file", + "SKILL.md 为空", + path=relative, + line=1, + ) + ) + return entry, issues + + fields, closing_line, frontmatter_error = _parse_frontmatter(text) + if frontmatter_error is not None: + line = 1 if frontmatter_error == "frontmatter_missing" else max(1, len(text.splitlines())) + messages = { + "frontmatter_missing": "缺少以 --- 开始的 frontmatter", + "frontmatter_unclosed": "frontmatter 未闭合", + } + issues.append( + _issue( + frontmatter_error, + messages[frontmatter_error], + path=relative, + line=line, + ) + ) + else: + entry["frontmatter_present"] = True + name = fields.get("name") if fields is not None else None + entry["name"] = name or None + if not name: + issues.append( + _issue( + "frontmatter_name_missing", + "frontmatter 缺少非空 name", + path=relative, + line=1, + ) + ) + elif name != path.parent.name: + issues.append( + _issue( + "name_directory_mismatch", + "frontmatter name 与 Skill 目录名不一致", + path=relative, + line=2, + expected=path.parent.name, + actual=name, + ) + ) + + development_section_level: Optional[int] = None + for line_number, line in enumerate(text.splitlines(), start=1): + heading_match = _DEVELOPMENT_HEADING_PATTERN.match(line) + if heading_match: + heading_level = len(heading_match.group(1)) + heading_title = re.sub(r"\s+#+\s*$", "", heading_match.group(2)).strip() + if ( + development_section_level is not None + and heading_level <= development_section_level + ): + development_section_level = None + if ( + 2 <= heading_level <= 6 + and _DEVELOPMENT_SECTION_TITLE_PATTERN.search(heading_title) + ): + development_section_level = heading_level + issues.append( + _issue( + "development_test_heading", + "发现开发测试章节标题", + path=relative, + line=line_number, + snippet=_snippet(line), + ) + ) + + for code, pattern, message in _POLLUTION_RULES: + if pattern.search(line): + issues.append( + _issue( + code, + message, + path=relative, + line=line_number, + snippet=_snippet(line), + ) + ) + if ( + development_section_level is not None + and _DEVELOPMENT_COMMAND_PATTERN.search(line) + ): + issues.append( + _issue( + "development_test_command_reference", + "发现开发测试章节中的合同测试命令", + path=relative, + line=line_number, + snippet=_snippet(line), + ) + ) + if _is_development_coverage_reference(line): + issues.append( + _issue( + "coverage_reference", + "发现开发测试语境中的覆盖率声明或工具名", + path=relative, + line=line_number, + snippet=_snippet(line), + ) + ) + + return entry, issues + + +def _scan_skill_directories(root: Path) -> tuple[list[dict[str, Any]], list[dict[str, Any]]]: + skills_root = root / SKILLS_ROOT + relative_skills_root = _relative_path(skills_root, root) + entries: list[dict[str, Any]] = [] + issues: list[dict[str, Any]] = [] + + if not skills_root.exists(): + issues.append( + _issue( + "skills_directory_missing", + "运行时 Skill 目录不存在", + path=relative_skills_root, + ) + ) + return entries, issues + if not skills_root.is_dir(): + issues.append( + _issue( + "skills_directory_not_directory", + "运行时 Skill 路径不是目录", + path=relative_skills_root, + ) + ) + return entries, issues + + try: + children = sorted(skills_root.iterdir(), key=lambda item: item.name) + except OSError as exc: + issues.append( + _issue( + "skills_directory_unreadable", + "无法列出运行时 Skill 目录", + path=relative_skills_root, + error=str(exc), + ) + ) + return entries, issues + + directories = [child for child in children if child.is_dir()] + if not directories: + issues.append( + _issue( + "no_skills_found", + "运行时 Skill 目录中未发现 Skill 子目录", + path=relative_skills_root, + ) + ) + return entries, issues + + for directory in directories: + skill_file = directory / "SKILL.md" + relative_skill_file = _relative_path(skill_file, root) + disk_entry: dict[str, Any] = { + "directory": directory.name, + "skill_path": relative_skill_file, + "name": None, + "frontmatter_present": False, + } + if not skill_file.exists(): + entries.append(disk_entry) + issues.append( + _issue( + "missing_skill_file", + "Skill 目录缺少 SKILL.md", + path=relative_skill_file, + ) + ) + continue + if not skill_file.is_file(): + entries.append(disk_entry) + issues.append( + _issue( + "skill_file_not_file", + "SKILL.md 路径不是普通文件", + path=relative_skill_file, + ) + ) + continue + + entry, file_issues = _inspect_skill_file(skill_file, root) + entries.append(entry) + issues.extend(file_issues) + + return entries, issues + + +def _resolve_manifest_path(root: Path, manifest: str | Path | None) -> Path: + path = DEFAULT_MANIFEST_PATH if manifest is None else Path(manifest).expanduser() + if not path.is_absolute(): + path = root / path + return path.resolve() + + +def _normalise_manifest_skill_path( + root: Path, value: str +) -> tuple[Optional[str], Optional[Path]]: + declared = Path(value) + if declared.is_absolute(): + return None, None + candidate = (root / declared).resolve() + try: + relative = candidate.relative_to(root).as_posix() + except ValueError: + return None, None + return relative, candidate + + +def _validate_manifest( + root: Path, + entries: Sequence[dict[str, Any]], + manifest_path: Path, +) -> tuple[dict[str, Any], list[dict[str, Any]]]: + """验证 Skill 登记表,并要求它与磁盘目录一一对应。""" + + relative_manifest = _relative_path(manifest_path, root) + metadata: dict[str, Any] = { + "path": relative_manifest, + "loaded": False, + "skills_declared": 0, + } + issues: list[dict[str, Any]] = [] + + if not manifest_path.exists(): + issues.append( + _issue( + "manifest_missing", + "Skill manifest 不存在", + path=relative_manifest, + ) + ) + return metadata, issues + if not manifest_path.is_file(): + issues.append( + _issue( + "manifest_not_file", + "Skill manifest 路径不是普通文件", + path=relative_manifest, + ) + ) + return metadata, issues + + try: + manifest_text = manifest_path.read_text(encoding="utf-8") + manifest = json.loads(manifest_text) + except UnicodeError as exc: + issues.append( + _issue( + "manifest_unreadable", + "无法按 UTF-8 读取 Skill manifest", + path=relative_manifest, + error=str(exc), + ) + ) + return metadata, issues + except OSError as exc: + issues.append( + _issue( + "manifest_unreadable", + "无法读取 Skill manifest", + path=relative_manifest, + error=str(exc), + ) + ) + return metadata, issues + except json.JSONDecodeError as exc: + issues.append( + _issue( + "manifest_invalid_json", + "Skill manifest 不是合法 JSON", + path=relative_manifest, + line=exc.lineno, + column=exc.colno, + error=exc.msg, + ) + ) + return metadata, issues + + metadata["loaded"] = True + if not isinstance(manifest, dict): + issues.append( + _issue( + "manifest_invalid_structure", + "Skill manifest 顶层必须是对象", + path=relative_manifest, + ) + ) + return metadata, issues + + raw_skills = manifest.get("skills") + if not isinstance(raw_skills, list): + issues.append( + _issue( + "manifest_skills_invalid", + "Skill manifest.skills 必须是数组", + path=relative_manifest, + ) + ) + return metadata, issues + + metadata["skills_declared"] = len(raw_skills) + disk_paths = { + str(entry["skill_path"]) + for entry in entries + if isinstance(entry.get("skill_path"), str) + } + declared_paths: set[str] = set() + path_occurrences: dict[str, list[int]] = {} + name_occurrences: dict[str, list[int]] = {} + skills_root = (root / SKILLS_ROOT).resolve() + + for index, item in enumerate(raw_skills): + if not isinstance(item, dict): + issues.append( + _issue( + "manifest_entry_invalid", + "Skill manifest 条目必须是对象", + path=relative_manifest, + entry_index=index, + ) + ) + continue + + name = item.get("name") + if ( + not isinstance(name, str) + or not name.strip() + or "/" in name + or "\\" in name + or name in {".", ".."} + ): + issues.append( + _issue( + "manifest_name_invalid", + "manifest 条目的 name 必须是非空目录名", + path=relative_manifest, + entry_index=index, + actual=name, + ) + ) + else: + name_occurrences.setdefault(name, []).append(index) + + contract_owner = item.get("contract_owner") + if not isinstance(contract_owner, str) or not contract_owner.strip(): + issues.append( + _issue( + "manifest_contract_owner_invalid", + "manifest 条目的 contract_owner 必须非空", + path=relative_manifest, + entry_index=index, + ) + ) + + collaborations = item.get("collaborates_with") + if ( + not isinstance(collaborations, list) + or any(not isinstance(value, str) or not value.strip() for value in collaborations) + ): + issues.append( + _issue( + "manifest_collaborates_with_invalid", + "manifest 条目的 collaborates_with 必须是非空字符串数组", + path=relative_manifest, + entry_index=index, + ) + ) + + declared_skill_path = item.get("skill_path") + if not isinstance(declared_skill_path, str) or not declared_skill_path.strip(): + issues.append( + _issue( + "manifest_skill_path_invalid", + "manifest 条目的 skill_path 必须是非空相对路径", + path=relative_manifest, + entry_index=index, + actual=declared_skill_path, + ) + ) + continue + + normalised_path, candidate = _normalise_manifest_skill_path( + root, declared_skill_path + ) + if normalised_path is None or candidate is None: + issues.append( + _issue( + "manifest_skill_path_invalid", + "manifest 条目的 skill_path 必须位于项目根目录内", + path=relative_manifest, + entry_index=index, + actual=declared_skill_path, + ) + ) + continue + + declared_paths.add(normalised_path) + path_occurrences.setdefault(normalised_path, []).append(index) + if candidate != candidate.parent / "SKILL.md" or not candidate.is_relative_to(skills_root): + issues.append( + _issue( + "manifest_skill_path_invalid", + "manifest 条目的 skill_path 必须指向 .claude/skills/*/SKILL.md", + path=normalised_path, + entry_index=index, + ) + ) + + if not candidate.exists(): + issues.append( + _issue( + "manifest_skill_path_missing", + "manifest 登记的 Skill 路径不存在", + path=normalised_path, + entry_index=index, + ) + ) + elif not candidate.is_file(): + issues.append( + _issue( + "manifest_skill_path_not_file", + "manifest 登记的 Skill 路径不是普通文件", + path=normalised_path, + entry_index=index, + ) + ) + + directory_name = candidate.parent.name + if isinstance(name, str) and name.strip() and name != directory_name: + issues.append( + _issue( + "manifest_name_directory_mismatch", + "manifest name 与 skill_path 的目录名不一致", + path=normalised_path, + entry_index=index, + expected=directory_name, + actual=name, + ) + ) + if isinstance(name, str) and name.strip(): + expected_path = (root / SKILLS_ROOT / name / "SKILL.md").resolve() + expected_relative = _relative_path(expected_path, root) + if normalised_path != expected_relative: + issues.append( + _issue( + "manifest_skill_path_mismatch", + "manifest skill_path 与 name 对应的 Skill 路径不一致", + path=normalised_path, + entry_index=index, + expected=expected_relative, + actual=normalised_path, + ) + ) + + for path, indexes in sorted(path_occurrences.items()): + if len(indexes) > 1: + issues.append( + _issue( + "manifest_duplicate_skill_path", + "manifest 重复登记同一个 Skill 路径", + path=path, + entry_indexes=indexes, + ) + ) + for name, indexes in sorted(name_occurrences.items()): + if len(indexes) > 1: + issues.append( + _issue( + "manifest_duplicate_name", + "manifest 重复登记同一个 Skill name", + path=relative_manifest, + name=name, + entry_indexes=indexes, + ) + ) + + for path in sorted(disk_paths - declared_paths): + issues.append( + _issue( + "manifest_skill_missing", + "磁盘 Skill 未在 manifest 中登记", + path=path, + ) + ) + for path in sorted(declared_paths - disk_paths): + issues.append( + _issue( + "manifest_skill_extra", + "manifest 登记了磁盘中不存在的额外 Skill", + path=path, + ) + ) + + return metadata, issues + + +def _duplicate_name_issues(entries: Iterable[dict[str, Any]]) -> list[dict[str, Any]]: + by_name: dict[str, list[str]] = {} + for entry in entries: + name = entry.get("name") + if isinstance(name, str) and name: + by_name.setdefault(name, []).append(str(entry["skill_path"])) + + issues: list[dict[str, Any]] = [] + for name in sorted(by_name): + paths = sorted(by_name[name]) + if len(paths) > 1: + issues.append( + _issue( + "duplicate_name", + "多个 Skill 使用相同的 frontmatter name", + path=paths[0], + name=name, + paths=paths, + ) + ) + return issues + + +def _summarize(issues: Sequence[dict[str, Any]], skill_count: int) -> dict[str, Any]: + counts = Counter(str(issue["code"]) for issue in issues) + return { + "status": "passed" if not issues else "failed", + "skills_scanned": skill_count, + "issue_count": len(issues), + "issue_codes": {code: counts[code] for code in sorted(counts)}, + } + + +def audit_skills( + root: str | Path = ".", manifest: str | Path | None = None +) -> dict[str, Any]: + """审计 Skill 文档与 manifest,并返回 JSON 可序列化报告。""" + + root_path = _resolve_root(root) + manifest_path = _resolve_manifest_path(root_path, manifest) + report: dict[str, Any] = { + "schema_version": SCHEMA_VERSION, + "root": str(root_path), + "skills_root": SKILLS_ROOT.as_posix(), + "manifest": { + "path": _relative_path(manifest_path, root_path), + "loaded": False, + }, + "skills_scanned": 0, + "skills": [], + "issues": [], + } + + if not root_path.exists(): + report["issues"] = [ + _issue("root_missing", "审计根目录不存在", path=".") + ] + elif not root_path.is_dir(): + report["issues"] = [ + _issue("root_not_directory", "审计根路径不是目录", path=".") + ] + else: + entries, issues = _scan_skill_directories(root_path) + issues.extend(_duplicate_name_issues(entries)) + manifest_info, manifest_issues = _validate_manifest( + root_path, entries, manifest_path + ) + issues.extend(manifest_issues) + report["manifest"] = manifest_info + report["skills"] = entries + report["skills_scanned"] = len(entries) + report["issues"] = issues + + report["summary"] = _summarize(report["issues"], report["skills_scanned"]) + report["issue_count"] = len(report["issues"]) + report["ok"] = not report["issues"] + report["status"] = "passed" if report["ok"] else "failed" + return report + + +def human_summary(report: dict[str, Any]) -> str: + """把结构化报告渲染为简洁的人读摘要。""" + + status = "通过" if report["ok"] else "失败" + lines = [ + f"Skill 静态审计:{status}", + f"扫描 Skill:{report['skills_scanned']} 个;问题:{report['issue_count']} 个", + ] + issues = report["issues"] + if not issues: + lines.append("未发现 frontmatter、目录命名、空文件或开发测试污染问题。") + return "\n".join(lines) + + lines.append("问题明细:") + for issue in issues: + location = str(issue.get("path", "")) + if issue.get("line") is not None: + location += f":{issue['line']}" + lines.append(f"- {location} [{issue['code']}] {issue['message']}") + return "\n".join(lines) + + +def _build_parser() -> argparse.ArgumentParser: + parser = argparse.ArgumentParser( + description="只读扫描项目运行时 .claude/skills/*/SKILL.md 的静态卫生审计器。" + ) + parser.add_argument( + "--root", + default=".", + help="项目根目录,默认为当前目录", + ) + parser.add_argument( + "--manifest", + default=None, + help="Skill manifest 路径;相对路径按项目根目录解析,默认 harness/manifests/skills.json", + ) + parser.add_argument( + "--json", + action="store_true", + help="输出结构化 JSON 报告", + ) + parser.add_argument( + "--quiet", + action="store_true", + help="不输出人读摘要;与 --json 同用时仍输出 JSON", + ) + return parser + + +def main(argv: Optional[Sequence[str]] = None) -> int: + args = _build_parser().parse_args(argv) + report = audit_skills(args.root, args.manifest) + if args.json: + print(json.dumps(report, ensure_ascii=False, indent=2, sort_keys=True)) + elif not args.quiet: + print(human_summary(report)) + return 0 if report["ok"] else 1 + + +if __name__ == "__main__": + sys.exit(main()) diff --git a/harness/specs/skill-testing.md b/harness/specs/skill-testing.md new file mode 100644 index 0000000..2b76ecc --- /dev/null +++ b/harness/specs/skill-testing.md @@ -0,0 +1,115 @@ +# Skill 测试与评测规范 + +> 状态:生效规范 +> 适用范围:`agent-example/.claude/skills/*/SKILL.md`、其确定性工具和 Skill 行为评测。 +> Owner:`agent-example/harness/` +> 本文件是开发与评测支架规范,不属于任何 Skill 的运行时提示词。 + +## 1. 三层边界 + +### 1.1 Skill 运行时指令 + +`SKILL.md` 只描述 Agent 执行该 Skill 时必须知道的稳定合同:唯一目的、消费者、输入、输出、schema/version、允许读取、允许副作用、禁止动作、数据库读写、授权、预算、raw 和失败关闭规则,以及实际执行所需的业务工具与流程。 + +`SKILL.md` 不承载开发者测试说明、测试文件路径、测试框架命令、夹具细节或测试结论。 + +禁止在运行时 Skill 文本中出现: + +- `## 自测`、`## 测试`、`## 开发验证` 等开发测试章节; +- `test_*.py`、`*_test.py`、pytest、unittest 等实现测试入口; +- “本测试通过”“测试覆盖 N 项”等开发证据; +- 把测试文件、测试常量或测试输出说成事实源或业务合同。 + +评测 Skill 的生产/评测执行命令可以保留,但必须是用户请求该评测时实际执行的业务步骤,而不是验证实现代码的开发测试。 + +### 1.2 确定性工具验证 + +工具验证回答“Python、数据库或状态机实现是否守住机械合同”,不回答模型是否写得好,也不证明 Skill 的自然语言指令有效。 + +- 单元测试、离线合同测试和数据库集成测试属于这一层; +- 项目运行时 Skill 的实现测试统一放在 `tests/skills//`,领域包保留自己的 `tests/`;两者都必须在 harness 的开发验证清单中登记; +- 测试通过只能证明列出的机械行为,不得升级为模型质量或产品可用性结论; +- 真实数据库、网络、模型和额度探针必须显式标记并单独授权。 + +### 1.3 Skill 行为评测 + +行为评测回答“把 Skill 内容交给 Agent 后,Agent 是否按合同行动”。评测由 harness 从外部驱动,不能由 `SKILL.md` 自己宣布通过。 + +行为评测至少区分: + +- 正向触发:应该使用该 Skill 的任务; +- 负向触发:不应该使用该 Skill 的任务; +- 输入缺失与越界; +- 输出合同与失败关闭; +- 关键禁止动作; +- 多样例稳定性和已知混淆项。 + +行为评测结果必须保留输入、Skill 版本指纹、模型/运行配置、输出摘要、判定证据和失败原因。小样本合同回放不等于文学质量、通用效果或生产完成。 + +## 2. Harness 职责 + +Harness 只负责从外部发现、运行、收集和裁决测试/评测证据: + +1. 扫描运行时 Skill 文档中的开发测试污染; +2. 登记和分类确定性工具测试、集成测试、fake pipeline 测试和真实探针; +3. 驱动 Skill 行为评测案例; +4. 生成带证据类型和失败原因的结构化报告; +5. 防止空跑、静默跳过、宽泛异常吞错和证据级别越权。 + +Harness 不负责: + +- 修改 `SKILL.md` 或业务代码; +- 判断文学质量; +- 用字符串出现证明自然语言合同有效; +- 把 fake、离线回放或小样本结果升级成生产结论。 + +## 3. 证据分级 + +从低到高仅表示证据类型,不允许自动跨级: + +1. 静态结构检查:frontmatter、路径、禁用污染模式; +2. 确定性离线测试:纯函数、schema、状态机和失败分支; +3. 真实依赖集成测试:PostgreSQL、文件系统或外部服务; +4. Skill 行为评测:外部 Agent/模型运行与结构化裁决; +5. 人工内容评审:文学质量、声音和语义效果。 + +任一层通过都不能代替更高层证据。没有行为评测,不得声称 Skill 内容有效;没有真实依赖证据,不得声称生产链路可用。 + +## 4. 改动门禁 + +### 只改运行时合同 + +- 通过 harness 的运行时文档污染扫描; +- 更新受影响的行为评测案例,或记录只改措辞、不改变行为的理由; +- 不新增把测试细节塞回 `SKILL.md` 的说明。 + +### 只改确定性工具 + +- 更新针对变更不变量的离线测试; +- 对真实数据库、网络和模型测试单独标记; +- 保留失败关闭和副作用边界证据。 + +### 改变 Skill 意图或边界 + +- 更新行为评测清单和正/负向案例; +- 检查消费者、Agent 槽位和 `meta/chains` 映射; +- 重新执行相关层级的验证,不得只跑 Python 单测。 + +## 5. 明确禁止 + +- 用 `assertIn` 检查几个词出现,就宣称 Skill 内容正确; +- 用测试文件或测试输出作为运行时事实源; +- 用 `except Exception: pass` 把错误依赖、连接失败或实现错误当成预期拒绝; +- 用任意总入口的“零测试”结果当成通过; +- 把离线 fake、模型探针或小样本合同回放写成真实质量结论; +- 为了让 harness 变绿而修改业务合同、降低断言或静默跳过测试。 + +## 6. 完成定义 + +本治理任务只有同时满足以下条件,才可称为完成: + +- 运行时 `SKILL.md` 不再携带开发测试说明; +- harness 能机械发现并阻断明显的测试污染; +- 实现测试、集成测试和 Skill 行为评测的证据类型可区分; +- 改动范围内的回归测试真实执行,失败不会被吞掉; +- 报告明确区分已验证事实、推断和未验证的模型质量假设。 diff --git a/harness/test_run_selected.py b/harness/test_run_selected.py new file mode 100644 index 0000000..e5f7257 --- /dev/null +++ b/harness/test_run_selected.py @@ -0,0 +1,438 @@ +#!/usr/bin/env python3 +"""run_selected 的标准库离线回归测试。 + +测试项目、manifest 和 .venv/bin/python 均在临时目录中生成,不连接数据库、网络或模型。 +""" + +from __future__ import annotations + +import io +import json +import stat +import sys +import tempfile +import unittest +from contextlib import redirect_stdout +from pathlib import Path +from typing import Any, Iterable + +try: + from .run_selected import main +except ImportError: # 允许直接执行 `.venv/bin/python harness/test_run_selected.py` + from run_selected import main + + +class RunSelectedTests(unittest.TestCase): + def make_project( + self, + entries: Iterable[dict[str, Any]], + *, + generated_scope: str | None = None, + ) -> tuple[tempfile.TemporaryDirectory[str], Path]: + temporary = tempfile.TemporaryDirectory() + root = Path(temporary.name) + (root / ".venv" / "bin").mkdir(parents=True) + fake_python = root / ".venv" / "bin" / "python" + fake_python.write_text( + f"#!{sys.executable}\n" + "import pathlib\n" + "import sys\n" + "import time\n" + "script = pathlib.Path(sys.argv[1])\n" + "(pathlib.Path.cwd() / 'invocations.log').open('a', encoding='utf-8').write(script.name + '\\n')\n" + "if script.name == 'empty.py':\n" + " raise SystemExit(0)\n" + "if script.name == 'skip.py':\n" + " print('1 skipped')\n" + " raise SystemExit(0)\n" + "if script.name == 'fail.py':\n" + " print('failure stdout')\n" + " print('failure stderr', file=sys.stderr)\n" + " raise SystemExit(7)\n" + "if script.name == 'sleep.py':\n" + " time.sleep(2)\n" + "if script.name == 'noisy.py':\n" + " print('x' * 5000)\n" + "else:\n" + " print('ran ' + script.name)\n", + encoding="utf-8", + ) + fake_python.chmod(fake_python.stat().st_mode | stat.S_IXUSR | stat.S_IXGRP | stat.S_IXOTH) + + manifest_entries = [] + for entry in entries: + entry_copy = dict(entry) + path = root / str(entry_copy["path"]) + if entry_copy.pop("create", True): + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text("# temporary test script\n", encoding="utf-8") + manifest_entries.append(entry_copy) + + manifest_path = root / "manifest.json" + manifest_payload: dict[str, Any] = { + "schema_version": 1, + "entries": manifest_entries, + } + if generated_scope is not None: + manifest_payload["generated_scope"] = generated_scope + manifest_path.write_text( + json.dumps(manifest_payload, indent=2) + "\n", + encoding="utf-8", + ) + self.addCleanup(temporary.cleanup) + return temporary, manifest_path + + def invoke(self, root: Path, manifest: Path, *arguments: str) -> tuple[int, dict[str, Any]]: + output = io.StringIO() + with redirect_stdout(output): + return_code = main( + [ + "--root", + str(root), + "--manifest", + str(manifest), + *arguments, + "--json", + ] + ) + return return_code, json.loads(output.getvalue()) + + @staticmethod + def entry( + path: str, + *, + owner: str = "alpha", + kind: str = "tool_unit", + scope: str = "runtime_skill", + requires: list[str] | None = None, + create: bool = True, + ) -> dict[str, Any]: + return { + "path": path, + "owner_skill_or_domain": owner, + "kind": kind, + "scope": scope, + "requires": ["offline"] if requires is None else requires, + "create": create, + } + + def test_selector_alias_and_intersection_are_reported_as_json(self) -> None: + _, manifest = self.make_project( + [ + self.entry("pass.py", owner="alpha", kind="tool_unit"), + self.entry("other.py", owner="beta", kind="tool_contract"), + ] + ) + root = manifest.parent + + return_code, report = self.invoke( + root, + manifest, + "--skill", + "alpha", + "--kind", + "tool_unit", + "--path", + "pass.py", + ) + + self.assertEqual(return_code, 0) + self.assertEqual(report["status"], "passed") + self.assertEqual(report["selected_count"], 1) + self.assertEqual(report["entries"][0]["path"], "pass.py") + self.assertEqual(report["entries"][0]["returncode"], 0) + self.assertIn("ran pass.py", report["entries"][0]["stdout"]) + + def test_offline_dependency_is_blocked_without_running_child(self) -> None: + _, manifest = self.make_project( + [self.entry("db.py", requires=["offline", "postgresql"])] + ) + root = manifest.parent + + return_code, report = self.invoke(root, manifest, "--path", "db.py") + + result = report["entries"][0] + self.assertEqual(return_code, 1) + self.assertEqual(report["status"], "failed") + self.assertEqual(result["status"], "blocked_dependency") + self.assertEqual(result["requires"], ["offline", "postgresql"]) + self.assertEqual(result["blocked_requires"], ["postgresql"]) + self.assertFalse((root / "invocations.log").exists()) + + def test_allow_requires_runs_and_still_reports_dependency(self) -> None: + _, manifest = self.make_project( + [self.entry("db.py", requires=["offline", "postgresql"])] + ) + root = manifest.parent + + return_code, report = self.invoke( + root, + manifest, + "--path", + "db.py", + "--allow-requires", + "postgresql", + ) + + result = report["entries"][0] + self.assertEqual(return_code, 0) + self.assertEqual(result["status"], "passed") + self.assertEqual(result["requires"], ["offline", "postgresql"]) + self.assertEqual((root / "invocations.log").read_text(encoding="utf-8"), "db.py\n") + + def test_nonzero_and_timeout_have_distinct_structured_statuses(self) -> None: + _, manifest = self.make_project( + [self.entry("fail.py"), self.entry("sleep.py")] + ) + root = manifest.parent + + return_code, report = self.invoke( + root, + manifest, + "--path", + "fail.py", + "--path", + "sleep.py", + "--timeout-seconds", + "0.5", + ) + + self.assertEqual(return_code, 1) + results = {entry["path"]: entry for entry in report["entries"]} + self.assertEqual(results["fail.py"]["status"], "failed") + self.assertEqual(results["fail.py"]["returncode"], 7) + self.assertIn("failure stdout", results["fail.py"]["stdout"]) + self.assertIn("failure stderr", results["fail.py"]["stderr"]) + self.assertEqual(results["sleep.py"]["status"], "timeout") + self.assertIsNone(results["sleep.py"]["returncode"]) + self.assertEqual(results["sleep.py"]["error"]["code"], "timeout") + + def test_no_matches_and_missing_path_are_not_silent(self) -> None: + _, manifest = self.make_project( + [ + self.entry("missing.py", create=False), + self.entry("present.py"), + ] + ) + root = manifest.parent + + return_code, no_match = self.invoke(root, manifest, "--kind", "does_not_exist") + self.assertEqual(return_code, 1) + self.assertEqual(no_match["status"], "no_matches") + self.assertEqual(no_match["entries"], []) + self.assertEqual(no_match["issues"][0]["code"], "no_matches") + + return_code, missing = self.invoke(root, manifest, "--path", "missing.py") + self.assertEqual(return_code, 1) + self.assertEqual(missing["entries"][0]["status"], "not_found") + self.assertEqual(missing["entries"][0]["error"]["code"], "path_not_found") + + def test_invalid_manifest_and_selector_requirement_are_structured(self) -> None: + temporary = tempfile.TemporaryDirectory() + self.addCleanup(temporary.cleanup) + root = Path(temporary.name) + manifest = root / "manifest.json" + manifest.write_text("{broken", encoding="utf-8") + + return_code, invalid_manifest = self.invoke(root, manifest, "--path", "x.py") + self.assertEqual(return_code, 1) + self.assertEqual(invalid_manifest["status"], "manifest_invalid") + self.assertEqual(invalid_manifest["issues"][0]["code"], "manifest_invalid_json") + + valid_root, valid_manifest = self.make_project([self.entry("pass.py")]) + return_code, invalid_selector = self.invoke(valid_root, valid_manifest) + self.assertEqual(return_code, 1) + self.assertEqual(invalid_selector["status"], "invalid_selector") + self.assertEqual(invalid_selector["issues"][0]["code"], "selector_required") + + def test_all_offline_is_explicit_opt_in_and_output_is_summary_only(self) -> None: + _, manifest = self.make_project( + [ + self.entry("pass.py"), + self.entry("db.py", requires=["postgresql"]), + self.entry("noisy.py"), + ] + ) + root = manifest.parent + + return_code, report = self.invoke(root, manifest, "--all-offline") + + self.assertEqual(return_code, 1) + self.assertEqual(report["selected_count"], 3) + statuses = {entry["path"]: entry["status"] for entry in report["entries"]} + self.assertEqual(statuses["pass.py"], "passed") + self.assertEqual(statuses["db.py"], "blocked_dependency") + self.assertEqual(statuses["noisy.py"], "passed") + noisy = next(entry for entry in report["entries"] if entry["path"] == "noisy.py") + self.assertLess(len(noisy["stdout"]), 2001) + self.assertIn("output summary truncated", noisy["stdout"]) + + def test_generated_scope_rejects_unregistered_disk_asset(self) -> None: + _, manifest = self.make_project( + [self.entry("tests/skills/registered/test_registered.py")], + generated_scope="temporary test asset inventory", + ) + root = manifest.parent + unregistered = root / "tests" / "skills" / "new" / "test_unregistered.py" + unregistered.parent.mkdir(parents=True) + unregistered.write_text("# unregistered test asset\n", encoding="utf-8") + + return_code, report = self.invoke(root, manifest, "--all-offline") + + self.assertEqual(return_code, 1) + self.assertEqual(report["status"], "manifest_invalid") + self.assertEqual(report["entries"], []) + self.assertEqual( + [ + issue["path"] + for issue in report["issues"] + if issue["code"] == "manifest_test_asset_missing" + ], + ["tests/skills/new/test_unregistered.py"], + ) + self.assertFalse((root / "invocations.log").exists()) + + def test_generated_scope_rejects_manifest_extra_asset(self) -> None: + _, manifest = self.make_project( + [ + self.entry( + "tests/skills/removed/test_removed.py", + create=False, + ) + ], + generated_scope="temporary test asset inventory", + ) + root = manifest.parent + + return_code, report = self.invoke(root, manifest, "--all-offline") + + self.assertEqual(return_code, 1) + self.assertEqual(report["status"], "manifest_invalid") + self.assertEqual(report["entries"], []) + self.assertEqual( + [ + issue["path"] + for issue in report["issues"] + if issue["code"] == "manifest_test_asset_extra" + ], + ["tests/skills/removed/test_removed.py"], + ) + self.assertFalse((root / "invocations.log").exists()) + + def test_generated_scope_ignores_non_test_helpers_under_test_roots(self) -> None: + _, manifest = self.make_project( + [self.entry("tests/skills/registered/test_registered.py")], + generated_scope="temporary test asset inventory", + ) + root = manifest.parent + registered_test = root / "tests" / "skills" / "registered" / "test_registered.py" + registered_test.write_text("def test_registered():\n pass\n", encoding="utf-8") + (root / "tests" / "skills" / "registered" / "helper.py").write_text( + "VALUE = 1\n", encoding="utf-8" + ) + (root / "humanization" / "tests" / "helper.py").parent.mkdir( + parents=True, exist_ok=True + ) + (root / "humanization" / "tests" / "helper.py").write_text( + "VALUE = 2\n", encoding="utf-8" + ) + + return_code, report = self.invoke(root, manifest, "--all-offline") + + self.assertEqual(return_code, 0) + self.assertEqual(report["status"], "passed") + self.assertEqual(report["manifest"]["test_assets_scanned"], 1) + self.assertEqual(report["entries"][0]["path"], "tests/skills/registered/test_registered.py") + + def test_generated_scope_blocks_empty_and_print_only_scripts_before_child(self) -> None: + entries = [ + self.entry("tests/skills/empty/test_empty.py"), + self.entry("tests/skills/print_only/test_print_only.py"), + ] + _, manifest = self.make_project( + entries, + generated_scope="temporary test asset inventory", + ) + root = manifest.parent + (root / "tests" / "skills" / "empty" / "test_empty.py").write_text( + "", encoding="utf-8" + ) + (root / "tests" / "skills" / "print_only" / "test_print_only.py").write_text( + "print('not a test')\n", encoding="utf-8" + ) + + return_code, report = self.invoke(root, manifest, "--all-offline") + + self.assertEqual(return_code, 1) + self.assertEqual(report["status"], "failed") + results = {entry["path"]: entry for entry in report["entries"]} + for path in ( + "tests/skills/empty/test_empty.py", + "tests/skills/print_only/test_print_only.py", + ): + self.assertEqual(results[path]["status"], "failed") + self.assertIsNone(results[path]["returncode"]) + self.assertEqual(results[path]["error"]["code"], "test_shape_missing") + self.assertFalse((root / "invocations.log").exists()) + + def test_generated_scope_accepts_supported_python_shapes(self) -> None: + entries = [ + self.entry("tests/skills/function/test_function.py"), + self.entry("tests/skills/class/test_class.py"), + self.entry("harness/test_main_entry.py"), + ] + _, manifest = self.make_project( + entries, + generated_scope="temporary test asset inventory", + ) + root = manifest.parent + (root / "tests" / "skills" / "function" / "test_function.py").write_text( + "def test_function():\n pass\n", encoding="utf-8" + ) + (root / "tests" / "skills" / "class" / "test_class.py").write_text( + "import unittest\n\nclass Fixture(unittest.TestCase):\n pass\n", + encoding="utf-8", + ) + (root / "harness" / "test_main_entry.py").write_text( + "if __name__ == \"__main__\":\n print(\"script\")\n", + encoding="utf-8", + ) + + return_code, report = self.invoke(root, manifest, "--all-offline") + + self.assertEqual(return_code, 0) + self.assertEqual(report["status"], "passed") + self.assertEqual( + [entry["status"] for entry in report["entries"]], + ["passed", "passed", "passed"], + ) + self.assertEqual( + (root / "invocations.log").read_text(encoding="utf-8").splitlines(), + ["test_function.py", "test_class.py", "test_main_entry.py"], + ) + + def test_zero_exit_without_execution_evidence_fails_closed(self) -> None: + _, manifest = self.make_project( + [self.entry("empty.py"), self.entry("skip.py")] + ) + root = manifest.parent + + return_code, report = self.invoke( + root, + manifest, + "--path", + "empty.py", + "--path", + "skip.py", + ) + + self.assertEqual(return_code, 1) + results = {entry["path"]: entry for entry in report["entries"]} + for path in ("empty.py", "skip.py"): + self.assertEqual(results[path]["status"], "failed") + self.assertEqual(results[path]["returncode"], 0) + self.assertEqual(results[path]["error"]["code"], "no_execution_evidence") + + +if __name__ == "__main__": + unittest.main(verbosity=2) diff --git a/harness/test_skill_harness.py b/harness/test_skill_harness.py new file mode 100644 index 0000000..bafbb53 --- /dev/null +++ b/harness/test_skill_harness.py @@ -0,0 +1,420 @@ +#!/usr/bin/env python3 +"""skill_harness 的纯标准库离线回归测试。 + +所有夹具都在临时目录中构造,不读取当前仓库的 Skill,也不依赖数据库、网络或模型。 +""" + +from __future__ import annotations + +import io +import json +import tempfile +import unittest +from contextlib import redirect_stdout +from pathlib import Path +from typing import Optional + +try: + from .skill_harness import audit_skills, main +except ImportError: # 允许直接执行 `.venv/bin/python harness/test_skill_harness.py` + from skill_harness import audit_skills, main + + +class SkillHarnessTests(unittest.TestCase): + def make_skill( + self, + root: Path, + directory: str, + *, + name: Optional[str] = None, + body: str = "# 合同\n\n只描述运行时行为。\n", + ) -> Path: + skill_dir = root / ".claude" / "skills" / directory + skill_dir.mkdir(parents=True, exist_ok=True) + skill_path = skill_dir / "SKILL.md" + frontmatter_name = directory if name is None else name + skill_path.write_text( + f"---\nname: {frontmatter_name}\ndescription: 离线夹具\n---\n{body}", + encoding="utf-8", + ) + self.write_manifest(root) + return skill_path + + def write_manifest( + self, + root: Path, + entries: Optional[list[dict[str, object]]] = None, + *, + path: Optional[Path] = None, + ) -> Path: + if entries is None: + entries = [] + skills_root = root / ".claude" / "skills" + if skills_root.exists(): + for directory in sorted(skills_root.iterdir(), key=lambda item: item.name): + if directory.is_dir(): + entries.append( + { + "name": directory.name, + "contract_owner": "测试夹具", + "collaborates_with": [], + "skill_path": ( + Path(".claude") + / "skills" + / directory.name + / "SKILL.md" + ).as_posix(), + } + ) + manifest_path = path or root / "harness" / "manifests" / "skills.json" + manifest_path.parent.mkdir(parents=True, exist_ok=True) + manifest_path.write_text( + json.dumps( + {"schema_version": 1, "skills": entries}, + ensure_ascii=False, + indent=2, + ) + + "\n", + encoding="utf-8", + ) + return manifest_path + + def test_passing_project_is_independent_of_current_repository(self) -> None: + with tempfile.TemporaryDirectory() as temporary: + root = Path(temporary) + self.make_skill(root, "alpha") + + report = audit_skills(root) + + self.assertTrue(report["ok"]) + self.assertEqual(report["status"], "passed") + self.assertEqual(report["skills_scanned"], 1) + self.assertEqual(report["issues"], []) + + def test_pollution_is_reported_for_each_obvious_category(self) -> None: + with tempfile.TemporaryDirectory() as temporary: + root = Path(temporary) + self.make_skill( + root, + "polluted", + body=( + "## 自测\n" + "pytest -q\n" + "import unittest\n" + "run test_sample.py and sample_test.py\n" + "覆盖率达到 100%。\n" + "测试通过。\n" + ), + ) + + report = audit_skills(root) + codes = {issue["code"] for issue in report["issues"]} + + self.assertFalse(report["ok"]) + self.assertTrue( + { + "development_test_heading", + "pytest_reference", + "unittest_reference", + "test_file_reference", + "coverage_reference", + "test_pass_declaration", + }.issubset(codes) + ) + + def test_development_headings_support_multiple_levels(self) -> None: + with tempfile.TemporaryDirectory() as temporary: + root = Path(temporary) + self.make_skill( + root, + "nested-development-sections", + body=( + "### 自测\n" + "这里不能登记测试命令。\n" + "## 离线验证\n" + "这里也不能登记测试命令。\n" + ), + ) + + report = audit_skills(root) + headings = [ + issue + for issue in report["issues"] + if issue["code"] == "development_test_heading" + ] + + self.assertFalse(report["ok"]) + self.assertEqual(len(headings), 2) + self.assertEqual([issue["line"] for issue in headings], [5, 7]) + + def test_check_contract_command_requires_development_section_context(self) -> None: + with tempfile.TemporaryDirectory() as temporary: + root = Path(temporary) + self.make_skill( + root, + "contextual-contract-command", + body=( + "业务合同字段名可以写 check_contract.py。\n" + "## 离线验证\n" + "python check_contract.py\n" + "## 业务说明\n" + "普通业务 check_contract.py 不是测试入口。\n" + ), + ) + + report = audit_skills(root) + command_issues = [ + issue + for issue in report["issues"] + if issue["code"] == "development_test_command_reference" + ] + + self.assertFalse(report["ok"]) + self.assertEqual(len(command_issues), 1) + self.assertEqual(command_issues[0]["line"], 7) + + def test_business_coverage_terms_are_not_development_pollution(self) -> None: + with tempfile.TemporaryDirectory() as temporary: + root = Path(temporary) + self.make_skill( + root, + "business-coverage", + body=( + "细纲覆盖率达到 100%。\n" + "事件覆盖率为 80%。\n" + "硬约束覆盖率 100%。\n" + ), + ) + + report = audit_skills(root) + + self.assertTrue(report["ok"]) + self.assertNotIn( + "coverage_reference", + {issue["code"] for issue in report["issues"]}, + ) + + def test_development_coverage_terms_are_reported(self) -> None: + with tempfile.TemporaryDirectory() as temporary: + root = Path(temporary) + self.make_skill( + root, + "development-coverage", + body=( + "coverage.py\n" + "pytest-cov\n" + "pytest --cov=harness\n" + "coverage report\n" + "测试覆盖率达到 90%。\n" + "覆盖率达到 80%。\n" + ), + ) + + report = audit_skills(root) + + self.assertFalse(report["ok"]) + self.assertIn( + "coverage_reference", + {issue["code"] for issue in report["issues"]}, + ) + + def test_frontmatter_name_must_match_skill_directory(self) -> None: + with tempfile.TemporaryDirectory() as temporary: + root = Path(temporary) + self.make_skill(root, "alpha", name="beta") + + report = audit_skills(root) + + self.assertFalse(report["ok"]) + self.assertIn( + "name_directory_mismatch", + {issue["code"] for issue in report["issues"]}, + ) + + def test_duplicate_frontmatter_names_are_reported(self) -> None: + with tempfile.TemporaryDirectory() as temporary: + root = Path(temporary) + self.make_skill(root, "alpha", name="shared") + self.make_skill(root, "beta", name="shared") + + report = audit_skills(root) + duplicates = [ + issue for issue in report["issues"] if issue["code"] == "duplicate_name" + ] + + self.assertFalse(report["ok"]) + self.assertEqual(len(duplicates), 1) + self.assertEqual(duplicates[0]["name"], "shared") + self.assertEqual(len(duplicates[0]["paths"]), 2) + + def test_clean_manifest_can_be_selected_explicitly(self) -> None: + with tempfile.TemporaryDirectory() as temporary: + root = Path(temporary) + self.make_skill(root, "alpha") + manifest_path = self.write_manifest( + root, + path=root / "clean-manifest.json", + ) + output = io.StringIO() + + with redirect_stdout(output): + return_code = main( + [ + "--root", + str(root), + "--manifest", + str(manifest_path), + "--json", + "--quiet", + ] + ) + + payload = json.loads(output.getvalue()) + self.assertEqual(return_code, 0) + self.assertTrue(payload["ok"]) + self.assertTrue(payload["manifest"]["loaded"]) + self.assertEqual(payload["manifest"]["path"], "clean-manifest.json") + + def test_error_manifest_fails_with_structured_issues(self) -> None: + with tempfile.TemporaryDirectory() as temporary: + root = Path(temporary) + self.make_skill(root, "alpha") + self.make_skill(root, "beta") + broken_manifest = self.write_manifest( + root, + entries=[ + { + "name": "alpha", + "contract_owner": "", + "collaborates_with": ["上下文与知识检索", 3], + "skill_path": ".claude/skills/beta/SKILL.md", + }, + { + "name": "alpha", + "contract_owner": "测试夹具", + "collaborates_with": [], + "skill_path": ".claude/skills/beta/SKILL.md", + }, + { + "name": "extra", + "contract_owner": "测试夹具", + "collaborates_with": [], + "skill_path": ".claude/skills/extra/SKILL.md", + }, + ], + path=root / "broken-manifest.json", + ) + output = io.StringIO() + + with redirect_stdout(output): + return_code = main( + [ + "--root", + str(root), + "--manifest", + str(broken_manifest), + "--json", + ] + ) + + payload = json.loads(output.getvalue()) + codes = {issue["code"] for issue in payload["issues"]} + self.assertEqual(return_code, 1) + self.assertFalse(payload["ok"]) + self.assertTrue(all(isinstance(issue, dict) for issue in payload["issues"])) + self.assertTrue( + { + "manifest_contract_owner_invalid", + "manifest_collaborates_with_invalid", + "manifest_duplicate_skill_path", + "manifest_duplicate_name", + "manifest_skill_missing", + "manifest_skill_extra", + "manifest_skill_path_mismatch", + "manifest_name_directory_mismatch", + "manifest_skill_path_missing", + }.issubset(codes) + ) + + def test_missing_manifest_fails_closed(self) -> None: + with tempfile.TemporaryDirectory() as temporary: + root = Path(temporary) + self.make_skill(root, "alpha") + (root / "harness" / "manifests" / "skills.json").unlink() + + report = audit_skills(root) + + self.assertFalse(report["ok"]) + self.assertIn( + "manifest_missing", + {issue["code"] for issue in report["issues"]}, + ) + + def test_missing_skills_directory_fails_closed(self) -> None: + with tempfile.TemporaryDirectory() as temporary: + root = Path(temporary) + + report = audit_skills(root) + + self.assertFalse(report["ok"]) + self.assertIn( + "skills_directory_missing", + {issue["code"] for issue in report["issues"]}, + ) + + def test_skill_directory_without_skill_file_fails(self) -> None: + with tempfile.TemporaryDirectory() as temporary: + root = Path(temporary) + (root / ".claude" / "skills" / "missing").mkdir(parents=True) + + report = audit_skills(root) + + self.assertFalse(report["ok"]) + self.assertIn( + "missing_skill_file", + {issue["code"] for issue in report["issues"]}, + ) + + def test_empty_skill_file_is_reported(self) -> None: + with tempfile.TemporaryDirectory() as temporary: + root = Path(temporary) + skill_path = self.make_skill(root, "empty") + skill_path.write_text("", encoding="utf-8") + + report = audit_skills(root) + + self.assertFalse(report["ok"]) + self.assertIn( + "empty_skill_file", + {issue["code"] for issue in report["issues"]}, + ) + + def test_json_cli_and_exit_code(self) -> None: + with tempfile.TemporaryDirectory() as temporary: + root = Path(temporary) + self.make_skill(root, "alpha") + output = io.StringIO() + + with redirect_stdout(output): + return_code = main(["--root", str(root), "--json", "--quiet"]) + + payload = json.loads(output.getvalue()) + self.assertEqual(return_code, 0) + self.assertTrue(payload["ok"]) + self.assertEqual(payload["skills_scanned"], 1) + + (root / ".claude" / "skills" / "alpha" / "SKILL.md").write_text( + "---\nname: alpha\n---\n## 测试\n", + encoding="utf-8", + ) + output = io.StringIO() + with redirect_stdout(output): + return_code = main(["--root", str(root), "--json"]) + + failed_payload = json.loads(output.getvalue()) + self.assertEqual(return_code, 1) + self.assertFalse(failed_payload["ok"]) + + +if __name__ == "__main__": + unittest.main(verbosity=2) diff --git a/humanization/research/20-project-skill-coverage.yaml b/humanization/research/20-project-skill-coverage.yaml index f3b5a7f..28109ca 100644 --- a/humanization/research/20-project-skill-coverage.yaml +++ b/humanization/research/20-project-skill-coverage.yaml @@ -73,7 +73,7 @@ capabilities: evidence_sources: [no-ai-slop, neuro-book, humanizer] owner_skill: prevent-ai-flavor implementation: .claude/skills/prevent-ai-flavor/scripts/prevent_ai_flavor.py - test: .claude/skills/assemble-context/scripts/test_assemble_writer_context.py + test: tests/skills/assemble-context/test_assemble_writer_context.py status: implemented - id: sf_snf_boundary_regression evidence_sources: [speak-human-tw, shuorenhua, neuro-book] @@ -98,14 +98,14 @@ capabilities: evidence_sources: [inkos, oh-story-claudecode, neuro-book] owner_skill: revise-ai-flavor implementation: humanization/src/deai/pipeline.py - test: .claude/skills/revise-ai-flavor/scripts/test_revise_ai_flavor.py + test: tests/skills/revise-ai-flavor/test_revise_ai_flavor.py status: implemented note: 当前以最多3轮合同、复扫和pairwise no_gain实现;跨轮最佳快照编排仍由上层负责 - id: source_revalidation evidence_sources: [shuorenhua, neuro-book] owner_skill: capture-ai-flavor-cases implementation: .claude/skills/capture-ai-flavor-cases/scripts/capture_cases.py - test: .claude/skills/capture-ai-flavor-cases/scripts/test_capture_cases.py + test: tests/skills/capture-ai-flavor-cases/test_capture_cases.py status: implemented - id: case_to_rule_lifecycle evidence_sources: [unslop, shuorenhua, neuro-book] diff --git a/.claude/skills/access-database/scripts/test_authorization_snapshot_ddl.py b/tests/skills/access-database/test_authorization_snapshot_ddl.py similarity index 91% rename from .claude/skills/access-database/scripts/test_authorization_snapshot_ddl.py rename to tests/skills/access-database/test_authorization_snapshot_ddl.py index ca8bca2..9554592 100644 --- a/.claude/skills/access-database/scripts/test_authorization_snapshot_ddl.py +++ b/tests/skills/access-database/test_authorization_snapshot_ddl.py @@ -6,7 +6,8 @@ import re import unittest -DDL_PATH = pathlib.Path(__file__).resolve().parents[4] / "db" / "ddl" / "96-example参考作品授权快照.sql" +PROJECT_ROOT = pathlib.Path(__file__).resolve().parents[3] +DDL_PATH = PROJECT_ROOT / "db" / "ddl" / "96-example参考作品授权快照.sql" class AuthorizationSnapshotDdlTest(unittest.TestCase): diff --git a/.claude/skills/access-database/scripts/test_db_params.py b/tests/skills/access-database/test_db_params.py similarity index 81% rename from .claude/skills/access-database/scripts/test_db_params.py rename to tests/skills/access-database/test_db_params.py index b63650b..a836357 100644 --- a/.claude/skills/access-database/scripts/test_db_params.py +++ b/tests/skills/access-database/test_db_params.py @@ -1,14 +1,16 @@ #!/usr/bin/env python3 """db execparams 参数装载逻辑离线自测(不连库)。 -跑法(仓库根目录):.venv/bin/python .claude/skills/access-database/scripts/test_db_params.py +跑法(仓库根目录):.venv/bin/python tests/skills/access-database/test_db_params.py """ import io import json -import os import sys +from pathlib import Path -sys.path.insert(0, os.path.dirname(os.path.abspath(__file__))) +PROJECT_ROOT = Path(__file__).resolve().parents[3] +SCRIPTS_DIR = PROJECT_ROOT / ".claude" / "skills" / "access-database" / "scripts" +sys.path.insert(0, str(SCRIPTS_DIR)) import click # noqa: E402 diff --git a/.claude/skills/access-database/scripts/test_skill_catalog.py b/tests/skills/access-database/test_skill_catalog.py similarity index 91% rename from .claude/skills/access-database/scripts/test_skill_catalog.py rename to tests/skills/access-database/test_skill_catalog.py index d239f1d..f6fb724 100644 --- a/.claude/skills/access-database/scripts/test_skill_catalog.py +++ b/tests/skills/access-database/test_skill_catalog.py @@ -2,9 +2,14 @@ """Skill 目录命名与 frontmatter 一致性的离线测试。""" import pathlib +import sys import tempfile import unittest +PROJECT_ROOT = pathlib.Path(__file__).resolve().parents[3] +SCRIPTS_DIR = PROJECT_ROOT / ".claude" / "skills" / "access-database" / "scripts" +sys.path.insert(0, str(SCRIPTS_DIR)) + from sync_agent_registry import validate_skill_catalog diff --git a/.claude/skills/assemble-context/scripts/test_assemble_writer_context.py b/tests/skills/assemble-context/test_assemble_writer_context.py similarity index 99% rename from .claude/skills/assemble-context/scripts/test_assemble_writer_context.py rename to tests/skills/assemble-context/test_assemble_writer_context.py index c227b4f..34abf23 100644 --- a/.claude/skills/assemble-context/scripts/test_assemble_writer_context.py +++ b/tests/skills/assemble-context/test_assemble_writer_context.py @@ -9,7 +9,9 @@ import pathlib import sys import unittest -sys.path.insert(0, str(pathlib.Path(__file__).resolve().parent)) +PROJECT_ROOT = pathlib.Path(__file__).resolve().parents[3] +SCRIPT_DIR = PROJECT_ROOT / ".claude" / "skills" / "assemble-context" / "scripts" +sys.path.insert(0, str(SCRIPT_DIR)) from assemble_writer_context import ( # noqa: E402 AssemblyError, diff --git a/.claude/skills/assemble-context/scripts/test_fine_outline_reader.py b/tests/skills/assemble-context/test_fine_outline_reader.py similarity index 92% rename from .claude/skills/assemble-context/scripts/test_fine_outline_reader.py rename to tests/skills/assemble-context/test_fine_outline_reader.py index d85d4c2..3bdfb7f 100644 --- a/.claude/skills/assemble-context/scripts/test_fine_outline_reader.py +++ b/tests/skills/assemble-context/test_fine_outline_reader.py @@ -5,7 +5,9 @@ import pathlib import sys import unittest -sys.path.insert(0, str(pathlib.Path(__file__).resolve().parent)) +PROJECT_ROOT = pathlib.Path(__file__).resolve().parents[3] +SCRIPT_DIR = PROJECT_ROOT / ".claude" / "skills" / "assemble-context" / "scripts" +sys.path.insert(0, str(SCRIPT_DIR)) from retrieve_writer_sources import RetrievalError, load_confirmed_fine_outline # noqa: E402 diff --git a/.claude/skills/assemble-context/scripts/test_fine_outline_unification.py b/tests/skills/assemble-context/test_fine_outline_unification.py similarity index 95% rename from .claude/skills/assemble-context/scripts/test_fine_outline_unification.py rename to tests/skills/assemble-context/test_fine_outline_unification.py index d363139..3a2aae5 100644 --- a/.claude/skills/assemble-context/scripts/test_fine_outline_unification.py +++ b/tests/skills/assemble-context/test_fine_outline_unification.py @@ -14,7 +14,9 @@ import pathlib import sys import unittest -sys.path.insert(0, str(pathlib.Path(__file__).resolve().parent)) +PROJECT_ROOT = pathlib.Path(__file__).resolve().parents[3] +SCRIPT_DIR = PROJECT_ROOT / ".claude" / "skills" / "assemble-context" / "scripts" +sys.path.insert(0, str(SCRIPT_DIR)) from assemble_writer_context import AssemblyError, _outline_contract # noqa: E402 diff --git a/.claude/skills/assemble-context/scripts/test_pattern_binding_reader.py b/tests/skills/assemble-context/test_pattern_binding_reader.py similarity index 93% rename from .claude/skills/assemble-context/scripts/test_pattern_binding_reader.py rename to tests/skills/assemble-context/test_pattern_binding_reader.py index 68b5827..578adcd 100644 --- a/.claude/skills/assemble-context/scripts/test_pattern_binding_reader.py +++ b/tests/skills/assemble-context/test_pattern_binding_reader.py @@ -9,7 +9,9 @@ import pathlib import sys import unittest -sys.path.insert(0, str(pathlib.Path(__file__).resolve().parent)) +PROJECT_ROOT = pathlib.Path(__file__).resolve().parents[3] +SCRIPT_DIR = PROJECT_ROOT / ".claude" / "skills" / "assemble-context" / "scripts" +sys.path.insert(0, str(SCRIPT_DIR)) from retrieve_writer_sources import load_confirmed_pattern_bindings # noqa: E402 diff --git a/.claude/skills/assemble-context/scripts/test_retrieve_writer_sources.py b/tests/skills/assemble-context/test_retrieve_writer_sources.py similarity index 96% rename from .claude/skills/assemble-context/scripts/test_retrieve_writer_sources.py rename to tests/skills/assemble-context/test_retrieve_writer_sources.py index cd37871..69f70a8 100644 --- a/.claude/skills/assemble-context/scripts/test_retrieve_writer_sources.py +++ b/tests/skills/assemble-context/test_retrieve_writer_sources.py @@ -8,10 +8,18 @@ import pathlib import sys import unittest -SCRIPT_DIR = pathlib.Path(__file__).resolve().parent -sys.path.insert(0, str(SCRIPT_DIR)) -sys.path.insert(0, str(SCRIPT_DIR.parents[1] / "freeze-context" / "scripts")) -sys.path.insert(0, str(SCRIPT_DIR.parents[1] / "search-knowledge" / "scripts")) +PROJECT_ROOT = pathlib.Path(__file__).resolve().parents[3] +ASSEMBLE_CONTEXT_SCRIPTS = PROJECT_ROOT / ".claude" / "skills" / "assemble-context" / "scripts" +FREEZE_CONTEXT_SCRIPTS = PROJECT_ROOT / ".claude" / "skills" / "freeze-context" / "scripts" +SEARCH_KNOWLEDGE_SCRIPTS = PROJECT_ROOT / ".claude" / "skills" / "search-knowledge" / "scripts" +EMBED_KNOWLEDGE_SCRIPTS = PROJECT_ROOT / ".claude" / "skills" / "embed-knowledge" / "scripts" +for script_dir in ( + ASSEMBLE_CONTEXT_SCRIPTS, + FREEZE_CONTEXT_SCRIPTS, + SEARCH_KNOWLEDGE_SCRIPTS, + EMBED_KNOWLEDGE_SCRIPTS, +): + sys.path.insert(0, str(script_dir)) from load_reference_work import begin_read_snapshot # noqa: E402 from search import search_cards # noqa: E402 diff --git a/.claude/skills/assemble-context/scripts/test_style_loader.py b/tests/skills/assemble-context/test_style_loader.py similarity index 93% rename from .claude/skills/assemble-context/scripts/test_style_loader.py rename to tests/skills/assemble-context/test_style_loader.py index 28ce715..8620dad 100644 --- a/.claude/skills/assemble-context/scripts/test_style_loader.py +++ b/tests/skills/assemble-context/test_style_loader.py @@ -9,7 +9,9 @@ import pathlib import sys import unittest -sys.path.insert(0, str(pathlib.Path(__file__).resolve().parent)) +PROJECT_ROOT = pathlib.Path(__file__).resolve().parents[3] +SCRIPT_DIR = PROJECT_ROOT / ".claude" / "skills" / "assemble-context" / "scripts" +sys.path.insert(0, str(SCRIPT_DIR)) from retrieve_writer_sources import derive_style_constraints, load_confirmed_style # noqa: E402 diff --git a/.claude/skills/assemble-context/scripts/test_writer_contract.py b/tests/skills/assemble-context/test_writer_contract.py similarity index 99% rename from .claude/skills/assemble-context/scripts/test_writer_contract.py rename to tests/skills/assemble-context/test_writer_contract.py index dbb44f4..bee2602 100644 --- a/.claude/skills/assemble-context/scripts/test_writer_contract.py +++ b/tests/skills/assemble-context/test_writer_contract.py @@ -9,7 +9,9 @@ import pathlib import sys import unittest -sys.path.insert(0, str(pathlib.Path(__file__).resolve().parent)) +PROJECT_ROOT = pathlib.Path(__file__).resolve().parents[3] +SCRIPT_DIR = PROJECT_ROOT / ".claude" / "skills" / "assemble-context" / "scripts" +sys.path.insert(0, str(SCRIPT_DIR)) from writer_contract import ( # noqa: E402 ContractError, diff --git a/.claude/skills/call-content-model/scripts/test_call_persistence.py b/tests/skills/call-content-model/test_call_persistence.py similarity index 95% rename from .claude/skills/call-content-model/scripts/test_call_persistence.py rename to tests/skills/call-content-model/test_call_persistence.py index f6554ec..5260155 100644 --- a/.claude/skills/call-content-model/scripts/test_call_persistence.py +++ b/tests/skills/call-content-model/test_call_persistence.py @@ -9,7 +9,9 @@ import sys import types -sys.path.insert(0, str(pathlib.Path(__file__).resolve().parent)) +PROJECT_ROOT = pathlib.Path(__file__).resolve().parents[3] +SCRIPT_DIR = PROJECT_ROOT / ".claude" / "skills" / "call-content-model" / "scripts" +sys.path.insert(0, str(SCRIPT_DIR)) import llm # noqa: E402 diff --git a/.claude/skills/call-content-model/scripts/test_quota.py b/tests/skills/call-content-model/test_quota.py similarity index 98% rename from .claude/skills/call-content-model/scripts/test_quota.py rename to tests/skills/call-content-model/test_quota.py index 4dd3171..d2d311a 100644 --- a/.claude/skills/call-content-model/scripts/test_quota.py +++ b/tests/skills/call-content-model/test_quota.py @@ -10,7 +10,9 @@ import sys import types from datetime import datetime, timedelta -sys.path.insert(0, str(pathlib.Path(__file__).resolve().parent)) +PROJECT_ROOT = pathlib.Path(__file__).resolve().parents[3] +SCRIPT_DIR = PROJECT_ROOT / ".claude" / "skills" / "call-content-model" / "scripts" +sys.path.insert(0, str(SCRIPT_DIR)) import llm # noqa: E402 # 费率缓存预置为兜底表:cost_usd/chat_governed 记账时 get_pricing() 直接命中缓存,绝不触网 diff --git a/.claude/skills/capture-ai-flavor-cases/scripts/test_capture_cases.py b/tests/skills/capture-ai-flavor-cases/test_capture_cases.py similarity index 98% rename from .claude/skills/capture-ai-flavor-cases/scripts/test_capture_cases.py rename to tests/skills/capture-ai-flavor-cases/test_capture_cases.py index 595efa5..f64aa9e 100644 --- a/.claude/skills/capture-ai-flavor-cases/scripts/test_capture_cases.py +++ b/tests/skills/capture-ai-flavor-cases/test_capture_cases.py @@ -3,10 +3,16 @@ from __future__ import annotations +import sys import tempfile import unittest from pathlib import Path from unittest.mock import patch + +PROJECT_ROOT = Path(__file__).resolve().parents[3] +SCRIPT_DIR = PROJECT_ROOT / ".claude" / "skills" / "capture-ai-flavor-cases" / "scripts" +sys.path.insert(0, str(SCRIPT_DIR)) + import yaml from capture_cases import ( @@ -233,20 +239,20 @@ class CaptureCasesTest(unittest.TestCase): self.assertEqual("narration", sample["carrier"]) def test_shipped_fixtures_pass_the_same_validator(self): - root = Path(__file__).resolve().parents[1] / "references" / "fixtures" + root = SCRIPT_DIR.parent / "references" / "fixtures" for name in ("backfill-hash-only.yaml", "canonical-samples.yaml"): data = yaml.safe_load((root / name).read_text(encoding="utf-8")) for card in data["cards"]: validate_card(card) def test_shipped_rule_seed_is_candidate_only(self): - root = Path(__file__).resolve().parents[1] / "references" / "fixtures" + root = SCRIPT_DIR.parent / "references" / "fixtures" data = yaml.safe_load((root / "rule-candidates.yaml").read_text(encoding="utf-8")) self.assertTrue(data["rules"]) self.assertTrue(all(rule["status"] == "candidate" for rule in data["rules"])) def test_shipped_revalidation_report_is_structurally_usable(self): - root = Path(__file__).resolve().parents[1] / "references" / "fixtures" + root = SCRIPT_DIR.parent / "references" / "fixtures" data = yaml.safe_load((root / "revalidation-2026-08-14.json").read_text(encoding="utf-8")) self.assertEqual("ai-flavor-revalidation-v1", data["schema_version"]) self.assertTrue(data["usable"]) diff --git a/.claude/skills/check-content-consistency/scripts/test_build_semantic_input.py b/tests/skills/check-content-consistency/test_build_semantic_input.py similarity index 85% rename from .claude/skills/check-content-consistency/scripts/test_build_semantic_input.py rename to tests/skills/check-content-consistency/test_build_semantic_input.py index c29038e..49ad971 100644 --- a/.claude/skills/check-content-consistency/scripts/test_build_semantic_input.py +++ b/tests/skills/check-content-consistency/test_build_semantic_input.py @@ -4,7 +4,7 @@ 验证:WriterContext + 候选能被投影成通过闭集校验的 semantic-detector-input-v3; sourceRef 多余字段被清洗;身份字段严格绑定;哈希自洽。 -跑法:.venv/bin/python .claude/skills/check-content-consistency/scripts/test_build_semantic_input.py +跑法:.venv/bin/python tests/skills/check-content-consistency/test_build_semantic_input.py """ from __future__ import annotations @@ -14,12 +14,13 @@ import pathlib import sys import unittest -SCRIPT_DIR = pathlib.Path(__file__).resolve().parent -SKILLS_DIR = SCRIPT_DIR.parents[1] -for path in (SCRIPT_DIR, - SKILLS_DIR / "check-content-consistency" / "scripts", - SKILLS_DIR / "write-next-chapter" / "scripts", - SKILLS_DIR / "assemble-context" / "scripts"): +PROJECT_ROOT = pathlib.Path(__file__).resolve().parents[3] +SKILLS_DIR = PROJECT_ROOT / ".claude" / "skills" +TEST_DIR = PROJECT_ROOT / "tests" / "skills" / "check-content-consistency" +SCRIPT_DIR = SKILLS_DIR / "check-content-consistency" / "scripts" +CONTINUATION_DIR = SKILLS_DIR / "write-next-chapter" / "scripts" +READ_CONTEXT_DIR = SKILLS_DIR / "assemble-context" / "scripts" +for path in (TEST_DIR, SCRIPT_DIR, CONTINUATION_DIR, READ_CONTEXT_DIR): if str(path) not in sys.path: sys.path.insert(0, str(path)) diff --git a/.claude/skills/check-content-consistency/scripts/test_check_writer_candidate.py b/tests/skills/check-content-consistency/test_check_writer_candidate.py similarity index 91% rename from .claude/skills/check-content-consistency/scripts/test_check_writer_candidate.py rename to tests/skills/check-content-consistency/test_check_writer_candidate.py index ffb13b8..587a3fd 100644 --- a/.claude/skills/check-content-consistency/scripts/test_check_writer_candidate.py +++ b/tests/skills/check-content-consistency/test_check_writer_candidate.py @@ -9,12 +9,15 @@ import pathlib import sys import unittest -SCRIPT_DIR = pathlib.Path(__file__).resolve().parent -SKILLS_DIR = SCRIPT_DIR.parents[1] +PROJECT_ROOT = pathlib.Path(__file__).resolve().parents[3] +SKILLS_DIR = PROJECT_ROOT / ".claude" / "skills" +SCRIPT_DIR = SKILLS_DIR / "check-content-consistency" / "scripts" CONTINUATION_DIR = SKILLS_DIR / "write-next-chapter" / "scripts" READ_CONTEXT_DIR = SKILLS_DIR / "assemble-context" / "scripts" -for path in (SCRIPT_DIR, CONTINUATION_DIR, READ_CONTEXT_DIR): - sys.path.insert(0, str(path)) +WRITER_TEST_DIR = PROJECT_ROOT / "tests" / "skills" / "write-next-chapter" +for path in (SCRIPT_DIR, CONTINUATION_DIR, READ_CONTEXT_DIR, WRITER_TEST_DIR): + if str(path) not in sys.path: + sys.path.insert(0, str(path)) from test_run_writer import _bound_context # noqa: E402 from check_writer_candidate import check_writer_candidate # noqa: E402 diff --git a/.claude/skills/check-content-consistency/scripts/test_run_writer_semantic_detector.py b/tests/skills/check-content-consistency/test_run_writer_semantic_detector.py similarity index 99% rename from .claude/skills/check-content-consistency/scripts/test_run_writer_semantic_detector.py rename to tests/skills/check-content-consistency/test_run_writer_semantic_detector.py index 9f45b2d..0c70a40 100644 --- a/.claude/skills/check-content-consistency/scripts/test_run_writer_semantic_detector.py +++ b/tests/skills/check-content-consistency/test_run_writer_semantic_detector.py @@ -10,7 +10,10 @@ import sys import unittest from typing import Any, Mapping, Sequence -sys.path.insert(0, str(pathlib.Path(__file__).resolve().parent)) +PROJECT_ROOT = pathlib.Path(__file__).resolve().parents[3] +SCRIPT_DIR = PROJECT_ROOT / ".claude" / "skills" / "check-content-consistency" / "scripts" +if str(SCRIPT_DIR) not in sys.path: + sys.path.insert(0, str(SCRIPT_DIR)) from run_writer_semantic_detector import ( # noqa: E402 SEMANTIC_DETECTOR_REPORT_JSON_SCHEMA, diff --git a/.claude/skills/clean-book-text/scripts/test_clean_detect_offline.py b/tests/skills/clean-book-text/test_clean_detect_offline.py similarity index 98% rename from .claude/skills/clean-book-text/scripts/test_clean_detect_offline.py rename to tests/skills/clean-book-text/test_clean_detect_offline.py index f27ff09..8f4f561 100644 --- a/.claude/skills/clean-book-text/scripts/test_clean_detect_offline.py +++ b/tests/skills/clean-book-text/test_clean_detect_offline.py @@ -16,7 +16,9 @@ from unittest.mock import patch from click.testing import CliRunner -sys.path.insert(0, str(pathlib.Path(__file__).resolve().parent)) +PROJECT_ROOT = pathlib.Path(__file__).resolve().parents[3] +SCRIPT_DIR = PROJECT_ROOT / ".claude" / "skills" / "clean-book-text" / "scripts" +sys.path.insert(0, str(SCRIPT_DIR)) import clean_detect # noqa: E402 diff --git a/.claude/skills/decide-candidate/scripts/test_confirm_knowledge_offline.py b/tests/skills/decide-candidate/test_confirm_knowledge_offline.py similarity index 94% rename from .claude/skills/decide-candidate/scripts/test_confirm_knowledge_offline.py rename to tests/skills/decide-candidate/test_confirm_knowledge_offline.py index 28749ec..81539ea 100644 --- a/.claude/skills/decide-candidate/scripts/test_confirm_knowledge_offline.py +++ b/tests/skills/decide-candidate/test_confirm_knowledge_offline.py @@ -5,7 +5,8 @@ import sys import unittest -SCRIPT_DIR = pathlib.Path(__file__).resolve().parent +PROJECT_ROOT = pathlib.Path(__file__).resolve().parents[3] +SCRIPT_DIR = PROJECT_ROOT / ".claude" / "skills" / "decide-candidate" / "scripts" sys.path.insert(0, str(SCRIPT_DIR)) import confirm_knowledge as confirm # noqa: E402 diff --git a/.claude/skills/decide-candidate/scripts/test_fact_delta.py b/tests/skills/decide-candidate/test_fact_delta.py similarity index 96% rename from .claude/skills/decide-candidate/scripts/test_fact_delta.py rename to tests/skills/decide-candidate/test_fact_delta.py index 849abb2..51f867d 100644 --- a/.claude/skills/decide-candidate/scripts/test_fact_delta.py +++ b/tests/skills/decide-candidate/test_fact_delta.py @@ -4,7 +4,7 @@ 模型只能提类型化增量,证据引文必须真实出现在候选正文中(不得编造证据); 字段闭集、类型闭集、payload 合同、重复 ID 一律失败关闭。 -跑法:.venv/bin/python .claude/skills/decide-candidate/scripts/test_fact_delta.py +跑法:.venv/bin/python tests/skills/decide-candidate/test_fact_delta.py """ from __future__ import annotations @@ -14,8 +14,9 @@ import pathlib import sys import unittest -SCRIPT_DIR = pathlib.Path(__file__).resolve().parent -SKILLS_DIR = SCRIPT_DIR.parents[1] +PROJECT_ROOT = pathlib.Path(__file__).resolve().parents[3] +SKILLS_DIR = PROJECT_ROOT / ".claude" / "skills" +SCRIPT_DIR = SKILLS_DIR / "decide-candidate" / "scripts" for path in (SCRIPT_DIR, SKILLS_DIR / "assemble-context" / "scripts"): if str(path) not in sys.path: sys.path.insert(0, str(path)) diff --git a/.claude/skills/decide-candidate/scripts/test_fact_delta_db.py b/tests/skills/decide-candidate/test_fact_delta_db.py similarity index 97% rename from .claude/skills/decide-candidate/scripts/test_fact_delta_db.py rename to tests/skills/decide-candidate/test_fact_delta_db.py index 159151a..e45847c 100644 --- a/.claude/skills/decide-candidate/scripts/test_fact_delta_db.py +++ b/tests/skills/decide-candidate/test_fact_delta_db.py @@ -11,7 +11,7 @@ 测试数据 unittest-delta- 前缀隔离;清理时短暂禁用账本防删触发器(try/finally 恢复)。 跑法(需 Tailscale 内网可达 muse-example): -.venv/bin/python .claude/skills/decide-candidate/scripts/test_fact_delta_db.py +.venv/bin/python tests/skills/decide-candidate/test_fact_delta_db.py """ from __future__ import annotations @@ -21,8 +21,11 @@ import pathlib import sys import uuid -SCRIPT_DIR = pathlib.Path(__file__).resolve().parent -SKILLS_DIR = SCRIPT_DIR.parents[1] +from psycopg.errors import RaiseException + +PROJECT_ROOT = pathlib.Path(__file__).resolve().parents[3] +SKILLS_DIR = PROJECT_ROOT / ".claude" / "skills" +SCRIPT_DIR = SKILLS_DIR / "decide-candidate" / "scripts" for path in (SCRIPT_DIR, SKILLS_DIR / "access-database" / "scripts"): if str(path) not in sys.path: sys.path.insert(0, str(path)) @@ -266,7 +269,7 @@ def test_ledger_append_only(work_id: int) -> None: raise AssertionError("账本必须 append-only") except AssertionError: raise - except Exception: + except RaiseException: pass try: with connect() as conn: @@ -275,7 +278,7 @@ def test_ledger_append_only(work_id: int) -> None: raise AssertionError("账本必须 append-only") except AssertionError: raise - except Exception: + except RaiseException: pass diff --git a/.claude/skills/decide-candidate/scripts/test_projection_db.py b/tests/skills/decide-candidate/test_projection_db.py similarity index 96% rename from .claude/skills/decide-candidate/scripts/test_projection_db.py rename to tests/skills/decide-candidate/test_projection_db.py index f2cb4cc..3144230 100644 --- a/.claude/skills/decide-candidate/scripts/test_projection_db.py +++ b/tests/skills/decide-candidate/test_projection_db.py @@ -10,7 +10,7 @@ 测试数据 unittest-proj- 前缀隔离,结束物理清理(本表可变,直接 DELETE)。 跑法(需 Tailscale 内网可达 muse-example): -.venv/bin/python .claude/skills/decide-candidate/scripts/test_projection_db.py +.venv/bin/python tests/skills/decide-candidate/test_projection_db.py """ from __future__ import annotations @@ -20,8 +20,11 @@ import pathlib import sys import uuid -SCRIPT_DIR = pathlib.Path(__file__).resolve().parent -SKILLS_DIR = SCRIPT_DIR.parents[1] +from psycopg.errors import RaiseException + +PROJECT_ROOT = pathlib.Path(__file__).resolve().parents[3] +SKILLS_DIR = PROJECT_ROOT / ".claude" / "skills" +SCRIPT_DIR = SKILLS_DIR / "decide-candidate" / "scripts" for path in (SCRIPT_DIR, SKILLS_DIR / "access-database" / "scripts"): if str(path) not in sys.path: sys.path.insert(0, str(path)) @@ -175,7 +178,7 @@ def test_failed_never_masquerades_completed(work_id: int) -> None: raise AssertionError("触发器必须拒绝 failed→completed") except AssertionError: raise - except Exception: + except RaiseException: pass # 显式 retry 才能回 pending,且 attempt+1 retried = retry_projection(pending[0]) @@ -202,7 +205,7 @@ def test_stale_projections_cannot_report_outcomes(work_id: int) -> None: raise AssertionError("stale→completed 必须被拒绝") except AssertionError: raise - except Exception: + except RaiseException: pass try: with connect() as conn: @@ -212,7 +215,7 @@ def test_stale_projections_cannot_report_outcomes(work_id: int) -> None: raise AssertionError("stale→failed 必须被拒绝") except AssertionError: raise - except Exception: + except RaiseException: pass # stale 走 retry 恢复:attempt+1 回 pending,随后可以正常完成 retried = retry_projection(stale_row[0]) diff --git a/.claude/skills/decide-candidate/scripts/test_write_canonical_db.py b/tests/skills/decide-candidate/test_write_canonical_db.py similarity index 98% rename from .claude/skills/decide-candidate/scripts/test_write_canonical_db.py rename to tests/skills/decide-candidate/test_write_canonical_db.py index 7e0291b..4eb6a5e 100644 --- a/.claude/skills/decide-candidate/scripts/test_write_canonical_db.py +++ b/tests/skills/decide-candidate/test_write_canonical_db.py @@ -12,7 +12,7 @@ 清理时短暂禁用其防删触发器(try/finally 保证恢复)。 跑法(需 Tailscale 内网可达 muse-example): -.venv/bin/python .claude/skills/decide-candidate/scripts/test_write_canonical_db.py +.venv/bin/python tests/skills/decide-candidate/test_write_canonical_db.py """ from __future__ import annotations @@ -24,8 +24,9 @@ import uuid from typing import Any from unittest import mock -SCRIPT_DIR = pathlib.Path(__file__).resolve().parent -SKILLS_DIR = SCRIPT_DIR.parents[1] +PROJECT_ROOT = pathlib.Path(__file__).resolve().parents[3] +SKILLS_DIR = PROJECT_ROOT / ".claude" / "skills" +SCRIPT_DIR = SKILLS_DIR / "decide-candidate" / "scripts" for path in (SCRIPT_DIR, SKILLS_DIR / "access-database" / "scripts"): if str(path) not in sys.path: sys.path.insert(0, str(path)) diff --git a/.claude/skills/decide-candidate/scripts/test_writer_acceptance.py b/tests/skills/decide-candidate/test_writer_acceptance.py similarity index 99% rename from .claude/skills/decide-candidate/scripts/test_writer_acceptance.py rename to tests/skills/decide-candidate/test_writer_acceptance.py index 8f9b758..a5f7c81 100644 --- a/.claude/skills/decide-candidate/scripts/test_writer_acceptance.py +++ b/tests/skills/decide-candidate/test_writer_acceptance.py @@ -14,8 +14,9 @@ import pathlib import sys import unittest -SCRIPT_DIR = pathlib.Path(__file__).resolve().parent -SKILLS_DIR = SCRIPT_DIR.parents[1] +PROJECT_ROOT = pathlib.Path(__file__).resolve().parents[3] +SKILLS_DIR = PROJECT_ROOT / ".claude" / "skills" +SCRIPT_DIR = SKILLS_DIR / "decide-candidate" / "scripts" CONTINUATION_DIR = SKILLS_DIR / "write-next-chapter" / "scripts" READ_CONTEXT_DIR = SKILLS_DIR / "assemble-context" / "scripts" for path in (SCRIPT_DIR, CONTINUATION_DIR, READ_CONTEXT_DIR): diff --git a/.claude/skills/deconstruct-book/scripts/test_parse_llm_offline.py b/tests/skills/deconstruct-book/test_parse_llm_offline.py similarity index 98% rename from .claude/skills/deconstruct-book/scripts/test_parse_llm_offline.py rename to tests/skills/deconstruct-book/test_parse_llm_offline.py index a942b55..f100bd4 100644 --- a/.claude/skills/deconstruct-book/scripts/test_parse_llm_offline.py +++ b/tests/skills/deconstruct-book/test_parse_llm_offline.py @@ -11,7 +11,9 @@ import unittest from unittest.mock import Mock, patch -sys.path.insert(0, str(pathlib.Path(__file__).resolve().parent)) +PROJECT_ROOT = pathlib.Path(__file__).resolve().parents[3] +SCRIPT_DIR = PROJECT_ROOT / ".claude" / "skills" / "deconstruct-book" / "scripts" +sys.path.insert(0, str(SCRIPT_DIR)) import parse_llm as pll # noqa: E402 diff --git a/.claude/skills/deconstruct-book/scripts/test_parse_outline_offline.py b/tests/skills/deconstruct-book/test_parse_outline_offline.py similarity index 95% rename from .claude/skills/deconstruct-book/scripts/test_parse_outline_offline.py rename to tests/skills/deconstruct-book/test_parse_outline_offline.py index cb8fada..5d8419f 100644 --- a/.claude/skills/deconstruct-book/scripts/test_parse_outline_offline.py +++ b/tests/skills/deconstruct-book/test_parse_outline_offline.py @@ -6,8 +6,9 @@ import sys import click -HERE = pathlib.Path(__file__).resolve().parent -sys.path.insert(0, str(HERE)) +PROJECT_ROOT = pathlib.Path(__file__).resolve().parents[3] +SCRIPT_DIR = PROJECT_ROOT / ".claude" / "skills" / "deconstruct-book" / "scripts" +sys.path.insert(0, str(SCRIPT_DIR)) import parse_outline as po # noqa: E402 diff --git a/.claude/skills/diagnose-ai-flavor/scripts/test_diagnose_ai_flavor.py b/tests/skills/diagnose-ai-flavor/test_diagnose_ai_flavor.py similarity index 96% rename from .claude/skills/diagnose-ai-flavor/scripts/test_diagnose_ai_flavor.py rename to tests/skills/diagnose-ai-flavor/test_diagnose_ai_flavor.py index 979328b..27ac736 100644 --- a/.claude/skills/diagnose-ai-flavor/scripts/test_diagnose_ai_flavor.py +++ b/tests/skills/diagnose-ai-flavor/test_diagnose_ai_flavor.py @@ -7,7 +7,9 @@ import tempfile import unittest from unittest.mock import patch -sys.path.insert(0, str(pathlib.Path(__file__).resolve().parent)) +PROJECT_ROOT = pathlib.Path(__file__).resolve().parents[3] +SCRIPT_DIR = PROJECT_ROOT / ".claude" / "skills" / "diagnose-ai-flavor" / "scripts" +sys.path.insert(0, str(SCRIPT_DIR)) import diagnose_ai_flavor as diag # noqa: E402 diff --git a/.claude/skills/embed-knowledge/scripts/test_embed_drafts_offline.py b/tests/skills/embed-knowledge/test_embed_drafts_offline.py similarity index 99% rename from .claude/skills/embed-knowledge/scripts/test_embed_drafts_offline.py rename to tests/skills/embed-knowledge/test_embed_drafts_offline.py index 4ec0bd8..f711218 100644 --- a/.claude/skills/embed-knowledge/scripts/test_embed_drafts_offline.py +++ b/tests/skills/embed-knowledge/test_embed_drafts_offline.py @@ -11,7 +11,8 @@ from unittest.mock import MagicMock, Mock, patch from click.testing import CliRunner -SCRIPT_DIR = pathlib.Path(__file__).resolve().parent +PROJECT_ROOT = pathlib.Path(__file__).resolve().parents[3] +SCRIPT_DIR = PROJECT_ROOT / ".claude" / "skills" / "embed-knowledge" / "scripts" sys.path.insert(0, str(SCRIPT_DIR)) import embed_drafts as embed # noqa: E402 diff --git a/.claude/skills/establish-voice-baseline/scripts/test_establish_voice_baseline.py b/tests/skills/establish-voice-baseline/test_establish_voice_baseline.py similarity index 96% rename from .claude/skills/establish-voice-baseline/scripts/test_establish_voice_baseline.py rename to tests/skills/establish-voice-baseline/test_establish_voice_baseline.py index bcff26a..ec3160b 100644 --- a/.claude/skills/establish-voice-baseline/scripts/test_establish_voice_baseline.py +++ b/tests/skills/establish-voice-baseline/test_establish_voice_baseline.py @@ -6,7 +6,9 @@ import tempfile import unittest from unittest.mock import patch -sys.path.insert(0, str(pathlib.Path(__file__).resolve().parent)) +PROJECT_ROOT = pathlib.Path(__file__).resolve().parents[3] +SCRIPT_DIR = PROJECT_ROOT / ".claude" / "skills" / "establish-voice-baseline" / "scripts" +sys.path.insert(0, str(SCRIPT_DIR)) import establish_voice_baseline as base # noqa: E402 diff --git a/.claude/skills/evaluate-frozen-replay/scripts/test_fine_outline_detector.py b/tests/skills/evaluate-frozen-replay/test_fine_outline_detector.py similarity index 93% rename from .claude/skills/evaluate-frozen-replay/scripts/test_fine_outline_detector.py rename to tests/skills/evaluate-frozen-replay/test_fine_outline_detector.py index 83ca4c7..45bc458 100644 --- a/.claude/skills/evaluate-frozen-replay/scripts/test_fine_outline_detector.py +++ b/tests/skills/evaluate-frozen-replay/test_fine_outline_detector.py @@ -7,7 +7,11 @@ import pathlib import sys import unittest -sys.path.insert(0, str(pathlib.Path(__file__).resolve().parent)) +PROJECT_ROOT = pathlib.Path(__file__).resolve().parents[3] +SCRIPT_DIR = PROJECT_ROOT / ".claude" / "skills" / "evaluate-frozen-replay" / "scripts" +if str(SCRIPT_DIR) not in sys.path: + sys.path.insert(0, str(SCRIPT_DIR)) + from fine_outline_detector import validate_detector_report # noqa: E402 @@ -76,7 +80,9 @@ class FineOutlineDetectorTest(unittest.TestCase): self.assertTrue(validate_detector_report(report, "blind-1")["ok"]) skill_path = ( - pathlib.Path(__file__).resolve().parents[2] + PROJECT_ROOT + / ".claude" + / "skills" / "check-content-consistency" / "SKILL.md" ) diff --git a/.claude/skills/evaluate-frozen-replay/scripts/test_gate_input_builder.py b/tests/skills/evaluate-frozen-replay/test_gate_input_builder.py similarity index 97% rename from .claude/skills/evaluate-frozen-replay/scripts/test_gate_input_builder.py rename to tests/skills/evaluate-frozen-replay/test_gate_input_builder.py index e9721b9..46d0b93 100644 --- a/.claude/skills/evaluate-frozen-replay/scripts/test_gate_input_builder.py +++ b/tests/skills/evaluate-frozen-replay/test_gate_input_builder.py @@ -8,9 +8,12 @@ import pathlib import sys import unittest -SCRIPT_DIR = pathlib.Path(__file__).resolve().parent -QUALITY_GATE_DIR = SCRIPT_DIR.parents[1] / "score-content-quality" / "scripts" -for _import_dir in (SCRIPT_DIR, QUALITY_GATE_DIR): +PROJECT_ROOT = pathlib.Path(__file__).resolve().parents[3] +SKILLS_DIR = PROJECT_ROOT / ".claude" / "skills" +SCRIPT_DIR = SKILLS_DIR / "evaluate-frozen-replay" / "scripts" +TEST_DIR = PROJECT_ROOT / "tests" / "skills" / "evaluate-frozen-replay" +QUALITY_GATE_DIR = SKILLS_DIR / "score-content-quality" / "scripts" +for _import_dir in (SCRIPT_DIR, TEST_DIR, QUALITY_GATE_DIR): if str(_import_dir) not in sys.path: sys.path.insert(0, str(_import_dir)) diff --git a/.claude/skills/evaluate-frozen-replay/scripts/test_load_writer_reference_work.py b/tests/skills/evaluate-frozen-replay/test_load_writer_reference_work.py similarity index 98% rename from .claude/skills/evaluate-frozen-replay/scripts/test_load_writer_reference_work.py rename to tests/skills/evaluate-frozen-replay/test_load_writer_reference_work.py index 873a194..5436c4c 100644 --- a/.claude/skills/evaluate-frozen-replay/scripts/test_load_writer_reference_work.py +++ b/tests/skills/evaluate-frozen-replay/test_load_writer_reference_work.py @@ -14,12 +14,24 @@ import unittest from datetime import datetime, timedelta, timezone from unittest.mock import patch -SCRIPT_DIR = pathlib.Path(__file__).resolve().parent -READ_CONTEXT_SCRIPTS = SCRIPT_DIR.parents[1] / "assemble-context" / "scripts" -DETECT_SCRIPTS = SCRIPT_DIR.parents[1] / "check-content-consistency" / "scripts" -CONTINUATION_SCRIPTS = SCRIPT_DIR.parents[1] / "write-next-chapter" / "scripts" -for _path in (SCRIPT_DIR, READ_CONTEXT_SCRIPTS, DETECT_SCRIPTS, CONTINUATION_SCRIPTS): - sys.path.insert(0, str(_path)) +PROJECT_ROOT = pathlib.Path(__file__).resolve().parents[3] +SKILLS_DIR = PROJECT_ROOT / ".claude" / "skills" +SCRIPT_DIR = SKILLS_DIR / "evaluate-frozen-replay" / "scripts" +TEST_DIR = PROJECT_ROOT / "tests" / "skills" / "evaluate-frozen-replay" +READ_CONTEXT_SCRIPTS = SKILLS_DIR / "assemble-context" / "scripts" +DETECT_SCRIPTS = SKILLS_DIR / "check-content-consistency" / "scripts" +CONTINUATION_SCRIPTS = SKILLS_DIR / "write-next-chapter" / "scripts" +CONTINUATION_TESTS = PROJECT_ROOT / "tests" / "skills" / "write-next-chapter" +for _path in ( + SCRIPT_DIR, + TEST_DIR, + READ_CONTEXT_SCRIPTS, + DETECT_SCRIPTS, + CONTINUATION_SCRIPTS, + CONTINUATION_TESTS, +): + if str(_path) not in sys.path: + sys.path.insert(0, str(_path)) import load_writer_reference_work as loader # noqa: E402 import run_writer_replay as replay_module # noqa: E402 diff --git a/.claude/skills/evaluate-frozen-replay/scripts/test_pattern_reference_injection.py b/tests/skills/evaluate-frozen-replay/test_pattern_reference_injection.py similarity index 98% rename from .claude/skills/evaluate-frozen-replay/scripts/test_pattern_reference_injection.py rename to tests/skills/evaluate-frozen-replay/test_pattern_reference_injection.py index 607cf35..17fcebb 100644 --- a/.claude/skills/evaluate-frozen-replay/scripts/test_pattern_reference_injection.py +++ b/tests/skills/evaluate-frozen-replay/test_pattern_reference_injection.py @@ -13,8 +13,11 @@ import pathlib import sys import unittest -SCRIPT_DIR = pathlib.Path(__file__).resolve().parent -SKILLS_DIR = SCRIPT_DIR.parents[1] +PROJECT_ROOT = pathlib.Path(__file__).resolve().parents[3] +SKILLS_DIR = PROJECT_ROOT / ".claude" / "skills" +SCRIPT_DIR = SKILLS_DIR / "evaluate-frozen-replay" / "scripts" +TEST_DIR = PROJECT_ROOT / "tests" / "skills" / "evaluate-frozen-replay" +ASSEMBLE_CONTEXT_TESTS = PROJECT_ROOT / "tests" / "skills" / "assemble-context" READ_CONTEXT_SCRIPTS = SKILLS_DIR / "assemble-context" / "scripts" # 端到端链路测试要导入回放包(run_writer_replay.sample),其依赖执行、证据与 # 评分三个 Skill 的 scripts 目录,路径口径与 test_run_writer_replay 保持一致。 @@ -23,12 +26,15 @@ EVIDENCE_SCRIPTS = SKILLS_DIR / "record-run-evidence" / "scripts" QUALITY_GATE_SCRIPTS = SKILLS_DIR / "score-content-quality" / "scripts" for _path in ( SCRIPT_DIR, + TEST_DIR, + ASSEMBLE_CONTEXT_TESTS, READ_CONTEXT_SCRIPTS, EXECUTION_SCRIPTS, EVIDENCE_SCRIPTS, QUALITY_GATE_SCRIPTS, ): - sys.path.insert(0, str(_path)) + if str(_path) not in sys.path: + sys.path.insert(0, str(_path)) import load_writer_reference_work as loader # noqa: E402 from load_writer_reference_work import ( # noqa: E402 diff --git a/.claude/skills/evaluate-frozen-replay/scripts/test_refresh_runtime_probe.py b/tests/skills/evaluate-frozen-replay/test_refresh_runtime_probe.py similarity index 98% rename from .claude/skills/evaluate-frozen-replay/scripts/test_refresh_runtime_probe.py rename to tests/skills/evaluate-frozen-replay/test_refresh_runtime_probe.py index 3dd010d..38a272c 100644 --- a/.claude/skills/evaluate-frozen-replay/scripts/test_refresh_runtime_probe.py +++ b/tests/skills/evaluate-frozen-replay/test_refresh_runtime_probe.py @@ -21,15 +21,14 @@ import unittest from typing import Any, Mapping from unittest import mock -SCRIPT_DIR = pathlib.Path(__file__).resolve().parent -EXECUTION_DIR = SCRIPT_DIR.parents[1] / "execute-claude-task" / "scripts" -QUALITY_GATE_DIR = SCRIPT_DIR.parents[1] / "score-content-quality" / "scripts" -if str(SCRIPT_DIR) not in sys.path: - sys.path.insert(0, str(SCRIPT_DIR)) -if str(EXECUTION_DIR) not in sys.path: - sys.path.insert(0, str(EXECUTION_DIR)) -if str(QUALITY_GATE_DIR) not in sys.path: - sys.path.insert(0, str(QUALITY_GATE_DIR)) +PROJECT_ROOT = pathlib.Path(__file__).resolve().parents[3] +SKILLS_DIR = PROJECT_ROOT / ".claude" / "skills" +SCRIPT_DIR = SKILLS_DIR / "evaluate-frozen-replay" / "scripts" +EXECUTION_DIR = SKILLS_DIR / "execute-claude-task" / "scripts" +QUALITY_GATE_DIR = SKILLS_DIR / "score-content-quality" / "scripts" +for _import_dir in (SCRIPT_DIR, EXECUTION_DIR, QUALITY_GATE_DIR): + if str(_import_dir) not in sys.path: + sys.path.insert(0, str(_import_dir)) import refresh_runtime_probe as refresh_module # noqa: E402 from refresh_runtime_probe import ( # noqa: E402 diff --git a/.claude/skills/evaluate-frozen-replay/scripts/test_run_replay.py b/tests/skills/evaluate-frozen-replay/test_run_replay.py similarity index 98% rename from .claude/skills/evaluate-frozen-replay/scripts/test_run_replay.py rename to tests/skills/evaluate-frozen-replay/test_run_replay.py index d08f207..50637ea 100644 --- a/.claude/skills/evaluate-frozen-replay/scripts/test_run_replay.py +++ b/tests/skills/evaluate-frozen-replay/test_run_replay.py @@ -11,10 +11,14 @@ import tempfile import unittest from unittest.mock import patch -SCRIPT_DIR = pathlib.Path(__file__).resolve().parent -QUALITY_GATE_DIR = SCRIPT_DIR.parents[1] / "score-content-quality" / "scripts" -sys.path.insert(0, str(SCRIPT_DIR)) -sys.path.insert(0, str(QUALITY_GATE_DIR)) +PROJECT_ROOT = pathlib.Path(__file__).resolve().parents[3] +SKILLS_DIR = PROJECT_ROOT / ".claude" / "skills" +SCRIPT_DIR = SKILLS_DIR / "evaluate-frozen-replay" / "scripts" +TEST_DIR = PROJECT_ROOT / "tests" / "skills" / "evaluate-frozen-replay" +QUALITY_GATE_DIR = SKILLS_DIR / "score-content-quality" / "scripts" +for _import_dir in (SCRIPT_DIR, TEST_DIR, QUALITY_GATE_DIR): + if str(_import_dir) not in sys.path: + sys.path.insert(0, str(_import_dir)) from run_replay import _parse_args, _planner_prompt, run_replay as _run_replay # noqa: E402 from test_writer_gate import build_input, source_bundle # noqa: E402 from writer_gate import issue_gate_report_and_receipt # noqa: E402 diff --git a/.claude/skills/evaluate-frozen-replay/scripts/test_run_writer_replay.py b/tests/skills/evaluate-frozen-replay/test_run_writer_replay.py similarity index 99% rename from .claude/skills/evaluate-frozen-replay/scripts/test_run_writer_replay.py rename to tests/skills/evaluate-frozen-replay/test_run_writer_replay.py index 43fb931..3699714 100644 --- a/.claude/skills/evaluate-frozen-replay/scripts/test_run_writer_replay.py +++ b/tests/skills/evaluate-frozen-replay/test_run_writer_replay.py @@ -19,8 +19,10 @@ from dataclasses import replace from decimal import Decimal from datetime import datetime, timedelta, timezone -SCRIPT_DIR = pathlib.Path(__file__).resolve().parent -SKILLS_DIR = SCRIPT_DIR.parents[1] +TEST_DIR = pathlib.Path(__file__).resolve().parent +PROJECT_ROOT = pathlib.Path(__file__).resolve().parents[3] +SKILLS_DIR = PROJECT_ROOT / ".claude" / "skills" +SCRIPT_DIR = SKILLS_DIR / "evaluate-frozen-replay" / "scripts" READ_CONTEXT_DIR = SKILLS_DIR / "assemble-context" / "scripts" EXECUTION_DIR = SKILLS_DIR / "execute-claude-task" / "scripts" EVIDENCE_DIR = SKILLS_DIR / "record-run-evidence" / "scripts" @@ -32,7 +34,8 @@ for import_path in ( EVIDENCE_DIR, QUALITY_GATE_DIR, ): - sys.path.insert(0, str(import_path)) + if str(import_path) not in sys.path: + sys.path.insert(0, str(import_path)) import run_writer_replay as replay_module # noqa: E402 import run_writer_replay.execute as execute_module # noqa: E402 diff --git a/.claude/skills/evaluate-frozen-replay/scripts/test_writer_eval_preregister.py b/tests/skills/evaluate-frozen-replay/test_writer_eval_preregister.py similarity index 94% rename from .claude/skills/evaluate-frozen-replay/scripts/test_writer_eval_preregister.py rename to tests/skills/evaluate-frozen-replay/test_writer_eval_preregister.py index 3c5936a..8e4c285 100644 --- a/.claude/skills/evaluate-frozen-replay/scripts/test_writer_eval_preregister.py +++ b/tests/skills/evaluate-frozen-replay/test_writer_eval_preregister.py @@ -8,7 +8,10 @@ import pathlib import sys import unittest -sys.path.insert(0, str(pathlib.Path(__file__).resolve().parent)) +PROJECT_ROOT = pathlib.Path(__file__).resolve().parents[3] +SCRIPT_DIR = PROJECT_ROOT / ".claude" / "skills" / "evaluate-frozen-replay" / "scripts" +if str(SCRIPT_DIR) not in sys.path: + sys.path.insert(0, str(SCRIPT_DIR)) from writer_eval_preregister import ( # noqa: E402 PreregistrationError, diff --git a/.claude/skills/evaluate-frozen-replay/scripts/test_writer_gate.py b/tests/skills/evaluate-frozen-replay/test_writer_gate.py similarity index 98% rename from .claude/skills/evaluate-frozen-replay/scripts/test_writer_gate.py rename to tests/skills/evaluate-frozen-replay/test_writer_gate.py index 7e56d70..be03f41 100644 --- a/.claude/skills/evaluate-frozen-replay/scripts/test_writer_gate.py +++ b/tests/skills/evaluate-frozen-replay/test_writer_gate.py @@ -10,10 +10,13 @@ import sys import tempfile import unittest -SCRIPT_DIR = pathlib.Path(__file__).resolve().parent -QUALITY_GATE_DIR = SCRIPT_DIR.parents[1] / "score-content-quality" / "scripts" -EVIDENCE_DIR = SCRIPT_DIR.parents[1] / "record-run-evidence" / "scripts" -for _import_dir in (SCRIPT_DIR, QUALITY_GATE_DIR, EVIDENCE_DIR): +PROJECT_ROOT = pathlib.Path(__file__).resolve().parents[3] +SKILLS_DIR = PROJECT_ROOT / ".claude" / "skills" +SCRIPT_DIR = SKILLS_DIR / "evaluate-frozen-replay" / "scripts" +TEST_DIR = PROJECT_ROOT / "tests" / "skills" / "evaluate-frozen-replay" +QUALITY_GATE_DIR = SKILLS_DIR / "score-content-quality" / "scripts" +EVIDENCE_DIR = SKILLS_DIR / "record-run-evidence" / "scripts" +for _import_dir in (SCRIPT_DIR, TEST_DIR, QUALITY_GATE_DIR, EVIDENCE_DIR): if str(_import_dir) not in sys.path: sys.path.insert(0, str(_import_dir)) diff --git a/.claude/skills/execute-claude-task/scripts/test_claude_runtime.py b/tests/skills/execute-claude-task/test_claude_runtime.py similarity index 99% rename from .claude/skills/execute-claude-task/scripts/test_claude_runtime.py rename to tests/skills/execute-claude-task/test_claude_runtime.py index 9ffeed0..6989b88 100644 --- a/.claude/skills/execute-claude-task/scripts/test_claude_runtime.py +++ b/tests/skills/execute-claude-task/test_claude_runtime.py @@ -17,7 +17,8 @@ import unittest from unittest import mock from decimal import Decimal -SCRIPT_DIR = pathlib.Path(__file__).resolve().parent +PROJECT_ROOT = pathlib.Path(__file__).resolve().parents[3] +SCRIPT_DIR = PROJECT_ROOT / ".claude" / "skills" / "execute-claude-task" / "scripts" sys.path.insert(0, str(SCRIPT_DIR)) from claude_runtime import ( # noqa: E402 diff --git a/.claude/skills/extract-chapter-knowledge/scripts/test_extract_knowledge_offline.py b/tests/skills/extract-chapter-knowledge/test_extract_knowledge_offline.py similarity index 92% rename from .claude/skills/extract-chapter-knowledge/scripts/test_extract_knowledge_offline.py rename to tests/skills/extract-chapter-knowledge/test_extract_knowledge_offline.py index a25f5e9..6ca3e47 100644 --- a/.claude/skills/extract-chapter-knowledge/scripts/test_extract_knowledge_offline.py +++ b/tests/skills/extract-chapter-knowledge/test_extract_knowledge_offline.py @@ -3,7 +3,9 @@ import pathlib import sys -sys.path.insert(0, str(pathlib.Path(__file__).resolve().parent)) +PROJECT_ROOT = pathlib.Path(__file__).resolve().parents[3] +SCRIPT_DIR = PROJECT_ROOT / ".claude" / "skills" / "extract-chapter-knowledge" / "scripts" +sys.path.insert(0, str(SCRIPT_DIR)) from extract_knowledge import ExtractionContractError, normalize_extraction, salvage_extraction # noqa: E402 diff --git a/.claude/skills/extract-work-knowledge/scripts/test_parse_upgrade_offline.py b/tests/skills/extract-work-knowledge/test_parse_upgrade_offline.py similarity index 91% rename from .claude/skills/extract-work-knowledge/scripts/test_parse_upgrade_offline.py rename to tests/skills/extract-work-knowledge/test_parse_upgrade_offline.py index 523ac00..d1b9f7d 100644 --- a/.claude/skills/extract-work-knowledge/scripts/test_parse_upgrade_offline.py +++ b/tests/skills/extract-work-knowledge/test_parse_upgrade_offline.py @@ -13,13 +13,12 @@ 本测试**不触碰真实服务**——质量修复命令只通过 fake DB 验证 preview、CAS 和原子写入, 不在离线自测中执行真实 work/window。 -跑法:仓根 `.venv/bin/python .claude/skills/extract-work-knowledge/scripts/test_parse_upgrade_offline.py` +跑法:仓库根目录 `.venv/bin/python tests/skills/extract-work-knowledge/test_parse_upgrade_offline.py` """ import json import hashlib import inspect from io import StringIO -import os import pathlib import sys from contextlib import nullcontext @@ -29,8 +28,10 @@ from unittest.mock import patch from click.testing import CliRunner -# 与 upgrade 同目录:直接 import 触发其 sys.path 装配(含 parse-book/embed/llm scripts),随后可导 embed_drafts -sys.path.insert(0, str(pathlib.Path(__file__).resolve().parent)) +# 从生产 scripts 导入 upgrade;upgrade 自己装配 parse-book/embed/llm 的生产模块路径。 +PROJECT_ROOT = pathlib.Path(__file__).resolve().parents[3] +SCRIPT_DIR = PROJECT_ROOT / ".claude" / "skills" / "extract-work-knowledge" / "scripts" +sys.path.insert(0, str(SCRIPT_DIR)) import upgrade as pu # noqa: E402 import embed_drafts # noqa: E402 (型修正验证在 embed skill 本体) @@ -4724,17 +4725,6 @@ def test_recovery_and_redo_rejection_contracts(): check("redo-run首行机械拒绝", False, detail="_run 接受了 redo") -def test_real_pg_rollback_smoke_entry(): - """真实 PG 入口默认不运行;显式开启后验证 public 窗状态写入可回滚。""" - - if os.getenv("MUSE_REAL_PG_ROLLBACK_SMOKE") != "1": - check("real-pg-rollback-smoke默认跳过", True) - return - work_id = int(os.getenv("MUSE_REAL_PG_ROLLBACK_SMOKE_WORK_ID", "8")) - pu.real_pg_rollback_smoke(work_id) - check("real-pg-rollback-smoke真实public窗回滚", True) - - class _LegacyRecoveryConn: """恢复命令离线夹具:用安全墓碑及异常软删验证 SQL 过滤。""" @@ -5847,538 +5837,7 @@ def test_quality_repair_rejects_alias_and_presence_duplicates(): ) -# ── presence 冗余收口(work8 十组双行:保留 MIN(id) 软删 MAX(id))── -# 离线 fixture 用字符串时间戳代替数据库驱动的 datetime:窄合同只比较同组两行是否相等, -# _sha256_json 计算摘要时统一 default=str 转写,二者行为一致。 -_PRESENCE_DEDUPE_BASE_TIME = "2026-07-01 12:00:00" - - -def _presence_dedupe_rows(): - """构造满足窄合同的 work8 presence 行 fixture。 - - 十组双行:每组 observation 不同、create_time 相同、恰好一个待删 ID 命中 DELETE_IDS、 - 保留 ID 是组内较小者;另有四个单行,证明快照会跳过不构成冗余的 key。 - """ - - rows = [] - for index, delete_id in enumerate(sorted(pu.PRESENCE_DEDUPE_DELETE_IDS)): - keep_id = 11501 + index - entity_type = "location" if index % 2 == 0 else "item" - name = f"重复实体{index}" - for row_id, observation in ((keep_id, f"观察A{index}"), - (delete_id, f"观察B{index}")): - rows.append((row_id, 8, 57 + index, 403 + index, entity_type, name, - observation, "upgrade", _PRESENCE_DEDUPE_BASE_TIME, - False, pu.TENANT)) - for index in range(4): - rows.append((11400 + index, 8, 10 + index, 20 + index, "character", - f"单体实体{index}", f"单体观察{index}", - "upgrade", _PRESENCE_DEDUPE_BASE_TIME, False, pu.TENANT)) - return sorted(rows, key=lambda row: row[0]) - - -def _presence_dedupe_state(): - return { - "rows": _presence_dedupe_rows(), - "writes": [], - "queries": [], - "commits": 0, - } - - -def _mutate_presence_row(state, row_id, column, value): - """替换指定行的单列;元组不可变,整体重建后写回。""" - - state["rows"] = [ - row[:column] + (value,) + row[column + 1:] if row[0] == row_id else row - for row in state["rows"] - ] - - -class _PresenceDedupeConn: - """presence 冗余收口 fake DB:只实现本命令的读快照与软删 SQL。 - - 其他任何域(draft/window/alias/card_state/audit/embedding)的 SQL 会落到末尾 - AssertionError,因此「零副作用」无需逐条枚举禁写语句即可离线断言。fake 只做 - 「未软删」粗过滤;列数/work/tenant/deleted 类型等窄合同行像校验正是离线断言对象, - fake 不能代为过滤。 - """ - - def __init__(self, state): - self.state = state - - def __enter__(self): - return self - - def __exit__(self, exc_type, exc, tb): - return False - - def execute(self, query, params=()): - normalized = " ".join(query.split()) - self.state.setdefault("queries", []).append(normalized) - if normalized.startswith("SET TRANSACTION") or normalized.startswith("LOCK TABLE"): - return _CardResult() - if normalized.startswith( - "SELECT id, work_id, window_no, chapter_no, entity_type, name, observation,"): - rows = [ - row for row in sorted(self.state["rows"], key=lambda item: item[0]) - if len(row) > 9 and not row[9] - ] - return _CardResult(rows=rows) - if normalized.startswith("UPDATE example_upgrade_presence SET deleted=TRUE"): - tenant_id, work_id, delete_ids = params - assert tenant_id == pu.TENANT - delete_set = set(delete_ids) - hit = [] - new_rows = [] - for row in self.state["rows"]: - if row[1] == work_id and row[0] in delete_set and not row[9]: - row = row[:9] + (True,) + row[10:] - hit.append((row[0],)) - new_rows.append(row) - self.state["rows"] = new_rows - self.state["writes"].append("UPDATE example_upgrade_presence") - return _CardResult(rows=hit) - raise AssertionError(f"presence 冗余收口 fake DB 未覆盖 SQL:{normalized}") - - def commit(self): - self.state["commits"] += 1 - - -def _presence_dedupe_cli(state, args): - """离线执行 repair-presence-duplicates:注入 fake DB 与空放同书锁。""" - - conn = _PresenceDedupeConn(state) - with patch.object(pu.psycopg, "connect", return_value=conn), \ - patch.object(pu, "upgrade_work_lock", return_value=nullcontext()): - result = CliRunner().invoke(pu.maintenance_cli, ["repair-presence-duplicates"] + args) - return result, conn - - -def _presence_dedupe_preview_sha(state): - """离线取 preview 的 confirmation_sha,作为 execute 的合法输入。""" - - result, _ = _presence_dedupe_cli(state, ["--work-id", "8", "--preview"]) - assert result.exit_code == 0, result.output - return json.loads(result.output)["confirmation_sha"] - - -def _presence_snapshot_raises(label, state, work_id=pu.PRESENCE_DEDUPE_WORK_ID): - """断言快照以 CompensationFenceConflict 拒绝,且没有发出任何写入。""" - - conn = _PresenceDedupeConn(state) - try: - pu._presence_dedupe_capture_snapshot(conn, work_id) - except pu.CompensationFenceConflict as exc: - check(f"presence-dedupe-{label}失败关闭", True, detail=str(exc)) - else: - check(f"presence-dedupe-{label}失败关闭", False, - detail="未抛 CompensationFenceConflict") - check( - f"presence-dedupe-{label}零副作用", - state["writes"] == [] and state["commits"] == 0, - ) - - -def test_presence_dedupe_actions_planner_guard(): - """planner 级收口守卫:同观察幂等留首条、不同观察失败关闭、非 presence 原样透传。""" - - presence_a = ("presence", None, 403, "location", "重复星体", "观察一") - presence_a_dup = ("presence", None, 403, "location", "重复星体", "观察一") - presence_b = ("presence", None, 404, "item", "镜面护盾", "观察二") - alias_action = ("alias", 7, "规范名", "别名", "ai") - new_action = ("new", -1, {"名称": "新实体"}, set(), []) - actions = [presence_a, alias_action, presence_a_dup, new_action, presence_b, ()] - original = list(actions) - deduped = pu._dedupe_presence_actions(actions) - check( - "presence-dedupe-planner同观察去重保留首条", - deduped == [presence_a, alias_action, new_action, presence_b, ()] - and deduped[0] is presence_a, - detail=str(deduped), - ) - check("presence-dedupe-重复应用幂等", - pu._dedupe_presence_actions(deduped) == deduped) - check("presence-dedupe-不修改输入列表", actions == original) - - conflict = [presence_a, - ("presence", None, 403, "location", "重复星体", "另一个观察")] - try: - pu._dedupe_presence_actions(conflict) - except RuntimeError as exc: - check("presence-dedupe-不同观察失败关闭", - "observation 冲突" in str(exc), detail=str(exc)) - else: - check("presence-dedupe-不同观察失败关闭", False, detail="未抛 RuntimeError") - - for label, bad, keyword in ( - ("结构缺一元", ("presence", None, 403, "location", "重复星体"), - "结构非法"), - ("结构多一元", - ("presence", None, 403, "location", "重复星体", "观察一", "extra"), - "结构非法"), - ("observation非字符串", ("presence", None, 403, "location", "重复星体", None), - "必须是字符串"), - ): - try: - pu._dedupe_presence_actions([bad]) - except RuntimeError as exc: - check(f"presence-dedupe-{label}抛错", keyword in str(exc), detail=str(exc)) - else: - check(f"presence-dedupe-{label}抛错", False, detail="未抛 RuntimeError") - - -def test_presence_dedupe_snapshot_success(): - """快照成功路径:恰好 10 组/20 行,待删集合精确等于 DELETE_IDS,保留为组内较小者。""" - - state = _presence_dedupe_state() - snapshot = pu._presence_dedupe_capture_snapshot( - _PresenceDedupeConn(state), pu.PRESENCE_DEDUPE_WORK_ID, - ) - check( - "presence-dedupe-快照组数行数精确", - len(snapshot["groups"]) == pu.PRESENCE_DEDUPE_GROUP_COUNT - and len(snapshot["rows"]) == pu.PRESENCE_DEDUPE_ROW_COUNT - and snapshot["contract"] == pu.PRESENCE_DEDUPE_CONTRACT - and snapshot["work_id"] == pu.PRESENCE_DEDUPE_WORK_ID, - ) - check( - "presence-dedupe-待删集合精确等于DELETE_IDS", - snapshot["delete_ids"] == sorted(pu.PRESENCE_DEDUPE_DELETE_IDS), - detail=str(snapshot["delete_ids"]), - ) - check( - "presence-dedupe-保留为组内较小ID", - snapshot["keep_ids"] == list(range(11501, 11511)) - and all(group["keep_id"] < group["delete_id"] for group in snapshot["groups"]) - and all(len(group["rows"]) == 2 for group in snapshot["groups"]), - detail=str(snapshot["keep_ids"]), - ) - check( - "presence-dedupe-行按id排序且排除单体行", - [row[0] for row in snapshot["rows"]] - == sorted(row[0] for row in snapshot["rows"]) - and not any(row[0] in range(11400, 11404) for row in snapshot["rows"]), - ) - - -def test_presence_dedupe_snapshot_fail_closed(): - """各类窄合同漂移都必须在写入前以 CompensationFenceConflict 逐条拒绝。""" - - # 额外冗余组:组数 11 ≠ 10(新增组复用合同内待删 ID,确保失败点落在组数校验)。 - state = _presence_dedupe_state() - state["rows"].extend(( - (11581, 8, 90, 500, "event", "额外重复组", "观察C", - "upgrade", _PRESENCE_DEDUPE_BASE_TIME, False, pu.TENANT), - (11582, 8, 90, 500, "event", "额外重复组", "观察D", - "upgrade", _PRESENCE_DEDUPE_BASE_TIME, False, pu.TENANT), - )) - _presence_snapshot_raises("额外冗余组", state) - - # 三行组不是「恰好两行」。 - state = _presence_dedupe_state() - state["rows"].append( - (11601, 8, 57, 403, "location", "重复实体0", "观察E", - "upgrade", _PRESENCE_DEDUPE_BASE_TIME, False, pu.TENANT) - ) - _presence_snapshot_raises("三行冗余组", state) - - # 缺行:一组只剩单行被跳过 → 组数 9 ≠ 10。 - state = _presence_dedupe_state() - state["rows"] = [row for row in state["rows"] - if row[0] != sorted(pu.PRESENCE_DEDUPE_DELETE_IDS)[0]] - _presence_snapshot_raises("缺行", state) - - # 同组两行 create_time 漂移。 - state = _presence_dedupe_state() - _mutate_presence_row(state, 11501, 8, "2026-07-02 09:30:00") - _presence_snapshot_raises("create_time漂移", state) - - # 待删 ID 命中 0:组内两行都不在 DELETE_IDS。 - state = _presence_dedupe_state() - _mutate_presence_row(state, 11501, 0, 12001) - _mutate_presence_row(state, 11582, 0, 12002) - _presence_snapshot_raises("待删ID无命中", state) - - # 待删 ID 命中 2:组内两行都在 DELETE_IDS。 - state = _presence_dedupe_state() - _mutate_presence_row(state, 11501, 0, 11693) - _presence_snapshot_raises("待删ID双命中", state) - - # 保留不是组内较小者:待删 ID 反而是组内较小者。 - state = _presence_dedupe_state() - _mutate_presence_row(state, 11501, 0, 11900) - _presence_snapshot_raises("保留非组内较小", state) - - # 同组两行 observation 相同(byte-exact 重复违反窄合同)。 - state = _presence_dedupe_state() - _mutate_presence_row(state, 11582, 6, "观察A0") - _presence_snapshot_raises("observation相同", state) - - # work_id ≠ 8 拒绝。 - _presence_snapshot_raises("work不符", _presence_dedupe_state(), work_id=9) - - # 窄合同行像漂移:列数不足。 - state = _presence_dedupe_state() - state["rows"] = [row[:10] if row[0] == 11501 else row for row in state["rows"]] - _presence_snapshot_raises("列数不符", state) - - # 窄合同行像漂移:work 列与查询目标不符。 - state = _presence_dedupe_state() - _mutate_presence_row(state, 11501, 1, 9) - _presence_snapshot_raises("work列不符", state) - - # 窄合同行像漂移:deleted 不是严格 False(整数 0 也必须拒绝)。 - state = _presence_dedupe_state() - _mutate_presence_row(state, 11501, 9, 0) - _presence_snapshot_raises("deleted类型不符", state) - - # 窄合同行像漂移:tenant 列不符。 - state = _presence_dedupe_state() - _mutate_presence_row(state, 11501, 10, "other-tenant") - _presence_snapshot_raises("tenant不符", state) - - -def test_presence_dedupe_confirmation_sha_binds_rows_and_ids(): - """confirmation_sha 必须确定且绑定二十行完整内容与保留/删除 ID,任一来源漂移即变化。""" - - base_snapshot = pu._presence_dedupe_capture_snapshot( - _PresenceDedupeConn(_presence_dedupe_state()), pu.PRESENCE_DEDUPE_WORK_ID, - ) - base_sha = pu._presence_dedupe_confirmation_sha(base_snapshot) - again_snapshot = pu._presence_dedupe_capture_snapshot( - _PresenceDedupeConn(_presence_dedupe_state()), pu.PRESENCE_DEDUPE_WORK_ID, - ) - check( - "presence-dedupe-confirmation_sha确定性", - base_sha == pu._presence_dedupe_confirmation_sha(again_snapshot) - and len(base_sha) == 64, - detail=base_sha, - ) - for label, row_id, column, value in ( - ("observation", 11582, 6, "被篡改的观察"), - ("保留行ID", 11501, 0, 11001), - ("creator", 11501, 7, "篡改者"), - ): - state = _presence_dedupe_state() - _mutate_presence_row(state, row_id, column, value) - drifted = pu._presence_dedupe_capture_snapshot( - _PresenceDedupeConn(state), pu.PRESENCE_DEDUPE_WORK_ID, - ) - check( - f"presence-dedupe-sha感知{label}漂移", - pu._presence_dedupe_confirmation_sha(drifted) != base_sha, - ) - - -def test_presence_dedupe_cli_preview_read_only(): - """preview:RR READ ONLY 输出十组摘要与 SHA,不写任何表,不触嵌入/模型。""" - - state = _presence_dedupe_state() - result, _ = _presence_dedupe_cli(state, ["--work-id", "8", "--preview"]) - check("presence-dedupe-preview成功", result.exit_code == 0, detail=result.output) - data = json.loads(result.output) - check( - "presence-dedupe-preview摘要完整", - data["mode"] == "preview" - and data["work_id"] == pu.PRESENCE_DEDUPE_WORK_ID - and data["group_count"] == pu.PRESENCE_DEDUPE_GROUP_COUNT - and data["row_count"] == pu.PRESENCE_DEDUPE_ROW_COUNT - and data["delete_ids"] == sorted(pu.PRESENCE_DEDUPE_DELETE_IDS) - and data["keep_ids"] == list(range(11501, 11511)) - and len(data["confirmation_sha"]) == 64 - and data["deleted_ids"] is None - and data["audit_rows_written"] == 0 - and len(data["groups"]) == pu.PRESENCE_DEDUPE_GROUP_COUNT, - detail=result.output[:500], - ) - check( - "presence-dedupe-preview不输出observation原文", - "观察A0" not in result.output and "观察B0" not in result.output, - ) - check( - "presence-dedupe-preview零副作用", - state["writes"] == [] and state["commits"] == 0 - and all(not row[9] for row in state["rows"]), - ) - check( - "presence-dedupe-preview首条SQL为RR只读", - state["queries"][0] == "SET TRANSACTION ISOLATION LEVEL REPEATABLE READ, READ ONLY", - detail=str(state["queries"][:2]), - ) - check( - "presence-dedupe-preview读快照不加行锁", - not any(query.endswith("FOR UPDATE") for query in state["queries"]), - detail=str(state["queries"]), - ) - - -def test_presence_dedupe_cli_execute_soft_deletes_exactly_ten(): - """execute:锁内重算精确匹配后同事务只软删十个精确 ID;其他域零写入。""" - - state = _presence_dedupe_state() - sha = _presence_dedupe_preview_sha(_presence_dedupe_state()) - result, _ = _presence_dedupe_cli( - state, - ["--work-id", "8", "--execute", "--confirmation-sha", sha, - "--confirm-no-live-process"], - ) - check("presence-dedupe-execute成功", result.exit_code == 0, detail=result.output) - data = json.loads(result.output) - check( - "presence-dedupe-execute只软删十个精确ID", - data["mode"] == "execute" - and data["deleted_ids"] == sorted(pu.PRESENCE_DEDUPE_DELETE_IDS) - and sorted(row[0] for row in state["rows"] if row[9] is True) - == sorted(pu.PRESENCE_DEDUPE_DELETE_IDS), - detail=result.output[:500], - ) - check( - "presence-dedupe-execute保留行与单体行全部留存", - sorted(row[0] for row in state["rows"] if not row[9]) - == sorted(list(range(11501, 11511)) + list(range(11400, 11404))), - detail=str(sorted(row[0] for row in state["rows"] if not row[9])), - ) - check( - "presence-dedupe-execute同事务单次提交", - state["writes"] == ["UPDATE example_upgrade_presence"] and state["commits"] == 1, - detail=str(state["writes"]), - ) - mutating = [ - query for query in state["queries"] - if query.startswith(("UPDATE ", "INSERT ", "DELETE ")) - ] - check( - "presence-dedupe-execute写SQL仅presence软删", - len(mutating) == 1 - and mutating[0].startswith("UPDATE example_upgrade_presence SET deleted=TRUE") - and "RETURNING id" in mutating[0], - detail=str(mutating), - ) - check( - "presence-dedupe-execute锁内带行锁重算快照", - any( - query.startswith("SELECT id, work_id, window_no, chapter_no") - and query.endswith("FOR UPDATE") - for query in state["queries"] - ), - detail=str(state["queries"]), - ) - - -def test_presence_dedupe_execute_second_run_fails_closed(): - """execute 成功后二次运行:软删后每组只剩单行不成十组,必须失败关闭且无新写入。""" - - state = _presence_dedupe_state() - sha = _presence_dedupe_preview_sha(_presence_dedupe_state()) - args = ["--work-id", "8", "--execute", "--confirmation-sha", sha, - "--confirm-no-live-process"] - first, _ = _presence_dedupe_cli(state, args) - check("presence-dedupe-首次execute成功", first.exit_code == 0, - detail=first.output) - second, _ = _presence_dedupe_cli(state, args) - check( - "presence-dedupe-二次execute失败关闭", - second.exit_code != 0 and "冗余组数量不精确" in second.output, - detail=second.output, - ) - check( - "presence-dedupe-二次execute无新写入", - state["writes"] == ["UPDATE example_upgrade_presence"] and state["commits"] == 1, - detail=str(state["writes"]), - ) - - -def test_presence_dedupe_execute_rejects_stale_or_drifted_snapshot(): - """旧 SHA 与 preview/execute 间快照漂移都必须在写入前失败关闭。""" - - # 旧/伪 confirmation-sha:锁内重算后精确匹配拒绝。 - state = _presence_dedupe_state() - stale, _ = _presence_dedupe_cli( - state, - ["--work-id", "8", "--execute", "--confirmation-sha", "0" * 64, - "--confirm-no-live-process"], - ) - check("presence-dedupe-旧SHA失败关闭", stale.exit_code != 0, - detail=stale.output) - check( - "presence-dedupe-旧SHA零写入", - state["writes"] == [] and state["commits"] == 0 - and all(not row[9] for row in state["rows"]), - ) - - # preview 与 execute 之间漂移(observation 被改):锁内重算 SHA 不匹配。 - sha = _presence_dedupe_preview_sha(_presence_dedupe_state()) - drifted = _presence_dedupe_state() - _mutate_presence_row(drifted, 11582, 6, "被篡改的观察") - result, _ = _presence_dedupe_cli( - drifted, - ["--work-id", "8", "--execute", "--confirmation-sha", sha, - "--confirm-no-live-process"], - ) - check( - "presence-dedupe-锁内漂移失败关闭", - result.exit_code != 0 and "不匹配" in result.output, - detail=result.output, - ) - check( - "presence-dedupe-锁内漂移零写入", - drifted["writes"] == [] and drifted["commits"] == 0 - and all(not row[9] for row in drifted["rows"]), - ) - - -def test_presence_dedupe_cli_argument_contracts(): - """preview/execute 互斥;execute 必须同时提供 SHA 与无活进程确认。""" - - cases = ( - ("preview与execute同给", ["--work-id", "8", "--preview", "--execute"]), - ("两模式都不给", ["--work-id", "8"]), - ("execute缺SHA", ["--work-id", "8", "--execute", "--confirm-no-live-process"]), - ("execute缺进程确认", - ["--work-id", "8", "--execute", "--confirmation-sha", "ab" * 32]), - ("work不符", ["--work-id", "9", "--preview"]), - ) - for label, args in cases: - state = _presence_dedupe_state() - result, _ = _presence_dedupe_cli(state, args) - check(f"presence-dedupe-{label}拒绝", result.exit_code != 0, - detail=result.output) - check( - f"presence-dedupe-{label}零副作用", - state["writes"] == [] and state["commits"] == 0, - ) - - -def test_presence_dedupe_cli_preview_fail_closed_on_drift(): - """CLI 层各类快照漂移都在 preview 被拒,且零副作用。""" - - first_delete_id = sorted(pu.PRESENCE_DEDUPE_DELETE_IDS)[0] - cases = ( - ("额外冗余组", lambda state: state["rows"].extend(( - (11581, 8, 90, 500, "event", "额外重复组", "观察C", - "upgrade", _PRESENCE_DEDUPE_BASE_TIME, False, pu.TENANT), - (11582, 8, 90, 500, "event", "额外重复组", "观察D", - "upgrade", _PRESENCE_DEDUPE_BASE_TIME, False, pu.TENANT), - ))), - ("observation相同", - lambda state: _mutate_presence_row(state, first_delete_id, 6, "观察A0")), - ("缺行", lambda state: state["rows"].remove( - next(row for row in state["rows"] if row[0] == first_delete_id))), - ("tenant不符", - lambda state: _mutate_presence_row(state, 11501, 10, "other-tenant")), - ) - for label, mutate in cases: - state = _presence_dedupe_state() - mutate(state) - result, _ = _presence_dedupe_cli(state, ["--work-id", "8", "--preview"]) - check(f"presence-dedupe-preview{label}拒绝", result.exit_code != 0, - detail=result.output) - check( - f"presence-dedupe-preview{label}零副作用", - state["writes"] == [] and state["commits"] == 0, - ) if __name__ == "__main__": @@ -6443,7 +5902,7 @@ if __name__ == "__main__": test_prompts_disciplines, test_window_fence_hashes_and_markers, test_capture_window_input_rejects_incomplete_or_empty_content, test_two_phase_external_calls_and_atomic_embedding_contract, - test_recovery_and_redo_rejection_contracts, test_real_pg_rollback_smoke_entry, + test_recovery_and_redo_rejection_contracts, test_compensation_attempt_baseline_exact_allows_old_non_allowlist_tombstone, test_compensation_attempt_baseline_input_drift_blocks, test_compensation_attempt_baseline_state_drift_blocks, @@ -6458,21 +5917,8 @@ if __name__ == "__main__": test_quality_repair_rejects_final_snapshot_drift_and_bad_embedding, test_quality_repair_rejects_invalid_vector_dimension_and_values, test_quality_repair_rejects_alias_and_presence_duplicates, - test_presence_dedupe_actions_planner_guard, - test_presence_dedupe_snapshot_success, - test_presence_dedupe_snapshot_fail_closed, - test_presence_dedupe_confirmation_sha_binds_rows_and_ids, - test_presence_dedupe_cli_preview_read_only, - test_presence_dedupe_cli_execute_soft_deletes_exactly_ten, - test_presence_dedupe_execute_second_run_fails_closed, - test_presence_dedupe_execute_rejects_stale_or_drifted_snapshot, - test_presence_dedupe_cli_argument_contracts, - test_presence_dedupe_cli_preview_fail_closed_on_drift, test_capture_window_state_binds_current_window_and_preserves_other_status, test_load_window_material_orders_blocks_like_capture_input, test_run_defaults_enforce_safety_contract): fn() - if os.getenv("MUSE_REAL_PG_ROLLBACK_SMOKE") == "1": - print(f"\n全部自测通过:{_passed} 项(含真实 PG public 窗 rollback smoke;未发模型/嵌入调用)") - else: - print(f"\n全部离线自测通过:{_passed} 项(真实 PG smoke 未启用;未发网络/嵌入/LLM 调用)") + print(f"\n全部离线自测通过:{_passed} 项(未发网络/嵌入/LLM 调用)") diff --git a/tests/skills/extract-work-knowledge/test_parse_upgrade_pg_smoke.py b/tests/skills/extract-work-knowledge/test_parse_upgrade_pg_smoke.py new file mode 100644 index 0000000..0a4db4b --- /dev/null +++ b/tests/skills/extract-work-knowledge/test_parse_upgrade_pg_smoke.py @@ -0,0 +1,28 @@ +#!/usr/bin/env python3 +"""Opt-in real PostgreSQL rollback smoke entry point.""" + +import os +import pathlib +import sys + +PROJECT_ROOT = pathlib.Path(__file__).resolve().parents[3] +SCRIPT_DIR = PROJECT_ROOT / ".claude" / "skills" / "extract-work-knowledge" / "scripts" +sys.path.insert(0, str(SCRIPT_DIR)) +import upgrade as pu # noqa: E402 + + +def main(): + if os.getenv("MUSE_REAL_PG_ROLLBACK_SMOKE") != "1": + print( + "BLOCKED: set MUSE_REAL_PG_ROLLBACK_SMOKE=1 to run the real PostgreSQL rollback smoke" + ) + return 2 + + work_id = int(os.getenv("MUSE_REAL_PG_ROLLBACK_SMOKE_WORK_ID", "8")) + pu.real_pg_rollback_smoke(work_id) + print(f"PASS: real PostgreSQL rollback smoke completed for work_id={work_id}") + return 0 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/tests/skills/extract-work-knowledge/test_presence_dedupe.py b/tests/skills/extract-work-knowledge/test_presence_dedupe.py new file mode 100644 index 0000000..2584a65 --- /dev/null +++ b/tests/skills/extract-work-knowledge/test_presence_dedupe.py @@ -0,0 +1,599 @@ +#!/usr/bin/env python3 +"""presence 冗余收口(work8 十组双行)的纯逻辑离线自测。 + +红线:**不连真实库、不发任何网络/嵌入/LLM 调用**——只使用内存 fake DB 验证 +presence 去重的 planner、快照、确认摘要和 preview/execute 事务合同。 + +跑法:仓库根目录 `.venv/bin/python tests/skills/extract-work-knowledge/test_presence_dedupe.py`; +也支持从其他工作目录通过该文件的绝对路径运行。 +""" +import json +import pathlib +import sys +from contextlib import nullcontext +from unittest.mock import patch + +from click.testing import CliRunner + +# 从原生产 scripts 导入 upgrade;测试替身只隔离数据库连接,不替代生产逻辑。 +PROJECT_ROOT = pathlib.Path(__file__).resolve().parents[3] +SCRIPT_DIR = PROJECT_ROOT / ".claude" / "skills" / "extract-work-knowledge" / "scripts" +sys.path.insert(0, str(SCRIPT_DIR)) +import upgrade as pu # noqa: E402 + +_passed = 0 + + +def check(name, cond, detail=""): + """单项断言:通过打 [PASS],失败抛 AssertionError(带上下文,令 CI/人工一眼定位)。""" + global _passed + assert cond, f"[FAIL] {name} :: {detail}" + _passed += 1 + print(f"[PASS] {name}") + + +class _CardResult: + """为 presence 去重离线测试提供最小查询结果对象。""" + + def __init__(self, row=None, rows=None): + self.row = row + self.rows = rows or [] + + def fetchone(self): + """返回预置的单行结果。""" + + return self.row + + def fetchall(self): + """返回预置的多行结果。""" + + return self.rows + + +# ── presence 冗余收口(work8 十组双行:保留 MIN(id) 软删 MAX(id))── + +# 离线 fixture 用字符串时间戳代替数据库驱动的 datetime:窄合同只比较同组两行是否相等, +# _sha256_json 计算摘要时统一 default=str 转写,二者行为一致。 +_PRESENCE_DEDUPE_BASE_TIME = "2026-07-01 12:00:00" + + +def _presence_dedupe_rows(): + """构造满足窄合同的 work8 presence 行 fixture。 + + 十组双行:每组 observation 不同、create_time 相同、恰好一个待删 ID 命中 DELETE_IDS、 + 保留 ID 是组内较小者;另有四个单行,证明快照会跳过不构成冗余的 key。 + """ + + rows = [] + for index, delete_id in enumerate(sorted(pu.PRESENCE_DEDUPE_DELETE_IDS)): + keep_id = 11501 + index + entity_type = "location" if index % 2 == 0 else "item" + name = f"重复实体{index}" + for row_id, observation in ((keep_id, f"观察A{index}"), + (delete_id, f"观察B{index}")): + rows.append((row_id, 8, 57 + index, 403 + index, entity_type, name, + observation, "upgrade", _PRESENCE_DEDUPE_BASE_TIME, + False, pu.TENANT)) + for index in range(4): + rows.append((11400 + index, 8, 10 + index, 20 + index, "character", + f"单体实体{index}", f"单体观察{index}", + "upgrade", _PRESENCE_DEDUPE_BASE_TIME, False, pu.TENANT)) + return sorted(rows, key=lambda row: row[0]) + + +def _presence_dedupe_state(): + return { + "rows": _presence_dedupe_rows(), + "writes": [], + "queries": [], + "commits": 0, + } + + +def _mutate_presence_row(state, row_id, column, value): + """替换指定行的单列;元组不可变,整体重建后写回。""" + + state["rows"] = [ + row[:column] + (value,) + row[column + 1:] if row[0] == row_id else row + for row in state["rows"] + ] + + +class _PresenceDedupeConn: + """presence 冗余收口 fake DB:只实现本命令的读快照与软删 SQL。 + + 其他任何域(draft/window/alias/card_state/audit/embedding)的 SQL 会落到末尾 + AssertionError,因此「零副作用」无需逐条枚举禁写语句即可离线断言。fake 只做 + 「未软删」粗过滤;列数/work/tenant/deleted 类型等窄合同行像校验正是离线断言对象, + fake 不能代为过滤。 + """ + + def __init__(self, state): + self.state = state + + def __enter__(self): + return self + + def __exit__(self, exc_type, exc, tb): + return False + + def execute(self, query, params=()): + normalized = " ".join(query.split()) + self.state.setdefault("queries", []).append(normalized) + if normalized.startswith("SET TRANSACTION") or normalized.startswith("LOCK TABLE"): + return _CardResult() + if normalized.startswith( + "SELECT id, work_id, window_no, chapter_no, entity_type, name, observation,"): + rows = [ + row for row in sorted(self.state["rows"], key=lambda item: item[0]) + if len(row) > 9 and not row[9] + ] + return _CardResult(rows=rows) + if normalized.startswith("UPDATE example_upgrade_presence SET deleted=TRUE"): + tenant_id, work_id, delete_ids = params + assert tenant_id == pu.TENANT + delete_set = set(delete_ids) + hit = [] + new_rows = [] + for row in self.state["rows"]: + if row[1] == work_id and row[0] in delete_set and not row[9]: + row = row[:9] + (True,) + row[10:] + hit.append((row[0],)) + new_rows.append(row) + self.state["rows"] = new_rows + self.state["writes"].append("UPDATE example_upgrade_presence") + return _CardResult(rows=hit) + raise AssertionError(f"presence 冗余收口 fake DB 未覆盖 SQL:{normalized}") + + def commit(self): + self.state["commits"] += 1 + + +def _presence_dedupe_cli(state, args): + """离线执行 repair-presence-duplicates:注入 fake DB 与空放同书锁。""" + + conn = _PresenceDedupeConn(state) + with patch.object(pu.psycopg, "connect", return_value=conn), \ + patch.object(pu, "upgrade_work_lock", return_value=nullcontext()): + result = CliRunner().invoke(pu.maintenance_cli, ["repair-presence-duplicates"] + args) + return result, conn + + +def _presence_dedupe_preview_sha(state): + """离线取 preview 的 confirmation_sha,作为 execute 的合法输入。""" + + result, _ = _presence_dedupe_cli(state, ["--work-id", "8", "--preview"]) + assert result.exit_code == 0, result.output + return json.loads(result.output)["confirmation_sha"] + + +def _presence_snapshot_raises(label, state, work_id=pu.PRESENCE_DEDUPE_WORK_ID): + """断言快照以 CompensationFenceConflict 拒绝,且没有发出任何写入。""" + + conn = _PresenceDedupeConn(state) + try: + pu._presence_dedupe_capture_snapshot(conn, work_id) + except pu.CompensationFenceConflict as exc: + check(f"presence-dedupe-{label}失败关闭", True, detail=str(exc)) + else: + check(f"presence-dedupe-{label}失败关闭", False, + detail="未抛 CompensationFenceConflict") + check( + f"presence-dedupe-{label}零副作用", + state["writes"] == [] and state["commits"] == 0, + ) + + +def test_presence_dedupe_actions_planner_guard(): + """planner 级收口守卫:同观察幂等留首条、不同观察失败关闭、非 presence 原样透传。""" + + presence_a = ("presence", None, 403, "location", "重复星体", "观察一") + presence_a_dup = ("presence", None, 403, "location", "重复星体", "观察一") + presence_b = ("presence", None, 404, "item", "镜面护盾", "观察二") + alias_action = ("alias", 7, "规范名", "别名", "ai") + new_action = ("new", -1, {"名称": "新实体"}, set(), []) + actions = [presence_a, alias_action, presence_a_dup, new_action, presence_b, ()] + original = list(actions) + deduped = pu._dedupe_presence_actions(actions) + check( + "presence-dedupe-planner同观察去重保留首条", + deduped == [presence_a, alias_action, new_action, presence_b, ()] + and deduped[0] is presence_a, + detail=str(deduped), + ) + check("presence-dedupe-重复应用幂等", + pu._dedupe_presence_actions(deduped) == deduped) + check("presence-dedupe-不修改输入列表", actions == original) + + conflict = [presence_a, + ("presence", None, 403, "location", "重复星体", "另一个观察")] + try: + pu._dedupe_presence_actions(conflict) + except RuntimeError as exc: + check("presence-dedupe-不同观察失败关闭", + "observation 冲突" in str(exc), detail=str(exc)) + else: + check("presence-dedupe-不同观察失败关闭", False, detail="未抛 RuntimeError") + + for label, bad, keyword in ( + ("结构缺一元", ("presence", None, 403, "location", "重复星体"), + "结构非法"), + ("结构多一元", + ("presence", None, 403, "location", "重复星体", "观察一", "extra"), + "结构非法"), + ("observation非字符串", ("presence", None, 403, "location", "重复星体", None), + "必须是字符串"), + ): + try: + pu._dedupe_presence_actions([bad]) + except RuntimeError as exc: + check(f"presence-dedupe-{label}抛错", keyword in str(exc), detail=str(exc)) + else: + check(f"presence-dedupe-{label}抛错", False, detail="未抛 RuntimeError") + + +def test_presence_dedupe_snapshot_success(): + """快照成功路径:恰好 10 组/20 行,待删集合精确等于 DELETE_IDS,保留为组内较小者。""" + + state = _presence_dedupe_state() + snapshot = pu._presence_dedupe_capture_snapshot( + _PresenceDedupeConn(state), pu.PRESENCE_DEDUPE_WORK_ID, + ) + check( + "presence-dedupe-快照组数行数精确", + len(snapshot["groups"]) == pu.PRESENCE_DEDUPE_GROUP_COUNT + and len(snapshot["rows"]) == pu.PRESENCE_DEDUPE_ROW_COUNT + and snapshot["contract"] == pu.PRESENCE_DEDUPE_CONTRACT + and snapshot["work_id"] == pu.PRESENCE_DEDUPE_WORK_ID, + ) + check( + "presence-dedupe-待删集合精确等于DELETE_IDS", + snapshot["delete_ids"] == sorted(pu.PRESENCE_DEDUPE_DELETE_IDS), + detail=str(snapshot["delete_ids"]), + ) + check( + "presence-dedupe-保留为组内较小ID", + snapshot["keep_ids"] == list(range(11501, 11511)) + and all(group["keep_id"] < group["delete_id"] for group in snapshot["groups"]) + and all(len(group["rows"]) == 2 for group in snapshot["groups"]), + detail=str(snapshot["keep_ids"]), + ) + check( + "presence-dedupe-行按id排序且排除单体行", + [row[0] for row in snapshot["rows"]] + == sorted(row[0] for row in snapshot["rows"]) + and not any(row[0] in range(11400, 11404) for row in snapshot["rows"]), + ) + + +def test_presence_dedupe_snapshot_fail_closed(): + """各类窄合同漂移都必须在写入前以 CompensationFenceConflict 逐条拒绝。""" + + # 额外冗余组:组数 11 ≠ 10(新增组复用合同内待删 ID,确保失败点落在组数校验)。 + state = _presence_dedupe_state() + state["rows"].extend(( + (11581, 8, 90, 500, "event", "额外重复组", "观察C", + "upgrade", _PRESENCE_DEDUPE_BASE_TIME, False, pu.TENANT), + (11582, 8, 90, 500, "event", "额外重复组", "观察D", + "upgrade", _PRESENCE_DEDUPE_BASE_TIME, False, pu.TENANT), + )) + _presence_snapshot_raises("额外冗余组", state) + + # 三行组不是「恰好两行」。 + state = _presence_dedupe_state() + state["rows"].append( + (11601, 8, 57, 403, "location", "重复实体0", "观察E", + "upgrade", _PRESENCE_DEDUPE_BASE_TIME, False, pu.TENANT) + ) + _presence_snapshot_raises("三行冗余组", state) + + # 缺行:一组只剩单行被跳过 → 组数 9 ≠ 10。 + state = _presence_dedupe_state() + state["rows"] = [row for row in state["rows"] + if row[0] != sorted(pu.PRESENCE_DEDUPE_DELETE_IDS)[0]] + _presence_snapshot_raises("缺行", state) + + # 同组两行 create_time 漂移。 + state = _presence_dedupe_state() + _mutate_presence_row(state, 11501, 8, "2026-07-02 09:30:00") + _presence_snapshot_raises("create_time漂移", state) + + # 待删 ID 命中 0:组内两行都不在 DELETE_IDS。 + state = _presence_dedupe_state() + _mutate_presence_row(state, 11501, 0, 12001) + _mutate_presence_row(state, 11582, 0, 12002) + _presence_snapshot_raises("待删ID无命中", state) + + # 待删 ID 命中 2:组内两行都在 DELETE_IDS。 + state = _presence_dedupe_state() + _mutate_presence_row(state, 11501, 0, 11693) + _presence_snapshot_raises("待删ID双命中", state) + + # 保留不是组内较小者:待删 ID 反而是组内较小者。 + state = _presence_dedupe_state() + _mutate_presence_row(state, 11501, 0, 11900) + _presence_snapshot_raises("保留非组内较小", state) + + # 同组两行 observation 相同(byte-exact 重复违反窄合同)。 + state = _presence_dedupe_state() + _mutate_presence_row(state, 11582, 6, "观察A0") + _presence_snapshot_raises("observation相同", state) + + # work_id ≠ 8 拒绝。 + _presence_snapshot_raises("work不符", _presence_dedupe_state(), work_id=9) + + # 窄合同行像漂移:列数不足。 + state = _presence_dedupe_state() + state["rows"] = [row[:10] if row[0] == 11501 else row for row in state["rows"]] + _presence_snapshot_raises("列数不符", state) + + # 窄合同行像漂移:work 列与查询目标不符。 + state = _presence_dedupe_state() + _mutate_presence_row(state, 11501, 1, 9) + _presence_snapshot_raises("work列不符", state) + + # 窄合同行像漂移:deleted 不是严格 False(整数 0 也必须拒绝)。 + state = _presence_dedupe_state() + _mutate_presence_row(state, 11501, 9, 0) + _presence_snapshot_raises("deleted类型不符", state) + + # 窄合同行像漂移:tenant 列不符。 + state = _presence_dedupe_state() + _mutate_presence_row(state, 11501, 10, "other-tenant") + _presence_snapshot_raises("tenant不符", state) + + +def test_presence_dedupe_confirmation_sha_binds_rows_and_ids(): + """confirmation_sha 必须确定且绑定二十行完整内容与保留/删除 ID,任一来源漂移即变化。""" + + base_snapshot = pu._presence_dedupe_capture_snapshot( + _PresenceDedupeConn(_presence_dedupe_state()), pu.PRESENCE_DEDUPE_WORK_ID, + ) + base_sha = pu._presence_dedupe_confirmation_sha(base_snapshot) + again_snapshot = pu._presence_dedupe_capture_snapshot( + _PresenceDedupeConn(_presence_dedupe_state()), pu.PRESENCE_DEDUPE_WORK_ID, + ) + check( + "presence-dedupe-confirmation_sha确定性", + base_sha == pu._presence_dedupe_confirmation_sha(again_snapshot) + and len(base_sha) == 64, + detail=base_sha, + ) + for label, row_id, column, value in ( + ("observation", 11582, 6, "被篡改的观察"), + ("保留行ID", 11501, 0, 11001), + ("creator", 11501, 7, "篡改者"), + ): + state = _presence_dedupe_state() + _mutate_presence_row(state, row_id, column, value) + drifted = pu._presence_dedupe_capture_snapshot( + _PresenceDedupeConn(state), pu.PRESENCE_DEDUPE_WORK_ID, + ) + check( + f"presence-dedupe-sha感知{label}漂移", + pu._presence_dedupe_confirmation_sha(drifted) != base_sha, + ) + + +def test_presence_dedupe_cli_preview_read_only(): + """preview:RR READ ONLY 输出十组摘要与 SHA,不写任何表,不触嵌入/模型。""" + + state = _presence_dedupe_state() + result, _ = _presence_dedupe_cli(state, ["--work-id", "8", "--preview"]) + check("presence-dedupe-preview成功", result.exit_code == 0, detail=result.output) + data = json.loads(result.output) + check( + "presence-dedupe-preview摘要完整", + data["mode"] == "preview" + and data["work_id"] == pu.PRESENCE_DEDUPE_WORK_ID + and data["group_count"] == pu.PRESENCE_DEDUPE_GROUP_COUNT + and data["row_count"] == pu.PRESENCE_DEDUPE_ROW_COUNT + and data["delete_ids"] == sorted(pu.PRESENCE_DEDUPE_DELETE_IDS) + and data["keep_ids"] == list(range(11501, 11511)) + and len(data["confirmation_sha"]) == 64 + and data["deleted_ids"] is None + and data["audit_rows_written"] == 0 + and len(data["groups"]) == pu.PRESENCE_DEDUPE_GROUP_COUNT, + detail=result.output[:500], + ) + check( + "presence-dedupe-preview不输出observation原文", + "观察A0" not in result.output and "观察B0" not in result.output, + ) + check( + "presence-dedupe-preview零副作用", + state["writes"] == [] and state["commits"] == 0 + and all(not row[9] for row in state["rows"]), + ) + check( + "presence-dedupe-preview首条SQL为RR只读", + state["queries"][0] == "SET TRANSACTION ISOLATION LEVEL REPEATABLE READ, READ ONLY", + detail=str(state["queries"][:2]), + ) + check( + "presence-dedupe-preview读快照不加行锁", + not any(query.endswith("FOR UPDATE") for query in state["queries"]), + detail=str(state["queries"]), + ) + + +def test_presence_dedupe_cli_execute_soft_deletes_exactly_ten(): + """execute:锁内重算精确匹配后同事务只软删十个精确 ID;其他域零写入。""" + + state = _presence_dedupe_state() + sha = _presence_dedupe_preview_sha(_presence_dedupe_state()) + result, _ = _presence_dedupe_cli( + state, + ["--work-id", "8", "--execute", "--confirmation-sha", sha, + "--confirm-no-live-process"], + ) + check("presence-dedupe-execute成功", result.exit_code == 0, detail=result.output) + data = json.loads(result.output) + check( + "presence-dedupe-execute只软删十个精确ID", + data["mode"] == "execute" + and data["deleted_ids"] == sorted(pu.PRESENCE_DEDUPE_DELETE_IDS) + and sorted(row[0] for row in state["rows"] if row[9] is True) + == sorted(pu.PRESENCE_DEDUPE_DELETE_IDS), + detail=result.output[:500], + ) + check( + "presence-dedupe-execute保留行与单体行全部留存", + sorted(row[0] for row in state["rows"] if not row[9]) + == sorted(list(range(11501, 11511)) + list(range(11400, 11404))), + detail=str(sorted(row[0] for row in state["rows"] if not row[9])), + ) + check( + "presence-dedupe-execute同事务单次提交", + state["writes"] == ["UPDATE example_upgrade_presence"] and state["commits"] == 1, + detail=str(state["writes"]), + ) + mutating = [ + query for query in state["queries"] + if query.startswith(("UPDATE ", "INSERT ", "DELETE ")) + ] + check( + "presence-dedupe-execute写SQL仅presence软删", + len(mutating) == 1 + and mutating[0].startswith("UPDATE example_upgrade_presence SET deleted=TRUE") + and "RETURNING id" in mutating[0], + detail=str(mutating), + ) + check( + "presence-dedupe-execute锁内带行锁重算快照", + any( + query.startswith("SELECT id, work_id, window_no, chapter_no") + and query.endswith("FOR UPDATE") + for query in state["queries"] + ), + detail=str(state["queries"]), + ) + + +def test_presence_dedupe_execute_second_run_fails_closed(): + """execute 成功后二次运行:软删后每组只剩单行不成十组,必须失败关闭且无新写入。""" + + state = _presence_dedupe_state() + sha = _presence_dedupe_preview_sha(_presence_dedupe_state()) + args = ["--work-id", "8", "--execute", "--confirmation-sha", sha, + "--confirm-no-live-process"] + first, _ = _presence_dedupe_cli(state, args) + check("presence-dedupe-首次execute成功", first.exit_code == 0, + detail=first.output) + second, _ = _presence_dedupe_cli(state, args) + check( + "presence-dedupe-二次execute失败关闭", + second.exit_code != 0 and "冗余组数量不精确" in second.output, + detail=second.output, + ) + check( + "presence-dedupe-二次execute无新写入", + state["writes"] == ["UPDATE example_upgrade_presence"] and state["commits"] == 1, + detail=str(state["writes"]), + ) + + +def test_presence_dedupe_execute_rejects_stale_or_drifted_snapshot(): + """旧 SHA 与 preview/execute 间快照漂移都必须在写入前失败关闭。""" + + # 旧/伪 confirmation-sha:锁内重算后精确匹配拒绝。 + state = _presence_dedupe_state() + stale, _ = _presence_dedupe_cli( + state, + ["--work-id", "8", "--execute", "--confirmation-sha", "0" * 64, + "--confirm-no-live-process"], + ) + check("presence-dedupe-旧SHA失败关闭", stale.exit_code != 0, + detail=stale.output) + check( + "presence-dedupe-旧SHA零写入", + state["writes"] == [] and state["commits"] == 0 + and all(not row[9] for row in state["rows"]), + ) + + # preview 与 execute 之间漂移(observation 被改):锁内重算 SHA 不匹配。 + sha = _presence_dedupe_preview_sha(_presence_dedupe_state()) + drifted = _presence_dedupe_state() + _mutate_presence_row(drifted, 11582, 6, "被篡改的观察") + result, _ = _presence_dedupe_cli( + drifted, + ["--work-id", "8", "--execute", "--confirmation-sha", sha, + "--confirm-no-live-process"], + ) + check( + "presence-dedupe-锁内漂移失败关闭", + result.exit_code != 0 and "不匹配" in result.output, + detail=result.output, + ) + check( + "presence-dedupe-锁内漂移零写入", + drifted["writes"] == [] and drifted["commits"] == 0 + and all(not row[9] for row in drifted["rows"]), + ) + + +def test_presence_dedupe_cli_argument_contracts(): + """preview/execute 互斥;execute 必须同时提供 SHA 与无活进程确认。""" + + cases = ( + ("preview与execute同给", ["--work-id", "8", "--preview", "--execute"]), + ("两模式都不给", ["--work-id", "8"]), + ("execute缺SHA", ["--work-id", "8", "--execute", "--confirm-no-live-process"]), + ("execute缺进程确认", + ["--work-id", "8", "--execute", "--confirmation-sha", "ab" * 32]), + ("work不符", ["--work-id", "9", "--preview"]), + ) + for label, args in cases: + state = _presence_dedupe_state() + result, _ = _presence_dedupe_cli(state, args) + check(f"presence-dedupe-{label}拒绝", result.exit_code != 0, + detail=result.output) + check( + f"presence-dedupe-{label}零副作用", + state["writes"] == [] and state["commits"] == 0, + ) + + +def test_presence_dedupe_cli_preview_fail_closed_on_drift(): + """CLI 层各类快照漂移都在 preview 被拒,且零副作用。""" + + first_delete_id = sorted(pu.PRESENCE_DEDUPE_DELETE_IDS)[0] + cases = ( + ("额外冗余组", lambda state: state["rows"].extend(( + (11581, 8, 90, 500, "event", "额外重复组", "观察C", + "upgrade", _PRESENCE_DEDUPE_BASE_TIME, False, pu.TENANT), + (11582, 8, 90, 500, "event", "额外重复组", "观察D", + "upgrade", _PRESENCE_DEDUPE_BASE_TIME, False, pu.TENANT), + ))), + ("observation相同", + lambda state: _mutate_presence_row(state, first_delete_id, 6, "观察A0")), + ("缺行", lambda state: state["rows"].remove( + next(row for row in state["rows"] if row[0] == first_delete_id))), + ("tenant不符", + lambda state: _mutate_presence_row(state, 11501, 10, "other-tenant")), + ) + for label, mutate in cases: + state = _presence_dedupe_state() + mutate(state) + result, _ = _presence_dedupe_cli(state, ["--work-id", "8", "--preview"]) + check(f"presence-dedupe-preview{label}拒绝", result.exit_code != 0, + detail=result.output) + check( + f"presence-dedupe-preview{label}零副作用", + state["writes"] == [] and state["commits"] == 0, + ) + + +if __name__ == "__main__": + for fn in (test_presence_dedupe_actions_planner_guard, + test_presence_dedupe_snapshot_success, + test_presence_dedupe_snapshot_fail_closed, + test_presence_dedupe_confirmation_sha_binds_rows_and_ids, + test_presence_dedupe_cli_preview_read_only, + test_presence_dedupe_cli_execute_soft_deletes_exactly_ten, + test_presence_dedupe_execute_second_run_fails_closed, + test_presence_dedupe_execute_rejects_stale_or_drifted_snapshot, + test_presence_dedupe_cli_argument_contracts, + test_presence_dedupe_cli_preview_fail_closed_on_drift): + fn() + print(f"\n全部离线自测通过:{_passed} 项(未发网络/嵌入/LLM 调用)") diff --git a/.claude/skills/extract-work-knowledge/scripts/test_upgrade_work_lock_offline.py b/tests/skills/extract-work-knowledge/test_upgrade_work_lock_offline.py similarity index 96% rename from .claude/skills/extract-work-knowledge/scripts/test_upgrade_work_lock_offline.py rename to tests/skills/extract-work-knowledge/test_upgrade_work_lock_offline.py index 90f2004..17e1af0 100644 --- a/.claude/skills/extract-work-knowledge/scripts/test_upgrade_work_lock_offline.py +++ b/tests/skills/extract-work-knowledge/test_upgrade_work_lock_offline.py @@ -6,7 +6,8 @@ import sys import unittest -SCRIPT_DIR = pathlib.Path(__file__).resolve().parent +PROJECT_ROOT = pathlib.Path(__file__).resolve().parents[3] +SCRIPT_DIR = PROJECT_ROOT / ".claude" / "skills" / "extract-work-knowledge" / "scripts" sys.path.insert(0, str(SCRIPT_DIR)) from upgrade_work_lock import ( # noqa: E402 diff --git a/.claude/skills/freeze-context/scripts/test_audit_leakage.py b/tests/skills/freeze-context/test_audit_leakage.py similarity index 96% rename from .claude/skills/freeze-context/scripts/test_audit_leakage.py rename to tests/skills/freeze-context/test_audit_leakage.py index f169689..5d24b3b 100644 --- a/.claude/skills/freeze-context/scripts/test_audit_leakage.py +++ b/tests/skills/freeze-context/test_audit_leakage.py @@ -8,7 +8,9 @@ import pathlib import sys import unittest -sys.path.insert(0, str(pathlib.Path(__file__).resolve().parent)) +PROJECT_ROOT = pathlib.Path(__file__).resolve().parents[3] +SCRIPT_DIR = PROJECT_ROOT / ".claude" / "skills" / "freeze-context" / "scripts" +sys.path.insert(0, str(SCRIPT_DIR)) from audit_leakage import audit_snapshot # noqa: E402 diff --git a/.claude/skills/freeze-context/scripts/test_build_snapshot.py b/tests/skills/freeze-context/test_build_snapshot.py similarity index 97% rename from .claude/skills/freeze-context/scripts/test_build_snapshot.py rename to tests/skills/freeze-context/test_build_snapshot.py index 04182c5..1f541f7 100644 --- a/.claude/skills/freeze-context/scripts/test_build_snapshot.py +++ b/tests/skills/freeze-context/test_build_snapshot.py @@ -5,7 +5,9 @@ import pathlib import sys import unittest -sys.path.insert(0, str(pathlib.Path(__file__).resolve().parent)) +PROJECT_ROOT = pathlib.Path(__file__).resolve().parents[3] +SCRIPT_DIR = PROJECT_ROOT / ".claude" / "skills" / "freeze-context" / "scripts" +sys.path.insert(0, str(SCRIPT_DIR)) from build_snapshot import ( # noqa: E402 SnapshotError, build_snapshot, diff --git a/.claude/skills/freeze-context/scripts/test_check_snapshot.py b/tests/skills/freeze-context/test_check_snapshot.py similarity index 98% rename from .claude/skills/freeze-context/scripts/test_check_snapshot.py rename to tests/skills/freeze-context/test_check_snapshot.py index 39a0b8a..d22cf0b 100644 --- a/.claude/skills/freeze-context/scripts/test_check_snapshot.py +++ b/tests/skills/freeze-context/test_check_snapshot.py @@ -5,7 +5,9 @@ import pathlib import sys import unittest -sys.path.insert(0, str(pathlib.Path(__file__).resolve().parent)) +PROJECT_ROOT = pathlib.Path(__file__).resolve().parents[3] +SCRIPT_DIR = PROJECT_ROOT / ".claude" / "skills" / "freeze-context" / "scripts" +sys.path.insert(0, str(SCRIPT_DIR)) from check_snapshot import ( # noqa: E402 STATUS_BLOCKED_AUTHORIZATION, STATUS_INVALID_ARM_DIFF, diff --git a/.claude/skills/freeze-context/scripts/test_load_reference_work.py b/tests/skills/freeze-context/test_load_reference_work.py similarity index 98% rename from .claude/skills/freeze-context/scripts/test_load_reference_work.py rename to tests/skills/freeze-context/test_load_reference_work.py index 784d58f..dd95e37 100644 --- a/.claude/skills/freeze-context/scripts/test_load_reference_work.py +++ b/tests/skills/freeze-context/test_load_reference_work.py @@ -10,7 +10,9 @@ import sys import unittest from unittest.mock import patch -sys.path.insert(0, str(pathlib.Path(__file__).resolve().parent)) +PROJECT_ROOT = pathlib.Path(__file__).resolve().parents[3] +SCRIPT_DIR = PROJECT_ROOT / ".claude" / "skills" / "freeze-context" / "scripts" +sys.path.insert(0, str(SCRIPT_DIR)) import load_reference_work as loader # noqa: E402 from load_reference_work import ( # noqa: E402 AdapterError, diff --git a/.claude/skills/maintain-work-extraction/scripts/test_backup_upgrade_work_offline.py b/tests/skills/maintain-work-extraction/test_backup_upgrade_work_offline.py similarity index 99% rename from .claude/skills/maintain-work-extraction/scripts/test_backup_upgrade_work_offline.py rename to tests/skills/maintain-work-extraction/test_backup_upgrade_work_offline.py index 41d0b4f..6dd1976 100644 --- a/.claude/skills/maintain-work-extraction/scripts/test_backup_upgrade_work_offline.py +++ b/tests/skills/maintain-work-extraction/test_backup_upgrade_work_offline.py @@ -12,8 +12,9 @@ from decimal import Decimal from unittest import mock -HERE = pathlib.Path(__file__).resolve().parent -sys.path.insert(0, str(HERE)) +PROJECT_ROOT = pathlib.Path(__file__).resolve().parents[3] +SCRIPT_DIR = PROJECT_ROOT / ".claude" / "skills" / "maintain-work-extraction" / "scripts" +sys.path.insert(0, str(SCRIPT_DIR)) import backup_upgrade_work as backup # noqa: E402 diff --git a/.claude/skills/maintain-work-extraction/scripts/test_reset_upgrade_work_offline.py b/tests/skills/maintain-work-extraction/test_reset_upgrade_work_offline.py similarity index 99% rename from .claude/skills/maintain-work-extraction/scripts/test_reset_upgrade_work_offline.py rename to tests/skills/maintain-work-extraction/test_reset_upgrade_work_offline.py index 083e23e..4acd66f 100644 --- a/.claude/skills/maintain-work-extraction/scripts/test_reset_upgrade_work_offline.py +++ b/tests/skills/maintain-work-extraction/test_reset_upgrade_work_offline.py @@ -11,9 +11,11 @@ from unittest.mock import Mock, patch from click.testing import CliRunner -SCRIPT_DIR = pathlib.Path(__file__).resolve().parent +PROJECT_ROOT = pathlib.Path(__file__).resolve().parents[3] +SCRIPT_DIR = PROJECT_ROOT / ".claude" / "skills" / "maintain-work-extraction" / "scripts" +EXTRACTION_SCRIPTS = PROJECT_ROOT / ".claude" / "skills" / "extract-work-knowledge" / "scripts" sys.path.insert(0, str(SCRIPT_DIR)) -sys.path.insert(0, str(SCRIPT_DIR.parents[1] / "extract-work-knowledge" / "scripts")) +sys.path.insert(0, str(EXTRACTION_SCRIPTS)) import upgrade as parse # noqa: E402 import backup_upgrade_work as backup # noqa: E402 diff --git a/.claude/skills/plan-chapter/scripts/test_contract.py b/tests/skills/plan-chapter/test_contract.py similarity index 85% rename from .claude/skills/plan-chapter/scripts/test_contract.py rename to tests/skills/plan-chapter/test_contract.py index e3f09b7..fe062ba 100644 --- a/.claude/skills/plan-chapter/scripts/test_contract.py +++ b/tests/skills/plan-chapter/test_contract.py @@ -9,11 +9,15 @@ from pathlib import Path import unittest -ROOT = Path(__file__).resolve().parents[4] -SKILL = (ROOT / ".claude/skills/plan-chapter/SKILL.md").read_text(encoding="utf-8") -PLANNER = (ROOT / ".claude/agents/planner.md").read_text(encoding="utf-8") -CHAINS = (ROOT / "meta/chains/README.md").read_text(encoding="utf-8") -SCHEMA = (ROOT / "meta/schemas/fine_outline.yaml").read_text(encoding="utf-8") +PROJECT_ROOT = Path(__file__).resolve().parents[3] +SKILL = ( + PROJECT_ROOT / ".claude" / "skills" / "plan-chapter" / "SKILL.md" +).read_text(encoding="utf-8") +PLANNER = (PROJECT_ROOT / ".claude" / "agents" / "planner.md").read_text(encoding="utf-8") +CHAINS = (PROJECT_ROOT / "meta" / "chains" / "README.md").read_text(encoding="utf-8") +SCHEMA = ( + PROJECT_ROOT / "meta" / "schemas" / "fine_outline.yaml" +).read_text(encoding="utf-8") # 与 meta/schemas/fine_outline.yaml 的必填/推荐集保持一致。 REQUIRED_FIELDS = ( diff --git a/.claude/skills/plan-story/scripts/test_field_coverage.py b/tests/skills/plan-story/test_field_coverage.py similarity index 96% rename from .claude/skills/plan-story/scripts/test_field_coverage.py rename to tests/skills/plan-story/test_field_coverage.py index 1d9a18d..a7305f8 100644 --- a/.claude/skills/plan-story/scripts/test_field_coverage.py +++ b/tests/skills/plan-story/test_field_coverage.py @@ -10,8 +10,10 @@ import pathlib import sys import unittest -sys.path.insert(0, str(pathlib.Path(__file__).resolve().parent)) -sys.path.insert(0, str(pathlib.Path(__file__).resolve().parents[2] / "access-database" / "scripts")) + +PROJECT_ROOT = pathlib.Path(__file__).resolve().parents[3] +SCRIPT_DIR = PROJECT_ROOT / ".claude" / "skills" / "plan-story" / "scripts" +sys.path.insert(0, str(SCRIPT_DIR)) from persist_planning import check_field_coverage, write_section, _load_schema # noqa: E402 diff --git a/.claude/skills/plan-story/scripts/test_record_planning_execution.py b/tests/skills/plan-story/test_record_planning_execution.py similarity index 84% rename from .claude/skills/plan-story/scripts/test_record_planning_execution.py rename to tests/skills/plan-story/test_record_planning_execution.py index f45f548..142ab1b 100644 --- a/.claude/skills/plan-story/scripts/test_record_planning_execution.py +++ b/tests/skills/plan-story/test_record_planning_execution.py @@ -1,11 +1,13 @@ #!/usr/bin/env python3 """规划执行留痕脚本的离线合同测试。""" + import pathlib import sys import unittest -SCRIPT_DIR = pathlib.Path(__file__).resolve().parent +PROJECT_ROOT = pathlib.Path(__file__).resolve().parents[3] +SCRIPT_DIR = PROJECT_ROOT / ".claude" / "skills" / "plan-story" / "scripts" sys.path.insert(0, str(SCRIPT_DIR)) import record_planning_execution as recorder # noqa: E402 diff --git a/.claude/skills/plan-story/scripts/test_repair_deterministic_receipt.py b/tests/skills/plan-story/test_repair_deterministic_receipt.py similarity index 88% rename from .claude/skills/plan-story/scripts/test_repair_deterministic_receipt.py rename to tests/skills/plan-story/test_repair_deterministic_receipt.py index d715d9e..54b74f5 100644 --- a/.claude/skills/plan-story/scripts/test_repair_deterministic_receipt.py +++ b/tests/skills/plan-story/test_repair_deterministic_receipt.py @@ -1,11 +1,13 @@ #!/usr/bin/env python3 """确定性规划回执修正的离线合同测试。""" + import pathlib import sys import unittest -SCRIPT_DIR = pathlib.Path(__file__).resolve().parent +PROJECT_ROOT = pathlib.Path(__file__).resolve().parents[3] +SCRIPT_DIR = PROJECT_ROOT / ".claude" / "skills" / "plan-story" / "scripts" sys.path.insert(0, str(SCRIPT_DIR)) import repair_deterministic_receipt as repair # noqa: E402 diff --git a/.claude/skills/plan-story/scripts/test_select_patterns_offline.py b/tests/skills/plan-story/test_select_patterns_offline.py similarity index 93% rename from .claude/skills/plan-story/scripts/test_select_patterns_offline.py rename to tests/skills/plan-story/test_select_patterns_offline.py index 78d5083..fd752bc 100644 --- a/.claude/skills/plan-story/scripts/test_select_patterns_offline.py +++ b/tests/skills/plan-story/test_select_patterns_offline.py @@ -1,12 +1,14 @@ #!/usr/bin/env python3 """范式选择投影的离线测试,不调用向量服务。""" + import pathlib import sys import unittest from unittest.mock import patch -SCRIPT_DIR = pathlib.Path(__file__).resolve().parent +PROJECT_ROOT = pathlib.Path(__file__).resolve().parents[3] +SCRIPT_DIR = PROJECT_ROOT / ".claude" / "skills" / "plan-story" / "scripts" sys.path.insert(0, str(SCRIPT_DIR)) import select_patterns as selector # noqa: E402 diff --git a/.claude/skills/prevent-ai-flavor/scripts/test_prevent_ai_flavor.py b/tests/skills/prevent-ai-flavor/test_prevent_ai_flavor.py similarity index 95% rename from .claude/skills/prevent-ai-flavor/scripts/test_prevent_ai_flavor.py rename to tests/skills/prevent-ai-flavor/test_prevent_ai_flavor.py index bf239e2..b848d27 100644 --- a/.claude/skills/prevent-ai-flavor/scripts/test_prevent_ai_flavor.py +++ b/tests/skills/prevent-ai-flavor/test_prevent_ai_flavor.py @@ -5,7 +5,9 @@ import sys import unittest from unittest.mock import patch -sys.path.insert(0, str(pathlib.Path(__file__).resolve().parent)) +PROJECT_ROOT = pathlib.Path(__file__).resolve().parents[3] +SCRIPT_DIR = PROJECT_ROOT / ".claude" / "skills" / "prevent-ai-flavor" / "scripts" +sys.path.insert(0, str(SCRIPT_DIR)) import prevent_ai_flavor as prev # noqa: E402 diff --git a/.claude/skills/record-run-evidence/scripts/test_file_cas.py b/tests/skills/record-run-evidence/test_file_cas.py similarity index 98% rename from .claude/skills/record-run-evidence/scripts/test_file_cas.py rename to tests/skills/record-run-evidence/test_file_cas.py index bd047ca..137df66 100644 --- a/.claude/skills/record-run-evidence/scripts/test_file_cas.py +++ b/tests/skills/record-run-evidence/test_file_cas.py @@ -12,7 +12,8 @@ import threading import unittest from unittest import mock -SCRIPT_DIR = pathlib.Path(__file__).resolve().parent +PROJECT_ROOT = pathlib.Path(__file__).resolve().parents[3] +SCRIPT_DIR = PROJECT_ROOT / ".claude" / "skills" / "record-run-evidence" / "scripts" sys.path.insert(0, str(SCRIPT_DIR)) from file_cas import CasConflictError, CasRecoveryError, FileCasStore # noqa: E402 diff --git a/.claude/skills/record-run-evidence/scripts/test_persist_raw.py b/tests/skills/record-run-evidence/test_persist_raw.py similarity index 82% rename from .claude/skills/record-run-evidence/scripts/test_persist_raw.py rename to tests/skills/record-run-evidence/test_persist_raw.py index 600cdfa..c083cb4 100644 --- a/.claude/skills/record-run-evidence/scripts/test_persist_raw.py +++ b/tests/skills/record-run-evidence/test_persist_raw.py @@ -3,7 +3,9 @@ import pathlib import sys -sys.path.insert(0, str(pathlib.Path(__file__).resolve().parent)) +PROJECT_ROOT = pathlib.Path(__file__).resolve().parents[3] +SCRIPT_DIR = PROJECT_ROOT / ".claude" / "skills" / "record-run-evidence" / "scripts" +sys.path.insert(0, str(SCRIPT_DIR)) from persist_raw import _check_no_secrets # noqa: E402 diff --git a/.claude/skills/record-run-evidence/scripts/test_raw_vault.py b/tests/skills/record-run-evidence/test_raw_vault.py similarity index 98% rename from .claude/skills/record-run-evidence/scripts/test_raw_vault.py rename to tests/skills/record-run-evidence/test_raw_vault.py index 024a6a1..ac0999d 100644 --- a/.claude/skills/record-run-evidence/scripts/test_raw_vault.py +++ b/tests/skills/record-run-evidence/test_raw_vault.py @@ -13,7 +13,8 @@ import unittest from datetime import datetime, timedelta, timezone from unittest import mock -SCRIPT_DIR = pathlib.Path(__file__).resolve().parent +PROJECT_ROOT = pathlib.Path(__file__).resolve().parents[3] +SCRIPT_DIR = PROJECT_ROOT / ".claude" / "skills" / "record-run-evidence" / "scripts" sys.path.insert(0, str(SCRIPT_DIR)) import raw_vault # noqa: E402 diff --git a/.claude/skills/record-run-evidence/scripts/test_record_failed_run.py b/tests/skills/record-run-evidence/test_record_failed_run.py similarity index 84% rename from .claude/skills/record-run-evidence/scripts/test_record_failed_run.py rename to tests/skills/record-run-evidence/test_record_failed_run.py index 97d9546..c4a4a80 100644 --- a/.claude/skills/record-run-evidence/scripts/test_record_failed_run.py +++ b/tests/skills/record-run-evidence/test_record_failed_run.py @@ -5,7 +5,8 @@ import sys import unittest -SCRIPT_DIR = pathlib.Path(__file__).resolve().parent +PROJECT_ROOT = pathlib.Path(__file__).resolve().parents[3] +SCRIPT_DIR = PROJECT_ROOT / ".claude" / "skills" / "record-run-evidence" / "scripts" sys.path.insert(0, str(SCRIPT_DIR)) import record_failed_run as failure # noqa: E402 diff --git a/.claude/skills/record-run-evidence/scripts/test_repair_receipt_evidence.py b/tests/skills/record-run-evidence/test_repair_receipt_evidence.py similarity index 79% rename from .claude/skills/record-run-evidence/scripts/test_repair_receipt_evidence.py rename to tests/skills/record-run-evidence/test_repair_receipt_evidence.py index de056b9..db8613c 100644 --- a/.claude/skills/record-run-evidence/scripts/test_repair_receipt_evidence.py +++ b/tests/skills/record-run-evidence/test_repair_receipt_evidence.py @@ -5,7 +5,8 @@ import sys import unittest -SCRIPT_DIR = pathlib.Path(__file__).resolve().parent +PROJECT_ROOT = pathlib.Path(__file__).resolve().parents[3] +SCRIPT_DIR = PROJECT_ROOT / ".claude" / "skills" / "record-run-evidence" / "scripts" sys.path.insert(0, str(SCRIPT_DIR)) import repair_receipt_evidence as repair # noqa: E402 diff --git a/.claude/skills/record-run-evidence/scripts/test_run_registry.py b/tests/skills/record-run-evidence/test_run_registry.py similarity index 83% rename from .claude/skills/record-run-evidence/scripts/test_run_registry.py rename to tests/skills/record-run-evidence/test_run_registry.py index e9cc93d..0b460d3 100644 --- a/.claude/skills/record-run-evidence/scripts/test_run_registry.py +++ b/tests/skills/record-run-evidence/test_run_registry.py @@ -3,7 +3,9 @@ import pathlib import sys -sys.path.insert(0, str(pathlib.Path(__file__).resolve().parent)) +PROJECT_ROOT = pathlib.Path(__file__).resolve().parents[3] +SCRIPT_DIR = PROJECT_ROOT / ".claude" / "skills" / "record-run-evidence" / "scripts" +sys.path.insert(0, str(SCRIPT_DIR)) import run_registry # noqa: E402 diff --git a/.claude/skills/revise-ai-flavor/scripts/test_revise_ai_flavor.py b/tests/skills/revise-ai-flavor/test_revise_ai_flavor.py similarity index 96% rename from .claude/skills/revise-ai-flavor/scripts/test_revise_ai_flavor.py rename to tests/skills/revise-ai-flavor/test_revise_ai_flavor.py index 967cf5e..1acfbe0 100644 --- a/.claude/skills/revise-ai-flavor/scripts/test_revise_ai_flavor.py +++ b/tests/skills/revise-ai-flavor/test_revise_ai_flavor.py @@ -7,8 +7,11 @@ import tempfile import unittest from unittest.mock import patch -sys.path.insert(0, str(pathlib.Path(__file__).resolve().parent)) -sys.path.insert(0, str(pathlib.Path(__file__).resolve().parents[3] / "skills" / "diagnose-ai-flavor" / "scripts")) +PROJECT_ROOT = pathlib.Path(__file__).resolve().parents[3] +SCRIPT_DIR = PROJECT_ROOT / ".claude" / "skills" / "revise-ai-flavor" / "scripts" +DIAGNOSE_SCRIPT_DIR = PROJECT_ROOT / ".claude" / "skills" / "diagnose-ai-flavor" / "scripts" +sys.path.insert(0, str(SCRIPT_DIR)) +sys.path.insert(0, str(DIAGNOSE_SCRIPT_DIR)) import revise_ai_flavor as rev # noqa: E402 import diagnose_ai_flavor as diag # noqa: E402 diff --git a/.claude/skills/score-content-quality/scripts/test_lesson_registry_db.py b/tests/skills/score-content-quality/test_lesson_registry_db.py similarity index 93% rename from .claude/skills/score-content-quality/scripts/test_lesson_registry_db.py rename to tests/skills/score-content-quality/test_lesson_registry_db.py index e13731f..6789896 100644 --- a/.claude/skills/score-content-quality/scripts/test_lesson_registry_db.py +++ b/tests/skills/score-content-quality/test_lesson_registry_db.py @@ -10,7 +10,7 @@ 测试行用 unittest-lesson- 前缀的 run_id 隔离,结束物理清理(本表可变)。 跑法(需 Tailscale 内网可达 muse-example): -.venv/bin/python .claude/skills/score-content-quality/scripts/test_lesson_registry_db.py +.venv/bin/python tests/skills/score-content-quality/test_lesson_registry_db.py """ from __future__ import annotations @@ -19,9 +19,13 @@ import pathlib import sys import uuid -SCRIPT_DIR = pathlib.Path(__file__).resolve().parent -SKILLS_DIR = SCRIPT_DIR.parents[1] -for path in (SCRIPT_DIR, SKILLS_DIR / "access-database" / "scripts"): +from psycopg.errors import RaiseException + +PROJECT_ROOT = pathlib.Path(__file__).resolve().parents[3] +SKILLS_DIR = PROJECT_ROOT / ".claude" / "skills" +SCRIPT_DIR = SKILLS_DIR / "score-content-quality" / "scripts" +DB_SCRIPT_DIR = SKILLS_DIR / "access-database" / "scripts" +for path in (SCRIPT_DIR, DB_SCRIPT_DIR): if str(path) not in sys.path: sys.path.insert(0, str(path)) @@ -96,7 +100,7 @@ def test_no_auto_promotion() -> None: raise AssertionError("触发器必须拒绝自动升格") except AssertionError: raise - except Exception: + except RaiseException: pass rows = list_lessons(status="proposed") assert any(row["id"] == out["lesson_id"] for row in rows) diff --git a/.claude/skills/score-content-quality/scripts/test_rubric.py b/tests/skills/score-content-quality/test_rubric.py similarity index 95% rename from .claude/skills/score-content-quality/scripts/test_rubric.py rename to tests/skills/score-content-quality/test_rubric.py index 56a759a..071123c 100644 --- a/.claude/skills/score-content-quality/scripts/test_rubric.py +++ b/tests/skills/score-content-quality/test_rubric.py @@ -6,7 +6,10 @@ import inspect import sys import unittest -sys.path.insert(0, str(pathlib.Path(__file__).resolve().parent)) +PROJECT_ROOT = pathlib.Path(__file__).resolve().parents[3] +SCRIPT_DIR = PROJECT_ROOT / ".claude" / "skills" / "score-content-quality" / "scripts" +if str(SCRIPT_DIR) not in sys.path: + sys.path.insert(0, str(SCRIPT_DIR)) from fine_outline_rubric import ( # noqa: E402 DIMENSIONS, RUBRIC_PROFILE, diff --git a/.claude/skills/score-content-quality/scripts/test_run_writer_blind_judge.py b/tests/skills/score-content-quality/test_run_writer_blind_judge.py similarity index 99% rename from .claude/skills/score-content-quality/scripts/test_run_writer_blind_judge.py rename to tests/skills/score-content-quality/test_run_writer_blind_judge.py index ad2cee8..a794fca 100644 --- a/.claude/skills/score-content-quality/scripts/test_run_writer_blind_judge.py +++ b/tests/skills/score-content-quality/test_run_writer_blind_judge.py @@ -11,7 +11,8 @@ import sys import unittest from typing import Any, Mapping -SCRIPT_DIR = pathlib.Path(__file__).resolve().parent +PROJECT_ROOT = pathlib.Path(__file__).resolve().parents[3] +SCRIPT_DIR = PROJECT_ROOT / ".claude" / "skills" / "score-content-quality" / "scripts" if str(SCRIPT_DIR) not in sys.path: sys.path.insert(0, str(SCRIPT_DIR)) diff --git a/.claude/skills/score-content-quality/scripts/test_writer_rubric.py b/tests/skills/score-content-quality/test_writer_rubric.py similarity index 98% rename from .claude/skills/score-content-quality/scripts/test_writer_rubric.py rename to tests/skills/score-content-quality/test_writer_rubric.py index bf5f00e..f3aaedc 100644 --- a/.claude/skills/score-content-quality/scripts/test_writer_rubric.py +++ b/tests/skills/score-content-quality/test_writer_rubric.py @@ -7,7 +7,10 @@ import pathlib import sys import unittest -sys.path.insert(0, str(pathlib.Path(__file__).resolve().parent)) +PROJECT_ROOT = pathlib.Path(__file__).resolve().parents[3] +SCRIPT_DIR = PROJECT_ROOT / ".claude" / "skills" / "score-content-quality" / "scripts" +if str(SCRIPT_DIR) not in sys.path: + sys.path.insert(0, str(SCRIPT_DIR)) from writer_rubric import ( # noqa: E402 COMMON_DIMENSIONS, diff --git a/.claude/skills/search-knowledge/scripts/test_search.py b/tests/skills/search-knowledge/test_search.py similarity index 89% rename from .claude/skills/search-knowledge/scripts/test_search.py rename to tests/skills/search-knowledge/test_search.py index 84abf9d..db94599 100644 --- a/.claude/skills/search-knowledge/scripts/test_search.py +++ b/tests/skills/search-knowledge/test_search.py @@ -3,8 +3,16 @@ from __future__ import annotations +import pathlib +import sys import unittest +PROJECT_ROOT = pathlib.Path(__file__).resolve().parents[3] +SCRIPT_DIR = PROJECT_ROOT / ".claude" / "skills" / "search-knowledge" / "scripts" +EMBED_SCRIPT_DIR = PROJECT_ROOT / ".claude" / "skills" / "embed-knowledge" / "scripts" +sys.path.insert(0, str(SCRIPT_DIR)) +sys.path.insert(0, str(EMBED_SCRIPT_DIR)) + from search import search_cards diff --git a/.claude/skills/write-next-chapter/scripts/test_candidate_cas.py b/tests/skills/write-next-chapter/test_candidate_cas.py similarity index 95% rename from .claude/skills/write-next-chapter/scripts/test_candidate_cas.py rename to tests/skills/write-next-chapter/test_candidate_cas.py index a161509..3b02cff 100644 --- a/.claude/skills/write-next-chapter/scripts/test_candidate_cas.py +++ b/tests/skills/write-next-chapter/test_candidate_cas.py @@ -6,7 +6,7 @@ - transition/start_next 的守卫与 token 递增; - 该存储可直接替换 InMemoryCasStateStore 驱动 run_writer_pipeline 全链。 -跑法(仓库根目录):.venv/bin/python .claude/skills/write-next-chapter/scripts/test_candidate_cas.py +跑法(仓库根目录):.venv/bin/python tests/skills/write-next-chapter/test_candidate_cas.py """ from __future__ import annotations @@ -16,11 +16,14 @@ import pathlib import sys import unittest -SCRIPT_DIR = pathlib.Path(__file__).resolve().parent -SKILLS_DIR = SCRIPT_DIR.parents[1] +PROJECT_ROOT = pathlib.Path(__file__).resolve().parents[3] +SKILLS_DIR = PROJECT_ROOT / ".claude" / "skills" +TEST_DIR = PROJECT_ROOT / "tests" / "skills" / "write-next-chapter" +CHECK_TEST_DIR = PROJECT_ROOT / "tests" / "skills" / "check-content-consistency" +SCRIPT_DIR = SKILLS_DIR / "write-next-chapter" / "scripts" DETECT_DIR = SKILLS_DIR / "check-content-consistency" / "scripts" READ_CONTEXT_DIR = SKILLS_DIR / "assemble-context" / "scripts" -for path in (SCRIPT_DIR, DETECT_DIR, READ_CONTEXT_DIR): +for path in (READ_CONTEXT_DIR, DETECT_DIR, SCRIPT_DIR, TEST_DIR, CHECK_TEST_DIR): if str(path) not in sys.path: sys.path.insert(0, str(path)) diff --git a/.claude/skills/write-next-chapter/scripts/test_candidate_cas_db.py b/tests/skills/write-next-chapter/test_candidate_cas_db.py similarity index 94% rename from .claude/skills/write-next-chapter/scripts/test_candidate_cas_db.py rename to tests/skills/write-next-chapter/test_candidate_cas_db.py index cff61b4..e767f02 100644 --- a/.claude/skills/write-next-chapter/scripts/test_candidate_cas_db.py +++ b/tests/skills/write-next-chapter/test_candidate_cas_db.py @@ -6,7 +6,7 @@ PASSED 终态、身份不可变,以及条件 UPDATE 在真实并发语义下 测试行用 unittest-cas- 前缀的 run_id,结束前物理清理(本表是可变注册表,无 append-only 约束)。 跑法(需 Tailscale 内网可达 muse-example): -.venv/bin/python .claude/skills/write-next-chapter/scripts/test_candidate_cas_db.py +.venv/bin/python tests/skills/write-next-chapter/test_candidate_cas_db.py """ from __future__ import annotations @@ -15,8 +15,11 @@ import pathlib import sys import uuid -SCRIPT_DIR = pathlib.Path(__file__).resolve().parent -SKILLS_DIR = SCRIPT_DIR.parents[1] +from psycopg.errors import RaiseException + +PROJECT_ROOT = pathlib.Path(__file__).resolve().parents[3] +SKILLS_DIR = PROJECT_ROOT / ".claude" / "skills" +SCRIPT_DIR = SKILLS_DIR / "write-next-chapter" / "scripts" for path in (SCRIPT_DIR, SKILLS_DIR / "access-database" / "scripts"): if str(path) not in sys.path: sys.path.insert(0, str(path)) @@ -107,7 +110,7 @@ def test_db_trigger_rejects_malformed_direct_updates() -> None: raise AssertionError(f"触发器必须拒绝: {sql}") except AssertionError: raise - except Exception: + except RaiseException: pass latest = store.latest(run_id) assert (latest.state, latest.revision) == ("DRAFT", 1), "非法 UPDATE 不得改变链" @@ -130,7 +133,7 @@ def test_start_next_monotonic_at_db_level() -> None: raise AssertionError("REJECTED 开新轮必须递增 attempt/candidate_version") except AssertionError: raise - except Exception: + except RaiseException: pass assert store.latest(run_id).state == "REJECTED", "非法开新轮不得改变链" @@ -146,7 +149,7 @@ def test_insert_must_start_draft_revision_one() -> None: raise AssertionError("初始行必须是 DRAFT") except AssertionError: raise - except Exception: + except RaiseException: pass try: with connect() as conn: @@ -157,7 +160,7 @@ def test_insert_must_start_draft_revision_one() -> None: raise AssertionError("初始 revision 必须是 1") except AssertionError: raise - except Exception: + except RaiseException: pass diff --git a/.claude/skills/write-next-chapter/scripts/test_persist_writer_run.py b/tests/skills/write-next-chapter/test_persist_writer_run.py similarity index 82% rename from .claude/skills/write-next-chapter/scripts/test_persist_writer_run.py rename to tests/skills/write-next-chapter/test_persist_writer_run.py index 289c146..b6362b0 100644 --- a/.claude/skills/write-next-chapter/scripts/test_persist_writer_run.py +++ b/tests/skills/write-next-chapter/test_persist_writer_run.py @@ -5,7 +5,8 @@ import sys import unittest -SCRIPT_DIR = pathlib.Path(__file__).resolve().parent +PROJECT_ROOT = pathlib.Path(__file__).resolve().parents[3] +SCRIPT_DIR = PROJECT_ROOT / ".claude" / "skills" / "write-next-chapter" / "scripts" sys.path.insert(0, str(SCRIPT_DIR)) import persist_writer_run as writer_persist # noqa: E402 diff --git a/.claude/skills/write-next-chapter/scripts/test_run_writer.py b/tests/skills/write-next-chapter/test_run_writer.py similarity index 97% rename from .claude/skills/write-next-chapter/scripts/test_run_writer.py rename to tests/skills/write-next-chapter/test_run_writer.py index 5a05a46..aca9874 100644 --- a/.claude/skills/write-next-chapter/scripts/test_run_writer.py +++ b/tests/skills/write-next-chapter/test_run_writer.py @@ -12,10 +12,13 @@ import sys import unittest from decimal import Decimal -SCRIPT_DIR = pathlib.Path(__file__).resolve().parent -READ_CONTEXT_DIR = SCRIPT_DIR.parents[1] / "assemble-context" / "scripts" +PROJECT_ROOT = pathlib.Path(__file__).resolve().parents[3] +SCRIPT_DIR = PROJECT_ROOT / ".claude" / "skills" / "write-next-chapter" / "scripts" +READ_CONTEXT_DIR = PROJECT_ROOT / ".claude" / "skills" / "assemble-context" / "scripts" +TEST_CONTEXT_DIR = PROJECT_ROOT / "tests" / "skills" / "assemble-context" sys.path.insert(0, str(SCRIPT_DIR)) sys.path.insert(0, str(READ_CONTEXT_DIR)) +sys.path.insert(0, str(TEST_CONTEXT_DIR)) from test_writer_contract import valid_context, valid_draft # noqa: E402 from writer_contract import build_writer_creative_input, canonical_json, retrieval_identity # noqa: E402 diff --git a/.claude/skills/write-next-chapter/scripts/test_run_writer_pipeline.py b/tests/skills/write-next-chapter/test_run_writer_pipeline.py similarity index 96% rename from .claude/skills/write-next-chapter/scripts/test_run_writer_pipeline.py rename to tests/skills/write-next-chapter/test_run_writer_pipeline.py index 8f3d432..224b049 100644 --- a/.claude/skills/write-next-chapter/scripts/test_run_writer_pipeline.py +++ b/tests/skills/write-next-chapter/test_run_writer_pipeline.py @@ -12,11 +12,14 @@ import tempfile import unittest from unittest import mock -SCRIPT_DIR = pathlib.Path(__file__).resolve().parent -SKILLS_DIR = SCRIPT_DIR.parents[1] +PROJECT_ROOT = pathlib.Path(__file__).resolve().parents[3] +SKILLS_DIR = PROJECT_ROOT / ".claude" / "skills" +TEST_DIR = PROJECT_ROOT / "tests" / "skills" / "write-next-chapter" +CHECK_TEST_DIR = PROJECT_ROOT / "tests" / "skills" / "check-content-consistency" +SCRIPT_DIR = SKILLS_DIR / "write-next-chapter" / "scripts" DETECT_DIR = SKILLS_DIR / "check-content-consistency" / "scripts" READ_CONTEXT_DIR = SKILLS_DIR / "assemble-context" / "scripts" -for path in (SCRIPT_DIR, DETECT_DIR, READ_CONTEXT_DIR): +for path in (READ_CONTEXT_DIR, DETECT_DIR, SCRIPT_DIR, TEST_DIR, CHECK_TEST_DIR): sys.path.insert(0, str(path)) from test_check_writer_candidate import _candidate_body, _requirements, _valid_pair # noqa: E402 diff --git a/.claude/skills/write-next-chapter/scripts/test_semantic_verdict.py b/tests/skills/write-next-chapter/test_semantic_verdict.py similarity index 95% rename from .claude/skills/write-next-chapter/scripts/test_semantic_verdict.py rename to tests/skills/write-next-chapter/test_semantic_verdict.py index 89a737f..226ad2a 100644 --- a/.claude/skills/write-next-chapter/scripts/test_semantic_verdict.py +++ b/tests/skills/write-next-chapter/test_semantic_verdict.py @@ -4,7 +4,7 @@ 语义状态固化到候选行是接受通道 DB 兜底的依据,绑定/哈希校验必须失败关闭: 报告版本、候选哈希、候选版本、上下文哈希、运行 ID 任一不一致都拒绝落库。 -跑法:.venv/bin/python .claude/skills/write-next-chapter/scripts/test_semantic_verdict.py +跑法:.venv/bin/python tests/skills/write-next-chapter/test_semantic_verdict.py """ from __future__ import annotations @@ -14,7 +14,8 @@ import pathlib import sys import unittest -SCRIPT_DIR = pathlib.Path(__file__).resolve().parent +PROJECT_ROOT = pathlib.Path(__file__).resolve().parents[3] +SCRIPT_DIR = PROJECT_ROOT / ".claude" / "skills" / "write-next-chapter" / "scripts" if str(SCRIPT_DIR) not in sys.path: sys.path.insert(0, str(SCRIPT_DIR))