muse-agent-example/humanization/research/20-project-skill-coverage.yaml
zizi c9f69d9d6d 治理: Skill 测试治理第一阶段——harness 控制平面 + 实现测试迁出运行时目录
范围(不含 design-story-foundation、docs/、humanization/README.md 等进行中改动):

1. 新增 harness/ 控制平面
   - skill_harness.py 静态审计:32 个运行时 Skill 的 frontmatter/manifest/文档污染,当前 0 问题
   - run_selected.py 选择性执行器:manifest 与磁盘一一对账、依赖阻断、
     空跑与 skip-only 失败关闭、AST 测试形状门
   - manifests/skills.json:32 个 Skill 的合同责任方与协作领域登记
   - manifests/test-inventory.json:81 个测试资产登记
   - specs/skill-testing.md 与 README.md:测试分层、证据边界与 harness 职责

2. 实现测试从 .claude/skills/*/scripts/ 迁至 tests/skills/<skill>/
   - 71 个测试文件迁移并修复项目根与临时目录运行导入
   - 数据库触发器测试宽泛异常收窄为 psycopg.errors.RaiseException
   - 抽取离线大测试拆出真实 PG smoke(默认阻断,不计入离线通过)
   - 抽取 presence 去重边界拆出独立测试:493 + 78 = 571 项检查不变

3. 运行时文档清理
   - 13 个 SKILL.md 移除自测/离线验证段落、测试命令与测试文件事实源表述,
     只保留运行时合同;业务运行合同、额度、授权与离线模式均保留

4. SoT 同步
   - AGENTS.md:新增 Skill 领域索引(7 个合同责任方分组,覆盖 32 个运行时 Skill)
   - 领域 07:测试入口改由 harness/manifests/ 登记,SKILL.md 不承载测试命令
   - humanization 覆盖矩阵:活动测试路径同步迁移

验证证据: harness 自测 15 项 + runner 自测 13 项通过;静态审计 32 Skill / 0 问题;
73 个非数据库测试通过;8 个集成条目中 6 个 PostgreSQL 项被依赖门明确阻断;
py_compile 与 git diff --check 通过。未连接 PostgreSQL、网络、真实模型或额度。

已知边界: 真正 skill_behavior_eval 仍为 0,尚未验证任何 Skill 自然语言行为;
evaluate-frozen-replay 的 raw 存储边界冲突留待单独治理。
2026-08-19 01:50:20 +08:00

144 lines
6.6 KiB
YAML
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

schema_version: humanization-research-skill-coverage-v1
source: ../../../.agents/knowledge/ai-writing-humanization-open-source-research.md
scope: agent-example自治验证期;研究出处不等于效果证明
projects:
- id: humanizer
mapped_capabilities: [voice_first_baseline, minimal_unique_patch]
- id: stop-slop
mapped_capabilities: [five_layer_detection]
- id: Humanizer-zh
mapped_capabilities: [preservation_and_modality]
- id: inkos
mapped_capabilities: [minimal_unique_patch, bounded_revision_and_original_wins, cross_model_pairwise]
- id: oh-story-claudecode
mapped_capabilities: [five_layer_detection, bounded_revision_and_original_wins]
- id: no-ai-slop
mapped_capabilities: [detect_edit_separation, pre_generation_guidance]
- id: PaperSpine
mapped_capabilities: [detect_edit_separation, bounded_revision_and_original_wins]
- id: im-not-ai
mapped_capabilities: [five_layer_detection, holdout_effect_evaluation]
- id: avoid-ai-writing
mapped_capabilities: [carrier_scope_and_mask, preservation_and_modality]
- id: human-writing
mapped_capabilities: [detect_edit_separation]
- id: talk-normal
mapped_capabilities: [pre_generation_guidance, holdout_effect_evaluation]
- id: AIWriteX
mapped_capabilities: [minimal_unique_patch]
- id: humanize-text
mapped_capabilities: [five_layer_detection, holdout_effect_evaluation]
- id: shuorenhua
mapped_capabilities: [carrier_scope_and_mask, sf_snf_boundary_regression]
- id: academic-humanizer
mapped_capabilities: [preservation_and_modality]
- id: speak-human-tw
mapped_capabilities: [sf_snf_boundary_regression, holdout_effect_evaluation]
- id: AI_paper
mapped_capabilities: [holdout_effect_evaluation]
- id: Openwrite
mapped_capabilities: [minimal_unique_patch]
- id: unslop
mapped_capabilities: [case_to_rule_lifecycle]
- id: neuro-book
mapped_capabilities: [five_layer_detection, case_to_rule_lifecycle, holdout_effect_evaluation]
capabilities:
- id: detect_edit_separation
evidence_sources: [no-ai-slop, oh-story-claudecode, academic-humanizer]
owner_skill: diagnose-ai-flavor
implementation: humanization/src/deai/diagnose.py
test: humanization/tests/test_contracts.py
status: implemented
- id: five_layer_detection
evidence_sources: [neuro-book, oh-story-claudecode, avoid-ai-writing]
owner_skill: diagnose-ai-flavor
implementation: humanization/src/deai/diagnose.py
test: humanization/tests/test_humanization_v2.py
status: implemented
note: regex/handler/density机械执行;semantic仍需外部模型或人工
- id: carrier_scope_and_mask
evidence_sources: [shuorenhua, avoid-ai-writing, neuro-book]
owner_skill: diagnose-ai-flavor
implementation: humanization/src/deai/carriers.py
test: humanization/tests/test_humanization_v2.py
status: implemented
- id: voice_first_baseline
evidence_sources: [humanizer, shuorenhua, inkos]
owner_skill: establish-voice-baseline
implementation: .claude/skills/establish-voice-baseline/scripts/establish_voice_baseline.py
test: humanization/tests/test_humanization_v2.py
status: implemented
note: 角色策略仍需planner/作者补充和确认
- id: pre_generation_guidance
evidence_sources: [no-ai-slop, neuro-book, humanizer]
owner_skill: prevent-ai-flavor
implementation: .claude/skills/prevent-ai-flavor/scripts/prevent_ai_flavor.py
test: tests/skills/assemble-context/test_assemble_writer_context.py
status: implemented
- id: sf_snf_boundary_regression
evidence_sources: [speak-human-tw, shuorenhua, neuro-book]
owner_skill: capture-ai-flavor-cases
implementation: humanization/samples
test: humanization/src/deai/evaluation.py
status: implemented
- id: minimal_unique_patch
evidence_sources: [inkos, Openwrite, shuorenhua]
owner_skill: revise-ai-flavor
implementation: humanization/src/deai/patch.py
test: humanization/tests/test_contracts.py
status: implemented
- id: preservation_and_modality
evidence_sources: [humanizer, academic-humanizer, shuorenhua]
owner_skill: revise-ai-flavor
implementation: humanization/src/deai/gates.py
test: humanization/tests/test_humanization_v2.py
status: implemented
note: POV/因果/伏笔仍需事实快照和语义detector
- id: bounded_revision_and_original_wins
evidence_sources: [inkos, oh-story-claudecode, neuro-book]
owner_skill: revise-ai-flavor
implementation: humanization/src/deai/pipeline.py
test: tests/skills/revise-ai-flavor/test_revise_ai_flavor.py
status: implemented
note: 当前以最多3轮合同、复扫和pairwise no_gain实现;跨轮最佳快照编排仍由上层负责
- id: source_revalidation
evidence_sources: [shuorenhua, neuro-book]
owner_skill: capture-ai-flavor-cases
implementation: .claude/skills/capture-ai-flavor-cases/scripts/capture_cases.py
test: tests/skills/capture-ai-flavor-cases/test_capture_cases.py
status: implemented
- id: case_to_rule_lifecycle
evidence_sources: [unslop, shuorenhua, neuro-book]
owner_skill: capture-ai-flavor-cases
implementation: .claude/skills/capture-ai-flavor-cases/scripts/mine_ai_flavor.py
test: humanization/tests/test_humanization_v2.py
status: implemented
note: 状态写回数据库;规则文件激活需显式输出和人工审批
- id: cross_model_pairwise
evidence_sources: [inkos, neuro-book, oh-story-claudecode]
owner_skill: revise-ai-flavor
implementation: humanization/src/deai/pairwise.py
test: humanization/tests/test_contracts.py
status: implemented
- id: holdout_effect_evaluation
evidence_sources: [neuro-book, speak-human-tw, im-not-ai]
owner_skill: capture-ai-flavor-cases
implementation: humanization/src/deai/evaluation.py
test: humanization/tests/test_humanization_v2.py
status: partial
note: 合同回放和holdout计数门已建;真实多作者多题材holdout尚未形成
- id: model_drift_deprecation
evidence_sources: [inkos, im-not-ai, neuro-book]
owner_skill: capture-ai-flavor-cases
implementation: humanization/src/deai/evaluation.py
test: humanization/tests/test_humanization_v2.py
status: partial
note: 人工审批的deprecate输出已建;自动跨模型巡检尚未接入生产调度
- id: independent_reader_blind_eval
evidence_sources: [neuro-book, inkos, humanizer]
owner_skill: capture-ai-flavor-cases
implementation: humanization/eval
test: humanization/tests/test_humanization_v2.py
status: pending
note: pairwise合同存在,但独立读者规模化盲评未完成