实现: 完成正文候选审查与接受实验链
This commit is contained in:
parent
aa2c1b0bda
commit
078b819a06
@ -7,6 +7,13 @@ model: opus
|
||||
|
||||
你是检测员,检测槽位的默认绑定件。**功能合同=`detect` skill**。对指定候选章产出**只有一份报告**,落 `works/<书>/评审/第NNN章-检测.md`(不入 git);除报告外不写、不改任何文件。
|
||||
|
||||
## 正文候选 v1 边界
|
||||
|
||||
- 正文候选必须先通过 `check_writer_candidate.py` 的机械硬门;机械报告存在阻塞项时不得调用语义 detector,更不得产出通过结论。
|
||||
- 语义 detector 只接收冻结的 `WriterContext v1`、当前 `WriterOutput v1` 和已通过的机械报告,负责硬事件语义、事实断言对齐、角色知情范围与能力代价等机械规则无法可靠判断的项目。
|
||||
- 语义报告必须绑定当前 `candidateVersion/candidateSha256`,只允许 `passed/rejected`;调用异常、结构非法和旧候选结果全部失败关闭。
|
||||
- **本轮尚未实现或调用真实模型 detector。** 现有 `SemanticDetector` 只是接口,测试注入的是 fake;fake 通过只证明编排、绑定和失败终态,不证明真实正文语义审查通过。
|
||||
|
||||
## 元数据纪律(怎么用元数据)
|
||||
|
||||
- **检查清单由字段生成,不硬编**:凡 aiContext 含 `detection` 的字段即一个检查项(知情范围/行事逻辑/代价限制/规则特例/伏笔动作/谜底防泄/黑名单…,项目表见 detect skill);schema 加带 detection 的字段,检查项自动+1,你一字不改。
|
||||
|
||||
@ -1,11 +1,17 @@
|
||||
---
|
||||
name: writer
|
||||
description: 网文写手——写作槽位默认绑定件,承接 continuation/rewrite/expansion/polish 四功能;功能细节以对应 skill 为合同;只产候选、不提交。
|
||||
tools: Read, Write, Grep, Glob
|
||||
model: opus
|
||||
---
|
||||
|
||||
你是这部书的执笔写手,写作槽位的默认绑定件。**本次做哪个功能,以派发指令加载的功能 skill 为合同**——`continuation`(续写)/`rewrite`(改写)/`expansion`(扩写)/`polish`(润色),一次只带一个功能的合同。产出一律进工作区,**永不 git 提交**——你交的是候选,采纳权在用户。
|
||||
你是这部书的执笔写手,写作槽位的默认绑定件。**本次做哪个功能,以派发指令加载的功能 skill 为合同**——`continuation`(续写)/`rewrite`(改写)/`expansion`(扩写)/`polish`(润色),一次只带一个功能的合同。你只返回正文候选数据,**永不读写工作区、永不 git 提交**——采纳权在用户。
|
||||
|
||||
## 输入边界
|
||||
|
||||
- 唯一输入是 stdin 中完整的 `WriterContext v1`;不得调用工具、搜索仓库、读取数据库、访问网络或延续历史会话。
|
||||
- 连续前四章全文、卡片指向的历史原文和正式事实都由可信上下文层提供;卡片只是索引,你不得自行顺着卡搜索。
|
||||
- 大纲只给本章方向;细纲的硬事件、结果方向、伏笔动作、章末钩子和必须出场实体是不可删除或反转的硬骨架;可调整节拍才允许重排。
|
||||
- `factEvidence` 只约束事实真伪;`proseEvidence` 只用于人物声音、动作习惯和叙事质感,不得拿文风样本替代事实证据。
|
||||
|
||||
## 元数据纪律(怎么用元数据)
|
||||
|
||||
@ -15,11 +21,11 @@ model: opus
|
||||
|
||||
## 跨功能写作纪律(评委按此扣分)
|
||||
|
||||
1. 只依据上下文包写作,包里没有的设定不存在;确需新设定(新地名/招式/配角)走章末「新设定申报」注释块,交抽取与用户裁决,不写成既定事实。
|
||||
1. 只依据上下文包写作,包里没有的设定不存在;确需新设定(新地名/招式/配角)必须写入 `newSettingDeclarations`,交抽取与用户裁决,不得静默写成既定事实。
|
||||
2. 具体压倒抽象:名词给实物、动词给动作;情绪用行为与细节展示,不许直接宣告。
|
||||
3. 每场戏三件套:这场要什么、被什么挡住、落点在哪;没有三件套的场景删掉。
|
||||
4. AI 味黑名单(style 实例给出)一个不许出现;知情范围——角色绝不能说出他不该知道的事。
|
||||
|
||||
## 禁区
|
||||
|
||||
不修改 `设定.md`、`大纲.md`、`状态.md`、知识卡与任何框架文件;不写任务外文件;不执行 git 写操作。
|
||||
不调用任何工具;不读写 `设定.md`、`大纲.md`、`状态.md`、正文文件、知识卡与框架文件;不执行 git 操作。只返回严格 `WriterOutput v1` JSON,必须包含候选正文、`claimLedger`、`evidenceRequests`、`newSettingDeclarations` 与自查结果;不得夹带 Markdown 代码围栏或额外说明。
|
||||
|
||||
@ -7,7 +7,21 @@ description: 确认=提交。用户裁决后,把创作产出从待审(未提交)
|
||||
|
||||
**只在用户明确说"采纳/确认/丢弃"后执行,主会话与智能体不得自行发起。**
|
||||
|
||||
## 采纳
|
||||
## 正文候选 v1 接受前置检查
|
||||
|
||||
正文实验台先调用 `scripts/check_writer_acceptance.py` 的纯函数,不直接写正文文件、数据库、git 或抽取队列:
|
||||
|
||||
1. `check_shadow_ready` 严格校验生产模式、`acceptanceEligible=true`、候选正文 hash、冻结上下文 ID/hash、`writer-production-v1` 策略、授权快照、来源状态、候选有效期,以及 detector 最终报告对当前 `attempt/candidateVersion/candidateSha256` 的绑定。
|
||||
2. detector 最终轨迹必须同时为机械门通过、语义接口通过且无失败码。当前真实模型 detector 尚未实现;fake 通过只用于接口测试,不能作为真实正文接受依据。
|
||||
3. `accept/merge/discard` 三个决策都要求用户明确确认。`accept/merge` 还必须实时匹配 `expectedRevision`;冲突返回 `REVISION_CONFLICT`,不得覆盖 Canonical。
|
||||
4. `merge` 必须以编辑前候选为基线生成严格的 `candidateVersion+1`,正文 hash 必须变化,并用编辑后的版本重新跑完整 detector,再回到 Shadow。
|
||||
5. 实验纯函数通过后只返回提交后副作用意图,固定标明 `canonicalMutationPerformed=false`、`commandIntent.allowed=false` 和 `requiresCanonicalCommit=true`;正式提交层验证 Canonical 提交凭证后才可异步排队章后抽卡。本轮不执行 Canonical 写入或抽取。
|
||||
|
||||
`acceptanceEligible=false` 的诊断/评测候选在第一道门硬拒绝,不能通过修改参数、重跑 fake 或用户确认进入接受链。
|
||||
|
||||
## 既有创作文件采纳
|
||||
|
||||
以下步骤只适用于已经由正式写入层形成的既有创作文件,不是正文实验纯函数的一部分:
|
||||
|
||||
1. `git status` 核对待审清单,向用户复述本次将确认哪些文件(正文?知识卡?规划?),多类混杂时建议分次确认。
|
||||
2. 更新回算字段:章文件「字数」、`设定.md` frontmatter「已确认章数/总字数」、章 frontmatter「依据」(来源归因:写手/大纲版本/上下文回显文件名)。
|
||||
@ -21,7 +35,7 @@ description: 确认=提交。用户裁决后,把创作产出从待审(未提交)
|
||||
## 丢弃 / 部分采纳
|
||||
|
||||
- 丢弃:`git restore <文件>`(新文件用 `git clean` 前先向用户列清单);
|
||||
- 部分采纳:用户改完再走采纳;改后合并的来源归因记 `writer+用户修订`。
|
||||
- 正文候选部分采纳:用户编辑后形成新 candidateVersion,重新跑 detector、Shadow 准入与接受前置检查;不得修改后直接合并。正式写入层完成 Canonical 提交后,来源归因记 `writer+用户修订`。
|
||||
|
||||
## 推送
|
||||
|
||||
|
||||
357
.claude/skills/confirm/scripts/check_writer_acceptance.py
Normal file
357
.claude/skills/confirm/scripts/check_writer_acceptance.py
Normal file
@ -0,0 +1,357 @@
|
||||
#!/usr/bin/env python3
|
||||
"""正文候选 Shadow 准入与用户三决策的纯函数前置检查。"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import pathlib
|
||||
import sys
|
||||
from datetime import datetime
|
||||
from typing import Any, Mapping
|
||||
|
||||
SCRIPT_DIR = pathlib.Path(__file__).resolve().parent
|
||||
READ_CONTEXT_DIR = SCRIPT_DIR.parents[1] / "read-context" / "scripts"
|
||||
if str(READ_CONTEXT_DIR) not in sys.path:
|
||||
sys.path.insert(0, str(READ_CONTEXT_DIR))
|
||||
|
||||
from writer_contract import ( # noqa: E402
|
||||
ContractError,
|
||||
han_count,
|
||||
validate_writer_context,
|
||||
validate_writer_output,
|
||||
)
|
||||
|
||||
|
||||
PRODUCTION_POLICY = "writer-production-v1"
|
||||
ACTIVE_SOURCE_STATUS = "active"
|
||||
DECISIONS = frozenset({"accept", "merge", "discard"})
|
||||
_LIVE_STATE_FIELDS = frozenset(
|
||||
{
|
||||
"qualityPolicyVersion",
|
||||
"contextSnapshotId",
|
||||
"contextSnapshotSha256",
|
||||
"authorizationSnapshotId",
|
||||
"authorizationValid",
|
||||
"sourceStatus",
|
||||
"candidateExpiresAt",
|
||||
"checkedAt",
|
||||
"canonicalRevision",
|
||||
}
|
||||
)
|
||||
|
||||
|
||||
class AcceptanceError(RuntimeError):
|
||||
"""携带稳定失败码的接受前置检查错误,任何错误都不可接受。"""
|
||||
|
||||
def __init__(
|
||||
self, code: str, message: str, *, details: Mapping[str, Any] | None = None
|
||||
) -> None:
|
||||
super().__init__(message)
|
||||
self.code = code
|
||||
self.details = dict(details or {})
|
||||
self.acceptance_eligible = False
|
||||
|
||||
|
||||
def _require_mapping(value: Any, field: str) -> Mapping[str, Any]:
|
||||
"""拒绝非对象输入,避免宽松取值绕过实时检查。"""
|
||||
|
||||
if not isinstance(value, Mapping):
|
||||
raise AcceptanceError("ACCEPTANCE_STATE_INVALID", f"{field} 必须是对象")
|
||||
return value
|
||||
|
||||
|
||||
def _parse_timestamp(value: Any, field: str) -> datetime:
|
||||
"""解析带时区的 ISO-8601 时间;无时区时间失败关闭。"""
|
||||
|
||||
if not isinstance(value, str) or not value:
|
||||
raise AcceptanceError("ACCEPTANCE_STATE_INVALID", f"{field} 必须是时间字符串")
|
||||
try:
|
||||
parsed = datetime.fromisoformat(value.replace("Z", "+00:00"))
|
||||
except ValueError as exc:
|
||||
raise AcceptanceError("ACCEPTANCE_STATE_INVALID", f"{field} 格式非法") from exc
|
||||
if parsed.tzinfo is None or parsed.utcoffset() is None:
|
||||
raise AcceptanceError("ACCEPTANCE_STATE_INVALID", f"{field} 必须带时区")
|
||||
return parsed
|
||||
|
||||
|
||||
def _validate_live_state(value: Any) -> dict[str, Any]:
|
||||
"""严格校验接受时重新读取的可信实时状态。"""
|
||||
|
||||
live = dict(_require_mapping(value, "live_state"))
|
||||
if set(live) != _LIVE_STATE_FIELDS:
|
||||
missing = sorted(_LIVE_STATE_FIELDS - set(live))
|
||||
extra = sorted(set(live) - _LIVE_STATE_FIELDS)
|
||||
raise AcceptanceError(
|
||||
"ACCEPTANCE_STATE_INVALID",
|
||||
"live_state 字段不完整或含未知字段",
|
||||
details={"missing": missing, "extra": extra},
|
||||
)
|
||||
if not isinstance(live["authorizationValid"], bool):
|
||||
raise AcceptanceError(
|
||||
"ACCEPTANCE_STATE_INVALID", "authorizationValid 必须是布尔值"
|
||||
)
|
||||
revision = live["canonicalRevision"]
|
||||
if isinstance(revision, bool) or not isinstance(revision, int) or revision < 0:
|
||||
raise AcceptanceError(
|
||||
"ACCEPTANCE_STATE_INVALID", "canonicalRevision 必须是非负整数"
|
||||
)
|
||||
checked_at = _parse_timestamp(live["checkedAt"], "checkedAt")
|
||||
expires_at = _parse_timestamp(live["candidateExpiresAt"], "candidateExpiresAt")
|
||||
live["_checkedAt"] = checked_at
|
||||
live["_expiresAt"] = expires_at
|
||||
return live
|
||||
|
||||
|
||||
def _validate_detector_result(
|
||||
detector_result: Any, candidate: Mapping[str, Any]
|
||||
) -> None:
|
||||
"""要求最终报告同时证明机械门和语义接口通过并绑定当前候选。"""
|
||||
|
||||
report = _require_mapping(detector_result, "detector_result")
|
||||
if report.get("schemaVersion") != "writer-pipeline-result-v1":
|
||||
raise AcceptanceError(
|
||||
"DETECTOR_REPORT_INVALID", "detector 终态报告版本不受支持"
|
||||
)
|
||||
if report.get("status") != "PASSED" or report.get("failureCode") is not None:
|
||||
raise AcceptanceError("DETECTOR_NOT_PASSED", "detector 尚未得到无阻塞通过终态")
|
||||
expected = {
|
||||
"runId": candidate["runId"],
|
||||
"attempt": candidate["attempt"],
|
||||
"candidateVersion": candidate["candidateVersion"],
|
||||
"candidateSha256": candidate["candidateSha256"],
|
||||
}
|
||||
mismatches = [field for field, value in expected.items() if report.get(field) != value]
|
||||
if mismatches:
|
||||
raise AcceptanceError(
|
||||
"DETECTOR_BINDING_MISMATCH",
|
||||
"detector 终态未绑定当前候选",
|
||||
details={"fields": mismatches},
|
||||
)
|
||||
trace = report.get("trace")
|
||||
if not isinstance(trace, list) or not trace or not isinstance(trace[-1], Mapping):
|
||||
raise AcceptanceError("DETECTOR_REPORT_INVALID", "detector 终态缺少审查轨迹")
|
||||
final_check = trace[-1]
|
||||
final_expected = {
|
||||
"attempt": candidate["attempt"],
|
||||
"candidateVersion": candidate["candidateVersion"],
|
||||
"candidateSha256": candidate["candidateSha256"],
|
||||
}
|
||||
if any(final_check.get(field) != value for field, value in final_expected.items()):
|
||||
raise AcceptanceError(
|
||||
"DETECTOR_BINDING_MISMATCH", "detector 最终审查轨迹绑定了旧候选"
|
||||
)
|
||||
if (
|
||||
final_check.get("mechanicalPassed") is not True
|
||||
or final_check.get("semanticStatus") != "passed"
|
||||
or final_check.get("failureCodes") != []
|
||||
):
|
||||
raise AcceptanceError(
|
||||
"DETECTOR_NOT_PASSED", "候选没有同时通过机械硬门和语义审查接口"
|
||||
)
|
||||
|
||||
|
||||
def _validate_context_binding(
|
||||
context: Mapping[str, Any], candidate: Mapping[str, Any]
|
||||
) -> None:
|
||||
"""保证候选没有逃离本次冻结上下文、运行身份和生产策略。"""
|
||||
|
||||
expected = {
|
||||
"runId": context["runId"],
|
||||
"attempt": context["attempt"],
|
||||
"mode": context["mode"],
|
||||
"qualityPolicyVersion": context["qualityPolicyVersion"],
|
||||
"contextSnapshotId": context["contextSnapshot"]["manifestId"],
|
||||
"contextSnapshotSha256": context["contextSnapshot"]["contextSha256"],
|
||||
"acceptanceEligible": context["acceptanceEligible"],
|
||||
}
|
||||
mismatches = [field for field, value in expected.items() if candidate.get(field) != value]
|
||||
if mismatches:
|
||||
raise AcceptanceError(
|
||||
"CONTEXT_BINDING_MISMATCH",
|
||||
"候选未绑定当前 WriterContext",
|
||||
details={"fields": mismatches},
|
||||
)
|
||||
|
||||
|
||||
def check_shadow_ready(
|
||||
*,
|
||||
context: Mapping[str, Any],
|
||||
candidate: Mapping[str, Any],
|
||||
detector_result: Mapping[str, Any],
|
||||
live_state: Mapping[str, Any],
|
||||
) -> dict[str, Any]:
|
||||
"""纯函数校验候选是否可进入 Shadow 待接受态,不执行任何写入。"""
|
||||
|
||||
raw_candidate = _require_mapping(candidate, "candidate")
|
||||
if raw_candidate.get("acceptanceEligible") is not True:
|
||||
raise AcceptanceError(
|
||||
"ACCEPTANCE_NOT_ELIGIBLE", "评测、诊断或显式不可接受候选不得进入接受链"
|
||||
)
|
||||
try:
|
||||
normalized_context = validate_writer_context(context)
|
||||
except ContractError as exc:
|
||||
raise AcceptanceError("CONTEXT_CONTRACT_INVALID", str(exc)) from exc
|
||||
try:
|
||||
normalized_candidate = validate_writer_output(candidate)
|
||||
except ContractError as exc:
|
||||
raise AcceptanceError("CANDIDATE_CONTRACT_INVALID", str(exc)) from exc
|
||||
if (
|
||||
normalized_context["mode"] != "production"
|
||||
or normalized_candidate["mode"] != "production"
|
||||
):
|
||||
raise AcceptanceError("ACCEPTANCE_NOT_ELIGIBLE", "只有生产候选可进入接受链")
|
||||
_validate_context_binding(normalized_context, normalized_candidate)
|
||||
contract = normalized_context["outputContract"]
|
||||
actual_han_chars = han_count(normalized_candidate["candidateBody"])
|
||||
if not contract["minChars"] <= actual_han_chars <= contract["maxChars"]:
|
||||
raise AcceptanceError(
|
||||
"CANDIDATE_LENGTH_OUT_OF_RANGE",
|
||||
"候选正文汉字数超出动态篇幅合同",
|
||||
details={
|
||||
"actualHanChars": actual_han_chars,
|
||||
"minChars": contract["minChars"],
|
||||
"maxChars": contract["maxChars"],
|
||||
"targetChars": contract["targetChars"],
|
||||
},
|
||||
)
|
||||
live = _validate_live_state(live_state)
|
||||
if (
|
||||
normalized_context["qualityPolicyVersion"] != PRODUCTION_POLICY
|
||||
or normalized_candidate["qualityPolicyVersion"] != PRODUCTION_POLICY
|
||||
or live["qualityPolicyVersion"] != PRODUCTION_POLICY
|
||||
):
|
||||
raise AcceptanceError("QUALITY_POLICY_STALE", "生产质量策略已变化或绑定错误")
|
||||
if (
|
||||
live["contextSnapshotId"] != normalized_candidate["contextSnapshotId"]
|
||||
or live["contextSnapshotSha256"]
|
||||
!= normalized_candidate["contextSnapshotSha256"]
|
||||
):
|
||||
raise AcceptanceError("CONTEXT_STALE", "冻结上下文已失效或哈希变化")
|
||||
if (
|
||||
live["authorizationSnapshotId"]
|
||||
!= normalized_context["authorizationSnapshot"]["snapshotId"]
|
||||
or live["authorizationValid"] is not True
|
||||
):
|
||||
raise AcceptanceError("AUTHORIZATION_STALE", "授权快照已失效或变化")
|
||||
if (
|
||||
live["sourceStatus"] != ACTIVE_SOURCE_STATUS
|
||||
or normalized_context["sourceStatus"] != ACTIVE_SOURCE_STATUS
|
||||
):
|
||||
raise AcceptanceError("SOURCE_STALE", "来源状态已不允许接受")
|
||||
if live["_checkedAt"] >= live["_expiresAt"]:
|
||||
raise AcceptanceError("CANDIDATE_EXPIRED", "候选接受窗口已过期")
|
||||
_validate_detector_result(detector_result, normalized_candidate)
|
||||
return {
|
||||
"schemaVersion": "writer-acceptance-result-v1",
|
||||
"status": "SHADOW_READY",
|
||||
"runId": normalized_candidate["runId"],
|
||||
"candidateVersion": normalized_candidate["candidateVersion"],
|
||||
"candidateSha256": normalized_candidate["candidateSha256"],
|
||||
"canonicalRevision": live["canonicalRevision"],
|
||||
"canonicalMutationPerformed": False,
|
||||
}
|
||||
|
||||
|
||||
def _validate_edited_candidate(
|
||||
previous_candidate: Any, candidate: Mapping[str, Any]
|
||||
) -> None:
|
||||
"""要求用户编辑形成严格下一版本,并保持同一冻结运行身份。"""
|
||||
|
||||
if previous_candidate is None:
|
||||
raise AcceptanceError("EDIT_BASE_REQUIRED", "修改后合并必须提供编辑前候选")
|
||||
try:
|
||||
previous = validate_writer_output(previous_candidate)
|
||||
except ContractError as exc:
|
||||
raise AcceptanceError("EDIT_BASE_INVALID", str(exc)) from exc
|
||||
if candidate["candidateVersion"] != previous["candidateVersion"] + 1:
|
||||
raise AcceptanceError(
|
||||
"EDIT_VERSION_INVALID", "用户编辑必须生成 candidateVersion 的严格下一版本"
|
||||
)
|
||||
identity_fields = (
|
||||
"runId",
|
||||
"attempt",
|
||||
"mode",
|
||||
"qualityPolicyVersion",
|
||||
"contextSnapshotId",
|
||||
"contextSnapshotSha256",
|
||||
"acceptanceEligible",
|
||||
)
|
||||
if any(candidate[field] != previous[field] for field in identity_fields):
|
||||
raise AcceptanceError("EDIT_BASE_MISMATCH", "编辑候选改变了冻结运行身份")
|
||||
if candidate["candidateSha256"] == previous["candidateSha256"]:
|
||||
raise AcceptanceError("EDIT_BODY_UNCHANGED", "修改后合并必须包含实际正文变更")
|
||||
|
||||
|
||||
def check_writer_acceptance(
|
||||
*,
|
||||
decision: str,
|
||||
confirmed: bool,
|
||||
context: Mapping[str, Any],
|
||||
candidate: Mapping[str, Any],
|
||||
detector_result: Mapping[str, Any],
|
||||
live_state: Mapping[str, Any],
|
||||
expected_revision: int,
|
||||
previous_candidate: Mapping[str, Any] | None = None,
|
||||
) -> dict[str, Any]:
|
||||
"""校验用户三决策并返回命令意图;实验台不写 Canonical。"""
|
||||
|
||||
if decision not in DECISIONS:
|
||||
raise AcceptanceError("DECISION_INVALID", "decision 只允许 accept、merge 或 discard")
|
||||
if confirmed is not True:
|
||||
raise AcceptanceError("CONFIRMATION_REQUIRED", f"{decision} 必须由用户明确确认")
|
||||
if decision == "discard":
|
||||
return {
|
||||
"schemaVersion": "writer-acceptance-result-v1",
|
||||
"status": "DISCARD_INTENT_READY",
|
||||
"decision": decision,
|
||||
"commandIntent": None,
|
||||
"canonicalMutationPerformed": False,
|
||||
}
|
||||
shadow = check_shadow_ready(
|
||||
context=context,
|
||||
candidate=candidate,
|
||||
detector_result=detector_result,
|
||||
live_state=live_state,
|
||||
)
|
||||
if isinstance(expected_revision, bool) or not isinstance(expected_revision, int):
|
||||
raise AcceptanceError("EXPECTED_REVISION_INVALID", "expectedRevision 必须是整数")
|
||||
if expected_revision != shadow["canonicalRevision"]:
|
||||
raise AcceptanceError(
|
||||
"REVISION_CONFLICT",
|
||||
"expectedRevision 与当前 Canonical revision 不一致",
|
||||
details={
|
||||
"expectedRevision": expected_revision,
|
||||
"canonicalRevision": shadow["canonicalRevision"],
|
||||
},
|
||||
)
|
||||
normalized_candidate = validate_writer_output(candidate)
|
||||
if decision == "merge":
|
||||
_validate_edited_candidate(previous_candidate, normalized_candidate)
|
||||
# detector 绑定已由 check_shadow_ready 针对编辑后的新版本重新校验。
|
||||
return {
|
||||
"schemaVersion": "writer-acceptance-result-v1",
|
||||
"status": "ACCEPTANCE_INTENT_READY",
|
||||
"decision": decision,
|
||||
"runId": normalized_candidate["runId"],
|
||||
"candidateVersion": normalized_candidate["candidateVersion"],
|
||||
"candidateSha256": normalized_candidate["candidateSha256"],
|
||||
"expectedRevision": expected_revision,
|
||||
"canonicalMutationPerformed": False,
|
||||
"commandIntent": {
|
||||
"command": "queue_chapter_extraction",
|
||||
"dispatch": "async",
|
||||
"executeAfter": "canonical_commit",
|
||||
# 本函数只是 preflight,尚无 Canonical 提交凭证;正式提交层验证凭证后才能放行。
|
||||
"allowed": False,
|
||||
"requiresCanonicalCommit": True,
|
||||
"workId": context["workId"],
|
||||
"chapter": context["targetChapter"],
|
||||
"candidateSha256": normalized_candidate["candidateSha256"],
|
||||
},
|
||||
}
|
||||
|
||||
|
||||
__all__ = [
|
||||
"AcceptanceError",
|
||||
"check_shadow_ready",
|
||||
"check_writer_acceptance",
|
||||
]
|
||||
419
.claude/skills/confirm/scripts/test_writer_acceptance.py
Normal file
419
.claude/skills/confirm/scripts/test_writer_acceptance.py
Normal file
@ -0,0 +1,419 @@
|
||||
#!/usr/bin/env python3
|
||||
"""正文候选 Shadow 准入、用户三决策与接受前置检查测试。"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import copy
|
||||
import hashlib
|
||||
import pathlib
|
||||
import sys
|
||||
import unittest
|
||||
|
||||
SCRIPT_DIR = pathlib.Path(__file__).resolve().parent
|
||||
SKILLS_DIR = SCRIPT_DIR.parents[1]
|
||||
CONTINUATION_DIR = SKILLS_DIR / "continuation" / "scripts"
|
||||
DETECT_DIR = SKILLS_DIR / "detect" / "scripts"
|
||||
READ_CONTEXT_DIR = SKILLS_DIR / "read-context" / "scripts"
|
||||
for path in (SCRIPT_DIR, CONTINUATION_DIR, DETECT_DIR, READ_CONTEXT_DIR):
|
||||
if str(path) not in sys.path:
|
||||
sys.path.insert(0, str(path))
|
||||
|
||||
from test_check_writer_candidate import _valid_pair # noqa: E402
|
||||
from test_check_writer_candidate import _requirements # noqa: E402
|
||||
from test_run_writer_pipeline import _semantic_pass, _writer_output # noqa: E402
|
||||
from check_writer_acceptance import ( # noqa: E402
|
||||
AcceptanceError,
|
||||
check_shadow_ready,
|
||||
check_writer_acceptance,
|
||||
)
|
||||
from run_writer_pipeline import InMemoryCasStateStore, run_writer_pipeline # noqa: E402
|
||||
|
||||
|
||||
def _detector_result(candidate: dict) -> dict:
|
||||
"""构造同时通过机械门和语义接口的最终 detector 结果。"""
|
||||
|
||||
return {
|
||||
"schemaVersion": "writer-pipeline-result-v1",
|
||||
"runId": candidate["runId"],
|
||||
"status": "PASSED",
|
||||
"attempt": candidate["attempt"],
|
||||
"candidateVersion": candidate["candidateVersion"],
|
||||
"candidateSha256": candidate["candidateSha256"],
|
||||
"failureCode": None,
|
||||
"evidenceRequestCount": 0,
|
||||
"rewriteCount": 0,
|
||||
"trace": [
|
||||
{
|
||||
"attempt": candidate["attempt"],
|
||||
"candidateVersion": candidate["candidateVersion"],
|
||||
"candidateSha256": candidate["candidateSha256"],
|
||||
"mechanicalPassed": True,
|
||||
"semanticStatus": "passed",
|
||||
"failureCodes": [],
|
||||
}
|
||||
],
|
||||
}
|
||||
|
||||
|
||||
def _live_state(context: dict) -> dict:
|
||||
"""构造接受时重新读取的可信实时状态。"""
|
||||
|
||||
return {
|
||||
"qualityPolicyVersion": context["qualityPolicyVersion"],
|
||||
"contextSnapshotId": context["contextSnapshot"]["manifestId"],
|
||||
"contextSnapshotSha256": context["contextSnapshot"]["contextSha256"],
|
||||
"authorizationSnapshotId": context["authorizationSnapshot"]["snapshotId"],
|
||||
"authorizationValid": True,
|
||||
"sourceStatus": "active",
|
||||
"candidateExpiresAt": "2026-07-20T12:00:00Z",
|
||||
"checkedAt": "2026-07-20T11:00:00Z",
|
||||
"canonicalRevision": 7,
|
||||
}
|
||||
|
||||
|
||||
def _valid_inputs() -> tuple[dict, dict, dict, dict]:
|
||||
"""返回合法上下文、候选、detector 终态和实时状态。"""
|
||||
|
||||
context, candidate = _valid_pair()
|
||||
return context, candidate, _detector_result(candidate), _live_state(context)
|
||||
|
||||
|
||||
def _edited_candidate(candidate: dict) -> dict:
|
||||
"""模拟用户编辑后生成严格递增、重新哈希的新候选。"""
|
||||
|
||||
edited = copy.deepcopy(candidate)
|
||||
edited["candidateVersion"] += 1
|
||||
edited["candidateBody"] = "章" + edited["candidateBody"][1:]
|
||||
digest = "sha256:" + hashlib.sha256(edited["candidateBody"].encode("utf-8")).hexdigest()
|
||||
edited["candidateSha256"] = digest
|
||||
for claim in edited["claimLedger"]:
|
||||
claim["candidateSha256"] = digest
|
||||
return edited
|
||||
|
||||
|
||||
class WriterAcceptanceTest(unittest.TestCase):
|
||||
def _assert_error(self, expected_code: str, function, **kwargs) -> None:
|
||||
"""断言纯函数失败关闭并返回稳定错误码。"""
|
||||
|
||||
with self.assertRaises(AcceptanceError) as raised:
|
||||
function(**kwargs)
|
||||
self.assertEqual(raised.exception.code, expected_code)
|
||||
self.assertFalse(raised.exception.acceptance_eligible)
|
||||
|
||||
def test_acceptance_eligible_false_is_hard_rejected(self):
|
||||
context, candidate, detector, live = _valid_inputs()
|
||||
candidate["acceptanceEligible"] = False
|
||||
|
||||
self._assert_error(
|
||||
"ACCEPTANCE_NOT_ELIGIBLE",
|
||||
check_shadow_ready,
|
||||
context=context,
|
||||
candidate=candidate,
|
||||
detector_result=detector,
|
||||
live_state=live,
|
||||
)
|
||||
|
||||
def test_shadow_requires_detector_context_hash_policy_authorization_and_freshness(self):
|
||||
context, candidate, detector, live = _valid_inputs()
|
||||
cases = []
|
||||
|
||||
rejected_detector = copy.deepcopy(detector)
|
||||
rejected_detector["status"] = "REJECTED"
|
||||
rejected_detector["failureCode"] = "REWRITE_LIMIT_REACHED"
|
||||
cases.append(("DETECTOR_NOT_PASSED", candidate, rejected_detector, live))
|
||||
|
||||
wrong_detector_schema = copy.deepcopy(detector)
|
||||
wrong_detector_schema["schemaVersion"] = "writer-detector-report-v1"
|
||||
cases.append(("DETECTOR_REPORT_INVALID", candidate, wrong_detector_schema, live))
|
||||
|
||||
stale_detector = copy.deepcopy(detector)
|
||||
stale_detector["candidateSha256"] = "sha256:" + "f" * 64
|
||||
cases.append(("DETECTOR_BINDING_MISMATCH", candidate, stale_detector, live))
|
||||
|
||||
mechanical_red = copy.deepcopy(detector)
|
||||
mechanical_red["trace"][-1]["mechanicalPassed"] = False
|
||||
cases.append(("DETECTOR_NOT_PASSED", candidate, mechanical_red, live))
|
||||
|
||||
semantic_pending = copy.deepcopy(detector)
|
||||
semantic_pending["trace"][-1]["semanticStatus"] = "pending"
|
||||
cases.append(("DETECTOR_NOT_PASSED", candidate, semantic_pending, live))
|
||||
|
||||
trace_has_failure = copy.deepcopy(detector)
|
||||
trace_has_failure["trace"][-1]["failureCodes"] = ["HARD_EVENT_MISSING"]
|
||||
cases.append(("DETECTOR_NOT_PASSED", candidate, trace_has_failure, live))
|
||||
|
||||
stale_trace = copy.deepcopy(detector)
|
||||
stale_trace["trace"][-1]["candidateVersion"] = 0
|
||||
cases.append(("DETECTOR_BINDING_MISMATCH", candidate, stale_trace, live))
|
||||
|
||||
bad_hash = copy.deepcopy(candidate)
|
||||
bad_hash["candidateSha256"] = "sha256:" + "e" * 64
|
||||
cases.append(("CANDIDATE_CONTRACT_INVALID", bad_hash, detector, live))
|
||||
|
||||
too_short = copy.deepcopy(candidate)
|
||||
too_short["candidateBody"] = "文" * 1000
|
||||
digest = "sha256:" + hashlib.sha256(too_short["candidateBody"].encode("utf-8")).hexdigest()
|
||||
too_short["candidateSha256"] = digest
|
||||
for claim in too_short["claimLedger"]:
|
||||
claim["candidateSha256"] = digest
|
||||
cases.append(
|
||||
(
|
||||
"CANDIDATE_LENGTH_OUT_OF_RANGE",
|
||||
too_short,
|
||||
_detector_result(too_short),
|
||||
live,
|
||||
)
|
||||
)
|
||||
|
||||
stale_context = copy.deepcopy(live)
|
||||
stale_context["contextSnapshotSha256"] = "sha256:" + "d" * 64
|
||||
cases.append(("CONTEXT_STALE", candidate, detector, stale_context))
|
||||
|
||||
stale_policy = copy.deepcopy(live)
|
||||
stale_policy["qualityPolicyVersion"] = "writer-production-v2"
|
||||
cases.append(("QUALITY_POLICY_STALE", candidate, detector, stale_policy))
|
||||
|
||||
stale_authorization = copy.deepcopy(live)
|
||||
stale_authorization["authorizationValid"] = False
|
||||
cases.append(("AUTHORIZATION_STALE", candidate, detector, stale_authorization))
|
||||
|
||||
stale_source = copy.deepcopy(live)
|
||||
stale_source["sourceStatus"] = "revoked"
|
||||
cases.append(("SOURCE_STALE", candidate, detector, stale_source))
|
||||
|
||||
expired = copy.deepcopy(live)
|
||||
expired["checkedAt"] = expired["candidateExpiresAt"]
|
||||
cases.append(("CANDIDATE_EXPIRED", candidate, detector, expired))
|
||||
|
||||
for expected_code, current_candidate, current_detector, current_live in cases:
|
||||
with self.subTest(expected_code=expected_code):
|
||||
self._assert_error(
|
||||
expected_code,
|
||||
check_shadow_ready,
|
||||
context=copy.deepcopy(context),
|
||||
candidate=copy.deepcopy(current_candidate),
|
||||
detector_result=copy.deepcopy(current_detector),
|
||||
live_state=copy.deepcopy(current_live),
|
||||
)
|
||||
|
||||
def test_valid_candidate_enters_shadow_without_writing_canonical(self):
|
||||
context, candidate, detector, live = _valid_inputs()
|
||||
|
||||
result = check_shadow_ready(
|
||||
context=context,
|
||||
candidate=candidate,
|
||||
detector_result=detector,
|
||||
live_state=live,
|
||||
)
|
||||
|
||||
self.assertEqual(result["status"], "SHADOW_READY")
|
||||
self.assertEqual(result["candidateVersion"], candidate["candidateVersion"])
|
||||
self.assertFalse(result["canonicalMutationPerformed"])
|
||||
|
||||
def test_pipeline_passed_result_is_accepted_by_shadow_contract(self):
|
||||
"""直接验证 Task 5 生产结果与 Task 6 消费合同一致。"""
|
||||
|
||||
context, candidate, _detector, live = _valid_inputs()
|
||||
pipeline_result = run_writer_pipeline(
|
||||
context=context,
|
||||
requirements=_requirements(),
|
||||
writer=lambda current, version, _repair: _writer_output(current, version),
|
||||
evidence_provider=lambda *_args: self.fail("通过候选不得补证"),
|
||||
semantic_detector=_semantic_pass,
|
||||
state_store=InMemoryCasStateStore(),
|
||||
)
|
||||
pipeline_candidate = pipeline_result["candidateArtifact"]
|
||||
|
||||
result = check_shadow_ready(
|
||||
context=context,
|
||||
candidate=pipeline_candidate,
|
||||
detector_result=pipeline_result,
|
||||
live_state=live,
|
||||
)
|
||||
|
||||
self.assertEqual(pipeline_candidate, candidate)
|
||||
self.assertEqual(
|
||||
pipeline_result["trace"][-1]["mechanicalReport"]["schemaVersion"],
|
||||
"writer-detector-report-v1",
|
||||
)
|
||||
self.assertEqual(
|
||||
pipeline_result["trace"][-1]["semanticReport"]["status"], "passed"
|
||||
)
|
||||
self.assertEqual(result["status"], "SHADOW_READY")
|
||||
|
||||
def test_accept_merge_and_discard_all_require_explicit_confirmation(self):
|
||||
context, candidate, detector, live = _valid_inputs()
|
||||
for decision in ("accept", "merge", "discard"):
|
||||
kwargs = {
|
||||
"decision": decision,
|
||||
"confirmed": False,
|
||||
"context": context,
|
||||
"candidate": candidate,
|
||||
"detector_result": detector,
|
||||
"live_state": live,
|
||||
"expected_revision": 7,
|
||||
}
|
||||
if decision == "merge":
|
||||
edited = _edited_candidate(candidate)
|
||||
kwargs["candidate"] = edited
|
||||
kwargs["detector_result"] = _detector_result(edited)
|
||||
kwargs["previous_candidate"] = candidate
|
||||
with self.subTest(decision=decision):
|
||||
self._assert_error("CONFIRMATION_REQUIRED", check_writer_acceptance, **kwargs)
|
||||
|
||||
def test_expected_revision_conflict_never_returns_command_intent(self):
|
||||
context, candidate, detector, live = _valid_inputs()
|
||||
|
||||
self._assert_error(
|
||||
"REVISION_CONFLICT",
|
||||
check_writer_acceptance,
|
||||
decision="accept",
|
||||
confirmed=True,
|
||||
context=context,
|
||||
candidate=candidate,
|
||||
detector_result=detector,
|
||||
live_state=live,
|
||||
expected_revision=6,
|
||||
)
|
||||
|
||||
def test_merge_requires_next_candidate_version_and_rerun_detector(self):
|
||||
context, candidate, _detector, live = _valid_inputs()
|
||||
edited = _edited_candidate(candidate)
|
||||
|
||||
same_version = copy.deepcopy(edited)
|
||||
same_version["candidateVersion"] = candidate["candidateVersion"]
|
||||
self._assert_error(
|
||||
"EDIT_VERSION_INVALID",
|
||||
check_writer_acceptance,
|
||||
decision="merge",
|
||||
confirmed=True,
|
||||
context=context,
|
||||
candidate=same_version,
|
||||
previous_candidate=candidate,
|
||||
detector_result=_detector_result(same_version),
|
||||
live_state=live,
|
||||
expected_revision=7,
|
||||
)
|
||||
|
||||
stale_detector = _detector_result(candidate)
|
||||
self._assert_error(
|
||||
"DETECTOR_BINDING_MISMATCH",
|
||||
check_writer_acceptance,
|
||||
decision="merge",
|
||||
confirmed=True,
|
||||
context=context,
|
||||
candidate=edited,
|
||||
previous_candidate=candidate,
|
||||
detector_result=stale_detector,
|
||||
live_state=live,
|
||||
expected_revision=7,
|
||||
)
|
||||
|
||||
def test_merge_rejects_version_jump_unchanged_body_and_invalid_base(self):
|
||||
"""编辑候选必须是有实际变更的严格下一版本,并绑定合法基线。"""
|
||||
|
||||
context, candidate, _detector, live = _valid_inputs()
|
||||
|
||||
jumped = _edited_candidate(candidate)
|
||||
jumped["candidateVersion"] += 1
|
||||
unchanged = copy.deepcopy(candidate)
|
||||
unchanged["candidateVersion"] += 1
|
||||
invalid_base = copy.deepcopy(candidate)
|
||||
del invalid_base["selfCheck"]
|
||||
cases = (
|
||||
("EDIT_VERSION_INVALID", jumped, candidate),
|
||||
("EDIT_BODY_UNCHANGED", unchanged, candidate),
|
||||
("EDIT_BASE_REQUIRED", _edited_candidate(candidate), None),
|
||||
("EDIT_BASE_INVALID", _edited_candidate(candidate), invalid_base),
|
||||
)
|
||||
for expected_code, current, previous in cases:
|
||||
with self.subTest(expected_code=expected_code):
|
||||
self._assert_error(
|
||||
expected_code,
|
||||
check_writer_acceptance,
|
||||
decision="merge",
|
||||
confirmed=True,
|
||||
context=context,
|
||||
candidate=current,
|
||||
previous_candidate=previous,
|
||||
detector_result=_detector_result(current),
|
||||
live_state=live,
|
||||
expected_revision=7,
|
||||
)
|
||||
|
||||
def test_accept_and_merge_only_return_async_extraction_command_intent(self):
|
||||
context, candidate, detector, live = _valid_inputs()
|
||||
edited = _edited_candidate(candidate)
|
||||
cases = (
|
||||
("accept", candidate, detector, None),
|
||||
("merge", edited, _detector_result(edited), candidate),
|
||||
)
|
||||
for decision, current, current_detector, previous in cases:
|
||||
with self.subTest(decision=decision):
|
||||
result = check_writer_acceptance(
|
||||
decision=decision,
|
||||
confirmed=True,
|
||||
context=context,
|
||||
candidate=current,
|
||||
previous_candidate=previous,
|
||||
detector_result=current_detector,
|
||||
live_state=live,
|
||||
expected_revision=7,
|
||||
)
|
||||
self.assertEqual(result["status"], "ACCEPTANCE_INTENT_READY")
|
||||
self.assertFalse(result["canonicalMutationPerformed"])
|
||||
self.assertEqual(result["commandIntent"]["command"], "queue_chapter_extraction")
|
||||
self.assertEqual(result["commandIntent"]["dispatch"], "async")
|
||||
self.assertEqual(result["commandIntent"]["executeAfter"], "canonical_commit")
|
||||
self.assertFalse(result["commandIntent"]["allowed"])
|
||||
self.assertTrue(result["commandIntent"]["requiresCanonicalCommit"])
|
||||
|
||||
def test_discard_returns_no_extraction_intent_and_performs_no_write(self):
|
||||
context, candidate, detector, live = _valid_inputs()
|
||||
|
||||
result = check_writer_acceptance(
|
||||
decision="discard",
|
||||
confirmed=True,
|
||||
context=context,
|
||||
candidate=candidate,
|
||||
detector_result=detector,
|
||||
live_state=live,
|
||||
expected_revision=7,
|
||||
)
|
||||
|
||||
self.assertEqual(result["status"], "DISCARD_INTENT_READY")
|
||||
self.assertEqual(result["runId"], candidate["runId"])
|
||||
self.assertEqual(result["candidateVersion"], candidate["candidateVersion"])
|
||||
self.assertEqual(result["candidateSha256"], candidate["candidateSha256"])
|
||||
self.assertEqual(result["commandIntent"]["command"], "close_shadow_candidate")
|
||||
self.assertEqual(
|
||||
result["commandIntent"]["expectedCandidateVersion"],
|
||||
candidate["candidateVersion"],
|
||||
)
|
||||
self.assertEqual(
|
||||
result["commandIntent"]["expectedCandidateSha256"],
|
||||
candidate["candidateSha256"],
|
||||
)
|
||||
self.assertFalse(result["canonicalMutationPerformed"])
|
||||
|
||||
def test_discard_rejects_malformed_candidate_instead_of_closing_unknown_target(self):
|
||||
"""丢弃也必须绑定合法候选身份,不能返回无目标关闭意图。"""
|
||||
|
||||
context, candidate, detector, live = _valid_inputs()
|
||||
del candidate["selfCheck"]
|
||||
|
||||
self._assert_error(
|
||||
"CANDIDATE_CONTRACT_INVALID",
|
||||
check_writer_acceptance,
|
||||
decision="discard",
|
||||
confirmed=True,
|
||||
context=context,
|
||||
candidate=candidate,
|
||||
detector_result=detector,
|
||||
live_state=live,
|
||||
expected_revision=7,
|
||||
)
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
unittest.main()
|
||||
@ -10,24 +10,28 @@ disable-model-invocation: true
|
||||
|
||||
## 元数据驱动(schema 怎么控制本功能)
|
||||
|
||||
- **章结构=字段合同**:frontmatter 按 `meta/schemas/chapter.yaml` 逐字段产出(章号/标题/本章目标/前情衔接/出场角色/场景列表/伏笔动作/章末钩子),场景列表逐项按 `scene.yaml`。**schema 加一个字段,frontmatter 立刻多一项,本 skill 与 writer 一字不改**(元数据驱动验收判据)。
|
||||
- **章结构=字段合同**:`candidateBody` 只承载正文,不包含 frontmatter;未来正式写入层在 Canonical 提交时依据 `meta/schemas/chapter.yaml` 与 `scene.yaml` 组装章结构,本轮不实现该写入层。
|
||||
- **行为约束=字段值**:文风(style 实例逐条)、人物言行(character.行事逻辑/说话方式/知情范围)、能力边界(power_system.代价限制)、地点规则(location.规则特例)——全部从上下文包字段读,不硬编。
|
||||
|
||||
## 输入(read-context 的 generation 视图)
|
||||
## 输入(`WriterContext v1`)
|
||||
|
||||
章级基线包(底牌已裁)+ 本章细纲。包里没有的设定不存在。
|
||||
adapter 只通过 stdin 传入已冻结的 `WriterContext v1`:连续前四章全文基线、卡索引回读原文、事实证据、文风证据、大纲定位和本章细纲。writer 不得自行检索;包里没有的设定不存在。
|
||||
|
||||
## 功能约束
|
||||
|
||||
1. 一次一整章,正文 2500–3500 字;场景以「## 场景N」分节,不另拆文件。
|
||||
1. 一次一整章;篇幅以 `outputContract.targetChars/minChars/maxChars` 为唯一口径,由目标章之前有效 Canonical 章长中位数和细纲密度确定性计算,并受 2000–10000 汉字硬边界约束;不得再使用固定字数范围。
|
||||
2. 伏笔只按细纲动作执行:说埋就埋、说推就推、说收就收;不擅自提前回收,不新开大坑。
|
||||
3. 前情衔接与上一章末场景无缝;章末钩子按文风画像的钩子风格。
|
||||
4. 缺细纲时停下报缺,不自编细纲(那是 planning 的活)。
|
||||
5. 事实断言逐项登记 `claimLedger`:绑定当前候选 hash、Unicode code point 左闭右开偏移、事实类型和事实证据;文风依据可另绑 `proseEvidenceId`。
|
||||
6. 证据不足时只填 `evidenceRequests`,最多由外层循环处理 3 次;新增 Canonical 事实必须填 `newSettingDeclarations`,不得静默补设定。
|
||||
|
||||
动态篇幅当前入口合同:`scripts/run_writer.py` 中的 `calculate_dynamic_output_contract()` 已实现并有纯函数测试;同一脚本的 `run_writer()` 只消费并校验上游写入 `WriterContext.outputContract` 的冻结结果。`scripts/run_writer_pipeline.py` 负责有限补证、重写、机械门、语义接口和 CAS 终态。当前任务没有修改 Task 3 上下文组装器,也没有实现 Task 8 真实回放编排调用;在上游显式调用动态篇幅函数并冻结合同前,不得声称端到端动态篇幅已接通。
|
||||
|
||||
## 输出合同
|
||||
|
||||
`works/<书>/manuscript/第NNN章-标题.md` 新文件,不提交。完成后报:依据卡清单/伏笔动作执行情况/新设定申报清单/自查结论。
|
||||
只返回严格 `WriterOutput v1` JSON,不写文件。正文放 `candidateBody`;同时返回候选版本/hash、`claimLedger`、`evidenceRequests`、`newSettingDeclarations` 和 `selfCheck`。CLI adapter 校验全部字段、上下文绑定、动态篇幅与证据引用后只形成待检测候选;机械门与语义 detector 最终通过后才进入 Shadow。
|
||||
|
||||
## 红线
|
||||
|
||||
不改既有章文件与规划文件;新设定走章末申报块(见 writer 身份纪律),不写成既定事实。
|
||||
不调用工具,不读写文件,不持久化会话;细纲硬骨架不得删除、反转或提前回收;新设定只申报,不冒充已确认事实。
|
||||
|
||||
267
.claude/skills/continuation/scripts/run_writer.py
Normal file
267
.claude/skills/continuation/scripts/run_writer.py
Normal file
@ -0,0 +1,267 @@
|
||||
#!/usr/bin/env python3
|
||||
"""以无工具、无会话方式运行正文写手,并严格校验 WriterOutput v1。"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import json
|
||||
import pathlib
|
||||
import subprocess
|
||||
import sys
|
||||
from decimal import Decimal, ROUND_HALF_UP
|
||||
from typing import Any, Callable, Mapping, Sequence
|
||||
|
||||
SCRIPT_DIR = pathlib.Path(__file__).resolve().parent
|
||||
READ_CONTEXT_DIR = SCRIPT_DIR.parents[1] / "read-context" / "scripts"
|
||||
if str(READ_CONTEXT_DIR) not in sys.path:
|
||||
sys.path.insert(0, str(READ_CONTEXT_DIR))
|
||||
|
||||
from writer_contract import ( # noqa: E402
|
||||
ContractError,
|
||||
calculate_target_chars,
|
||||
canonical_json,
|
||||
han_count,
|
||||
validate_writer_context,
|
||||
validate_writer_output,
|
||||
)
|
||||
|
||||
|
||||
class WriterAdapterError(RuntimeError):
|
||||
"""携带稳定失败码的写手 adapter 错误,所有错误都不可接受。"""
|
||||
|
||||
def __init__(self, code: str, message: str, *, details: Mapping[str, Any] | None = None):
|
||||
super().__init__(message)
|
||||
self.code = code
|
||||
self.details = dict(details or {})
|
||||
self.acceptance_eligible = False
|
||||
|
||||
|
||||
def _array_length(value: Any, field: str) -> int:
|
||||
"""读取细纲数组长度;错误类型失败关闭,避免密度被静默低估。"""
|
||||
|
||||
if value is None:
|
||||
return 0
|
||||
if not isinstance(value, list) or any(not isinstance(item, (str, Mapping)) for item in value):
|
||||
raise WriterAdapterError("dynamic_length_input_invalid", f"{field} 必须是字符串或对象数组")
|
||||
return len(value)
|
||||
|
||||
|
||||
def _round_half_up(value: Decimal) -> int:
|
||||
"""以十进制半入规则计算篇幅区间端点。"""
|
||||
|
||||
return int(value.quantize(Decimal("1"), rounding=ROUND_HALF_UP))
|
||||
|
||||
|
||||
def calculate_dynamic_output_contract(
|
||||
*,
|
||||
fine_outline: Mapping[str, Any],
|
||||
recent_chapter_bodies: Sequence[str],
|
||||
default_target_chars: int = 4000,
|
||||
hard_min_chars: int = 2000,
|
||||
hard_max_chars: int = 10000,
|
||||
) -> dict[str, Any]:
|
||||
"""按细纲密度和冻结历史中位章长计算确定性输出篇幅合同。
|
||||
|
||||
目标值复用 WriterContext 合同的唯一算法。允许区间固定为目标值上下 10%,
|
||||
端点按十进制半入取整后再受 2000-10000 的硬边界限制。
|
||||
"""
|
||||
|
||||
if not isinstance(fine_outline, Mapping):
|
||||
raise WriterAdapterError("dynamic_length_input_invalid", "fine_outline 必须是对象")
|
||||
if isinstance(recent_chapter_bodies, (str, bytes)):
|
||||
raise WriterAdapterError("dynamic_length_input_invalid", "recent_chapter_bodies 必须是正文数组")
|
||||
counts: list[int] = []
|
||||
for index, body in enumerate(recent_chapter_bodies):
|
||||
if not isinstance(body, str):
|
||||
raise WriterAdapterError(
|
||||
"dynamic_length_input_invalid",
|
||||
f"recent_chapter_bodies[{index}] 必须是字符串",
|
||||
)
|
||||
count = han_count(body)
|
||||
# 只有达到合同定义的有效章节才进入历史中位数,短章不会污染基线。
|
||||
if count >= 500:
|
||||
counts.append(count)
|
||||
explicit_target = fine_outline.get("targetChars")
|
||||
try:
|
||||
target = calculate_target_chars(
|
||||
explicit_target_chars=explicit_target,
|
||||
recent_chapter_han_counts=counts,
|
||||
default_target_chars=default_target_chars,
|
||||
hard_event_count=_array_length(
|
||||
fine_outline.get("hardEvents", fine_outline.get("hardConstraints", [])),
|
||||
"fine_outline.hardEvents",
|
||||
),
|
||||
foreshadowing_action_count=_array_length(
|
||||
fine_outline.get("foreshadowingActions", []),
|
||||
"fine_outline.foreshadowingActions",
|
||||
),
|
||||
required_scene_count=_array_length(
|
||||
fine_outline.get("requiredScenes", []),
|
||||
"fine_outline.requiredScenes",
|
||||
),
|
||||
min_chars=hard_min_chars,
|
||||
max_chars=hard_max_chars,
|
||||
)
|
||||
except ContractError as exc:
|
||||
raise WriterAdapterError("dynamic_length_input_invalid", str(exc)) from exc
|
||||
lower = max(hard_min_chars, _round_half_up(Decimal(target) * Decimal("0.90")))
|
||||
upper = min(hard_max_chars, _round_half_up(Decimal(target) * Decimal("1.10")))
|
||||
return {
|
||||
"targetChars": target,
|
||||
"minChars": lower,
|
||||
"maxChars": upper,
|
||||
"frontmatterRequired": False,
|
||||
"newSettingDeclarationRequired": True,
|
||||
}
|
||||
|
||||
|
||||
def build_writer_command(executable: str = "claude") -> list[str]:
|
||||
"""基于本机 Claude Code 2.1.211 已验证参数构造无工具命令。"""
|
||||
|
||||
if not isinstance(executable, str) or not executable:
|
||||
raise WriterAdapterError("writer_cli_invalid", "Claude CLI 路径不能为空")
|
||||
return [
|
||||
executable,
|
||||
"--print",
|
||||
"--output-format",
|
||||
"json",
|
||||
"--agent",
|
||||
"writer",
|
||||
"--model",
|
||||
"opus",
|
||||
"--tools",
|
||||
"",
|
||||
"--no-session-persistence",
|
||||
"--disable-slash-commands",
|
||||
"--strict-mcp-config",
|
||||
"--mcp-config",
|
||||
'{"mcpServers":{}}',
|
||||
]
|
||||
|
||||
|
||||
def _parse_cli_output(stdout: str) -> Mapping[str, Any]:
|
||||
"""解析 Claude JSON 包装和其中的 WriterOutput JSON,拒绝宽松猜测。"""
|
||||
|
||||
try:
|
||||
envelope = json.loads(stdout)
|
||||
except (TypeError, json.JSONDecodeError) as exc:
|
||||
raise WriterAdapterError("writer_cli_json_invalid", "Claude CLI 未返回合法 JSON") from exc
|
||||
if not isinstance(envelope, Mapping) or envelope.get("type") != "result":
|
||||
raise WriterAdapterError("writer_cli_json_invalid", "Claude CLI 返回结构不是 result")
|
||||
if envelope.get("is_error") is True:
|
||||
raise WriterAdapterError("writer_cli_reported_error", "Claude CLI 报告执行失败")
|
||||
raw_result = envelope.get("result")
|
||||
if not isinstance(raw_result, str):
|
||||
raise WriterAdapterError("writer_cli_json_invalid", "Claude CLI result 必须是 JSON 字符串")
|
||||
try:
|
||||
parsed = json.loads(raw_result)
|
||||
except json.JSONDecodeError as exc:
|
||||
raise WriterAdapterError("writer_output_json_invalid", "写手结果不是合法 JSON") from exc
|
||||
if not isinstance(parsed, Mapping):
|
||||
raise WriterAdapterError("writer_output_json_invalid", "WriterOutput 必须是对象")
|
||||
return parsed
|
||||
|
||||
|
||||
def _validate_context_binding(context: Mapping[str, Any], output: Mapping[str, Any]) -> None:
|
||||
"""保证输出身份、快照和接受资格不能脱离本次 stdin 上下文。"""
|
||||
|
||||
expected = {
|
||||
"runId": context["runId"],
|
||||
"attempt": context["attempt"],
|
||||
"mode": context["mode"],
|
||||
"qualityPolicyVersion": context["qualityPolicyVersion"],
|
||||
"contextSnapshotId": context["contextSnapshot"]["manifestId"],
|
||||
"contextSnapshotSha256": context["contextSnapshot"]["contextSha256"],
|
||||
"acceptanceEligible": context["acceptanceEligible"],
|
||||
}
|
||||
mismatches = [field for field, value in expected.items() if output.get(field) != value]
|
||||
if mismatches:
|
||||
raise WriterAdapterError(
|
||||
"writer_output_context_mismatch",
|
||||
"WriterOutput 未绑定当前 WriterContext",
|
||||
details={"fields": mismatches},
|
||||
)
|
||||
|
||||
|
||||
def _validate_candidate_semantics(context: Mapping[str, Any], output: Mapping[str, Any]) -> None:
|
||||
"""校验动态篇幅和证据 ID 引用,补足结构合同无法看到的上下文绑定。"""
|
||||
|
||||
contract = context["outputContract"]
|
||||
actual_han_chars = han_count(output["candidateBody"])
|
||||
if not contract["minChars"] <= actual_han_chars <= contract["maxChars"]:
|
||||
raise WriterAdapterError(
|
||||
"candidate_length_out_of_range",
|
||||
"候选正文汉字数超出动态篇幅区间",
|
||||
details={
|
||||
"actualHanChars": actual_han_chars,
|
||||
"minChars": contract["minChars"],
|
||||
"maxChars": contract["maxChars"],
|
||||
"targetChars": contract["targetChars"],
|
||||
},
|
||||
)
|
||||
fact_ids = {item["evidenceId"] for item in context["factEvidence"]}
|
||||
prose_ids = {item["evidenceId"] for item in context["proseEvidence"]}
|
||||
coverage_states = {"supported", "declared_new", "card_gap", "style_gap", "unsupported", "conflict"}
|
||||
for index, claim in enumerate(output["claimLedger"]):
|
||||
invalid = claim["factEvidenceId"] not in fact_ids or claim["coverageState"] not in coverage_states
|
||||
prose_id = claim.get("proseEvidenceId")
|
||||
if prose_id is not None and prose_id not in prose_ids:
|
||||
invalid = True
|
||||
if invalid:
|
||||
raise WriterAdapterError(
|
||||
"claim_evidence_reference_invalid",
|
||||
f"claimLedger[{index}] 引用了当前上下文之外的证据",
|
||||
details={"claimId": claim["claimId"]},
|
||||
)
|
||||
|
||||
|
||||
def run_writer(
|
||||
context: Mapping[str, Any],
|
||||
*,
|
||||
runner: Callable[..., subprocess.CompletedProcess[str]] = subprocess.run,
|
||||
executable: str = "claude",
|
||||
timeout_seconds: float = 600,
|
||||
) -> dict[str, Any]:
|
||||
"""通过 stdin 调用无工具写手;任一错误均抛出不可接受的稳定失败码。"""
|
||||
|
||||
try:
|
||||
normalized_context = validate_writer_context(context)
|
||||
except ContractError as exc:
|
||||
raise WriterAdapterError("writer_context_contract_invalid", str(exc)) from exc
|
||||
command = build_writer_command(executable)
|
||||
try:
|
||||
completed = runner(
|
||||
command,
|
||||
input=canonical_json(normalized_context),
|
||||
text=True,
|
||||
capture_output=True,
|
||||
timeout=timeout_seconds,
|
||||
check=False,
|
||||
)
|
||||
except subprocess.TimeoutExpired as exc:
|
||||
raise WriterAdapterError("writer_timeout", "正文写手调用超时") from exc
|
||||
except FileNotFoundError as exc:
|
||||
raise WriterAdapterError("writer_cli_unavailable", "找不到 Claude CLI") from exc
|
||||
except OSError as exc:
|
||||
raise WriterAdapterError("writer_cli_start_failed", "Claude CLI 启动失败") from exc
|
||||
if completed.returncode != 0:
|
||||
raise WriterAdapterError(
|
||||
"writer_nonzero_exit",
|
||||
"Claude CLI 非零退出",
|
||||
details={"returnCode": completed.returncode, "stderr": (completed.stderr or "")[:1000]},
|
||||
)
|
||||
raw_output = _parse_cli_output(completed.stdout)
|
||||
try:
|
||||
normalized_output = validate_writer_output(raw_output)
|
||||
except ContractError as exc:
|
||||
raise WriterAdapterError("writer_output_contract_invalid", str(exc)) from exc
|
||||
_validate_context_binding(normalized_context, normalized_output)
|
||||
_validate_candidate_semantics(normalized_context, normalized_output)
|
||||
return normalized_output
|
||||
|
||||
|
||||
__all__ = [
|
||||
"WriterAdapterError",
|
||||
"build_writer_command",
|
||||
"calculate_dynamic_output_contract",
|
||||
"run_writer",
|
||||
]
|
||||
705
.claude/skills/continuation/scripts/run_writer_pipeline.py
Normal file
705
.claude/skills/continuation/scripts/run_writer_pipeline.py
Normal file
@ -0,0 +1,705 @@
|
||||
#!/usr/bin/env python3
|
||||
"""正文候选的有限补证、重写、机械审查与 CAS 编排。"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import copy
|
||||
import os
|
||||
import pathlib
|
||||
import sys
|
||||
import tempfile
|
||||
from dataclasses import dataclass
|
||||
from typing import Any, Callable, Mapping, Protocol, Sequence
|
||||
|
||||
SCRIPT_DIR = pathlib.Path(__file__).resolve().parent
|
||||
SKILLS_DIR = SCRIPT_DIR.parents[1]
|
||||
DETECT_DIR = SKILLS_DIR / "detect" / "scripts"
|
||||
READ_CONTEXT_DIR = SKILLS_DIR / "read-context" / "scripts"
|
||||
for path in (DETECT_DIR, READ_CONTEXT_DIR):
|
||||
if str(path) not in sys.path:
|
||||
sys.path.insert(0, str(path))
|
||||
|
||||
from check_writer_candidate import check_writer_candidate # noqa: E402
|
||||
from writer_contract import ( # noqa: E402
|
||||
ContractError,
|
||||
canonical_json,
|
||||
retrieval_identity,
|
||||
validate_writer_context,
|
||||
validate_writer_output,
|
||||
)
|
||||
|
||||
|
||||
MAX_EVIDENCE_REQUESTS = 3
|
||||
MAX_REWRITES = 2
|
||||
|
||||
|
||||
class PipelineError(RuntimeError):
|
||||
"""携带稳定失败码和终态结果的编排错误。"""
|
||||
|
||||
def __init__(
|
||||
self,
|
||||
code: str,
|
||||
message: str,
|
||||
*,
|
||||
details: Mapping[str, Any] | None = None,
|
||||
result: Mapping[str, Any] | None = None,
|
||||
):
|
||||
super().__init__(message)
|
||||
self.code = code
|
||||
self.details = dict(details or {})
|
||||
self.result = dict(result or {})
|
||||
self.acceptance_eligible = False
|
||||
|
||||
|
||||
@dataclass(frozen=True)
|
||||
class CasToken:
|
||||
"""一次候选状态的不可变 CAS 身份。"""
|
||||
|
||||
run_id: str
|
||||
attempt: int
|
||||
candidate_version: int
|
||||
state: str
|
||||
revision: int
|
||||
|
||||
|
||||
class CasStateStore(Protocol):
|
||||
"""生产存储必须实现的最小 CAS 接口;本轮测试只用内存 fake。"""
|
||||
|
||||
def create(self, run_id: str, *, attempt: int, candidate_version: int) -> CasToken: ...
|
||||
|
||||
def transition(self, expected: CasToken, target_state: str) -> CasToken | None: ...
|
||||
|
||||
def start_next(
|
||||
self, expected: CasToken, *, attempt: int, candidate_version: int
|
||||
) -> CasToken | None: ...
|
||||
|
||||
def latest(self, run_id: str) -> CasToken | None: ...
|
||||
|
||||
|
||||
class SemanticDetector(Protocol):
|
||||
"""语义 detector 的明确边界;实现方不得读写状态或候选文件。"""
|
||||
|
||||
def __call__(
|
||||
self,
|
||||
context: Mapping[str, Any],
|
||||
candidate: Mapping[str, Any],
|
||||
mechanical_report: Mapping[str, Any],
|
||||
) -> Mapping[str, Any]: ...
|
||||
|
||||
|
||||
class InMemoryCasStateStore:
|
||||
"""仅供实验和测试使用的线程外内存 CAS fake。"""
|
||||
|
||||
_ALLOWED = {
|
||||
"DRAFT": frozenset({"CHECKING"}),
|
||||
"CHECKING": frozenset({"PASSED", "REJECTED"}),
|
||||
"PASSED": frozenset(),
|
||||
"REJECTED": frozenset(),
|
||||
}
|
||||
|
||||
def __init__(self) -> None:
|
||||
self._latest: dict[str, CasToken] = {}
|
||||
|
||||
def create(self, run_id: str, *, attempt: int, candidate_version: int) -> CasToken:
|
||||
"""只允许为未存在的 run 创建第一条 DRAFT。"""
|
||||
|
||||
if run_id in self._latest:
|
||||
raise PipelineError("CAS_CONFLICT", "run 已存在,不能重复创建初始状态")
|
||||
token = CasToken(run_id, attempt, candidate_version, "DRAFT", 1)
|
||||
self._latest[run_id] = token
|
||||
return token
|
||||
|
||||
def transition(self, expected: CasToken, target_state: str) -> CasToken | None:
|
||||
"""仅当最新 token 完全匹配且转换合法时更新。"""
|
||||
|
||||
current = self._latest.get(expected.run_id)
|
||||
if current != expected or target_state not in self._ALLOWED.get(expected.state, frozenset()):
|
||||
return None
|
||||
updated = CasToken(
|
||||
expected.run_id,
|
||||
expected.attempt,
|
||||
expected.candidate_version,
|
||||
target_state,
|
||||
expected.revision + 1,
|
||||
)
|
||||
self._latest[expected.run_id] = updated
|
||||
return updated
|
||||
|
||||
def start_next(
|
||||
self, expected: CasToken, *, attempt: int, candidate_version: int
|
||||
) -> CasToken | None:
|
||||
"""只允许从最新 REJECTED 创建单调递增的新 DRAFT。"""
|
||||
|
||||
current = self._latest.get(expected.run_id)
|
||||
if (
|
||||
current != expected
|
||||
or expected.state != "REJECTED"
|
||||
or attempt <= expected.attempt
|
||||
or candidate_version <= expected.candidate_version
|
||||
):
|
||||
return None
|
||||
updated = CasToken(
|
||||
expected.run_id,
|
||||
attempt,
|
||||
candidate_version,
|
||||
"DRAFT",
|
||||
expected.revision + 1,
|
||||
)
|
||||
self._latest[expected.run_id] = updated
|
||||
return updated
|
||||
|
||||
def latest(self, run_id: str) -> CasToken | None:
|
||||
"""返回 run 的最新不可变状态。"""
|
||||
|
||||
return self._latest.get(run_id)
|
||||
|
||||
|
||||
def atomic_write_json(path: str | pathlib.Path, value: Mapping[str, Any]) -> None:
|
||||
"""先写同目录临时文件并 fsync,再以原子替换发布正式结果。"""
|
||||
|
||||
target = pathlib.Path(path)
|
||||
target.parent.mkdir(parents=True, exist_ok=True)
|
||||
descriptor, temporary_name = tempfile.mkstemp(
|
||||
prefix=f".{target.name}.", suffix=".tmp", dir=target.parent
|
||||
)
|
||||
temporary = pathlib.Path(temporary_name)
|
||||
try:
|
||||
with os.fdopen(descriptor, "w", encoding="utf-8", newline="\n") as handle:
|
||||
handle.write(canonical_json(value))
|
||||
handle.write("\n")
|
||||
handle.flush()
|
||||
os.fsync(handle.fileno())
|
||||
os.replace(str(temporary), str(target))
|
||||
except BaseException:
|
||||
# replace 成功后临时路径已不存在;失败时只清理本次临时文件。
|
||||
try:
|
||||
temporary.unlink(missing_ok=True)
|
||||
finally:
|
||||
raise
|
||||
|
||||
|
||||
def _cas_or_fail(token: CasToken | None, action: str) -> CasToken:
|
||||
"""把所有 CAS 竞争统一为稳定失败码。"""
|
||||
|
||||
if token is None:
|
||||
raise PipelineError("CAS_CONFLICT", f"状态 CAS 失败: {action}")
|
||||
return token
|
||||
|
||||
|
||||
def _next_context_without_new_evidence(
|
||||
context: Mapping[str, Any], *, attempt: int
|
||||
) -> dict[str, Any]:
|
||||
"""为修复重写创建新 attempt,并重新绑定上下文 hash。"""
|
||||
|
||||
advanced = copy.deepcopy(dict(context))
|
||||
advanced["attempt"] = attempt
|
||||
advanced["contextSnapshot"]["contextSha256"] = retrieval_identity(advanced)
|
||||
try:
|
||||
return validate_writer_context(advanced)
|
||||
except ContractError as exc:
|
||||
raise PipelineError("REASSEMBLED_CONTEXT_INVALID", str(exc)) from exc
|
||||
|
||||
|
||||
def _validate_semantic_report(
|
||||
report: Mapping[str, Any], candidate: Mapping[str, Any]
|
||||
) -> list[dict[str, Any]]:
|
||||
"""校验 fake/未来模型 detector 的候选绑定和阻塞项结构。"""
|
||||
|
||||
if not isinstance(report, Mapping):
|
||||
raise PipelineError("SEMANTIC_DETECTOR_INVALID", "语义 detector 必须返回对象")
|
||||
if report.get("status") not in {"passed", "rejected"}:
|
||||
raise PipelineError("SEMANTIC_DETECTOR_INVALID", "语义 detector status 非法")
|
||||
if (
|
||||
report.get("candidateVersion") != candidate["candidateVersion"]
|
||||
or report.get("candidateSha256") != candidate["candidateSha256"]
|
||||
):
|
||||
raise PipelineError("SEMANTIC_DETECTOR_STALE", "语义 detector 结果绑定了旧候选")
|
||||
failures = report.get("blockingFailures")
|
||||
if not isinstance(failures, list) or any(not isinstance(item, Mapping) for item in failures):
|
||||
raise PipelineError("SEMANTIC_DETECTOR_INVALID", "语义 detector blockingFailures 非法")
|
||||
if any(not isinstance(item.get("code"), str) or not item["code"] for item in failures):
|
||||
raise PipelineError("SEMANTIC_DETECTOR_INVALID", "语义 detector 阻塞项缺少稳定 code")
|
||||
if report["status"] == "passed" and failures:
|
||||
raise PipelineError("SEMANTIC_DETECTOR_INVALID", "语义 detector 通过时不得含阻塞项")
|
||||
if report["status"] == "rejected" and not failures:
|
||||
raise PipelineError("SEMANTIC_DETECTOR_INVALID", "语义 detector 拒绝时必须包含阻塞项")
|
||||
return [dict(item) for item in failures]
|
||||
|
||||
|
||||
def _validate_mechanical_report(
|
||||
report: Any, candidate: Mapping[str, Any]
|
||||
) -> list[dict[str, Any]]:
|
||||
"""严格校验机械 detector 报告版本、候选绑定和通过状态。"""
|
||||
|
||||
if not isinstance(report, Mapping):
|
||||
raise PipelineError("MECHANICAL_DETECTOR_INVALID", "机械 detector 必须返回对象")
|
||||
if report.get("schemaVersion") != "writer-detector-report-v1":
|
||||
raise PipelineError("MECHANICAL_DETECTOR_INVALID", "机械 detector 报告版本非法")
|
||||
expected = {
|
||||
"runId": candidate.get("runId"),
|
||||
"attempt": candidate.get("attempt"),
|
||||
"candidateVersion": candidate.get("candidateVersion"),
|
||||
"candidateSha256": candidate.get("candidateSha256"),
|
||||
}
|
||||
if any(report.get(field) != value for field, value in expected.items()):
|
||||
raise PipelineError("MECHANICAL_DETECTOR_INVALID", "机械 detector 报告绑定了旧候选")
|
||||
passed = report.get("passed")
|
||||
failures = report.get("blockingFailures")
|
||||
if not isinstance(passed, bool):
|
||||
raise PipelineError("MECHANICAL_DETECTOR_INVALID", "机械 detector passed 必须是布尔值")
|
||||
if not isinstance(failures, list) or any(not isinstance(item, Mapping) for item in failures):
|
||||
raise PipelineError("MECHANICAL_DETECTOR_INVALID", "机械 detector blockingFailures 非法")
|
||||
if any(not isinstance(item.get("code"), str) or not item["code"] for item in failures):
|
||||
raise PipelineError("MECHANICAL_DETECTOR_INVALID", "机械 detector 阻塞项缺少稳定 code")
|
||||
if passed != (not failures):
|
||||
raise PipelineError("MECHANICAL_DETECTOR_INVALID", "机械 detector 通过状态与阻塞项矛盾")
|
||||
return [dict(item) for item in failures]
|
||||
|
||||
|
||||
def _terminal_result(
|
||||
*,
|
||||
run_id: str,
|
||||
status: str,
|
||||
token: CasToken,
|
||||
evidence_request_count: int,
|
||||
rewrite_count: int,
|
||||
trace: Sequence[Mapping[str, Any]],
|
||||
candidate: Mapping[str, Any] | None,
|
||||
failure_code: str | None = None,
|
||||
) -> dict[str, Any]:
|
||||
"""生成可原子落盘的稳定终态结果。"""
|
||||
|
||||
return {
|
||||
"schemaVersion": "writer-pipeline-result-v1",
|
||||
"runId": run_id,
|
||||
"status": status,
|
||||
"attempt": token.attempt,
|
||||
"candidateVersion": token.candidate_version,
|
||||
"candidateSha256": candidate.get("candidateSha256") if candidate else None,
|
||||
"failureCode": failure_code,
|
||||
"evidenceRequestCount": evidence_request_count,
|
||||
"rewriteCount": rewrite_count,
|
||||
"trace": [dict(item) for item in trace],
|
||||
}
|
||||
|
||||
|
||||
def _publish_if_requested(
|
||||
result_path: str | pathlib.Path | None, result: Mapping[str, Any]
|
||||
) -> None:
|
||||
"""仅在调用方显式给出路径时发布终态文件。"""
|
||||
|
||||
if result_path is not None:
|
||||
atomic_write_json(result_path, result)
|
||||
|
||||
|
||||
def _terminal_failure(
|
||||
*,
|
||||
code: str,
|
||||
message: str,
|
||||
token: CasToken,
|
||||
state_store: CasStateStore,
|
||||
run_id: str,
|
||||
evidence_request_count: int,
|
||||
rewrite_count: int,
|
||||
trace: Sequence[Mapping[str, Any]],
|
||||
candidate: Mapping[str, Any] | None,
|
||||
result_path: str | pathlib.Path | None,
|
||||
details: Mapping[str, Any] | None = None,
|
||||
) -> PipelineError:
|
||||
"""将已进入检查的失败统一收敛为可发布的 REJECTED 终态。"""
|
||||
|
||||
if token.state != "CHECKING":
|
||||
raise PipelineError("CAS_CONFLICT", f"失败收敛不接受状态: {token.state}")
|
||||
rejected = _cas_or_fail(
|
||||
state_store.transition(token, "REJECTED"), "CHECKING -> REJECTED"
|
||||
)
|
||||
terminal_trace = [dict(item) for item in trace]
|
||||
terminal_trace.append(
|
||||
{
|
||||
"attempt": rejected.attempt,
|
||||
"candidateVersion": rejected.candidate_version,
|
||||
"candidateSha256": candidate.get("candidateSha256") if candidate else None,
|
||||
"status": "pipeline_failed",
|
||||
"failureCodes": [code],
|
||||
}
|
||||
)
|
||||
result = _terminal_result(
|
||||
run_id=run_id,
|
||||
status="REJECTED",
|
||||
token=rejected,
|
||||
evidence_request_count=evidence_request_count,
|
||||
rewrite_count=rewrite_count,
|
||||
trace=terminal_trace,
|
||||
candidate=candidate,
|
||||
failure_code=code,
|
||||
)
|
||||
_publish_if_requested(result_path, result)
|
||||
return PipelineError(code, message, details=details, result=result)
|
||||
|
||||
|
||||
def run_writer_pipeline(
|
||||
*,
|
||||
context: Mapping[str, Any],
|
||||
requirements: Mapping[str, Any],
|
||||
writer: Callable[[Mapping[str, Any], int, list[dict[str, Any]]], Mapping[str, Any]],
|
||||
evidence_provider: Callable[
|
||||
[Mapping[str, Any], list[dict[str, Any]], int], Mapping[str, Any]
|
||||
],
|
||||
semantic_detector: SemanticDetector,
|
||||
state_store: CasStateStore,
|
||||
result_path: str | pathlib.Path | None = None,
|
||||
) -> dict[str, Any]:
|
||||
"""执行最多 3 次补证、2 次重写的正文候选有限收敛循环。"""
|
||||
|
||||
try:
|
||||
current_context = validate_writer_context(context)
|
||||
except ContractError as exc:
|
||||
raise PipelineError("WRITER_CONTEXT_INVALID", str(exc)) from exc
|
||||
run_id = current_context["runId"]
|
||||
candidate_version = 1
|
||||
evidence_request_count = 0
|
||||
rewrite_count = 0
|
||||
repair_failures: list[dict[str, Any]] = []
|
||||
trace: list[dict[str, Any]] = []
|
||||
draft = state_store.create(
|
||||
run_id,
|
||||
attempt=current_context["attempt"],
|
||||
candidate_version=candidate_version,
|
||||
)
|
||||
|
||||
while True:
|
||||
checking = _cas_or_fail(state_store.transition(draft, "CHECKING"), "DRAFT -> CHECKING")
|
||||
try:
|
||||
raw_candidate = writer(current_context, candidate_version, repair_failures)
|
||||
except PipelineError as exc:
|
||||
raise _terminal_failure(
|
||||
code=exc.code,
|
||||
message=str(exc),
|
||||
token=checking,
|
||||
state_store=state_store,
|
||||
run_id=run_id,
|
||||
evidence_request_count=evidence_request_count,
|
||||
rewrite_count=rewrite_count,
|
||||
trace=trace,
|
||||
candidate=None,
|
||||
result_path=result_path,
|
||||
details=exc.details,
|
||||
) from exc
|
||||
except Exception as exc:
|
||||
raise _terminal_failure(
|
||||
code="WRITER_FAILED",
|
||||
message="fake/adapter 写手调用失败",
|
||||
token=checking,
|
||||
state_store=state_store,
|
||||
run_id=run_id,
|
||||
evidence_request_count=evidence_request_count,
|
||||
rewrite_count=rewrite_count,
|
||||
trace=trace,
|
||||
candidate=None,
|
||||
result_path=result_path,
|
||||
) from exc
|
||||
if not isinstance(raw_candidate, Mapping):
|
||||
raise _terminal_failure(
|
||||
code="WRITER_OUTPUT_INVALID",
|
||||
message="写手必须返回对象",
|
||||
token=checking,
|
||||
state_store=state_store,
|
||||
run_id=run_id,
|
||||
evidence_request_count=evidence_request_count,
|
||||
rewrite_count=rewrite_count,
|
||||
trace=trace,
|
||||
candidate=None,
|
||||
result_path=result_path,
|
||||
)
|
||||
candidate = dict(raw_candidate)
|
||||
if (
|
||||
candidate.get("runId") != run_id
|
||||
or candidate.get("attempt") != current_context["attempt"]
|
||||
or candidate.get("candidateVersion") != candidate_version
|
||||
):
|
||||
raise _terminal_failure(
|
||||
code="WRITER_OUTPUT_STALE",
|
||||
message="写手返回了旧 attempt 或 candidateVersion",
|
||||
token=checking,
|
||||
state_store=state_store,
|
||||
run_id=run_id,
|
||||
evidence_request_count=evidence_request_count,
|
||||
rewrite_count=rewrite_count,
|
||||
trace=trace,
|
||||
candidate=candidate,
|
||||
result_path=result_path,
|
||||
)
|
||||
try:
|
||||
candidate = validate_writer_output(candidate)
|
||||
except ContractError as exc:
|
||||
# 合同错误作为可定位审查失败进入有限重写,而不是绕过状态机。
|
||||
try:
|
||||
mechanical_report = check_writer_candidate(current_context, candidate, requirements)
|
||||
candidate_failures = _validate_mechanical_report(mechanical_report, candidate)
|
||||
except PipelineError as detector_exc:
|
||||
raise _terminal_failure(
|
||||
code=detector_exc.code,
|
||||
message=str(detector_exc),
|
||||
token=checking,
|
||||
state_store=state_store,
|
||||
run_id=run_id,
|
||||
evidence_request_count=evidence_request_count,
|
||||
rewrite_count=rewrite_count,
|
||||
trace=trace,
|
||||
candidate=candidate,
|
||||
result_path=result_path,
|
||||
details=detector_exc.details,
|
||||
) from detector_exc
|
||||
except Exception as detector_exc:
|
||||
raise _terminal_failure(
|
||||
code="MECHANICAL_DETECTOR_FAILED",
|
||||
message="机械 detector 调用失败",
|
||||
token=checking,
|
||||
state_store=state_store,
|
||||
run_id=run_id,
|
||||
evidence_request_count=evidence_request_count,
|
||||
rewrite_count=rewrite_count,
|
||||
trace=trace,
|
||||
candidate=candidate,
|
||||
result_path=result_path,
|
||||
) from detector_exc
|
||||
if not candidate_failures:
|
||||
candidate_failures = [{"code": "OUTPUT_CONTRACT_INVALID", "message": str(exc)}]
|
||||
else:
|
||||
candidate_failures = []
|
||||
mechanical_report = {}
|
||||
|
||||
requests = candidate.get("evidenceRequests", [])
|
||||
if not isinstance(requests, list):
|
||||
requests = []
|
||||
if evidence_request_count + len(requests) > MAX_EVIDENCE_REQUESTS:
|
||||
raise _terminal_failure(
|
||||
code="EVIDENCE_REQUEST_LIMIT_REACHED",
|
||||
message="补证请求累计超过 3 次",
|
||||
token=checking,
|
||||
state_store=state_store,
|
||||
run_id=run_id,
|
||||
evidence_request_count=evidence_request_count,
|
||||
rewrite_count=rewrite_count,
|
||||
trace=trace,
|
||||
candidate=candidate,
|
||||
result_path=result_path,
|
||||
)
|
||||
if requests:
|
||||
evidence_request_count += len(requests)
|
||||
if rewrite_count >= MAX_REWRITES:
|
||||
raise _terminal_failure(
|
||||
code="REWRITE_LIMIT_REACHED",
|
||||
message="补证后重写已达到 2 次",
|
||||
token=checking,
|
||||
state_store=state_store,
|
||||
run_id=run_id,
|
||||
evidence_request_count=evidence_request_count,
|
||||
rewrite_count=rewrite_count,
|
||||
trace=trace,
|
||||
candidate=candidate,
|
||||
result_path=result_path,
|
||||
)
|
||||
next_attempt = current_context["attempt"] + 1
|
||||
try:
|
||||
provided = evidence_provider(current_context, [dict(item) for item in requests], next_attempt)
|
||||
current_context = validate_writer_context(provided)
|
||||
except Exception as exc:
|
||||
raise _terminal_failure(
|
||||
code="REASSEMBLED_CONTEXT_INVALID",
|
||||
message=str(exc),
|
||||
token=checking,
|
||||
state_store=state_store,
|
||||
run_id=run_id,
|
||||
evidence_request_count=evidence_request_count,
|
||||
rewrite_count=rewrite_count,
|
||||
trace=trace,
|
||||
candidate=candidate,
|
||||
result_path=result_path,
|
||||
) from exc
|
||||
if current_context["runId"] != run_id or current_context["attempt"] != next_attempt:
|
||||
raise _terminal_failure(
|
||||
code="REASSEMBLED_CONTEXT_INVALID",
|
||||
message="补证上下文 runId/attempt 不匹配",
|
||||
token=checking,
|
||||
state_store=state_store,
|
||||
run_id=run_id,
|
||||
evidence_request_count=evidence_request_count,
|
||||
rewrite_count=rewrite_count,
|
||||
trace=trace,
|
||||
candidate=candidate,
|
||||
result_path=result_path,
|
||||
)
|
||||
rejected = _cas_or_fail(
|
||||
state_store.transition(checking, "REJECTED"), "CHECKING -> REJECTED"
|
||||
)
|
||||
rewrite_count += 1
|
||||
candidate_version += 1
|
||||
trace.append(
|
||||
{
|
||||
"attempt": checking.attempt,
|
||||
"candidateVersion": checking.candidate_version,
|
||||
"status": "evidence_requested",
|
||||
"requestIds": [item.get("requestId") for item in requests],
|
||||
}
|
||||
)
|
||||
repair_failures = []
|
||||
draft = _cas_or_fail(
|
||||
state_store.start_next(
|
||||
rejected,
|
||||
attempt=next_attempt,
|
||||
candidate_version=candidate_version,
|
||||
),
|
||||
"REJECTED -> next DRAFT",
|
||||
)
|
||||
continue
|
||||
|
||||
if not mechanical_report:
|
||||
try:
|
||||
mechanical_report = check_writer_candidate(current_context, candidate, requirements)
|
||||
candidate_failures = _validate_mechanical_report(mechanical_report, candidate)
|
||||
except PipelineError as detector_exc:
|
||||
raise _terminal_failure(
|
||||
code=detector_exc.code,
|
||||
message=str(detector_exc),
|
||||
token=checking,
|
||||
state_store=state_store,
|
||||
run_id=run_id,
|
||||
evidence_request_count=evidence_request_count,
|
||||
rewrite_count=rewrite_count,
|
||||
trace=trace,
|
||||
candidate=candidate,
|
||||
result_path=result_path,
|
||||
details=detector_exc.details,
|
||||
) from detector_exc
|
||||
except Exception as exc:
|
||||
raise _terminal_failure(
|
||||
code="MECHANICAL_DETECTOR_FAILED",
|
||||
message="机械 detector 调用失败",
|
||||
token=checking,
|
||||
state_store=state_store,
|
||||
run_id=run_id,
|
||||
evidence_request_count=evidence_request_count,
|
||||
rewrite_count=rewrite_count,
|
||||
trace=trace,
|
||||
candidate=candidate,
|
||||
result_path=result_path,
|
||||
) from exc
|
||||
semantic_report: Mapping[str, Any] | None = None
|
||||
if not candidate_failures:
|
||||
try:
|
||||
semantic_report = semantic_detector(current_context, candidate, mechanical_report)
|
||||
candidate_failures.extend(_validate_semantic_report(semantic_report, candidate))
|
||||
except PipelineError as exc:
|
||||
raise _terminal_failure(
|
||||
code=exc.code,
|
||||
message=str(exc),
|
||||
token=checking,
|
||||
state_store=state_store,
|
||||
run_id=run_id,
|
||||
evidence_request_count=evidence_request_count,
|
||||
rewrite_count=rewrite_count,
|
||||
trace=trace,
|
||||
candidate=candidate,
|
||||
result_path=result_path,
|
||||
details=exc.details,
|
||||
) from exc
|
||||
except Exception as exc:
|
||||
raise _terminal_failure(
|
||||
code="SEMANTIC_DETECTOR_FAILED",
|
||||
message="fake/语义 detector 调用失败",
|
||||
token=checking,
|
||||
state_store=state_store,
|
||||
run_id=run_id,
|
||||
evidence_request_count=evidence_request_count,
|
||||
rewrite_count=rewrite_count,
|
||||
trace=trace,
|
||||
candidate=candidate,
|
||||
result_path=result_path,
|
||||
) from exc
|
||||
trace.append(
|
||||
{
|
||||
"attempt": checking.attempt,
|
||||
"candidateVersion": checking.candidate_version,
|
||||
"candidateSha256": candidate.get("candidateSha256"),
|
||||
"mechanicalPassed": mechanical_report.get("passed", False),
|
||||
"semanticStatus": semantic_report.get("status") if semantic_report else "not_run",
|
||||
"failureCodes": [item.get("code") for item in candidate_failures],
|
||||
}
|
||||
)
|
||||
if not candidate_failures:
|
||||
passed = _cas_or_fail(
|
||||
state_store.transition(checking, "PASSED"), "CHECKING -> PASSED"
|
||||
)
|
||||
result = _terminal_result(
|
||||
run_id=run_id,
|
||||
status="PASSED",
|
||||
token=passed,
|
||||
evidence_request_count=evidence_request_count,
|
||||
rewrite_count=rewrite_count,
|
||||
trace=trace,
|
||||
candidate=candidate,
|
||||
)
|
||||
_publish_if_requested(result_path, result)
|
||||
return result
|
||||
|
||||
if rewrite_count >= MAX_REWRITES:
|
||||
raise _terminal_failure(
|
||||
code="REWRITE_LIMIT_REACHED",
|
||||
message="候选两次重写后仍未通过",
|
||||
token=checking,
|
||||
state_store=state_store,
|
||||
run_id=run_id,
|
||||
evidence_request_count=evidence_request_count,
|
||||
rewrite_count=rewrite_count,
|
||||
trace=trace,
|
||||
candidate=candidate,
|
||||
result_path=result_path,
|
||||
)
|
||||
rewrite_count += 1
|
||||
candidate_version += 1
|
||||
next_attempt = current_context["attempt"] + 1
|
||||
try:
|
||||
next_context = _next_context_without_new_evidence(
|
||||
current_context, attempt=next_attempt
|
||||
)
|
||||
except PipelineError as exc:
|
||||
raise _terminal_failure(
|
||||
code=exc.code,
|
||||
message=str(exc),
|
||||
token=checking,
|
||||
state_store=state_store,
|
||||
run_id=run_id,
|
||||
evidence_request_count=evidence_request_count,
|
||||
rewrite_count=rewrite_count,
|
||||
trace=trace,
|
||||
candidate=candidate,
|
||||
result_path=result_path,
|
||||
details=exc.details,
|
||||
) from exc
|
||||
rejected = _cas_or_fail(
|
||||
state_store.transition(checking, "REJECTED"), "CHECKING -> REJECTED"
|
||||
)
|
||||
current_context = next_context
|
||||
repair_failures = [dict(item) for item in candidate_failures]
|
||||
draft = _cas_or_fail(
|
||||
state_store.start_next(
|
||||
rejected,
|
||||
attempt=next_attempt,
|
||||
candidate_version=candidate_version,
|
||||
),
|
||||
"REJECTED -> next DRAFT",
|
||||
)
|
||||
|
||||
|
||||
__all__ = [
|
||||
"MAX_EVIDENCE_REQUESTS",
|
||||
"MAX_REWRITES",
|
||||
"CasToken",
|
||||
"CasStateStore",
|
||||
"InMemoryCasStateStore",
|
||||
"PipelineError",
|
||||
"SemanticDetector",
|
||||
"atomic_write_json",
|
||||
"run_writer_pipeline",
|
||||
]
|
||||
211
.claude/skills/continuation/scripts/test_run_writer.py
Normal file
211
.claude/skills/continuation/scripts/test_run_writer.py
Normal file
@ -0,0 +1,211 @@
|
||||
#!/usr/bin/env python3
|
||||
"""无工具正文写手 adapter 的隔离、合同与失败关闭测试。"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import copy
|
||||
import hashlib
|
||||
import json
|
||||
import pathlib
|
||||
import subprocess
|
||||
import sys
|
||||
import unittest
|
||||
|
||||
SCRIPT_DIR = pathlib.Path(__file__).resolve().parent
|
||||
READ_CONTEXT_DIR = SCRIPT_DIR.parents[1] / "read-context" / "scripts"
|
||||
sys.path.insert(0, str(SCRIPT_DIR))
|
||||
sys.path.insert(0, str(READ_CONTEXT_DIR))
|
||||
|
||||
from test_writer_contract import valid_context, valid_output # noqa: E402
|
||||
from writer_contract import canonical_json, retrieval_identity # noqa: E402
|
||||
from run_writer import ( # noqa: E402
|
||||
WriterAdapterError,
|
||||
calculate_dynamic_output_contract,
|
||||
run_writer,
|
||||
)
|
||||
|
||||
|
||||
def _bound_context() -> dict:
|
||||
"""构造带一条事实证据且可供 adapter 绑定检查的合法上下文。"""
|
||||
|
||||
context = valid_context()
|
||||
context["factEvidence"] = [
|
||||
{
|
||||
"evidenceId": "fact-1",
|
||||
"fact": "林澈仍在圣蒂曼城内",
|
||||
"sourceType": "canonical_state",
|
||||
"sourceRef": {
|
||||
"sourceId": "state:8",
|
||||
"sourceVersion": "state-v1",
|
||||
"sourceType": "canonical_state",
|
||||
},
|
||||
"contentSha256": "sha256:" + "a" * 64,
|
||||
"riskLevel": "high",
|
||||
}
|
||||
]
|
||||
context["contextSnapshot"]["contextSha256"] = retrieval_identity(context)
|
||||
return context
|
||||
|
||||
|
||||
def _bound_output(context: dict, *, body: str | None = None) -> dict:
|
||||
"""构造与给定上下文、篇幅和事实证据完全绑定的合法输出。"""
|
||||
|
||||
candidate_body = body if body is not None else "文" * 4000
|
||||
candidate_hash = "sha256:" + hashlib.sha256(candidate_body.encode("utf-8")).hexdigest()
|
||||
output = valid_output()
|
||||
output.update(
|
||||
{
|
||||
"runId": context["runId"],
|
||||
"attempt": context["attempt"],
|
||||
"mode": context["mode"],
|
||||
"qualityPolicyVersion": context["qualityPolicyVersion"],
|
||||
"contextSnapshotId": context["contextSnapshot"]["manifestId"],
|
||||
"contextSnapshotSha256": context["contextSnapshot"]["contextSha256"],
|
||||
"candidateSha256": candidate_hash,
|
||||
"acceptanceEligible": context["acceptanceEligible"],
|
||||
"candidateBody": candidate_body,
|
||||
"claimLedger": [
|
||||
{
|
||||
"claimId": "claim-1",
|
||||
"candidateSha256": candidate_hash,
|
||||
"startCodePoint": 0,
|
||||
"endCodePoint": 2,
|
||||
"factType": "character_location",
|
||||
"factEvidenceId": "fact-1",
|
||||
"coverageState": "supported",
|
||||
}
|
||||
],
|
||||
}
|
||||
)
|
||||
return output
|
||||
|
||||
|
||||
class FakeRunner:
|
||||
"""记录参数并返回预设结果,确保测试不会调用真实模型。"""
|
||||
|
||||
def __init__(self, completed: subprocess.CompletedProcess[str] | BaseException):
|
||||
self.completed = completed
|
||||
self.calls: list[tuple[list[str], dict]] = []
|
||||
|
||||
def __call__(self, command: list[str], **kwargs):
|
||||
self.calls.append((command, kwargs))
|
||||
if isinstance(self.completed, BaseException):
|
||||
raise self.completed
|
||||
return self.completed
|
||||
|
||||
|
||||
def _success_runner(context: dict, output: dict | None = None) -> FakeRunner:
|
||||
"""返回模拟 Claude `--output-format json` 包装结构的成功 runner。"""
|
||||
|
||||
payload = output if output is not None else _bound_output(context)
|
||||
stdout = json.dumps({"type": "result", "result": json.dumps(payload, ensure_ascii=False)})
|
||||
return FakeRunner(subprocess.CompletedProcess(["claude"], 0, stdout=stdout, stderr=""))
|
||||
|
||||
|
||||
class RunWriterTest(unittest.TestCase):
|
||||
def test_adapter_uses_verified_isolation_flags_and_stdin_only(self):
|
||||
context = _bound_context()
|
||||
runner = _success_runner(context)
|
||||
|
||||
result = run_writer(context, runner=runner, executable="claude", timeout_seconds=30)
|
||||
|
||||
self.assertEqual(result["candidateVersion"], 1)
|
||||
command, kwargs = runner.calls[0]
|
||||
required_pairs = {
|
||||
"--output-format": "json",
|
||||
"--agent": "writer",
|
||||
"--model": "opus",
|
||||
"--tools": "",
|
||||
}
|
||||
self.assertIn("--print", command)
|
||||
self.assertIn("--no-session-persistence", command)
|
||||
self.assertIn("--disable-slash-commands", command)
|
||||
self.assertIn("--strict-mcp-config", command)
|
||||
for flag, value in required_pairs.items():
|
||||
self.assertEqual(command[command.index(flag) + 1], value)
|
||||
self.assertEqual(command[0], "claude")
|
||||
self.assertEqual(kwargs["input"], canonical_json(context))
|
||||
self.assertTrue(kwargs["text"])
|
||||
self.assertTrue(kwargs["capture_output"])
|
||||
self.assertNotIn("shell", kwargs)
|
||||
self.assertNotIn(canonical_json(context), command)
|
||||
|
||||
def test_timeout_nonzero_non_json_and_schema_error_fail_closed(self):
|
||||
context = _bound_context()
|
||||
invalid_schema = _bound_output(context)
|
||||
del invalid_schema["selfCheck"]
|
||||
cases = {
|
||||
"writer_timeout": FakeRunner(subprocess.TimeoutExpired(["claude"], 1)),
|
||||
"writer_nonzero_exit": FakeRunner(
|
||||
subprocess.CompletedProcess(["claude"], 7, stdout="", stderr="拒绝执行")
|
||||
),
|
||||
"writer_cli_json_invalid": FakeRunner(
|
||||
subprocess.CompletedProcess(["claude"], 0, stdout="not-json", stderr="")
|
||||
),
|
||||
"writer_output_contract_invalid": _success_runner(context, invalid_schema),
|
||||
}
|
||||
|
||||
for expected_code, runner in cases.items():
|
||||
with self.subTest(expected_code=expected_code), self.assertRaises(WriterAdapterError) as raised:
|
||||
run_writer(context, runner=runner, timeout_seconds=1)
|
||||
self.assertEqual(raised.exception.code, expected_code)
|
||||
self.assertFalse(raised.exception.acceptance_eligible)
|
||||
|
||||
def test_candidate_length_must_fit_dynamic_context_range(self):
|
||||
context = _bound_context()
|
||||
too_short = _bound_output(context, body="短" * 3599)
|
||||
|
||||
with self.assertRaises(WriterAdapterError) as raised:
|
||||
run_writer(context, runner=_success_runner(context, too_short))
|
||||
|
||||
self.assertEqual(raised.exception.code, "candidate_length_out_of_range")
|
||||
self.assertEqual(raised.exception.details["actualHanChars"], 3599)
|
||||
|
||||
def test_claim_must_reference_evidence_in_the_same_context(self):
|
||||
context = _bound_context()
|
||||
output = _bound_output(context)
|
||||
output["claimLedger"][0]["factEvidenceId"] = "fact-from-another-context"
|
||||
|
||||
with self.assertRaises(WriterAdapterError) as raised:
|
||||
run_writer(context, runner=_success_runner(context, output))
|
||||
|
||||
self.assertEqual(raised.exception.code, "claim_evidence_reference_invalid")
|
||||
|
||||
def test_dynamic_output_contract_uses_outline_density_and_frozen_history(self):
|
||||
contract = calculate_dynamic_output_contract(
|
||||
fine_outline={
|
||||
"hardConstraints": ["事件一", "事件二", "事件三"],
|
||||
"foreshadowingActions": [],
|
||||
"requiredScenes": [],
|
||||
},
|
||||
recent_chapter_bodies=["文" * 2501, "文" * 2502, "文" * 2503, "文" * 2504],
|
||||
)
|
||||
|
||||
self.assertEqual(contract["targetChars"], 2100)
|
||||
self.assertEqual(contract["minChars"], 2000)
|
||||
self.assertEqual(contract["maxChars"], 2310)
|
||||
self.assertFalse(contract["frontmatterRequired"])
|
||||
self.assertTrue(contract["newSettingDeclarationRequired"])
|
||||
|
||||
def test_output_metadata_cannot_escape_the_input_context(self):
|
||||
context = _bound_context()
|
||||
for field, value in (
|
||||
("attempt", 2),
|
||||
("candidateVersion", 0),
|
||||
("contextSnapshotSha256", "sha256:" + "f" * 64),
|
||||
("acceptanceEligible", False),
|
||||
):
|
||||
output = _bound_output(context)
|
||||
output[field] = value
|
||||
if field == "candidateVersion":
|
||||
# candidateVersion=0 会先被严格 WriterOutput 合同拒绝。
|
||||
expected_code = "writer_output_contract_invalid"
|
||||
else:
|
||||
expected_code = "writer_output_context_mismatch"
|
||||
with self.subTest(field=field), self.assertRaises(WriterAdapterError) as raised:
|
||||
run_writer(copy.deepcopy(context), runner=_success_runner(context, output))
|
||||
self.assertEqual(raised.exception.code, expected_code)
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
unittest.main()
|
||||
489
.claude/skills/continuation/scripts/test_run_writer_pipeline.py
Normal file
489
.claude/skills/continuation/scripts/test_run_writer_pipeline.py
Normal file
@ -0,0 +1,489 @@
|
||||
#!/usr/bin/env python3
|
||||
"""正文补证、重写、CAS 和原子结果写入的有限收敛测试。"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import copy
|
||||
import json
|
||||
import pathlib
|
||||
import sys
|
||||
import tempfile
|
||||
import unittest
|
||||
from unittest import mock
|
||||
|
||||
SCRIPT_DIR = pathlib.Path(__file__).resolve().parent
|
||||
SKILLS_DIR = SCRIPT_DIR.parents[1]
|
||||
DETECT_DIR = SKILLS_DIR / "detect" / "scripts"
|
||||
READ_CONTEXT_DIR = SKILLS_DIR / "read-context" / "scripts"
|
||||
for path in (SCRIPT_DIR, DETECT_DIR, READ_CONTEXT_DIR):
|
||||
sys.path.insert(0, str(path))
|
||||
|
||||
from test_check_writer_candidate import ( # noqa: E402
|
||||
_candidate_body,
|
||||
_requirements,
|
||||
_valid_pair,
|
||||
_rehash,
|
||||
)
|
||||
from run_writer_pipeline import ( # noqa: E402
|
||||
CasToken,
|
||||
InMemoryCasStateStore,
|
||||
PipelineError,
|
||||
atomic_write_json,
|
||||
run_writer_pipeline,
|
||||
)
|
||||
from writer_contract import retrieval_identity # noqa: E402
|
||||
|
||||
|
||||
def _advance_context(context: dict, attempt: int) -> dict:
|
||||
"""模拟可信上下文层为下一次重写生成新 attempt 和快照。"""
|
||||
|
||||
advanced = copy.deepcopy(context)
|
||||
advanced["attempt"] = attempt
|
||||
advanced["contextSnapshot"]["contextSha256"] = retrieval_identity(advanced)
|
||||
return advanced
|
||||
|
||||
|
||||
def _writer_output(context: dict, candidate_version: int, *, body: str | None = None) -> dict:
|
||||
"""生成与本次 attempt、快照和候选版本绑定的 fake 写手结果。"""
|
||||
|
||||
from test_run_writer import _bound_output
|
||||
|
||||
output = _bound_output(context, body=body or _candidate_body())
|
||||
output["candidateVersion"] = candidate_version
|
||||
return output
|
||||
|
||||
|
||||
def _semantic_pass(_context: dict, candidate: dict, _mechanical_report: dict) -> dict:
|
||||
"""模拟只审语义项且无阻塞问题的模型 detector。"""
|
||||
|
||||
return {
|
||||
"status": "passed",
|
||||
"candidateVersion": candidate["candidateVersion"],
|
||||
"candidateSha256": candidate["candidateSha256"],
|
||||
"blockingFailures": [],
|
||||
"suggestions": [],
|
||||
}
|
||||
|
||||
|
||||
class RunWriterPipelineTest(unittest.TestCase):
|
||||
def _assert_terminal_failure(
|
||||
self,
|
||||
*,
|
||||
context: dict,
|
||||
writer,
|
||||
evidence_provider,
|
||||
semantic_detector,
|
||||
expected_code: str,
|
||||
) -> None:
|
||||
"""断言编排错误不会遗留 CHECKING,并且发布稳定拒绝结果。"""
|
||||
|
||||
store = InMemoryCasStateStore()
|
||||
with tempfile.TemporaryDirectory() as directory:
|
||||
result_path = pathlib.Path(directory) / "result.json"
|
||||
with self.assertRaises(PipelineError) as raised:
|
||||
run_writer_pipeline(
|
||||
context=context,
|
||||
requirements=_requirements(),
|
||||
writer=writer,
|
||||
evidence_provider=evidence_provider,
|
||||
semantic_detector=semantic_detector,
|
||||
state_store=store,
|
||||
result_path=result_path,
|
||||
)
|
||||
|
||||
self.assertEqual(raised.exception.code, expected_code)
|
||||
self.assertEqual(store.latest(context["runId"]).state, "REJECTED")
|
||||
self.assertEqual(raised.exception.result["status"], "REJECTED")
|
||||
self.assertEqual(raised.exception.result["failureCode"], expected_code)
|
||||
self.assertEqual(
|
||||
json.loads(result_path.read_text(encoding="utf-8")),
|
||||
raised.exception.result,
|
||||
)
|
||||
|
||||
def test_writer_error_paths_publish_rejected_terminal_result(self):
|
||||
"""写手边界错误必须收敛为拒绝终态,不能把状态留在 CHECKING。"""
|
||||
|
||||
context, _ = _valid_pair()
|
||||
|
||||
def stale_writer(current: dict, candidate_version: int, _repair: list[dict]) -> dict:
|
||||
output = _writer_output(current, candidate_version)
|
||||
output["attempt"] = current["attempt"] - 1
|
||||
return output
|
||||
|
||||
def adapter_failure(*_args) -> dict:
|
||||
raise PipelineError("WRITER_ADAPTER_TIMEOUT", "写手 adapter 超时")
|
||||
|
||||
cases = (
|
||||
("WRITER_OUTPUT_INVALID", lambda *_args: ["not-an-object"]),
|
||||
("WRITER_OUTPUT_STALE", stale_writer),
|
||||
("WRITER_ADAPTER_TIMEOUT", adapter_failure),
|
||||
)
|
||||
for expected_code, writer in cases:
|
||||
with self.subTest(expected_code=expected_code):
|
||||
self._assert_terminal_failure(
|
||||
context=copy.deepcopy(context),
|
||||
writer=writer,
|
||||
evidence_provider=lambda *_args: self.fail("写手失败后不得补证"),
|
||||
semantic_detector=_semantic_pass,
|
||||
expected_code=expected_code,
|
||||
)
|
||||
|
||||
def test_invalid_reassembled_context_publishes_rejected_terminal_result(self):
|
||||
"""补证层返回非法上下文时,当前候选必须保持 REJECTED 并发布原因。"""
|
||||
|
||||
context, _ = _valid_pair()
|
||||
|
||||
def writer(current: dict, candidate_version: int, _repair: list[dict]) -> dict:
|
||||
output = _writer_output(current, candidate_version)
|
||||
output["evidenceRequests"] = [
|
||||
{"requestId": "request-1", "query": "旧徽章", "reason": "需要表现证据", "priority": "high"}
|
||||
]
|
||||
return output
|
||||
|
||||
self._assert_terminal_failure(
|
||||
context=context,
|
||||
writer=writer,
|
||||
evidence_provider=lambda *_args: {"schemaVersion": "invalid-context"},
|
||||
semantic_detector=_semantic_pass,
|
||||
expected_code="REASSEMBLED_CONTEXT_INVALID",
|
||||
)
|
||||
|
||||
def test_semantic_detector_error_paths_publish_rejected_terminal_result(self):
|
||||
"""语义 detector 的调用、结构和旧结果错误都必须失败关闭。"""
|
||||
|
||||
context, _ = _valid_pair()
|
||||
|
||||
def writer(current: dict, candidate_version: int, _repair: list[dict]) -> dict:
|
||||
return _writer_output(current, candidate_version)
|
||||
|
||||
def detector_failure(*_args):
|
||||
raise RuntimeError("detector unavailable")
|
||||
|
||||
def invalid_shape(*_args):
|
||||
return ["not-an-object"]
|
||||
|
||||
def invalid_status(_context: dict, candidate: dict, _mechanical_report: dict) -> dict:
|
||||
report = _semantic_pass(_context, candidate, _mechanical_report)
|
||||
report["status"] = "unknown"
|
||||
return report
|
||||
|
||||
def rejected_without_failures(
|
||||
_context: dict, candidate: dict, _mechanical_report: dict
|
||||
) -> dict:
|
||||
report = _semantic_pass(_context, candidate, _mechanical_report)
|
||||
report["status"] = "rejected"
|
||||
return report
|
||||
|
||||
def passed_with_failures(
|
||||
_context: dict, candidate: dict, _mechanical_report: dict
|
||||
) -> dict:
|
||||
report = _semantic_pass(_context, candidate, _mechanical_report)
|
||||
report["blockingFailures"] = [
|
||||
{"code": "KNOWLEDGE_SCOPE_VIOLATION", "message": "角色知情范围冲突"}
|
||||
]
|
||||
return report
|
||||
|
||||
def invalid_failure_shape(
|
||||
_context: dict, candidate: dict, _mechanical_report: dict
|
||||
) -> dict:
|
||||
report = _semantic_pass(_context, candidate, _mechanical_report)
|
||||
report["status"] = "rejected"
|
||||
report["blockingFailures"] = ["not-an-object"]
|
||||
return report
|
||||
|
||||
def stale_report(_context: dict, candidate: dict, _mechanical_report: dict) -> dict:
|
||||
report = _semantic_pass(_context, candidate, _mechanical_report)
|
||||
report["candidateVersion"] = candidate["candidateVersion"] - 1
|
||||
return report
|
||||
|
||||
cases = (
|
||||
("SEMANTIC_DETECTOR_FAILED", detector_failure),
|
||||
("SEMANTIC_DETECTOR_INVALID", invalid_shape),
|
||||
("SEMANTIC_DETECTOR_INVALID", invalid_status),
|
||||
("SEMANTIC_DETECTOR_INVALID", rejected_without_failures),
|
||||
("SEMANTIC_DETECTOR_INVALID", passed_with_failures),
|
||||
("SEMANTIC_DETECTOR_INVALID", invalid_failure_shape),
|
||||
("SEMANTIC_DETECTOR_STALE", stale_report),
|
||||
)
|
||||
for expected_code, semantic_detector in cases:
|
||||
with self.subTest(expected_code=expected_code, detector=semantic_detector.__name__):
|
||||
self._assert_terminal_failure(
|
||||
context=copy.deepcopy(context),
|
||||
writer=writer,
|
||||
evidence_provider=lambda *_args: self.fail("detector 失败后不得补证"),
|
||||
semantic_detector=semantic_detector,
|
||||
expected_code=expected_code,
|
||||
)
|
||||
|
||||
def test_invalid_mechanical_report_publishes_rejected_terminal_result(self):
|
||||
"""机械 detector 非法结构不得裸抛或把当前状态留在 CHECKING。"""
|
||||
|
||||
context, _ = _valid_pair()
|
||||
|
||||
def writer(current: dict, candidate_version: int, _repair: list[dict]) -> dict:
|
||||
return _writer_output(current, candidate_version)
|
||||
|
||||
invalid_reports = (
|
||||
["not-an-object"],
|
||||
{"schemaVersion": "wrong-version", "passed": True, "blockingFailures": []},
|
||||
{
|
||||
"schemaVersion": "writer-detector-report-v1",
|
||||
"runId": context["runId"],
|
||||
"attempt": context["attempt"],
|
||||
"candidateVersion": 1,
|
||||
"candidateSha256": _writer_output(context, 1)["candidateSha256"],
|
||||
"passed": True,
|
||||
"blockingFailures": [{"code": "HARD_EVENT_MISSING"}],
|
||||
},
|
||||
)
|
||||
for report in invalid_reports:
|
||||
with self.subTest(report=report):
|
||||
with mock.patch(
|
||||
"run_writer_pipeline.check_writer_candidate", return_value=report
|
||||
):
|
||||
self._assert_terminal_failure(
|
||||
context=copy.deepcopy(context),
|
||||
writer=writer,
|
||||
evidence_provider=lambda *_args: self.fail("机械报告非法后不得补证"),
|
||||
semantic_detector=_semantic_pass,
|
||||
expected_code="MECHANICAL_DETECTOR_INVALID",
|
||||
)
|
||||
|
||||
def test_concurrent_new_attempt_blocks_old_evidence_failure_publication(self):
|
||||
"""补证等待期被并发取消时,旧 attempt 不得覆盖新 attempt 的结果。"""
|
||||
|
||||
context, _ = _valid_pair()
|
||||
store = InMemoryCasStateStore()
|
||||
|
||||
def writer(current: dict, candidate_version: int, _repair: list[dict]) -> dict:
|
||||
output = _writer_output(current, candidate_version)
|
||||
output["evidenceRequests"] = [
|
||||
{"requestId": "request-1", "query": "旧徽章", "reason": "需要表现证据", "priority": "high"}
|
||||
]
|
||||
return output
|
||||
|
||||
def concurrent_provider(_current: dict, _requests: list[dict], _attempt: int) -> dict:
|
||||
checking = store.latest(context["runId"])
|
||||
rejected = store.transition(checking, "REJECTED")
|
||||
self.assertIsNotNone(rejected)
|
||||
next_draft = store.start_next(rejected, attempt=2, candidate_version=2)
|
||||
self.assertIsNotNone(next_draft)
|
||||
return {"schemaVersion": "invalid-context"}
|
||||
|
||||
with tempfile.TemporaryDirectory() as directory:
|
||||
result_path = pathlib.Path(directory) / "result.json"
|
||||
with self.assertRaises(PipelineError) as raised:
|
||||
run_writer_pipeline(
|
||||
context=context,
|
||||
requirements=_requirements(),
|
||||
writer=writer,
|
||||
evidence_provider=concurrent_provider,
|
||||
semantic_detector=_semantic_pass,
|
||||
state_store=store,
|
||||
result_path=result_path,
|
||||
)
|
||||
|
||||
self.assertEqual(raised.exception.code, "CAS_CONFLICT")
|
||||
self.assertEqual(store.latest(context["runId"]).state, "DRAFT")
|
||||
self.assertEqual(store.latest(context["runId"]).attempt, 2)
|
||||
self.assertFalse(result_path.exists())
|
||||
|
||||
def test_concurrent_new_attempt_blocks_old_rewrite_context_failure_publication(self):
|
||||
"""普通重写组装失败也只能凭当前 CHECKING token 发布终态。"""
|
||||
|
||||
context, _ = _valid_pair()
|
||||
store = InMemoryCasStateStore()
|
||||
|
||||
def writer(current: dict, candidate_version: int, _repair: list[dict]) -> dict:
|
||||
body = _candidate_body().replace("完成围攻突围", "完成撤离")
|
||||
return _writer_output(current, candidate_version, body=body)
|
||||
|
||||
def concurrent_context_failure(*_args, **_kwargs):
|
||||
current = store.latest(context["runId"])
|
||||
rejected = (
|
||||
store.transition(current, "REJECTED")
|
||||
if current.state == "CHECKING"
|
||||
else current
|
||||
)
|
||||
self.assertIsNotNone(rejected)
|
||||
next_draft = store.start_next(rejected, attempt=2, candidate_version=2)
|
||||
self.assertIsNotNone(next_draft)
|
||||
raise PipelineError("REASSEMBLED_CONTEXT_INVALID", "模拟上下文组装失败")
|
||||
|
||||
with tempfile.TemporaryDirectory() as directory:
|
||||
result_path = pathlib.Path(directory) / "result.json"
|
||||
with mock.patch(
|
||||
"run_writer_pipeline._next_context_without_new_evidence",
|
||||
side_effect=concurrent_context_failure,
|
||||
), self.assertRaises(PipelineError) as raised:
|
||||
run_writer_pipeline(
|
||||
context=context,
|
||||
requirements=_requirements(),
|
||||
writer=writer,
|
||||
evidence_provider=lambda *_args: self.fail("普通重写不得补证"),
|
||||
semantic_detector=_semantic_pass,
|
||||
state_store=store,
|
||||
result_path=result_path,
|
||||
)
|
||||
|
||||
self.assertEqual(raised.exception.code, "CAS_CONFLICT")
|
||||
self.assertEqual(store.latest(context["runId"]).state, "DRAFT")
|
||||
self.assertEqual(store.latest(context["runId"]).attempt, 2)
|
||||
self.assertFalse(result_path.exists())
|
||||
|
||||
def test_dynamic_length_failure_never_reaches_semantic_detector(self):
|
||||
"""短候选只能进入机械返修,不能靠语义 detector 放行。"""
|
||||
|
||||
context, _ = _valid_pair()
|
||||
|
||||
def writer(current: dict, candidate_version: int, _repair: list[dict]) -> dict:
|
||||
return _writer_output(current, candidate_version, body="文" * 1000)
|
||||
|
||||
self._assert_terminal_failure(
|
||||
context=context,
|
||||
writer=writer,
|
||||
evidence_provider=lambda *_args: self.fail("篇幅失败不得补证"),
|
||||
semantic_detector=lambda *_args: self.fail("机械失败不得调用语义 detector"),
|
||||
expected_code="REWRITE_LIMIT_REACHED",
|
||||
)
|
||||
|
||||
def test_evidence_requests_are_limited_to_three(self):
|
||||
context, _ = _valid_pair()
|
||||
|
||||
def writer(current: dict, candidate_version: int, _repair: list[dict]) -> dict:
|
||||
output = _writer_output(current, candidate_version)
|
||||
output["evidenceRequests"] = [
|
||||
{"requestId": f"request-{index}", "query": "补证", "reason": "缺口", "priority": "high"}
|
||||
for index in range(4)
|
||||
]
|
||||
return output
|
||||
|
||||
self._assert_terminal_failure(
|
||||
context=context,
|
||||
writer=writer,
|
||||
evidence_provider=lambda *_args: self.fail("超过上限后不得取证"),
|
||||
semantic_detector=_semantic_pass,
|
||||
expected_code="EVIDENCE_REQUEST_LIMIT_REACHED",
|
||||
)
|
||||
|
||||
def test_rewrite_is_limited_to_two_and_returns_stable_failure(self):
|
||||
context, _ = _valid_pair()
|
||||
calls: list[int] = []
|
||||
|
||||
def writer(current: dict, candidate_version: int, _repair: list[dict]) -> dict:
|
||||
calls.append(candidate_version)
|
||||
body = _candidate_body().replace("完成围攻突围", "完成撤离")
|
||||
return _writer_output(current, candidate_version, body=body)
|
||||
|
||||
self._assert_terminal_failure(
|
||||
context=context,
|
||||
writer=writer,
|
||||
evidence_provider=lambda current, _requests, attempt: _advance_context(
|
||||
current, attempt
|
||||
),
|
||||
semantic_detector=_semantic_pass,
|
||||
expected_code="REWRITE_LIMIT_REACHED",
|
||||
)
|
||||
self.assertEqual(calls, [1, 2, 3])
|
||||
|
||||
def test_evidence_rewrite_limit_publishes_stable_terminal_result(self):
|
||||
"""连续补证触及重写上限时也必须原子发布 REJECTED 结果。"""
|
||||
|
||||
context, _ = _valid_pair()
|
||||
|
||||
def writer(current: dict, candidate_version: int, _repair: list[dict]) -> dict:
|
||||
output = _writer_output(current, candidate_version)
|
||||
output["evidenceRequests"] = [
|
||||
{
|
||||
"requestId": f"request-{candidate_version}",
|
||||
"query": "旧徽章",
|
||||
"reason": "需要表现证据",
|
||||
"priority": "high",
|
||||
}
|
||||
]
|
||||
return output
|
||||
|
||||
self._assert_terminal_failure(
|
||||
context=context,
|
||||
writer=writer,
|
||||
evidence_provider=lambda current, _requests, attempt: _advance_context(
|
||||
current, attempt
|
||||
),
|
||||
semantic_detector=_semantic_pass,
|
||||
expected_code="REWRITE_LIMIT_REACHED",
|
||||
)
|
||||
|
||||
def test_evidence_request_reassembles_context_then_rewrites_once(self):
|
||||
context, _ = _valid_pair()
|
||||
calls: list[tuple[int, int]] = []
|
||||
evidence_calls: list[list[str]] = []
|
||||
|
||||
def writer(current: dict, candidate_version: int, _repair: list[dict]) -> dict:
|
||||
calls.append((current["attempt"], candidate_version))
|
||||
output = _writer_output(current, candidate_version)
|
||||
if candidate_version == 1:
|
||||
output["evidenceRequests"] = [
|
||||
{"requestId": "request-1", "query": "旧徽章", "reason": "需要表现证据", "priority": "high"}
|
||||
]
|
||||
return output
|
||||
|
||||
def provide(current: dict, requests: list[dict], attempt: int) -> dict:
|
||||
evidence_calls.append([item["requestId"] for item in requests])
|
||||
return _advance_context(current, attempt)
|
||||
|
||||
store = InMemoryCasStateStore()
|
||||
result = run_writer_pipeline(
|
||||
context=context,
|
||||
requirements=_requirements(),
|
||||
writer=writer,
|
||||
evidence_provider=provide,
|
||||
semantic_detector=_semantic_pass,
|
||||
state_store=store,
|
||||
)
|
||||
|
||||
self.assertEqual(result["status"], "PASSED")
|
||||
self.assertEqual(result["evidenceRequestCount"], 1)
|
||||
self.assertEqual(result["rewriteCount"], 1)
|
||||
self.assertEqual(calls, [(1, 1), (2, 2)])
|
||||
self.assertEqual(evidence_calls, [["request-1"]])
|
||||
|
||||
def test_cas_rejects_late_result_from_old_attempt_and_candidate_version(self):
|
||||
store = InMemoryCasStateStore()
|
||||
first = store.create("run-cas", attempt=1, candidate_version=1)
|
||||
checking = store.transition(first, "CHECKING")
|
||||
rejected = store.transition(checking, "REJECTED")
|
||||
second = store.start_next(rejected, attempt=2, candidate_version=2)
|
||||
|
||||
stale_success = store.transition(checking, "PASSED")
|
||||
|
||||
self.assertIsNone(stale_success)
|
||||
self.assertEqual(store.latest("run-cas"), second)
|
||||
self.assertEqual(store.latest("run-cas").state, "DRAFT")
|
||||
|
||||
def test_atomic_result_fsyncs_before_replace(self):
|
||||
events: list[str] = []
|
||||
with tempfile.TemporaryDirectory() as directory:
|
||||
target = pathlib.Path(directory) / "result.json"
|
||||
real_fsync = __import__("os").fsync
|
||||
real_replace = __import__("os").replace
|
||||
|
||||
def recording_fsync(fd: int) -> None:
|
||||
events.append("fsync")
|
||||
real_fsync(fd)
|
||||
|
||||
def recording_replace(source: str, destination: str) -> None:
|
||||
events.append("replace")
|
||||
real_replace(source, destination)
|
||||
|
||||
with mock.patch("run_writer_pipeline.os.fsync", side_effect=recording_fsync), mock.patch(
|
||||
"run_writer_pipeline.os.replace", side_effect=recording_replace
|
||||
):
|
||||
atomic_write_json(target, {"status": "PASSED"})
|
||||
|
||||
self.assertTrue(target.exists())
|
||||
self.assertEqual(json.loads(target.read_text(encoding="utf-8")), {"status": "PASSED"})
|
||||
self.assertLess(events.index("fsync"), events.index("replace"))
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
unittest.main()
|
||||
@ -8,6 +8,17 @@ disable-model-invocation: true
|
||||
|
||||
何时用:候选章生成后、用户三决策前的伴随检查(C4);也可对既有章回溯检查。
|
||||
|
||||
## 正文候选 v1 执行顺序
|
||||
|
||||
1. 先运行 `scripts/check_writer_candidate.py`;它只做确定性机械硬门,不调用模型。
|
||||
2. 机械硬门阻塞时直接返回结构化失败码,保留审查轨迹,不调用语义 detector。
|
||||
3. 机械硬门通过后,编排层才可调用 `SemanticDetector(context, candidate, mechanicalReport)` 接口。
|
||||
4. 语义报告必须绑定当前候选版本和正文哈希;异常、非对象、非法状态、阻塞项结构错误或旧结果都 CAS 到 `REJECTED`,并原子发布稳定 `failureCode`。
|
||||
|
||||
当前机械硬门覆盖:WriterContext/WriterOutput 严格合同、上下文与候选哈希绑定、claim Unicode 偏移与证据引用、细纲事件/角色/伏笔/章末钩子锚点、新设定申报及冻结事实冲突。`semanticReview.status=pending` 表示还需要真实语义审查,不能解释为通过。
|
||||
|
||||
**实现状态:本轮只有机械硬门、语义接口和 fake 编排测试,没有真实模型 detector adapter,也没有任何真实模型调用。** 因此测试中的 `PASSED` 只证明接口与有限状态机按合同工作,不构成真实正文语义验收证据。
|
||||
|
||||
## 元数据驱动(本功能的核心机制)
|
||||
|
||||
**检查清单由字段生成,不硬编**:凡 aiContext 含 `detection` 的字段,自动成为一个检查项——
|
||||
@ -29,6 +40,12 @@ schema 给字段加上 detection 用途,检查项自动+1,本 skill 与 detector
|
||||
|
||||
## 输出合同
|
||||
|
||||
正文候选 v1 的机械报告为 `writer-detector-report-v1`,至少包含候选身份、`passed`、`blockingFailures[]` 和状态为 `pending` 的 `semanticReview`。未来真实语义报告只允许:
|
||||
|
||||
`status(passed/rejected) / candidateVersion / candidateSha256 / blockingFailures[] / suggestions[]`
|
||||
|
||||
以下是既有创作文件检测与回溯检查的人工报告合同,不替代上述结构化接口:
|
||||
|
||||
唯一一份报告落 `works/<书>/评审/第NNN章-检测.md`(不入 git)。每条问题:
|
||||
|
||||
`[严重度 高/中/低] 位置(场景N·引原句) | 类型 | 依据(哪张卡哪个字段) | 建议改法`
|
||||
|
||||
339
.claude/skills/detect/scripts/check_writer_candidate.py
Normal file
339
.claude/skills/detect/scripts/check_writer_candidate.py
Normal file
@ -0,0 +1,339 @@
|
||||
#!/usr/bin/env python3
|
||||
"""正文候选机械硬门;不调用模型、数据库或外部服务。"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import hashlib
|
||||
import pathlib
|
||||
import sys
|
||||
from typing import Any, Mapping, Sequence
|
||||
|
||||
SCRIPT_DIR = pathlib.Path(__file__).resolve().parent
|
||||
READ_CONTEXT_DIR = SCRIPT_DIR.parents[1] / "read-context" / "scripts"
|
||||
if str(READ_CONTEXT_DIR) not in sys.path:
|
||||
sys.path.insert(0, str(READ_CONTEXT_DIR))
|
||||
|
||||
from writer_contract import ( # noqa: E402
|
||||
ContractError,
|
||||
han_count,
|
||||
normalize_text,
|
||||
validate_writer_context,
|
||||
validate_writer_output,
|
||||
)
|
||||
|
||||
|
||||
SEMANTIC_CHECKS = (
|
||||
"hard_event_semantics",
|
||||
"claim_truth_alignment",
|
||||
"knowledge_scope",
|
||||
"ability_cost_consistency",
|
||||
)
|
||||
|
||||
|
||||
def _failure(code: str, message: str, **details: Any) -> dict[str, Any]:
|
||||
"""生成稳定且可追踪的阻塞项。"""
|
||||
|
||||
return {"code": code, "message": message, **details}
|
||||
|
||||
|
||||
def _anchors_present(body: str, anchors: Any) -> bool:
|
||||
"""只执行确定性的文本锚点匹配,语义等价判断留给模型 detector。"""
|
||||
|
||||
return isinstance(anchors, list) and bool(anchors) and any(
|
||||
isinstance(anchor, str) and anchor and normalize_text(anchor) in body for anchor in anchors
|
||||
)
|
||||
|
||||
|
||||
def _safe_sequence(value: Any) -> Sequence[Any]:
|
||||
"""把非法数组降为空序列,随后由输入合同阻塞项统一报告。"""
|
||||
|
||||
return value if isinstance(value, list) else ()
|
||||
|
||||
|
||||
def _precheck_candidate_integrity(candidate: Mapping[str, Any]) -> list[dict[str, Any]]:
|
||||
"""在严格 schema 前检查候选与 claim 身份,保留精确失败码。"""
|
||||
|
||||
failures: list[dict[str, Any]] = []
|
||||
body = candidate.get("candidateBody")
|
||||
candidate_hash = candidate.get("candidateSha256")
|
||||
if not isinstance(body, str):
|
||||
return [_failure("OUTPUT_CONTRACT_INVALID", "candidateBody 必须是字符串")]
|
||||
try:
|
||||
normalized_body = normalize_text(body)
|
||||
except ContractError as exc:
|
||||
return [_failure("OUTPUT_CONTRACT_INVALID", str(exc))]
|
||||
expected_hash = "sha256:" + hashlib.sha256(normalized_body.encode("utf-8")).hexdigest()
|
||||
if candidate_hash != expected_hash:
|
||||
failures.append(
|
||||
_failure(
|
||||
"CANDIDATE_HASH_MISMATCH",
|
||||
"candidateSha256 与规范化候选正文不一致",
|
||||
expectedSha256=expected_hash,
|
||||
)
|
||||
)
|
||||
claims = candidate.get("claimLedger")
|
||||
if not isinstance(claims, list):
|
||||
failures.append(_failure("OUTPUT_CONTRACT_INVALID", "claimLedger 必须是数组"))
|
||||
return failures
|
||||
for index, claim in enumerate(claims):
|
||||
if not isinstance(claim, Mapping):
|
||||
failures.append(_failure("OUTPUT_CONTRACT_INVALID", "claim 必须是对象", claimIndex=index))
|
||||
continue
|
||||
if claim.get("candidateSha256") != expected_hash:
|
||||
failures.append(
|
||||
_failure(
|
||||
"CLAIM_HASH_MISMATCH",
|
||||
"claim 未绑定当前候选正文哈希",
|
||||
claimIndex=index,
|
||||
claimId=claim.get("claimId"),
|
||||
)
|
||||
)
|
||||
start = claim.get("startCodePoint")
|
||||
end = claim.get("endCodePoint")
|
||||
if (
|
||||
isinstance(start, bool)
|
||||
or isinstance(end, bool)
|
||||
or not isinstance(start, int)
|
||||
or not isinstance(end, int)
|
||||
or start < 0
|
||||
or end <= start
|
||||
or end > len(normalized_body)
|
||||
):
|
||||
failures.append(
|
||||
_failure(
|
||||
"CLAIM_UNICODE_OFFSET_INVALID",
|
||||
"claim 偏移必须是规范正文的 Unicode code point 左闭右开区间",
|
||||
claimIndex=index,
|
||||
claimId=claim.get("claimId"),
|
||||
)
|
||||
)
|
||||
return failures
|
||||
|
||||
|
||||
def _check_context_binding(
|
||||
context: Mapping[str, Any], candidate: Mapping[str, Any]
|
||||
) -> list[dict[str, Any]]:
|
||||
"""检查候选元数据和所有证据 ID 均来自本次冻结上下文。"""
|
||||
|
||||
failures: list[dict[str, Any]] = []
|
||||
expected = {
|
||||
"runId": context["runId"],
|
||||
"attempt": context["attempt"],
|
||||
"mode": context["mode"],
|
||||
"qualityPolicyVersion": context["qualityPolicyVersion"],
|
||||
"contextSnapshotId": context["contextSnapshot"]["manifestId"],
|
||||
"contextSnapshotSha256": context["contextSnapshot"]["contextSha256"],
|
||||
"acceptanceEligible": context["acceptanceEligible"],
|
||||
}
|
||||
mismatches = [field for field, value in expected.items() if candidate.get(field) != value]
|
||||
if mismatches:
|
||||
failures.append(
|
||||
_failure("CONTEXT_BINDING_MISMATCH", "候选未绑定当前冻结上下文", fields=mismatches)
|
||||
)
|
||||
fact_ids = {item["evidenceId"] for item in context["factEvidence"]}
|
||||
prose_ids = {item["evidenceId"] for item in context["proseEvidence"]}
|
||||
for index, claim in enumerate(_safe_sequence(candidate.get("claimLedger"))):
|
||||
if not isinstance(claim, Mapping):
|
||||
continue
|
||||
fact_id = claim.get("factEvidenceId")
|
||||
prose_id = claim.get("proseEvidenceId")
|
||||
if fact_id not in fact_ids or (prose_id is not None and prose_id not in prose_ids):
|
||||
failures.append(
|
||||
_failure(
|
||||
"CLAIM_EVIDENCE_MISSING",
|
||||
"claim 引用了当前上下文之外的事实或文风证据",
|
||||
claimIndex=index,
|
||||
claimId=claim.get("claimId"),
|
||||
)
|
||||
)
|
||||
if claim.get("coverageState") in {"unsupported", "conflict"}:
|
||||
failures.append(
|
||||
_failure(
|
||||
"FROZEN_FACT_CONFLICT",
|
||||
"claim 使用了不受支持或互相冲突的冻结事实",
|
||||
claimIndex=index,
|
||||
claimId=claim.get("claimId"),
|
||||
)
|
||||
)
|
||||
return failures
|
||||
|
||||
|
||||
def _check_outline_anchors(body: str, requirements: Mapping[str, Any]) -> list[dict[str, Any]]:
|
||||
"""机械检查硬事件、必须出场角色、伏笔动作和章末钩子。"""
|
||||
|
||||
failures: list[dict[str, Any]] = []
|
||||
for item in _safe_sequence(requirements.get("requiredEvents")):
|
||||
if not isinstance(item, Mapping) or not _anchors_present(body, item.get("anchors")):
|
||||
failures.append(
|
||||
_failure(
|
||||
"HARD_EVENT_MISSING",
|
||||
"候选未命中细纲硬事件锚点",
|
||||
requirementId=item.get("requirementId") if isinstance(item, Mapping) else None,
|
||||
)
|
||||
)
|
||||
for character in _safe_sequence(requirements.get("requiredCharacters")):
|
||||
if not isinstance(character, str) or not character or normalize_text(character) not in body:
|
||||
failures.append(
|
||||
_failure("REQUIRED_CHARACTER_MISSING", "细纲要求的角色未出场", character=character)
|
||||
)
|
||||
for item in _safe_sequence(requirements.get("foreshadowingActions")):
|
||||
if not isinstance(item, Mapping) or not _anchors_present(body, item.get("anchors")):
|
||||
failures.append(
|
||||
_failure(
|
||||
"FORESHADOWING_ACTION_MISSING",
|
||||
"候选未命中细纲伏笔动作锚点",
|
||||
requirementId=item.get("requirementId") if isinstance(item, Mapping) else None,
|
||||
)
|
||||
)
|
||||
hook = requirements.get("chapterEndHook")
|
||||
if isinstance(hook, Mapping):
|
||||
max_distance = hook.get("maxDistanceFromEnd")
|
||||
if isinstance(max_distance, bool) or not isinstance(max_distance, int) or max_distance <= 0:
|
||||
failures.append(_failure("DETECTOR_INPUT_INVALID", "章末钩子距离必须是正整数"))
|
||||
else:
|
||||
tail = body[-max_distance:]
|
||||
if not _anchors_present(tail, hook.get("anchors")):
|
||||
failures.append(
|
||||
_failure(
|
||||
"CHAPTER_END_HOOK_MISSING",
|
||||
"候选结尾未命中细纲章末钩子",
|
||||
requirementId=hook.get("requirementId"),
|
||||
)
|
||||
)
|
||||
else:
|
||||
failures.append(_failure("DETECTOR_INPUT_INVALID", "chapterEndHook 必须是对象"))
|
||||
return failures
|
||||
|
||||
|
||||
def _spans_overlap(left_start: int, left_end: int, right_start: int, right_end: int) -> bool:
|
||||
"""判断两个 Unicode code point 左闭右开区间是否相交。"""
|
||||
|
||||
return left_start < right_end and right_start < left_end
|
||||
|
||||
|
||||
def _check_new_settings_and_conflicts(
|
||||
candidate: Mapping[str, Any], requirements: Mapping[str, Any]
|
||||
) -> list[dict[str, Any]]:
|
||||
"""检查可信实体识别结果是否有对应申报,并回显冻结冲突来源。"""
|
||||
|
||||
failures: list[dict[str, Any]] = []
|
||||
declarations = [
|
||||
item
|
||||
for item in _safe_sequence(candidate.get("newSettingDeclarations"))
|
||||
if isinstance(item, Mapping)
|
||||
]
|
||||
for setting in _safe_sequence(requirements.get("detectedNewSettings")):
|
||||
if not isinstance(setting, Mapping):
|
||||
failures.append(_failure("DETECTOR_INPUT_INVALID", "detectedNewSettings 项必须是对象"))
|
||||
continue
|
||||
start = setting.get("startCodePoint")
|
||||
end = setting.get("endCodePoint")
|
||||
matched = False
|
||||
if isinstance(start, int) and not isinstance(start, bool) and isinstance(end, int) and not isinstance(end, bool):
|
||||
matched = any(
|
||||
declaration.get("factType") == setting.get("factType")
|
||||
and isinstance(declaration.get("startCodePoint"), int)
|
||||
and isinstance(declaration.get("endCodePoint"), int)
|
||||
and _spans_overlap(
|
||||
start,
|
||||
end,
|
||||
declaration["startCodePoint"],
|
||||
declaration["endCodePoint"],
|
||||
)
|
||||
for declaration in declarations
|
||||
)
|
||||
if not matched:
|
||||
failures.append(
|
||||
_failure(
|
||||
"NEW_SETTING_UNDECLARED",
|
||||
"候选中的新设定没有对应申报",
|
||||
settingId=setting.get("settingId"),
|
||||
)
|
||||
)
|
||||
for conflict in _safe_sequence(requirements.get("frozenConflicts")):
|
||||
if not isinstance(conflict, Mapping):
|
||||
failures.append(_failure("DETECTOR_INPUT_INVALID", "frozenConflicts 项必须是对象"))
|
||||
continue
|
||||
failures.append(
|
||||
_failure(
|
||||
"FROZEN_FACT_CONFLICT",
|
||||
str(conflict.get("message") or "候选与冻结事实冲突"),
|
||||
conflictId=conflict.get("conflictId"),
|
||||
sourceRef=conflict.get("sourceRef"),
|
||||
)
|
||||
)
|
||||
return failures
|
||||
|
||||
|
||||
def check_writer_candidate(
|
||||
context: Mapping[str, Any], candidate: Mapping[str, Any], requirements: Mapping[str, Any]
|
||||
) -> dict[str, Any]:
|
||||
"""运行不含模型调用的机械硬门,并返回结构化 detector 报告。"""
|
||||
|
||||
failures: list[dict[str, Any]] = []
|
||||
if not isinstance(context, Mapping) or not isinstance(candidate, Mapping) or not isinstance(requirements, Mapping):
|
||||
failures.append(_failure("DETECTOR_INPUT_INVALID", "context、candidate、requirements 必须是对象"))
|
||||
return {
|
||||
"schemaVersion": "writer-detector-report-v1",
|
||||
"passed": False,
|
||||
"blockingFailures": failures,
|
||||
"semanticReview": {"status": "pending", "checks": list(SEMANTIC_CHECKS)},
|
||||
}
|
||||
try:
|
||||
normalized_context = validate_writer_context(context)
|
||||
except ContractError as exc:
|
||||
failures.append(_failure("CONTEXT_CONTRACT_INVALID", str(exc)))
|
||||
normalized_context = None
|
||||
failures.extend(_precheck_candidate_integrity(candidate))
|
||||
try:
|
||||
normalized_candidate = validate_writer_output(candidate)
|
||||
except ContractError as exc:
|
||||
failures.append(_failure("OUTPUT_CONTRACT_INVALID", str(exc)))
|
||||
normalized_candidate = None
|
||||
if normalized_context is not None:
|
||||
failures.extend(_check_context_binding(normalized_context, candidate))
|
||||
body = candidate.get("candidateBody")
|
||||
if isinstance(body, str):
|
||||
normalized_body = normalize_text(body)
|
||||
if normalized_context is not None:
|
||||
contract = normalized_context["outputContract"]
|
||||
actual_han_chars = han_count(normalized_body)
|
||||
if not contract["minChars"] <= actual_han_chars <= contract["maxChars"]:
|
||||
failures.append(
|
||||
_failure(
|
||||
"CANDIDATE_LENGTH_OUT_OF_RANGE",
|
||||
"候选正文汉字数超出动态篇幅合同",
|
||||
actualHanChars=actual_han_chars,
|
||||
minChars=contract["minChars"],
|
||||
maxChars=contract["maxChars"],
|
||||
targetChars=contract["targetChars"],
|
||||
)
|
||||
)
|
||||
failures.extend(_check_outline_anchors(normalized_body, requirements))
|
||||
failures.extend(_check_new_settings_and_conflicts(candidate, requirements))
|
||||
# 按 code 与定位去重,避免严格合同与预检重复报告同一处问题。
|
||||
unique: list[dict[str, Any]] = []
|
||||
seen: set[tuple[Any, ...]] = set()
|
||||
for item in failures:
|
||||
key = (item.get("code"), item.get("claimIndex"), item.get("requirementId"), item.get("settingId"), item.get("conflictId"))
|
||||
if key not in seen:
|
||||
seen.add(key)
|
||||
unique.append(item)
|
||||
source = normalized_candidate or candidate
|
||||
return {
|
||||
"schemaVersion": "writer-detector-report-v1",
|
||||
"runId": source.get("runId"),
|
||||
"attempt": source.get("attempt"),
|
||||
"candidateVersion": source.get("candidateVersion"),
|
||||
"candidateSha256": source.get("candidateSha256"),
|
||||
"passed": not unique,
|
||||
"blockingFailures": unique,
|
||||
"semanticReview": {
|
||||
"status": "pending",
|
||||
"checks": list(SEMANTIC_CHECKS),
|
||||
"inputContract": "context + candidate + mechanicalReport -> semanticDetectorReport",
|
||||
},
|
||||
}
|
||||
|
||||
|
||||
__all__ = ["SEMANTIC_CHECKS", "check_writer_candidate"]
|
||||
199
.claude/skills/detect/scripts/test_check_writer_candidate.py
Normal file
199
.claude/skills/detect/scripts/test_check_writer_candidate.py
Normal file
@ -0,0 +1,199 @@
|
||||
#!/usr/bin/env python3
|
||||
"""正文候选机械硬门的细纲、证据和 Unicode 绑定测试。"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import copy
|
||||
import hashlib
|
||||
import pathlib
|
||||
import sys
|
||||
import unittest
|
||||
|
||||
SCRIPT_DIR = pathlib.Path(__file__).resolve().parent
|
||||
SKILLS_DIR = SCRIPT_DIR.parents[1]
|
||||
CONTINUATION_DIR = SKILLS_DIR / "continuation" / "scripts"
|
||||
READ_CONTEXT_DIR = SKILLS_DIR / "read-context" / "scripts"
|
||||
for path in (SCRIPT_DIR, CONTINUATION_DIR, READ_CONTEXT_DIR):
|
||||
sys.path.insert(0, str(path))
|
||||
|
||||
from test_run_writer import _bound_context, _bound_output # noqa: E402
|
||||
from check_writer_candidate import check_writer_candidate # noqa: E402
|
||||
from writer_contract import han_count # noqa: E402
|
||||
|
||||
|
||||
def _candidate_body() -> str:
|
||||
"""构造同时满足四类细纲机械锚点和动态篇幅的候选正文。"""
|
||||
|
||||
prefix = "林澈完成围攻突围,又把旧徽章压回掌心。"
|
||||
hook = "城门忽然打开"
|
||||
filler_count = 4000 - han_count(prefix) - han_count(hook)
|
||||
return prefix + "文" * filler_count + hook
|
||||
|
||||
|
||||
def _requirements() -> dict:
|
||||
"""构造可信上下文层派生的机械硬约束。"""
|
||||
|
||||
return {
|
||||
"requiredEvents": [
|
||||
{"requirementId": "event-1", "anchors": ["完成围攻突围"]},
|
||||
],
|
||||
"requiredCharacters": ["林澈"],
|
||||
"foreshadowingActions": [
|
||||
{"requirementId": "foreshadow-1", "anchors": ["旧徽章"]},
|
||||
],
|
||||
"chapterEndHook": {
|
||||
"requirementId": "hook-1",
|
||||
"anchors": ["城门忽然打开"],
|
||||
"maxDistanceFromEnd": 20,
|
||||
},
|
||||
"detectedNewSettings": [],
|
||||
"frozenConflicts": [],
|
||||
}
|
||||
|
||||
|
||||
def _rehash(output: dict) -> None:
|
||||
"""正文变更后同步候选和 claim 哈希,供单项硬门测试使用。"""
|
||||
|
||||
digest = "sha256:" + hashlib.sha256(output["candidateBody"].encode("utf-8")).hexdigest()
|
||||
output["candidateSha256"] = digest
|
||||
for claim in output["claimLedger"]:
|
||||
claim["candidateSha256"] = digest
|
||||
|
||||
|
||||
def _valid_pair() -> tuple[dict, dict]:
|
||||
"""构造一组能通过严格合同和机械硬门的上下文与候选。"""
|
||||
|
||||
context = _bound_context()
|
||||
output = _bound_output(context, body=_candidate_body())
|
||||
return context, output
|
||||
|
||||
|
||||
class CheckWriterCandidateTest(unittest.TestCase):
|
||||
def test_missing_event_character_foreshadow_or_hook_blocks(self):
|
||||
context, output = _valid_pair()
|
||||
mutations = {
|
||||
"HARD_EVENT_MISSING": ("完成围攻突围", "完成撤离"),
|
||||
"REQUIRED_CHARACTER_MISSING": ("林澈", "无名人"),
|
||||
"FORESHADOWING_ACTION_MISSING": ("旧徽章", "旧物件"),
|
||||
"CHAPTER_END_HOOK_MISSING": ("城门忽然打开", "风声渐渐平息"),
|
||||
}
|
||||
for expected_code, (old, new) in mutations.items():
|
||||
candidate = copy.deepcopy(output)
|
||||
candidate["candidateBody"] = candidate["candidateBody"].replace(old, new)
|
||||
_rehash(candidate)
|
||||
report = check_writer_candidate(context, candidate, _requirements())
|
||||
with self.subTest(expected_code=expected_code):
|
||||
self.assertFalse(report["passed"])
|
||||
self.assertIn(expected_code, {item["code"] for item in report["blockingFailures"]})
|
||||
|
||||
def test_candidate_claim_hash_unicode_offset_and_evidence_reference_block(self):
|
||||
context, output = _valid_pair()
|
||||
cases = {}
|
||||
|
||||
bad_candidate_hash = copy.deepcopy(output)
|
||||
bad_candidate_hash["candidateSha256"] = "sha256:" + "f" * 64
|
||||
cases["CANDIDATE_HASH_MISMATCH"] = bad_candidate_hash
|
||||
|
||||
bad_claim_hash = copy.deepcopy(output)
|
||||
bad_claim_hash["claimLedger"][0]["candidateSha256"] = "sha256:" + "e" * 64
|
||||
cases["CLAIM_HASH_MISMATCH"] = bad_claim_hash
|
||||
|
||||
bad_unicode_offset = copy.deepcopy(output)
|
||||
bad_unicode_offset["candidateBody"] = "林澈🙂" + bad_unicode_offset["candidateBody"][2:]
|
||||
_rehash(bad_unicode_offset)
|
||||
bad_unicode_offset["claimLedger"][0]["endCodePoint"] = len(bad_unicode_offset["candidateBody"]) + 1
|
||||
cases["CLAIM_UNICODE_OFFSET_INVALID"] = bad_unicode_offset
|
||||
|
||||
bad_evidence = copy.deepcopy(output)
|
||||
bad_evidence["claimLedger"][0]["factEvidenceId"] = "fact-missing"
|
||||
cases["CLAIM_EVIDENCE_MISSING"] = bad_evidence
|
||||
|
||||
for expected_code, candidate in cases.items():
|
||||
report = check_writer_candidate(context, candidate, _requirements())
|
||||
with self.subTest(expected_code=expected_code):
|
||||
self.assertFalse(report["passed"])
|
||||
self.assertIn(expected_code, {item["code"] for item in report["blockingFailures"]})
|
||||
|
||||
def test_candidate_body_outside_dynamic_length_contract_blocks(self):
|
||||
"""机械门必须独立复核动态篇幅,不能只信 writer adapter。"""
|
||||
|
||||
context, output = _valid_pair()
|
||||
output["candidateBody"] = "文" * 1000
|
||||
_rehash(output)
|
||||
|
||||
report = check_writer_candidate(context, output, _requirements())
|
||||
|
||||
self.assertFalse(report["passed"])
|
||||
self.assertIn(
|
||||
"CANDIDATE_LENGTH_OUT_OF_RANGE",
|
||||
{item["code"] for item in report["blockingFailures"]},
|
||||
)
|
||||
|
||||
def test_new_setting_must_have_overlapping_declaration(self):
|
||||
context, output = _valid_pair()
|
||||
marker = "赤霄城"
|
||||
insertion_start = len("林澈完成围攻突围,又把旧徽章压回掌心。")
|
||||
output["candidateBody"] = (
|
||||
output["candidateBody"][:insertion_start]
|
||||
+ marker
|
||||
+ output["candidateBody"][insertion_start + len(marker):]
|
||||
)
|
||||
_rehash(output)
|
||||
requirements = _requirements()
|
||||
requirements["detectedNewSettings"] = [
|
||||
{
|
||||
"settingId": "setting-1",
|
||||
"factType": "location",
|
||||
"startCodePoint": insertion_start,
|
||||
"endCodePoint": insertion_start + len(marker),
|
||||
}
|
||||
]
|
||||
|
||||
blocked = check_writer_candidate(context, output, requirements)
|
||||
self.assertIn("NEW_SETTING_UNDECLARED", {item["code"] for item in blocked["blockingFailures"]})
|
||||
|
||||
output["newSettingDeclarations"] = [
|
||||
{
|
||||
"declarationId": "declaration-1",
|
||||
"factType": "location",
|
||||
"text": marker,
|
||||
"startCodePoint": insertion_start,
|
||||
"endCodePoint": insertion_start + len(marker),
|
||||
}
|
||||
]
|
||||
passed = check_writer_candidate(context, output, requirements)
|
||||
self.assertTrue(passed["passed"])
|
||||
|
||||
def test_frozen_fact_conflict_blocks_with_traceable_source(self):
|
||||
context, output = _valid_pair()
|
||||
requirements = _requirements()
|
||||
requirements["frozenConflicts"] = [
|
||||
{
|
||||
"conflictId": "conflict-1",
|
||||
"message": "候选把已毁地点写成完好",
|
||||
"sourceRef": "state:8@state-v1",
|
||||
}
|
||||
]
|
||||
|
||||
report = check_writer_candidate(context, output, requirements)
|
||||
|
||||
self.assertFalse(report["passed"])
|
||||
failure = next(item for item in report["blockingFailures"] if item["code"] == "FROZEN_FACT_CONFLICT")
|
||||
self.assertEqual(failure["sourceRef"], "state:8@state-v1")
|
||||
|
||||
def test_valid_candidate_passes_mechanical_gate_and_exposes_semantic_interface(self):
|
||||
context, output = _valid_pair()
|
||||
|
||||
report = check_writer_candidate(context, output, _requirements())
|
||||
|
||||
self.assertTrue(report["passed"])
|
||||
self.assertEqual(report["candidateSha256"], output["candidateSha256"])
|
||||
self.assertEqual(report["semanticReview"]["status"], "pending")
|
||||
self.assertEqual(
|
||||
set(report["semanticReview"]["checks"]),
|
||||
{"hard_event_semantics", "claim_truth_alignment", "knowledge_scope", "ability_cost_consistency"},
|
||||
)
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
unittest.main()
|
||||
16
README.md
16
README.md
@ -77,7 +77,7 @@ agent-example/
|
||||
|
||||
| agent | SoT 身份(专题-06 §3 / 专题-03 §2.3) | 做什么 | 用在 |
|
||||
|---|---|---|---|
|
||||
| writer | 一层 · 写作(Writing 槽位) | 整章正文候选:2500–3500 字、AI 味黑名单零命中、新设定申报注释块;只进待审 | C4 |
|
||||
| writer | 一层 · 写作(Writing 槽位) | 无工具、无会话,只消费 stdin 的 `WriterContext v1`;动态篇幅算法和入口合同已实现,上游编排接线待 Task 8;返回严格 `WriterOutput v1`,本轮只验收 continuation | C4 |
|
||||
| planner | 一层 · 规划(规划链槽位) | 设定包/大纲/细纲候选;字段全覆盖或标「字段存疑」;未确认不进后续生成上下文 | C2 |
|
||||
| extractor | 一层 · 拆书/分析(Analysis 槽位) | 两用场:**拆书**(B2,按库内 schema 产知识行草稿+出处链;系统级只产抽象范式、不留原文)与**章后抽取**(C5);冲突标「⚠ 冲突待裁决」不覆盖 | B2 / C5 |
|
||||
| detector | 一层 · 检测(Detection 槽位) | 一致性/伏笔/设定违背检查,只产解释与定位落 `评审/`,不改正文与知识 | C4 伴随 |
|
||||
@ -87,14 +87,14 @@ SoT 第三层「知识基座」(检索/切块/入库,明文不是 app)在
|
||||
|
||||
**prompt 三段式**(登记表见 `meta/chains/README.md`):身份段(agent .md,跨功能共享)× 功能指令段(**功能 skill**,每 scenario 一只,派发只加载本次一只)× L0 任务特化——不做一个 agent 一个大 prompt(防指令稀释),也不按功能裂 agent(防身份复制)。**skill 辅助映射**:extractor 最重(db 读 schema 与既有知识、产出经 db 写入 draft、embed 嵌入);writer/planner/detector 经 read-context 取数、PG 版加 search 召回;judge 零外部依赖(基线包+rubric);import 是纯确定性工具,LLM 不参与。
|
||||
|
||||
**LLM 能力功能全覆盖对照**(SoT 功能 → 实验落点,防漏配):写作 7 功能(续写/改写/扩写/润色/纠错/去AI味/角色声音)→writer;规划 5 动作(生成/补全/整理/检查/多方案,产品-03 §3.5)→planner;分析 6 功能(目标分析/实体解析/候选草稿提取/全书解析/章节抽取/结构拆解)→extractor;检测 5 检查(风险/一致性/角色声音/文风/语义偏离)→detector;质量评分 8 维(专题-04:3 叙事关键+5 非关键)→judge;**有限重写**(≤2 次,专题-04 §5)→eval 环承担(judge 诊断+writer 重生成,人在环);**合规与语义围栏**→不镜像(阶段一由平台安全层承担,阶段二真后端承接);**组装内摘要化**(超预算先摘要低层)→read-context 流程内动作;**嵌入/重排**→embed/search skill(第三层模型能力,非 agent);**硬阻断 3 维**(来源安全/输出合规/静态结构)→规则层检查,skill/流程承担,不进评委职权。
|
||||
**LLM 能力功能全覆盖对照**(SoT 功能 → 实验落点,防漏配):写作 7 功能(续写/改写/扩写/润色/纠错/去AI味/角色声音)共享 writer 身份,但正文实验台 v1 **只验收续写**,其余场景不得据此宣称已优化;规划 5 动作(生成/补全/整理/检查/多方案,产品-03 §3.5)→planner;分析 6 功能(目标分析/实体解析/候选草稿提取/全书解析/章节抽取/结构拆解)→extractor;检测 5 检查(风险/一致性/角色声音/文风/语义偏离)→detector;质量评分 8 维(专题-04:3 叙事关键+5 非关键)→judge;**有限重写**(≤2 次,专题-04 §5)→eval 环承担(judge 诊断+writer 重生成,人在环);**合规与语义围栏**→不镜像(阶段一由平台安全层承担,阶段二真后端承接);**组装内摘要化**(超预算先摘要低层)→read-context 流程内动作;**嵌入/重排**→embed/search skill(第三层模型能力,非 agent);**硬阻断 3 维**(来源安全/输出合规/静态结构)→规则层检查,skill/流程承担,不进评委职权。
|
||||
|
||||
### 上下文组装结构(SoT=[专题-03 §4 四层上下文](../design-docs/专题-03-AI编排上下文与质量评测实现规范.md) + [专题-06 §7 统一读取器](../design-docs/专题-06-元数据驱动的智能体架构.md);操作细节以 read-context skill 为准)
|
||||
|
||||
```
|
||||
四层上下文(专题-03 §4.2;请求带 scenario 用途与授权语义)
|
||||
├── Layer 0 当前输入 任务目标+本章细纲+当前文本 ——不可省略
|
||||
├── Layer 1 近邻正文 上一章末尾场景原文+再前一章摘要+近期叙事状态 ——续写不可省略
|
||||
├── Layer 1 近邻正文 目标章之前连续四章全文+近期叙事状态 ——正文实验 v1 不可省略
|
||||
├── Layer 2 作品事实 设定四节+正式规划+状态台账+知识卡(仅出场+关联) ——只读已确认(Canonical)
|
||||
└── Layer 3 授权资料 绑定的公共范式库/全局知识 ——按场景可省;必须有绑定授权
|
||||
预算顺序:先留输出预算 → L0 完整保留 → L1 优先 L2 优先 L3;超限先摘要化低层再截断
|
||||
@ -115,11 +115,11 @@ SoT 第三层「知识基座」(检索/切块/入库,明文不是 app)在
|
||||
## 五、一次续写怎么走
|
||||
|
||||
1. 主会话读 `装配.yaml`(写作槽位绑了哪个写手、绑定了哪些公共库);
|
||||
2. 走 `read-context` 组装上下文:设定与知识卡按 schema 的 `aiContext` 裁剪(如「结局方向」续写时不给)、近两章正文尾部、`状态.md`、本章细纲;被裁掉的字段记入回显,落 `评审/`;
|
||||
3. 写手产出整章,直接写进 `manuscript/` 新章文件——**不提交**;
|
||||
4. 检测/评委只读产出报告与评分,落 `评审/`;
|
||||
5. 你读章 + 看报告:满意 → 走 `confirm` commit;要改 → 提意见重生成;不要 → restore;
|
||||
6. 确认后抽取员按 schema 从新章抽新实体/事件 → 知识卡(状态:草稿)落 `知识/`,与既有卡冲突时标冲突留你裁决;下次续写即可被读取器用上。
|
||||
2. 走 `read-context` 组装 `WriterContext v1`:知识卡只作索引并触发冻结原文回读,连续前四章全文作为基线,事实证据与文风证据分开;被裁掉的字段和来源记入回显;
|
||||
3. 无工具 adapter 只经 stdin 调用 writer;writer 返回严格 `WriterOutput v1`,校验通过后形成待检测候选,不直接写 `manuscript/`;
|
||||
4. 先运行机械硬门;机械通过后才允许调用语义 detector。当前真实模型 detector 尚未实现,fake 只验证接口与失败终态;两层最终通过后候选才进入 Shadow;
|
||||
5. 用户对 Shadow 明确选择接受、修改后合并或丢弃;修改生成严格下一 candidateVersion 并重跑 detector,接受/合并还需通过 revision、上下文、授权、来源、策略和过期检查;
|
||||
6. 本轮 accept_preflight 是纯函数,不提交 Canonical、不调用抽取进程;只返回要求 Canonical 成功提交后才可放行的异步抽卡命令意图。
|
||||
|
||||
「角色卡长什么样」由 schema 声明——加一个字段,抽取与生成的产出立刻多这个字段,智能体一行不改;换绑写手只改 `装配.yaml`。这两条是「元数据驱动」的活体证明(场景 A7 专门验收)。
|
||||
|
||||
|
||||
@ -16,7 +16,7 @@ agent 的提示词按**变化轴**拆三段,不做"一个 agent 一个大 prom
|
||||
|
||||
| scenario | 功能 skill | 槽位/agent | purpose | 保护节点序列(简化) | 状态 |
|
||||
|---|---|---|---|---|---|
|
||||
| continuation 续写 | `continuation` | 写作→writer | generation | read-context→槽位→detect 伴随→quality-gate→用户三决策(confirm) | 已建(2026-07-09) |
|
||||
| continuation 续写 | `continuation` | 写作→writer | generation | read-context→无工具 writer→机械门→语义 detector→Shadow→用户三决策→accept_preflight→正式写入→异步抽取 | v1 接口与机械门已建;真实语义 detector、正式写入和抽取接线未建 |
|
||||
| rewrite 改写 | `rewrite` | 写作→writer | generation | 同上+expectedRevision 核对 | 已建 |
|
||||
| expansion 扩写 | `expansion` | 写作→writer | generation | 同 continuation | 已建 |
|
||||
| polish 润色 | `polish` | 写作→writer | generation | 同 continuation(只动表达层) | 已建 |
|
||||
@ -34,3 +34,10 @@ agent 的提示词按**变化轴**拆三段,不做"一个 agent 一个大 prom
|
||||
- 拼装位置:身份段之后、增量段之前(见 read-context 拼装序);同功能多回合间稳定,吃缓存;
|
||||
- purpose 由 scenario 映射(生成类→generation、抽取类→extraction、规划→planning、检测类→detection),purpose 定字段可见集(aiContext),scenario 定功能 skill 与 L0 形态,两者不混;
|
||||
- 链登记与实况漂移=G2 治理缺陷,发现即修。
|
||||
|
||||
## 正文候选 v1 三决策合同
|
||||
|
||||
- `accept`:用户明确确认后,对当前候选运行实时 `accept_preflight`;`expectedRevision`、detector、上下文、策略、授权、来源或有效期任一不满足即失败关闭。
|
||||
- `merge`:用户编辑必须生成严格下一 `candidateVersion` 和新正文 hash,重新运行 detector 后再进入 Shadow;禁止“修改后直接合并”。
|
||||
- `discard`:用户明确确认后只关闭候选,不返回抽取命令意图。
|
||||
- 实验台 `check_writer_acceptance.py` 是无副作用纯函数,不写 Canonical。`accept/merge` 通过时只返回当前不可派发的提交后副作用意图(`allowed=false`、`requiresCanonicalCommit=true`);正式提交层验证 Canonical 提交凭证后才可异步排队章后抽取。本轮没有接入正式正文库或抽取进程。
|
||||
|
||||
Loading…
x
Reference in New Issue
Block a user