框架: 收敛 Agent 证据链并迁移生产写作入口
This commit is contained in:
parent
091b66a9bb
commit
1c05b8df7b
@ -2,6 +2,7 @@
|
||||
|
||||
- [项目长期文档](docs/_index.md)
|
||||
- [Skill 发现总索引](skills/_index.md)
|
||||
- [角色定义](agents/)(writer / planner / extractor / detector / judge,派发合同见 [07-Agent与Skill领域 §2](docs/architecture/domains/07-Agent与Skill领域.md))
|
||||
- [角色身份提示](agents/)(writer / planner / extractor / detector / judge)
|
||||
- [角色合同](docs/architecture/角色合同.md)(稳定角色边界、模型策略与派发合同唯一事实源)
|
||||
|
||||
项目入口与协作规则仍由根目录 [`AGENTS.md`](../AGENTS.md) 拥有;本目录只保存跨任务稳定知识。
|
||||
|
||||
@ -1,35 +1,18 @@
|
||||
---
|
||||
name: detector
|
||||
description: 检测员——只负责候选语义核查;机械合同、身份绑定和状态推进由可信编排器负责。
|
||||
tools: Read, Grep, Glob
|
||||
description: 检测员——负责对当前一个候选执行证据约束下的语义核查。
|
||||
skills: check-content-consistency
|
||||
tools: read, grep, find, ls
|
||||
---
|
||||
|
||||
你是检测员。功能合同以 `check-content-consistency` Skill 为准。你只返回当前调用要求的结构化检测草稿,不修改正文、规划、知识卡、运行状态或任何文件。
|
||||
你是检测员。只根据当前输入核查成立性,区分通过、冲突和证据不足。
|
||||
|
||||
## 模型口径
|
||||
## Skill 路由
|
||||
|
||||
模型不属于角色身份;交互式调用由主代理选择宿主子代理,自动化调用由版本化治理策略选择实际模型并写入回执。
|
||||
候选进入用户决策或独立评分前,用 `check-content-consistency` 做结构、事实、角色状态、能力代价、伏笔和证据缺口检查。
|
||||
|
||||
## 正文候选边界
|
||||
## 工具提示
|
||||
|
||||
- 可信编排器先完成 CandidateEnvelope v2 的 schema、hash、篇幅、细纲锚点和上下文绑定检查;机械失败时不会调用你。
|
||||
- 你只接收当前候选正文、冻结细纲与硬约束、事实证据、历史原文证据和 `asOf`,不接收运行身份、真实实验臂、raw 路径、oracle、其他候选或其他评审结果。
|
||||
- 每次调用只核查一个候选,不继承会话,不负责重写候选。
|
||||
- 返回 `semantic-detection-draft-v3`:`claims`、`findings`、`assertionVerdicts`、`hardConstraintVerdicts`、`newSettingCandidates`、`evidenceGaps`。
|
||||
- 引用候选时提供唯一可定位的 `candidateQuote`;证据只能引用输入登记的 ID。证据不足必须使用 `unknown` 和 `gapReason`,不得猜成 `pass`。
|
||||
- `runId`、候选与上下文 hash、字符 offset、模型回执、最终状态和报告 hash 都由 adapter 计算并形成 SemanticDetection v3,不由你生成。
|
||||
工具是否实际开放以本次任务包白名单为准。只读当前授权材料,不改候选、不补写事实、不推进状态。
|
||||
|
||||
## 语义检查范围
|
||||
|
||||
- 细纲硬事件、结果方向、出场实体、伏笔动作与章末钩子是否在语义上成立。
|
||||
- 冻结事实、角色知情范围、人物行为逻辑、能力代价、地点规则、物品边界与叙事状态是否冲突。
|
||||
- 候选是否引入需要登记的新设定,或是否存在必须补检索才能判断的证据缺口。
|
||||
- 谜底、真相和未来信息只能用于防止提前泄露,不能写入候选可见内容。
|
||||
|
||||
## 细纲回放
|
||||
|
||||
细纲回放只接收冻结到 `as_of` 的匿名候选与公共规划上下文,不读取目标章 proxy,不推断实验臂。报告类别使用 `check-content-consistency` 登记的闭集;任一高严重度问题由编排器阻断,检测员不改候选、不裁决卡效用。
|
||||
|
||||
## 禁区
|
||||
|
||||
不调用工具,不读写文件,不执行 git 操作,不修改候选,不生成身份或可信绑定字段,不把主观观感伪装成事实结论。
|
||||
稳定输入边界、输出字段、模型和失败规则见角色合同与对应 Skill。
|
||||
|
||||
@ -1,29 +1,19 @@
|
||||
---
|
||||
name: extractor
|
||||
description: 抽取员——分析槽位默认绑定件,承接 full_parse 与 extraction;分别加载 deconstruct-book 与 extract-chapter-knowledge,产出全为草稿。
|
||||
tools: Read, Write, Grep, Glob
|
||||
description: 知识抽取员——负责从当前任务材料中提取可核验的知识草稿。
|
||||
skills: deconstruct-book, extract-chapter-knowledge
|
||||
tools: read, grep, find, ls
|
||||
---
|
||||
|
||||
你是知识抽取员,分析槽位的默认绑定件。两用场各有功能合同:**拆书=`deconstruct-book`,章后抽取=`extract-chapter-knowledge`**,一次只带本次功能的合同。产出全部是草稿(文件版不提交;PG 版 status=draft)。
|
||||
你是知识抽取员。以正文证据为准,区分草稿与正式事实。
|
||||
|
||||
## 模型口径(抽取有两条路径)
|
||||
## Skill 路由
|
||||
|
||||
- **拆书 / 导入侧抽取**:经 `call-content-model` / `deconstruct-book` 的内容模型治理入口调用,不继承角色会话。
|
||||
- **创作期章后抽取**:由主代理派发 extractor 子代理;模型由宿主或自动化治理策略选择,不写进角色身份。
|
||||
- 参考书或旧稿的逆向拆解用 `deconstruct-book`。
|
||||
- 已接受章节的章后增量抽取用 `extract-chapter-knowledge`。
|
||||
|
||||
## 元数据纪律(怎么用元数据)
|
||||
## 工具提示
|
||||
|
||||
- **schema 有什么字段,你就抽什么;schema 没有的不抽**——字段合同就是抽取 checklist,不自造结构。
|
||||
- 归型走各 schema 的「判据」字段;归不进任何型的候选=枚举缺口,如实报,不硬塞。
|
||||
- 每字段要有正文证据;置信度低标「?」;字段不合用报「设计发现」(写进对应 schema yaml 的设计发现节)。
|
||||
工具是否实际开放以本次任务包白名单为准。只读取当前授权材料,不把抽取结果直接写成 Canonical,不执行 Git 写操作。
|
||||
|
||||
## 抽取通则(跨两用场)
|
||||
|
||||
1. 以正文为准,不脑补正文没写的;
|
||||
2. 与既有知识冲突时**不覆盖**——「⚠ 冲突待裁决」双版本留档并升级用户;
|
||||
3. 基础字段规范填:来源(抽取@第N章 / 拆书@书名)、状态(草稿);
|
||||
4. **采纳正文≠确认知识**:确认另走 `decide-candidate`,自动确认条件的判定不归你。
|
||||
|
||||
## 禁区
|
||||
|
||||
不动正文、大纲、框架文件;不执行 git 写操作;PG 版不直接写库(产结构化清单,经主会话走 `access-database` 入库)。
|
||||
稳定输入边界、输出字段、模型和失败规则见角色合同与对应 Skill。
|
||||
|
||||
@ -1,29 +1,18 @@
|
||||
---
|
||||
name: judge
|
||||
description: 质量评委——只负责一次匿名独立评分;盲化、身份绑定、稳定性和去盲由可信编排器负责。
|
||||
tools: Read, Grep, Glob
|
||||
description: 质量评委——负责对当前匿名候选执行独立、逐维、可复核的质量评审。
|
||||
skills: score-content-quality
|
||||
tools: read, grep, find, ls
|
||||
---
|
||||
|
||||
你是质量评委。功能合同以 `score-content-quality` Skill 为准。每次调用只完成一个独立评审,不修改候选、规划、知识卡、运行状态或任何文件。
|
||||
你是质量评委。保持独立和可复核,只按当前输入中的量表与证据判断。
|
||||
|
||||
## 正文回放边界
|
||||
## Skill 路由
|
||||
|
||||
- 只接收匿名候选正文、所有候选共同的细纲、oracle 断言和 rubric;不接收真实 A/B/C、卡注入、证据策略、候选 hash、raw 路径、提示词差异、其他评委结果或历史会话。
|
||||
- 返回 `blind-judge-draft-v3`。逐候选、逐维给出 0-10 的 0.5 步长分数、具体 `reason`、唯一可定位的 `candidateQuote` 和受控 `evidenceRefs`。
|
||||
- 证据类型只允许 `candidate`、`fine_outline`、`oracle_assertion`、`judge_inference`,引用 ID 必须来自当前输入。
|
||||
- 返回每个维度的匿名候选完整排序,以及每个匿名候选对全部 oracle 断言和硬约束的 verdict;不得省略或增加对象。
|
||||
- reviewer 身份、候选 hash、字符 offset、输入/报告 hash 和模型回执由 adapter 绑定,不由你输出。
|
||||
需要独立质量分数时用 `score-content-quality`;集合级通过或不通过由 `adjudicate-quality-gate` 负责,不由你代替。
|
||||
|
||||
## 独立性与稳定性
|
||||
## 工具提示
|
||||
|
||||
- 不根据“需要收敛”调整严格度,不尝试猜测另一评委分数。
|
||||
- 编排器以 fresh 无会话调用产生第二评;第二评的候选顺序反转。只有同维差异超过 0.5 时,编排器才启动一次 fresh 第三评。
|
||||
- 第三评后仍不存在稳定配对时,编排器标记 `invalid_unstable`;评委不得自行去盲、强行决定实验臂输赢或计算卡效用。
|
||||
工具是否实际开放以本次任务包白名单为准。只读当前授权材料,不猜测实验臂、不修改候选、不推进状态。
|
||||
|
||||
## 细纲回放
|
||||
|
||||
细纲回放只评价结构完整性、方向因果、事件顺序、实体状态、伏笔动作和承接钩子,不使用正文文风、文笔或可读性维度。输入只含冻结快照与匿名结构候选,不读取目标章全文或完整目标章细纲。
|
||||
|
||||
## 禁区
|
||||
|
||||
不调用工具,不读写文件,不执行 git 操作,不修改候选,不推断真实臂,不生成可信绑定字段,不使用当前输入之外的事实。
|
||||
稳定输入边界、输出字段、模型和失败规则见角色合同与对应 Skill。
|
||||
|
||||
@ -1,31 +1,20 @@
|
||||
---
|
||||
name: planner
|
||||
description: 规划师——规划槽位默认绑定件,承接 setting_init、planning 与 fine_outline;分别加载 design-story-foundation、plan-story 与 plan-chapter,产出全为草稿。
|
||||
tools: Read, Write, Grep, Glob
|
||||
description: 规划师——负责把当前任务中的创作要求转成设定、规划或细纲候选。
|
||||
skills: design-story-foundation, plan-story, plan-chapter
|
||||
tools: read, grep, find, ls
|
||||
---
|
||||
|
||||
你是这部书的总规划,规划槽位的默认绑定件。每次只执行一个功能合同:
|
||||
你是规划师。用结构、因果和可执行性组织规划,不把规划语言代替正文。
|
||||
|
||||
- `setting_init`:遵守 `design-story-foundation` Skill,独立完成一份用户挑选前的前期设定候选;
|
||||
- `planning`:遵守 `plan-story` Skill,负责立项与规划修订;
|
||||
- `fine_outline`:遵守 `plan-chapter` Skill,只产结构细纲,不写正文。
|
||||
## Skill 路由
|
||||
|
||||
产出全部不提交;未确认的规划不进生成上下文。回放任务中,`plan-chapter` 的冻结边界优先于本身份段里面向正式创作的全局规划能力。
|
||||
- 只有模糊 idea 或引擎未成形时用 `design-story-foundation`。
|
||||
- 正式作品设定、大纲和规划修订用 `plan-story`。
|
||||
- `fine_outline` 场景生成下一章结构细纲,用 `plan-chapter`。
|
||||
|
||||
## 元数据纪律(怎么用元数据)
|
||||
## 工具提示
|
||||
|
||||
- `setting_init` 的结构由 `design-story-foundation` 冻结的候选合同控制;下面的 schema 纪律只用于 `planning` 与 `fine_outline`。
|
||||
- **产出结构=schema 字段清单本身**:设定包/大纲/知识卡/状态的每一节每一卡,都按对应 schema 逐字段产出(落点表见 `plan-story`);**字段全覆盖**,写不出=设计问题,标「字段存疑:原因」——这是验证元数据设计的一等产出,不许静默跳过。
|
||||
- **你是唯一看全底牌的生成型角色**(谜底与真相/结局方向/未来卷粗纲):底牌管理是规划职责——底牌写进对应 aiContext 受限字段,绝不散进人人可见的字段。
|
||||
- schema 加字段,设定包立刻多一节,你一字不改。
|
||||
工具是否实际开放以本次任务包白名单为准。优先使用冻结输入,不自行扫描未授权资料,不写数据库或框架文件。
|
||||
|
||||
## 规划方法论(跨立项与修订)
|
||||
|
||||
1. **设定互相咬合**:势力实力用力量体系阶梯表述;人物境界有座标;地点归属对得上势力地盘;主角起点与第一卷冲突强度匹配。写完自查,咬不合当场改。
|
||||
2. **伏笔成网**:核心悬念拆进分卷粗纲,每条有埋设章与计划回收章,登记状态台账。
|
||||
3. **变奏自查**:与品类烂大街套路的差异点写进题材定位;没有差异点推倒重来。
|
||||
4. 「说话方式」必须给可执行语言指纹(口头禅/句长/称呼习惯),不许"豪爽""高冷"空词。
|
||||
|
||||
## 禁区
|
||||
|
||||
不写正文;不动 `meta/` 与框架文件;不执行 git 写操作、不写数据库(规划落库由主会话经 `plan-story` 的 `persist_planning.py` 做)。
|
||||
稳定输入边界、输出字段、模型和失败规则见角色合同与对应 Skill。
|
||||
|
||||
@ -1,30 +1,21 @@
|
||||
---
|
||||
name: writer
|
||||
description: 网文写手——写作槽位默认绑定件,承接 continuation/rewrite/expansion/polish;分别加载 write-next-chapter、rewrite-selection、expand-scene 与 polish-prose,只产候选。
|
||||
description: 网文写手——负责把当前任务中的创作要求转成正文候选。
|
||||
skills: write-next-chapter, rewrite-selection, expand-scene, polish-prose
|
||||
tools: read, grep, find, ls
|
||||
---
|
||||
|
||||
你是这部书的执笔写手,写作槽位的默认绑定件。派发指令按 scenario 只加载一个 Skill:`continuation`→`write-next-chapter`,`rewrite`→`rewrite-selection`,`expansion`→`expand-scene`,`polish`→`polish-prose`。你只返回当前合同要求的正文草稿,**永不读写工作区、永不 git 提交**——采纳权在用户。
|
||||
你是网文写手。保持具体、可读、有场景动作的中文表达。
|
||||
|
||||
## 输入边界
|
||||
## Skill 路由
|
||||
|
||||
- 唯一输入是 stdin 中的 `WriterCreativeInput v2`;不得调用工具、搜索仓库、读取数据库、访问网络或延续历史会话。
|
||||
- `fineOutline`、`narrativeState`、`factConstraints`、`proseExcerpts`、`patternReferences`、`lengthContract` 和 `styleConstraints` 都由可信上下文层投影;卡片只是索引,你不得自行顺着卡搜索。
|
||||
- 大纲只给本章方向;细纲的硬事件、结果方向、伏笔动作、章末钩子和必须出场实体是不可删除或反转的硬骨架;可调整节拍才允许重排。
|
||||
- `factConstraints` 只约束事实真伪;`proseExcerpts` 只用于人物声音、动作习惯和叙事质感,不得拿文风样本替代事实约束。
|
||||
- `patternReferences` 是可参考的写作范式:每条含名字(name)、一句话摘要(summary)和写法要点(writingPoints);只借鉴其写法节奏与技巧,不当作事实约束,不照抄。
|
||||
- 续写下一章用 `write-next-chapter`。
|
||||
- 用户点名改范围或翻案用 `rewrite-selection`。
|
||||
- 场景变薄但落点不变用 `expand-scene`。
|
||||
- 只修表达、语病、标点和节奏用 `polish-prose`。
|
||||
|
||||
## 元数据纪律(怎么用元数据)
|
||||
## 工具提示
|
||||
|
||||
- **结构合同来自 schema**:你只输出 `WriterDraft v2`,其唯一业务字段是 `candidateBody`;正文不包含 frontmatter。
|
||||
- **行为约束来自字段值**:文风画像(style)、人物卡「行事逻辑/说话方式/知情范围」、力量体系「代价限制」、地点「规则特例」——逐字段对照,它们是硬约束不是参考。
|
||||
工具是否实际开放以本次任务包白名单为准。默认把输入当作唯一材料,不写工作区、不执行 Git、不修改 Canonical。
|
||||
|
||||
## 跨功能写作纪律(评委按此扣分)
|
||||
|
||||
1. 只依据创作输入写作;不得把未声明的新地名、能力、组织、身份、战绩或关系写成已确认 Canonical 事实。
|
||||
2. 具体压倒抽象:名词给实物、动词给动作;情绪用行为与细节展示,不许直接宣告。
|
||||
3. 每场戏三件套:这场要什么、被什么挡住、落点在哪;没有三件套的场景删掉。
|
||||
4. AI 味黑名单(style 实例给出)一个不许出现;知情范围——角色绝不能说出他不该知道的事。
|
||||
|
||||
## 禁区
|
||||
|
||||
不调用任何工具;不读写 `设定.md`、`大纲.md`、`状态.md`、正文文件、知识卡与框架文件;不执行 git 操作。只返回严格 `WriterDraft v2` JSON:`{"candidateBody":"一章完整正文"}`;不得输出其他字段、Markdown 代码围栏或额外说明。
|
||||
稳定输入边界、输出字段、模型和失败规则见角色合同与对应 Skill。
|
||||
|
||||
@ -1,4 +1,4 @@
|
||||
# 长期文档索引
|
||||
|
||||
- [架构](architecture/_index.md)
|
||||
- [架构](architecture/_index.md),包含 [角色合同](architecture/角色合同.md)
|
||||
- [创作周期与 Skill 导读](architecture/创作周期与Skill导读.md)(教学地图,不是合同权威)
|
||||
|
||||
@ -1,4 +1,5 @@
|
||||
# 架构文档索引
|
||||
|
||||
- [单用户本地优先领域设计](domains/_index.md)
|
||||
- [角色合同](角色合同.md)(五个角色的唯一稳定合同事实源)
|
||||
- [创作周期与 Skill 导读](创作周期与Skill导读.md)(教学地图,不是合同权威)
|
||||
|
||||
@ -54,7 +54,7 @@
|
||||
- 规划期 `select_patterns` 从公共范式卡(`work_id=0`)选定本作范式,绑定 `pattern_bindings`;范式内容与消费尺寸合同见 [03-范式领域](03-范式领域.md)。
|
||||
- 装配 `assembly`:把已选范式与本作事实组织为待装配集合,落 `section_type='assembly'`(装配此前无领域 owner,本节收编,见下"归属收编")。
|
||||
- `style` 文风画像八字段(叙事人称视角 / 句式 / 叙述配比 / 用词质感 / AI 味黑名单 / 对话风格 / 章末钩子风格 / 达标样张,字段权威见 `meta/schemas/style.yaml`)。
|
||||
- 机械门禁:**`assembly` 经用户确认(`confirmed`)才进正文上下文**、**`assemble` 只消费已绑定范式**(对齐 [专题-07](../../../../../design-docs/专题-07-知识消费契约与质量闭环.md):公共范式只走规划期决策、写作期引用,不作写作期临场海选)、**`style` 真注入 writer**——**取数端与生产接线已建**:assemble-context 三个一等取数端(`load_confirmed_fine_outline` / `load_confirmed_pattern_bindings` / `load_confirmed_style`)已建并由生产编排(`docs/write-chapter/step2_write_chapter.py`)接线;范式只读已确认 assembly 绑定注入(实验仓承载,见 [03-范式领域 §6](03-范式领域.md)),确认文风投影为 `styleConstraints` 随冻结上下文注入 writer(不再写死为空)。**待建**:当前注入的是设定行的一句话文风(书12 现状),结构化 `style` 八字段画像的书级抽取尚未建;写作期范式须可回指 confirmed 绑定的门禁尚未机械强制。评测 A/B/C 臂走独立冻结注入,不读生产绑定。
|
||||
- 机械门禁:**`assembly` 经用户确认(`confirmed`)才进正文上下文**、**`assemble` 只消费已绑定范式**(对齐 [专题-07](../../../../../design-docs/专题-07-知识消费契约与质量闭环.md):公共范式只走规划期决策、写作期引用,不作写作期临场海选)、**`style` 真注入 writer**——**取数端与生产接线已建**:assemble-context 三个一等取数端(`load_confirmed_fine_outline` / `load_confirmed_pattern_bindings` / `load_confirmed_style`)已建并由生产编排(`.agent/skills/write-next-chapter/scripts/produce_next_chapter.py`)接线;范式只读已确认 assembly 绑定注入(实验仓承载,见 [03-范式领域 §6](03-范式领域.md)),确认文风投影为 `styleConstraints` 随冻结上下文注入 writer(不再写死为空)。**待建**:当前注入的是设定行的一句话文风(书12 现状),结构化 `style` 八字段画像的书级抽取尚未建;写作期范式须可回指 confirmed 绑定的门禁尚未机械强制。评测 A/B/C 臂走独立冻结注入,不读生产绑定。
|
||||
- 退出条件:`assembly` 已 `confirmed` 并完成 `pattern_bindings` 绑定,`style` 八字段齐备。
|
||||
|
||||
**阶段 4 · 细纲**
|
||||
@ -88,7 +88,7 @@
|
||||
|
||||
### 现状与待建
|
||||
|
||||
- **已建成(引用即可)**:五阶段顺序与产出落点、`example_planning_section` 的 `section_type` 白名单与 `fine_outline` 必带章号、`shadow → confirmed` 单通道、公共范式卡与 `select_patterns`;以及本轮落地的——细纲唯一字段合同(`meta/schemas/fine_outline.yaml`,产出形与装配消费形统一)、落库字段覆盖门禁(`fine_outline` 型已强制失败关闭)、assemble-context 三个一等取数端(`load_confirmed_fine_outline` / `load_confirmed_pattern_bindings` / `load_confirmed_style`)并由生产编排 `step2_write_chapter.py` 接线(细纲统一消费、范式只读已确认 assembly 绑定、文风投影为 `styleConstraints` 注入 writer)。
|
||||
- **已建成(引用即可)**:五阶段顺序与产出落点、`example_planning_section` 的 `section_type` 白名单与 `fine_outline` 必带章号、`shadow → confirmed` 单通道、公共范式卡与 `select_patterns`;以及本轮落地的——细纲唯一字段合同(`meta/schemas/fine_outline.yaml`,产出形与装配消费形统一)、落库字段覆盖门禁(`fine_outline` 型已强制失败关闭)、assemble-context 三个一等取数端(`load_confirmed_fine_outline` / `load_confirmed_pattern_bindings` / `load_confirmed_style`)并由生产编排 `produce_next_chapter.py` 接线(细纲统一消费、范式只读已确认 assembly 绑定、文风投影为 `styleConstraints` 注入 writer)。
|
||||
- **决策已定、机械落地待建(三个架构空白)**:合同见 02/01/03 三域决策记录——
|
||||
1. 设定全书闭环校验 + 全书设定台账(把「演变历程」从只追加日志升级为闭环义务,新增设定 × 章消费矩阵视图;见 [02-实体领域 §8](02-实体领域.md))。
|
||||
2. 卷数合同:`novel_work` 篇幅目标增「分卷数」,`outline` 分卷粗纲卷数须与之一致并机械校验(见 [01-作品领域 §6](01-作品领域.md))。
|
||||
|
||||
@ -15,13 +15,13 @@ Agent 与 Skill 领域拥有角色职责、可调用能力合同、确定性工
|
||||
| Skill | 定义一个可复用能力的输入、输出、允许动作、失败和验收 | 同时承担多个无关意图 |
|
||||
| Tool | 执行确定性解析、校验、计算、读写或报告 | 主观创作和质量裁决 |
|
||||
|
||||
角色至少包括 planner、writer、extractor、detector、judge。角色身份由 `.agent/agents/*.md` 拥有;具体功能步骤由 Skill 拥有,不复制进角色提示词。
|
||||
角色至少包括 planner、writer、extractor、detector、judge。角色短身份提示由 `.agent/agents/*.md` 拥有;稳定角色合同由 [角色合同](../角色合同.md) 统一拥有,具体功能步骤由 Skill 拥有,不复制进角色文件。
|
||||
|
||||
角色提示词用“声明不做什么”守边界:说清本角色不碰哪些支架职责(适配、hash、状态机、持久化),把该做的判断留给模型,把该走的步骤交给 Skill。
|
||||
派发器把身份提示、对应角色合同、功能合同和输出 Schema 按固定顺序装配。角色文件可以登记推荐 Skill 和工具能力供 Agent 路由,但不登记实际权限、模型、输入输出字段、hash、状态机或持久化规则。
|
||||
|
||||
角色是主 ReAct Agent 派发的子代理,按统一派发合同执行:
|
||||
|
||||
1. 派发方把角色文件全文作系统提示词注入,不裁剪改写;角色文件哈希进证据。
|
||||
1. 派发方把角色身份提示与 [角色合同](../角色合同.md) 对应章节作系统提示词注入,不裁剪合同;身份和合同哈希进证据。
|
||||
2. 输入是冻结的结构化 JSON,原样传入;输入哈希进证据。
|
||||
3. 会话全新,不携带历史上下文;角色不得访问未声明的工具。
|
||||
4. 输出是结构化 JSON,由派发方按 schema 校验后才可进入下游;校验失败按失败关闭。
|
||||
@ -91,7 +91,7 @@ ReAct Agent 不能把“扫全库、随意写表”当作通用工具。每次
|
||||
|
||||
## 7. 模型与运行
|
||||
|
||||
模型运行时可替换;角色合同不绑定某个 CLI 的私有状态,也不绑定任何宿主的原生角色装载机制。角色执行走 §2 派发合同:全新会话、角色提示词全文注入、冻结输入、输出校验、证据落库。冻结 profile 绑定角色版本、运行时版本、模型策略版本、prompt/schema 哈希、预算、deadline 与上下文上限;回执分别记录请求策略别名和链内实际模型。模型不可用只终止当次调用,不损坏库里已有的正式内容,也不让 Skill 跳过检查。
|
||||
模型运行时可替换;角色合同不绑定某个 CLI 的私有状态,也不绑定任何宿主的原生角色装载机制。角色执行走 §2 派发合同:全新会话、身份提示与中心合同注入、冻结输入、输出校验、证据落库。冻结 profile 绑定角色合同版本/哈希、运行时版本、模型策略版本、prompt/schema 哈希、预算、deadline 与上下文上限;每次框架调用显式记录 provider、请求模型和实际模型。模型不可用只终止当次调用,不损坏库里已有的正式内容,也不让 Skill 跳过检查。
|
||||
|
||||
运行底座分两层:宿主子代理机制承载交互式派发;`muse_role -> muse_llm.chat_governed` 承载自动化管线。工具隔离由会话授权实现:角色只用角色文件声明的工具,无工具角色不给任何工具。自动化角色调用不启动模型 CLI 子进程、不读取本机客户端配置。运行回执一经写入不可篡改;依赖只向下(上层调下层,不反向)。完整问答、完整原文这类 raw 进库可看全文;仓外保险库是可选备份,不是默认权威。
|
||||
|
||||
|
||||
@ -3,7 +3,7 @@
|
||||
> 这篇是**教学地图**,方便在编辑器里点目录、点链接往下读。
|
||||
> 它**不是**合同权威。阶段顺序、门禁、人机分界以 [05-创作流程领域](domains/05-创作流程领域.md) 为准;某只 Skill 此刻允许做什么,以对应 [`SKILL.md`](../../../.agent/skills/_index.md) 为准。两边若有出入,以那两处为准,回改本文。
|
||||
|
||||
磁盘上现有 **57** 只 Skill([`skills.json`](../../../harness/manifests/skills.json) 与 `.agent/skills/` 一致)。分类索引进度可能滞后,以 57 为准。
|
||||
磁盘上现有 **58** 只 Skill([`skills.json`](../../../harness/manifests/skills.json) 与 `.agent/skills/` 一致)。分类索引进度可能滞后,以 58 为准。
|
||||
|
||||
---
|
||||
|
||||
@ -33,7 +33,7 @@
|
||||
|
||||
### 2. 角色文件不写「这次做什么」
|
||||
|
||||
五种角色身份在 [`.agent/agents/`](../../../.agent/agents/):`planner` / `writer` / `extractor` / `detector` / `judge`。它们只写人设、知情范围和禁区。
|
||||
五种角色身份在 [`.agent/agents/`](../../../.agent/agents/):`planner` / `writer` / `extractor` / `detector` / `judge`;稳定合同集中在 [角色合同](角色合同.md),角色文件只写短身份提示。
|
||||
|
||||
这次做什么、按什么步骤、输出什么,写在功能 Skill 里。换功能不换角色;换角色不换功能指令。
|
||||
|
||||
@ -70,12 +70,13 @@
|
||||
|
||||
## 一次任务怎么加载 Skill
|
||||
|
||||
提示词拆成三段,不做成「一个角色一份大 prompt」:
|
||||
提示词拆成四段,不做成「一个角色一份大 prompt」:
|
||||
|
||||
```text
|
||||
身份段 ← .agent/agents/*.md 这次不变
|
||||
功能指令段 ← 一只功能 Skill 的 SKILL.md 由 scenario 决定
|
||||
L0 任务段 ← assemble-context 组装 本回合参数
|
||||
身份段 ← .agent/agents/*.md Agent-facing 身份、Skill 路由和工具提示
|
||||
角色合同段 ← architecture/角色合同.md 稳定边界、模型策略和失败规则
|
||||
功能指令段 ← 一只功能 Skill 的 SKILL.md 由 scenario 决定
|
||||
L0 任务段 ← assemble-context 组装 本回合参数
|
||||
```
|
||||
|
||||
`scenario` 是这条功能链的名字,例如续写 `continuation`、细纲 `fine_outline`。它映射到哪只 Skill,看 [meta/chains/README.md](../../../meta/chains/README.md)。
|
||||
@ -628,22 +629,23 @@ Gate A/B、A/B/C 三臂、参考书标准答案,只属于离线验收。评测
|
||||
| [execute-role-task](../../../.agent/skills/execute-role-task/SKILL.md) | 冻结 prompt/schema/profile 下的受治理角色调用 |
|
||||
| [record-run-evidence](../../../.agent/skills/record-run-evidence/SKILL.md) | 运行登记、不可变回执、原文证据、经验待审记录 |
|
||||
|
||||
角色模型合同:规划师、写手和评委固定 Opus,不得降级到内容模型;抽取员、检测员只有在 profile 明确登记时才可走内容治理链。内容链按 5 小时窗治理:MiniMax 累计花费上限 24 美元,全模型成功调用上限 6000 次。交互式角色由宿主子代理承载同一 prompt 与输入输出合同。
|
||||
角色模型合同见 [角色合同](角色合同.md):规划师、写手和评委固定 Opus,不得降级到内容模型;抽取员、检测员只有在 profile 明确登记时才可走内容治理链。每次框架调用显式指定 provider、model、thinking。内容链按 5 小时窗治理:MiniMax 累计花费上限 24 美元,全模型成功调用上限 6000 次。交互式角色由宿主子代理承载同一身份提示、角色合同与输入输出合同。
|
||||
|
||||
---
|
||||
|
||||
## 57 只 Skill 总表
|
||||
## 58 只 Skill 总表
|
||||
|
||||
分类取值对应 [索引九域](../../../.agent/skills/_index.md)。这是**主用阶段**,不是每次全加载。跨阶段取用见索引文末表。
|
||||
|
||||
`编排` = 只能被主会话或上游显式调用。`自路由` = 模型可读描述自行选用。
|
||||
|
||||
### 平台底座 · 5
|
||||
### 平台底座 · 6
|
||||
|
||||
| Skill | 调用 | 一句话 |
|
||||
|---|---|---|
|
||||
| [access-database](../../../.agent/skills/access-database/SKILL.md) | 自路由 | 受控查库、改库、跑可审计 DDL |
|
||||
| [call-content-model](../../../.agent/skills/call-content-model/SKILL.md) | 自路由 | New-API 治理入口 |
|
||||
| [dispatch-agent-task](../../../.agent/skills/dispatch-agent-task/SKILL.md) | 编排 | 显式模型策略下派发框架子代理并留痕 |
|
||||
| [execute-role-task](../../../.agent/skills/execute-role-task/SKILL.md) | 编排 | 跑一次受治理角色任务 |
|
||||
| [record-run-evidence](../../../.agent/skills/record-run-evidence/SKILL.md) | 编排 | 运行、回执、原文证据 |
|
||||
| [refresh-runtime-probe](../../../.agent/skills/refresh-runtime-probe/SKILL.md) | 编排 | 重签写手能力探针 |
|
||||
@ -773,7 +775,7 @@ Gate A/B、A/B/C 三臂、参考书标准答案,只属于离线验收。评测
|
||||
|---|---|
|
||||
| 阶段顺序、人机分界、节点菜单 | [05-创作流程领域](domains/05-创作流程领域.md) |
|
||||
| 这次该加载哪只功能 Skill | [meta/chains/README.md](../../../meta/chains/README.md) |
|
||||
| 57 只 Skill 挂在哪一段 | [.agent/skills/_index.md](../../../.agent/skills/_index.md) |
|
||||
| 58 只 Skill 挂在哪一段 | [.agent/skills/_index.md](../../../.agent/skills/_index.md) |
|
||||
| 某只 Skill 的输入、红线、工具 | `.agent/skills/<名字>/SKILL.md` |
|
||||
| 方法细节、案例、清单 | 同目录 `references/` |
|
||||
| 细纲字段 | [meta/schemas/fine_outline.yaml](../../../meta/schemas/fine_outline.yaml) |
|
||||
|
||||
112
.agent/docs/architecture/角色合同.md
Normal file
112
.agent/docs/architecture/角色合同.md
Normal file
@ -0,0 +1,112 @@
|
||||
---
|
||||
schemaVersion: role-contracts-v1
|
||||
roles:
|
||||
writer:
|
||||
displayName: 网文写手
|
||||
promptFile: .agent/agents/writer.md
|
||||
modelPolicy: fixed-opus
|
||||
modelPolicyVersion: fixed-opus-v1
|
||||
explicitModelRequired: true
|
||||
toolPolicy: task-spec-allowlist
|
||||
planner:
|
||||
displayName: 规划师
|
||||
promptFile: .agent/agents/planner.md
|
||||
modelPolicy: fixed-opus
|
||||
modelPolicyVersion: fixed-opus-v1
|
||||
explicitModelRequired: true
|
||||
toolPolicy: task-spec-allowlist
|
||||
extractor:
|
||||
displayName: 知识抽取员
|
||||
promptFile: .agent/agents/extractor.md
|
||||
modelPolicy: governed-chain-or-fixed
|
||||
modelPolicyVersion: muse-governed-chain-v1
|
||||
explicitModelRequired: true
|
||||
toolPolicy: task-spec-allowlist
|
||||
detector:
|
||||
displayName: 检测员
|
||||
promptFile: .agent/agents/detector.md
|
||||
modelPolicy: governed-chain-or-fixed
|
||||
modelPolicyVersion: muse-governed-chain-v1
|
||||
explicitModelRequired: true
|
||||
toolPolicy: task-spec-allowlist
|
||||
judge:
|
||||
displayName: 质量评委
|
||||
promptFile: .agent/agents/judge.md
|
||||
modelPolicy: fixed-opus
|
||||
modelPolicyVersion: fixed-opus-v1
|
||||
explicitModelRequired: true
|
||||
toolPolicy: task-spec-allowlist
|
||||
---
|
||||
|
||||
# 角色合同
|
||||
|
||||
本文件是五个角色的硬合同唯一事实源。`.agent/agents/*.md` 保存 Agent-facing 的身份、Skill 路由、推荐工具能力和工作方法,但不拥有实际权限、Schema、模型策略、证据字段或状态机。Skill 拥有具体功能步骤;本文件只规定角色在一次派发中的稳定责任和边界。
|
||||
|
||||
派发器必须同时装载:角色身份提示、对应本文件章节、冻结任务输入和本次输出 Schema。`provider`、`model`、`thinking` 属于执行策略;每次模型调用必须由调用方显式传入 `provider` 和 `model`,不得从环境变量隐式补全。
|
||||
|
||||
<!-- role-contract:detector -->
|
||||
## detector:检测员
|
||||
|
||||
**责任**:只对当前一个候选做语义核查,判断细纲硬事件、结果方向、出场实体、伏笔动作、章末钩子、冻结事实、角色知情范围、行为逻辑、能力代价、地点规则和物品边界是否成立。
|
||||
|
||||
**输入边界**:只接收当前候选、冻结细纲与硬约束、事实证据、历史原文证据和 `asOf`。不接收运行身份、真实实验臂、raw 路径、oracle、其他候选或其他评审结果。
|
||||
|
||||
**输出边界**:返回当前调用 Schema 要求的检测草稿。证据不足必须标记 `unknown` 并说明缺口;候选引文必须是输入正文中逐字相邻、可定位的片段。运行 ID、候选与上下文哈希、字符位置、模型回执和最终状态由编排器绑定。
|
||||
|
||||
**失败与禁区**:不改候选、不补写事实、不裁决知识卡效用、不生成可信绑定字段、不读取目标章未来信息、不调用未在任务包中开放的工具。
|
||||
|
||||
<!-- /role-contract:detector -->
|
||||
|
||||
<!-- role-contract:extractor -->
|
||||
## extractor:知识抽取员
|
||||
|
||||
**责任**:从给定正文或拆书材料中抽取实体、关系、事件和叙事状态草稿;拆书路径与章后抽取路径分别遵守调用方指定的 Skill 合同。
|
||||
|
||||
**输入边界**:以当前正文和冻结输入为唯一事实来源;schema 字段是抽取清单,不自行扩展字段。每个产出字段都应能回指正文证据,证据不足标记低置信度或设计发现。
|
||||
|
||||
**输出边界**:只产草稿和结构化清单,不把抽取结果当作 Canonical,不自行确认知识、不推进状态、不写正式数据库。与既有事实冲突时保留冲突信息并交由上层裁决。
|
||||
|
||||
**失败与禁区**:不改正文、大纲或框架文件,不执行 Git 写操作,不绕过受控模型入口,不把正文未写出的内容补成事实。
|
||||
|
||||
<!-- /role-contract:extractor -->
|
||||
|
||||
<!-- role-contract:judge -->
|
||||
## judge:质量评委
|
||||
|
||||
**责任**:对匿名候选做一次独立、逐维、可复核的质量评审;只按当前输入中的 rubric、细纲、oracle 断言和候选证据判断。
|
||||
|
||||
**输入边界**:只接收匿名候选、共同细纲、oracle 断言和 rubric。不接收真实 A/B/C 身份、卡注入策略、候选 hash、raw 路径、提示词差异、其他评委结果或历史会话。
|
||||
|
||||
**输出边界**:按调用 Schema 给出完整逐候选、逐维分数、理由、可定位引文和受控证据引用;不省略必需对象,不额外生成身份绑定字段。评委不计算实验臂胜负,稳定性由编排器比较多次 fresh 结果。
|
||||
|
||||
**失败与禁区**:不修改候选、规划、知识卡或运行状态,不猜测另一评委结论,不去盲、不强行裁决、不把主观感受伪装成事实。
|
||||
|
||||
<!-- /role-contract:judge -->
|
||||
|
||||
<!-- role-contract:planner -->
|
||||
## planner:规划师
|
||||
|
||||
**责任**:承接设定初始化、作品规划和单章细纲等规划任务;一次调用只执行任务包指定的一个功能合同,产出可比较或可校验的 Shadow 草稿。
|
||||
|
||||
**输入边界**:以任务包冻结输入和对应 Skill 合同为准。规划结构由 schema 字段控制;字段缺失或无法判断时明确标记设计问题,不静默跳过。底牌、未来信息和终局方向只进入被授权的受限字段。
|
||||
|
||||
**输出边界**:只返回调用 Schema 要求的规划结构,不写正文,不把未确认规划送入生成上下文,不生成运行身份、哈希、回执或数据库状态字段。
|
||||
|
||||
**失败与禁区**:不自行决定用户是否确认、不修改 `meta/` 或框架文件、不执行 Git 写操作、不写数据库、不读取未授权的目标章或未来信息。
|
||||
|
||||
<!-- /role-contract:planner -->
|
||||
|
||||
<!-- role-contract:writer -->
|
||||
## writer:网文写手
|
||||
|
||||
**责任**:按任务包指定的 continuation、rewrite、expansion 或 polish 合同生成正文候选;候选默认属于 Shadow,不直接进入 Canonical。
|
||||
|
||||
**输入边界**:唯一事实来源是冻结任务输入。细纲硬事件、结果方向、伏笔动作、章末钩子和必须出场实体不可删除、反转或提前回收;事实约束、声音样本、范式引用和篇幅合同各司其职,不互相替代。
|
||||
|
||||
**输出边界**:只返回调用 Schema 要求的正文草稿,不输出 frontmatter、运行身份、哈希、raw 路径、解释或额外字段。不得把未声明的新地名、能力、组织、身份、战绩或关系写成已确认 Canonical 事实。
|
||||
|
||||
**写作纪律**:具体名词和动作优先;情绪用行为和细节呈现;每场戏有目标、阻力和落点;遵守角色知情范围与声音指纹;不照抄范式或输入原文。
|
||||
|
||||
**失败与禁区**:不调用未授权工具,不读写工作区、正文、规划或知识卡,不执行 Git 操作,不自行提交候选,不绕过检测和用户决策。
|
||||
|
||||
<!-- /role-contract:writer -->
|
||||
@ -4,7 +4,7 @@
|
||||
|
||||
本索引只登记三字段:`skill_name`(目录名,即调用名)、`skill_file`(合同文件路径)、`skill_description`(适用与边界描述,与 SKILL.md frontmatter 逐字一致,frontmatter 是 SoT)。分类字段(`lifecycle` / `invocation` / `side_effects` / `compounding`)逐个登记在 [`harness/manifests/skills.json`](../../harness/manifests/skills.json),由 `harness/skill_harness.py` 机械校验,不在本索引重复。
|
||||
|
||||
本文件由 `harness/skills_index.py --write` 生成,手改会被覆盖;一致性由 `--check` 与 `tests/architecture/test_skills_index.py` 机械把关。按生命周期分域,共 57 个 skill。
|
||||
本文件由 `harness/skills_index.py --write` 生成,手改会被覆盖;一致性由 `--check` 与 `tests/architecture/test_skills_index.py` 机械把关。按生命周期分域,共 58 个 skill。
|
||||
|
||||
## 0 平台底座
|
||||
|
||||
@ -12,7 +12,8 @@
|
||||
|---|---|---|
|
||||
| access-database | `.agent/skills/access-database/SKILL.md` | 通过唯一受控入口查询或修改 muse-example PostgreSQL,并应用可审计 DDL。主会话或 Skill 需要通用数据库访问时使用;专用导入、嵌入和检索仍走各自 Skill,禁止裸连和一次性脚本。 |
|
||||
| call-content-model | `.agent/skills/call-content-model/SKILL.md` | 通过 New-API 的统一治理入口调用内容模型,执行额度窗口、模型降级、重试和 JSON 提取。清洗、拆书或知识审核需要 MiniMax 等内容模型时使用;不得裸调外部服务。 |
|
||||
| execute-role-task | `.agent/skills/execute-role-task/SKILL.md` | 以冻结 RoleExecutionProfile 运行一次受治理的提示词角色调用,校验模型策略、期限、预算、结构和输入输出哈希并返回 RoleExecutionReceipt。writer、planner、extractor、detector 或 judge 的自动化管线需要执行角色时使用;能力探针刷新交给 refresh-runtime-probe,本 Skill 不负责保存 raw、登记运行或裁决业务结果。 |
|
||||
| dispatch-agent-task | `.agent/skills/dispatch-agent-task/SKILL.md` | 把冻结角色任务包派发给 Agent 框架子代理执行并自动留痕:注入角色 prompt 与输出 Schema、按白名单开放工具、归一框架事件流写入代理事件账本,结构化输出经 Draft 2020-12 校验后返回回执。任何 Agent 框架(当前 pi)执行 writer/planner/detector/judge/extractor 角色任务时使用;不经框架的直接 HTTP 批处理走 execute-role-task;本 Skill 不做补证、重写等业务决策。 |
|
||||
| execute-role-task | `.agent/skills/execute-role-task/SKILL.md` | 以冻结 RoleExecutionProfile 运行一次不经框架的直接 HTTP 角色调用,校验模型策略、期限、预算、结构和输入输出哈希并返回 RoleExecutionReceipt。writer、planner、extractor、detector 或 judge 的无工具批处理需要直接模型调用时使用;需要框架原生 ReAct/工具循环的子代理执行走 dispatch-agent-task;能力探针刷新交给 refresh-runtime-probe,本 Skill 不负责保存 raw、登记运行或裁决业务结果。 |
|
||||
| record-run-evidence | `.agent/skills/record-run-evidence/SKILL.md` | 记录模型调用、运行登记、不可变回执、CAS revision 和受控 raw 证据。执行器或业务 Skill 需要持久化一次运行、追加失败证据、补回执引用或管理 raw 备份时使用;不负责调用模型或裁决内容质量。 |
|
||||
| refresh-runtime-probe | `.agent/skills/refresh-runtime-probe/SKILL.md` | 通过 execute-role-task 用当前 writer 提示词、结构和档案实跑一次极小合成角色任务,刷新运行探针记录与自哈希并把完整配置写到新文件。角色合同或运行时、模型策略版本变化导致执行门失败时使用;不就地覆盖原配置,不把离线预览伪装成成功证明。 |
|
||||
|
||||
|
||||
@ -2,8 +2,8 @@
|
||||
"""智能体/技能登记脚本——把 Git 侧的 agent/skill 元数据影子进库,供看板只读。
|
||||
|
||||
落库设计 §2.10 的 A 方案:
|
||||
- Git 侧(.agent/agents/*.md、.agent/skills/*/SKILL.md)仍是配置权威;
|
||||
- 本脚本抽取 frontmatter(name/description/model/tools)+ 尽力抽取涉及的表名,
|
||||
- Git 侧(角色合同文档、.agent/agents/*.md、.agent/skills/*/SKILL.md)仍是配置权威;
|
||||
- 本脚本从中心角色合同读取责任、模型策略和工具政策,角色 frontmatter 只提供 name/description;
|
||||
upsert 进 example_agent_role / example_skill;
|
||||
- 看板只读登记表、不读 Git。幂等可重跑(upsert),配置变更后重跑即同步。
|
||||
|
||||
@ -18,6 +18,7 @@ import re
|
||||
from pathlib import Path
|
||||
|
||||
from muse_db import connect
|
||||
from muse_role_contract import ROLE_CONTRACT_RELATIVE_PATH, load_role_contract_catalog
|
||||
|
||||
ROOT = Path(__file__).resolve().parents[4] # .agent/skills/access-database/scripts → 仓库根
|
||||
AGENTS_DIR = ROOT / ".agent" / "agents"
|
||||
@ -25,7 +26,7 @@ SKILLS_DIR = ROOT / ".agent" / "skills"
|
||||
TABLE_RE = re.compile(r"\b(muse_[a-z_]+|example_[a-z_]+)\b")
|
||||
SKILL_NAME_RE = re.compile(r"^[a-z][a-z0-9]*(?:-[a-z0-9]+)+$")
|
||||
SKILL_FRONTMATTER_KEYS = frozenset({"name", "description", "disable-model-invocation"})
|
||||
ROLE_FRONTMATTER_KEYS = frozenset({"name", "description", "tools"})
|
||||
ROLE_FRONTMATTER_KEYS = frozenset({"name", "description", "skills", "tools"})
|
||||
EXPECTED_ROLES = frozenset({"writer", "planner", "extractor", "detector", "judge"})
|
||||
|
||||
|
||||
@ -69,12 +70,19 @@ def validate_role_catalog(agents_dir: Path = AGENTS_DIR) -> list[tuple[Path, dic
|
||||
raise ValueError(f"{md}: 角色 name 重复: {role}")
|
||||
if not fm.get("description"):
|
||||
raise ValueError(f"{md}: frontmatter 缺少 description")
|
||||
for key in ("skills", "tools"):
|
||||
if key not in fm or not fm[key].strip():
|
||||
raise ValueError(f"{md}: frontmatter 缺少 {key}")
|
||||
seen.add(role)
|
||||
entries.append((md, fm))
|
||||
if seen != EXPECTED_ROLES:
|
||||
raise ValueError(
|
||||
f"角色目录必须精确包含 {sorted(EXPECTED_ROLES)},实际 {sorted(seen)}"
|
||||
)
|
||||
if agents_dir.resolve() == AGENTS_DIR.resolve():
|
||||
catalog = load_role_contract_catalog(ROOT)
|
||||
if set(catalog.roles) != EXPECTED_ROLES:
|
||||
raise ValueError("角色合同文档与角色目录不一致")
|
||||
return entries
|
||||
|
||||
|
||||
@ -114,24 +122,23 @@ def validate_skill_catalog(skills_dir: Path = SKILLS_DIR) -> list[tuple[Path, di
|
||||
|
||||
def sync_roles(conn, catalog: list[tuple[Path, dict]] | None = None) -> int:
|
||||
entries = catalog or validate_role_catalog()
|
||||
contracts = load_role_contract_catalog(ROOT)
|
||||
names = []
|
||||
for md, fm in entries:
|
||||
role = fm["name"]
|
||||
contract = contracts.for_role(role)
|
||||
names.append(role)
|
||||
desc = fm["description"]
|
||||
display = desc.split("——", 1)[0].strip() or None
|
||||
tools_raw = fm.get("tools")
|
||||
tools = json.dumps([t.strip() for t in tools_raw.split(",") if t.strip()],
|
||||
ensure_ascii=False) if tools_raw else None
|
||||
display = contract.display_name
|
||||
responsibility = contract.contract_prompt
|
||||
conn.execute(
|
||||
"""INSERT INTO example_agent_role
|
||||
(role, display_name, model, tools, responsibility, source_ref, synced_at, creator, updater)
|
||||
VALUES (%s,%s,NULL,%s::jsonb,%s,%s,CURRENT_TIMESTAMP,'sync_agent_registry','sync_agent_registry')
|
||||
VALUES (%s,%s,NULL,NULL,%s,%s,CURRENT_TIMESTAMP,'sync_agent_registry','sync_agent_registry')
|
||||
ON CONFLICT (tenant_id, role) DO UPDATE SET
|
||||
display_name=EXCLUDED.display_name, model=NULL, tools=EXCLUDED.tools,
|
||||
display_name=EXCLUDED.display_name, model=NULL, tools=NULL,
|
||||
responsibility=EXCLUDED.responsibility, source_ref=EXCLUDED.source_ref,
|
||||
synced_at=CURRENT_TIMESTAMP, updater='sync_agent_registry', deleted=FALSE""",
|
||||
(role, display, tools, desc, str(md.relative_to(ROOT))))
|
||||
(role, display, responsibility, ROLE_CONTRACT_RELATIVE_PATH.as_posix()))
|
||||
conn.execute(
|
||||
"""UPDATE example_agent_role
|
||||
SET deleted=TRUE, synced_at=CURRENT_TIMESTAMP, updater='sync_agent_registry'
|
||||
@ -177,13 +184,14 @@ def sync_skills(conn, catalog: list[tuple[Path, dict]] | None = None) -> int:
|
||||
def check_database(conn, roles, skills) -> None:
|
||||
"""把数据库当前活跃影子与 Git 目录逐项对账。"""
|
||||
|
||||
contracts = load_role_contract_catalog(ROOT)
|
||||
expected_roles = {
|
||||
fm["name"]: {
|
||||
"source_ref": str(path.relative_to(ROOT)),
|
||||
"responsibility": fm["description"],
|
||||
"source_ref": ROLE_CONTRACT_RELATIVE_PATH.as_posix(),
|
||||
"responsibility": contracts.for_role(fm["name"]).contract_prompt,
|
||||
"model": None,
|
||||
}
|
||||
for path, fm in roles
|
||||
for _path, fm in roles
|
||||
}
|
||||
actual_roles = {
|
||||
row[0]: {"source_ref": row[1], "responsibility": row[2], "model": row[3]}
|
||||
|
||||
50
.agent/skills/dispatch-agent-task/SKILL.md
Normal file
50
.agent/skills/dispatch-agent-task/SKILL.md
Normal file
@ -0,0 +1,50 @@
|
||||
---
|
||||
name: dispatch-agent-task
|
||||
description: 把冻结角色任务包派发给 Agent 框架子代理执行并自动留痕:注入角色 prompt 与输出 Schema、按白名单开放工具、归一框架事件流写入代理事件账本,结构化输出经 Draft 2020-12 校验后返回回执。任何 Agent 框架(当前 pi)执行 writer/planner/detector/judge/extractor 角色任务时使用;不经框架的直接 HTTP 批处理走 execute-role-task;本 Skill 不做补证、重写等业务决策。
|
||||
disable-model-invocation: true
|
||||
---
|
||||
|
||||
# 派发 Agent 框架任务
|
||||
|
||||
本 Skill 只拥有「框架派发」接缝:角色执行交给 Agent 框架(pi/codex/opencode…)的原生 ReAct 循环、工具调用与子代理机制,项目不自造编排。任务包可移植(角色 + 冻结输入 + 输出 Schema + 工具白名单,不含框架字段);框架适配器是全仓唯一直接调用框架二进制的位置(架构门禁白名单)。不经过框架、需要直接 HTTP 模型调用的无工具批处理走 `execute-role-task`,两者不共用执行路径。
|
||||
|
||||
## 入口
|
||||
|
||||
```
|
||||
.venv/bin/python .agent/skills/dispatch-agent-task/scripts/dispatch_agent_task.py \
|
||||
--spec task.json --provider P --model M [--thinking low] \
|
||||
[--repo-root .] [--run-id ID] [--run-dir DIR] [--trigger-source user]
|
||||
```
|
||||
|
||||
| 模块 | 职责 |
|
||||
|---|---|
|
||||
| `scripts/agent_task.py` | 可移植任务包合同:spec 加载校验、角色身份+中央角色合同+Schema 装配、结构化输出校验。 |
|
||||
| `scripts/pi_runner.py` | pi 框架适配器:argv 构造(`--system-prompt` 注入、`--tools` 白名单、`--no-context-files/--no-skills/--no-extensions` 隔离)、JSON 事件流消费、看门狗超时。 |
|
||||
| `scripts/dispatch_agent_task.py` | CLI:运行登记 -> 事件入账 -> 框架执行 -> 校验 -> 证据落库 -> 回执。 |
|
||||
|
||||
## 输入与输出
|
||||
|
||||
- 输入:`AgentTaskSpec` JSON(specVersion=agent-task-v1;role 限五个角色,角色合同来自 `.agent/docs/architecture/角色合同.md`;outputSchema 必须是合法 Draft 2020-12;inputSha256 可选校验)。`provider`、`model`、`thinking` 属于执行策略,其中 `provider` 和 `model` 必须由调用方显式传入并如实记账。
|
||||
- 输出:回执 JSON(runId、框架、请求/实际模型、逐回合用量与成本、哈希链、证据 ID)与退出码;结构化输出经 `muse_llm.extract_json` + 完整 Draft 2020-12 校验,失败关闭(退出码 4)。
|
||||
- 事件协议(九类闭集,逐条追加 `example_agent_event`):run.started / agent.started / model.completed / tool.started / tool.completed / agent.completed / agent.failed / run.completed / run.failed。
|
||||
- 退出码:0 成功;2 spec 非法;3 框架失败(超时/非零退出/流不可解析/无模型回合);4 输出不合 Schema;5 证据落库失败。
|
||||
|
||||
## 红线
|
||||
|
||||
- 适配器不含业务决策:补证、重写、下一步做什么属于框架里的模型与主代理,不属于本 Skill。
|
||||
- 只有 `pi_runner.py` 可以直接调用框架二进制;其余任何位置 shell 调模型 CLI 都被架构门禁阻断。
|
||||
- 不读取本机模型客户端配置文件;框架凭据走框架自身环境变量,本 Skill 不经手。
|
||||
- 角色文件只提供身份提示;中央角色合同是稳定边界唯一事实源。系统提示词由适配器按固定顺序装配,不裁剪角色合同;工具白名单外的能力不开放(空名单 = `--no-tools`)。
|
||||
- 失败一律关闭:框架异常、Schema 不符、证据落库失败都终止运行并记 run.failed,不部分成功。
|
||||
|
||||
## 数据边界
|
||||
|
||||
- `example_agent_event`(DDL-113,append-only):归一事件账本,只存身份、用量、成本与安全摘要。
|
||||
- `example_llm_call`:每个模型回合一条投影(无额度窗时 `window_key=NULL`,以 `run_id` 归属本次派发,raw 指针指向全量转录)。
|
||||
- raw 表:system prompt(prompt)、最终输出(response)、框架全量转录(supplier)经 `record-run-evidence/agent_trace.persist_agent_evidence` 单事务原子落库,写前密钥拦截。
|
||||
- `example_run`:start_run/finish_run 登记终态;运行目录(/tmp/muse-agent-runs/<run_id>)保留 task-spec、system-prompt、user-message、transcript、output、receipt 审计件。
|
||||
- 留痕是旁路义务:派发路径不提供「不留痕」选项,业务调用方不能决定是否记录。
|
||||
|
||||
## 复利合同
|
||||
|
||||
- **模式 C(平台底座)**:`lifecycle=platform`,D8 不适用;不登记创作经验 `example_lesson`。框架派发的效果信号由业务 Skill 在消费回执时归因。
|
||||
279
.agent/skills/dispatch-agent-task/scripts/agent_task.py
Normal file
279
.agent/skills/dispatch-agent-task/scripts/agent_task.py
Normal file
@ -0,0 +1,279 @@
|
||||
#!/usr/bin/env python3
|
||||
"""可移植的 Agent 任务包合同:角色 + 冻结输入 + 输出 Schema + 工具白名单。
|
||||
|
||||
任务包不含任何框架字段(provider/model/二进制路径都属于派发方 ExecutionPolicy),
|
||||
因此同一个任务包可以被 pi / codex / opencode 等任意框架适配器执行。
|
||||
系统提示词 = 角色身份文件 + 中央角色合同 + 结构化输出合同;
|
||||
用户消息 = 功能合同(taskPrompt)+ 冻结输入 JSON。输出按 Draft 2020-12 校验。
|
||||
"""
|
||||
from __future__ import annotations
|
||||
|
||||
import json
|
||||
import math
|
||||
import re
|
||||
from dataclasses import dataclass
|
||||
from pathlib import Path
|
||||
from typing import Any, Mapping
|
||||
|
||||
from jsonschema import Draft202012Validator
|
||||
from jsonschema.exceptions import SchemaError
|
||||
|
||||
from muse_role import canonical_json, format_schema_contract, sha256_json, sha256_text
|
||||
from muse_role_contract import (
|
||||
ROLE_CONTRACT_RELATIVE_PATH,
|
||||
ROLE_CONTRACT_VERSION,
|
||||
ROLE_NAMES,
|
||||
RoleContract,
|
||||
RoleContractError,
|
||||
load_role_contract_catalog,
|
||||
)
|
||||
|
||||
SPEC_VERSION = "agent-task-v1"
|
||||
SUPPORTED_AGENT_ROLES = ROLE_NAMES
|
||||
AGENT_TASK_SEPARATOR = "\n\n--- 冻结输入 ---\n"
|
||||
DEFAULT_MAX_DURATION_SECONDS = 600.0
|
||||
TOOL_NAME_PATTERN = re.compile(r"^[A-Za-z0-9_-]{1,64}$")
|
||||
_SPEC_REQUIRED_KEYS = frozenset(
|
||||
{"specVersion", "role", "taskPrompt", "input", "outputSchema", "outputSchemaId"}
|
||||
)
|
||||
_SPEC_OPTIONAL_KEYS = frozenset(
|
||||
{"toolAllowlist", "workId", "targetChapter", "maxDurationSeconds", "inputSha256"}
|
||||
)
|
||||
_SPEC_ALLOWED_KEYS = _SPEC_REQUIRED_KEYS | _SPEC_OPTIONAL_KEYS
|
||||
|
||||
|
||||
def _reject_nonstandard_json_constant(value: str) -> None:
|
||||
raise ValueError(f"JSON 不允许常量: {value}")
|
||||
|
||||
|
||||
class TaskSpecError(ValueError):
|
||||
"""任务包不合法:字段缺失、schema 非法、角色不受支持或哈希不符。"""
|
||||
|
||||
|
||||
@dataclass(frozen=True)
|
||||
class AgentTaskSpec:
|
||||
"""一次框架派发的可移植任务定义(无框架、无模型字段)。"""
|
||||
|
||||
role: str
|
||||
task_prompt: str
|
||||
input: Mapping[str, Any]
|
||||
output_schema: Mapping[str, Any]
|
||||
output_schema_id: str
|
||||
tool_allowlist: tuple[str, ...]
|
||||
work_id: int | None = None
|
||||
target_chapter: int | None = None
|
||||
max_duration_seconds: float = DEFAULT_MAX_DURATION_SECONDS
|
||||
input_sha256: str | None = None
|
||||
|
||||
def __post_init__(self) -> None:
|
||||
if not isinstance(self.role, str) or self.role not in SUPPORTED_AGENT_ROLES:
|
||||
raise TaskSpecError(
|
||||
f"role 不受支持: {self.role!r}(可选 {sorted(SUPPORTED_AGENT_ROLES)})"
|
||||
)
|
||||
if not isinstance(self.task_prompt, str) or not self.task_prompt.strip():
|
||||
raise TaskSpecError("taskPrompt 必须是非空字符串")
|
||||
if not isinstance(self.input, Mapping):
|
||||
raise TaskSpecError("input 必须是 JSON 对象")
|
||||
if not isinstance(self.output_schema_id, str) or not self.output_schema_id.strip():
|
||||
raise TaskSpecError("outputSchemaId 必须是非空字符串")
|
||||
if len(self.output_schema_id) > 128 or any(ord(char) < 32 for char in self.output_schema_id):
|
||||
raise TaskSpecError("outputSchemaId 超过 128 字符或含控制字符")
|
||||
if not isinstance(self.output_schema, Mapping):
|
||||
raise TaskSpecError("outputSchema 必须是 JSON 对象")
|
||||
try:
|
||||
Draft202012Validator.check_schema(self.output_schema)
|
||||
except SchemaError as exc:
|
||||
raise TaskSpecError(f"outputSchema 不符合 Draft 2020-12: {exc.message}") from exc
|
||||
if not isinstance(self.tool_allowlist, tuple):
|
||||
raise TaskSpecError("toolAllowlist 必须是字符串数组")
|
||||
for tool in self.tool_allowlist:
|
||||
if not isinstance(tool, str) or TOOL_NAME_PATTERN.fullmatch(tool) is None:
|
||||
raise TaskSpecError(f"toolAllowlist 含非法工具名: {tool!r}")
|
||||
if len(set(self.tool_allowlist)) != len(self.tool_allowlist):
|
||||
raise TaskSpecError("toolAllowlist 不得含重复工具名")
|
||||
if (
|
||||
isinstance(self.max_duration_seconds, bool)
|
||||
or not isinstance(self.max_duration_seconds, (int, float))
|
||||
or not math.isfinite(float(self.max_duration_seconds))
|
||||
or self.max_duration_seconds <= 0
|
||||
):
|
||||
raise TaskSpecError("maxDurationSeconds 必须是正的有限数")
|
||||
if self.work_id is not None and (
|
||||
isinstance(self.work_id, bool) or not isinstance(self.work_id, int) or self.work_id <= 0
|
||||
):
|
||||
raise TaskSpecError("workId 必须是正整数")
|
||||
if self.target_chapter is not None and (
|
||||
isinstance(self.target_chapter, bool)
|
||||
or not isinstance(self.target_chapter, int)
|
||||
or self.target_chapter <= 0
|
||||
):
|
||||
raise TaskSpecError("targetChapter 必须是正整数")
|
||||
|
||||
@property
|
||||
def canonical_input_sha256(self) -> str:
|
||||
return sha256_json(self.input)
|
||||
|
||||
|
||||
def load_spec(path: str | Path) -> AgentTaskSpec:
|
||||
"""从 JSON 文件加载并校验任务包;解析、字段或哈希异常统一失败关闭。"""
|
||||
|
||||
try:
|
||||
raw = json.loads(
|
||||
Path(path).read_text(encoding="utf-8"),
|
||||
parse_constant=_reject_nonstandard_json_constant,
|
||||
)
|
||||
except (OSError, UnicodeError, ValueError) as exc:
|
||||
raise TaskSpecError(f"spec 文件不可读或不是合法 JSON: {type(exc).__name__}") from exc
|
||||
if not isinstance(raw, Mapping):
|
||||
raise TaskSpecError("spec 文件必须是 JSON 对象")
|
||||
missing = sorted(_SPEC_REQUIRED_KEYS - set(raw))
|
||||
if missing:
|
||||
raise TaskSpecError(f"spec 缺少必填字段: {', '.join(missing)}")
|
||||
unknown = sorted(set(raw) - _SPEC_ALLOWED_KEYS)
|
||||
if unknown:
|
||||
raise TaskSpecError(f"spec 含未知字段: {', '.join(unknown)}")
|
||||
if raw.get("specVersion") != SPEC_VERSION:
|
||||
raise TaskSpecError(f"specVersion 必须是 {SPEC_VERSION}")
|
||||
tools = raw.get("toolAllowlist", [])
|
||||
if not isinstance(tools, list):
|
||||
raise TaskSpecError("toolAllowlist 必须是字符串数组")
|
||||
duration = raw.get("maxDurationSeconds", DEFAULT_MAX_DURATION_SECONDS)
|
||||
spec = AgentTaskSpec(
|
||||
role=raw["role"],
|
||||
task_prompt=raw["taskPrompt"],
|
||||
input=raw["input"],
|
||||
output_schema=raw["outputSchema"],
|
||||
output_schema_id=raw["outputSchemaId"],
|
||||
tool_allowlist=tuple(tools),
|
||||
work_id=raw.get("workId"),
|
||||
target_chapter=raw.get("targetChapter"),
|
||||
max_duration_seconds=duration,
|
||||
input_sha256=raw.get("inputSha256"),
|
||||
)
|
||||
if spec.input_sha256 is not None and spec.input_sha256 != spec.canonical_input_sha256:
|
||||
raise TaskSpecError("inputSha256 与 input 内容不一致")
|
||||
return spec
|
||||
|
||||
|
||||
@dataclass(frozen=True)
|
||||
class TaskPackage:
|
||||
"""框架适配器实际消费的执行材料(与框架无关)。"""
|
||||
|
||||
spec: AgentTaskSpec
|
||||
role_prompt: str
|
||||
role_contract: RoleContract
|
||||
system_prompt: str
|
||||
system_prompt_sha256: str
|
||||
user_message: str
|
||||
user_message_sha256: str
|
||||
input_sha256: str
|
||||
spec_sha256: str
|
||||
|
||||
def as_identity(self) -> dict[str, Any]:
|
||||
"""给回执/事件用的身份摘要(不含正文)。"""
|
||||
|
||||
return {
|
||||
"role": self.spec.role,
|
||||
"roleContractVersion": ROLE_CONTRACT_VERSION,
|
||||
"roleContractSha256": self.role_contract.contract_sha256,
|
||||
"roleContractSource": ROLE_CONTRACT_RELATIVE_PATH.as_posix(),
|
||||
"outputSchemaId": self.spec.output_schema_id,
|
||||
"outputSchemaSha256": sha256_json(self.spec.output_schema),
|
||||
"systemPromptSha256": self.system_prompt_sha256,
|
||||
"userMessageSha256": self.user_message_sha256,
|
||||
"inputSha256": self.input_sha256,
|
||||
"specSha256": self.spec_sha256,
|
||||
"toolAllowlist": list(self.spec.tool_allowlist),
|
||||
}
|
||||
|
||||
|
||||
def role_prompt_path(repo_root: str | Path, role: str) -> Path:
|
||||
"""角色文件路径由角色名单一决定,杜绝任意路径注入。"""
|
||||
|
||||
return Path(repo_root) / ".agent" / "agents" / f"{role}.md"
|
||||
|
||||
|
||||
def build_task_package(spec: AgentTaskSpec, repo_root: str | Path) -> TaskPackage:
|
||||
"""装配身份提示、中心角色合同、功能合同与冻结输入。"""
|
||||
|
||||
root = Path(repo_root)
|
||||
try:
|
||||
catalog = load_role_contract_catalog(root)
|
||||
role_contract = catalog.for_role(spec.role)
|
||||
except RoleContractError as exc:
|
||||
raise TaskSpecError(f"角色合同不可用: {type(exc).__name__}") from exc
|
||||
path = role_prompt_path(root, spec.role)
|
||||
if not path.is_file():
|
||||
raise TaskSpecError(f"角色文件不存在: {path}")
|
||||
role_prompt = path.read_text(encoding="utf-8")
|
||||
if not role_prompt.strip():
|
||||
raise TaskSpecError(f"角色文件为空: {path}")
|
||||
system_prompt = (
|
||||
role_prompt.rstrip()
|
||||
+ "\n\n--- 角色合同(唯一事实源) ---\n"
|
||||
+ role_contract.contract_prompt
|
||||
+ format_schema_contract(spec.output_schema)
|
||||
)
|
||||
user_message = spec.task_prompt.strip() + AGENT_TASK_SEPARATOR + canonical_json(spec.input)
|
||||
return TaskPackage(
|
||||
spec=spec,
|
||||
role_prompt=role_prompt,
|
||||
role_contract=role_contract,
|
||||
system_prompt=system_prompt,
|
||||
system_prompt_sha256=sha256_text(system_prompt),
|
||||
user_message=user_message,
|
||||
user_message_sha256=sha256_text(user_message),
|
||||
input_sha256=spec.canonical_input_sha256,
|
||||
spec_sha256=sha256_json(
|
||||
{
|
||||
"specVersion": SPEC_VERSION,
|
||||
"role": spec.role,
|
||||
"roleContractVersion": catalog.version,
|
||||
"roleContractSha256": role_contract.contract_sha256,
|
||||
"taskPrompt": spec.task_prompt.strip(),
|
||||
"input": spec.input,
|
||||
"outputSchema": spec.output_schema,
|
||||
"outputSchemaId": spec.output_schema_id,
|
||||
"toolAllowlist": list(spec.tool_allowlist),
|
||||
"workId": spec.work_id,
|
||||
"targetChapter": spec.target_chapter,
|
||||
"maxDurationSeconds": float(spec.max_duration_seconds),
|
||||
}
|
||||
),
|
||||
)
|
||||
|
||||
|
||||
class OutputInvalidError(ValueError):
|
||||
"""框架最终输出未通过结构化合同。"""
|
||||
|
||||
|
||||
def validate_structured_output(final_text: str, spec: AgentTaskSpec) -> dict[str, Any]:
|
||||
"""抽取 JSON 并按冻结 schema 校验;失败抛 OutputInvalidError(失败关闭)。"""
|
||||
|
||||
from muse_llm import extract_json
|
||||
|
||||
try:
|
||||
extracted = extract_json(final_text)
|
||||
Draft202012Validator(spec.output_schema).validate(extracted)
|
||||
except Exception as exc: # noqa: BLE001 - 任何解析/校验失败都统一失败关闭
|
||||
raise OutputInvalidError(f"结构化输出不满足 {spec.output_schema_id}: {type(exc).__name__}") from exc
|
||||
if not isinstance(extracted, Mapping):
|
||||
raise OutputInvalidError("结构化输出必须是 JSON 对象")
|
||||
return dict(extracted)
|
||||
|
||||
|
||||
__all__ = [
|
||||
"AGENT_TASK_SEPARATOR",
|
||||
"AgentTaskSpec",
|
||||
"DEFAULT_MAX_DURATION_SECONDS",
|
||||
"OutputInvalidError",
|
||||
"SPEC_VERSION",
|
||||
"SUPPORTED_AGENT_ROLES",
|
||||
"TaskPackage",
|
||||
"TaskSpecError",
|
||||
"TOOL_NAME_PATTERN",
|
||||
"build_task_package",
|
||||
"load_spec",
|
||||
"role_prompt_path",
|
||||
"validate_structured_output",
|
||||
]
|
||||
642
.agent/skills/dispatch-agent-task/scripts/dispatch_agent_task.py
Normal file
642
.agent/skills/dispatch-agent-task/scripts/dispatch_agent_task.py
Normal file
@ -0,0 +1,642 @@
|
||||
#!/usr/bin/env python3
|
||||
"""dispatch-agent-task CLI:把冻结任务包派发给 Agent 框架子代理并自动留痕。
|
||||
|
||||
流程(全部失败关闭):
|
||||
加载 spec -> 装配任务包 -> 登记 example_run -> 事件账本 run.started ->
|
||||
框架适配器执行(事件流逐条入账)-> 证据原子落库(raw + 逐回合 llm_call)
|
||||
-> 结构化输出校验 -> run.completed + 运行终态 -> 打印回执。
|
||||
|
||||
业务决策(补证、重写、下一步)不属于本入口:那是框架里模型的事。
|
||||
用法:
|
||||
.venv/bin/python dispatch_agent_task.py --spec task.json \
|
||||
[--provider P] [--model M] [--thinking low] [--run-id ID]
|
||||
"""
|
||||
from __future__ import annotations
|
||||
|
||||
import argparse
|
||||
import json
|
||||
import os
|
||||
import re
|
||||
import sys
|
||||
from dataclasses import replace
|
||||
from pathlib import Path
|
||||
from typing import Any, Callable, Iterable, Mapping
|
||||
|
||||
SCRIPT_DIR = Path(__file__).resolve().parent
|
||||
if str(SCRIPT_DIR) not in sys.path:
|
||||
sys.path.insert(0, str(SCRIPT_DIR))
|
||||
EVIDENCE_DIR = (SCRIPT_DIR.parent.parent / "record-run-evidence" / "scripts").resolve()
|
||||
if str(EVIDENCE_DIR) not in sys.path:
|
||||
sys.path.insert(0, str(EVIDENCE_DIR))
|
||||
|
||||
from agent_task import ( # noqa: E402
|
||||
AgentTaskSpec,
|
||||
OutputInvalidError,
|
||||
TaskSpecError,
|
||||
build_task_package,
|
||||
load_spec,
|
||||
validate_structured_output,
|
||||
)
|
||||
from agent_trace import AgentTraceWriter, persist_agent_evidence # noqa: E402
|
||||
from persist_raw import _check_no_secrets # noqa: E402
|
||||
from pi_runner import ( # noqa: E402
|
||||
DEFAULT_PI_BIN,
|
||||
AgentStreamOutcome,
|
||||
ExecutionPolicy,
|
||||
FrameworkError,
|
||||
PiAgentRunner,
|
||||
)
|
||||
from run_registry import finish_run, new_run_id, start_run # noqa: E402
|
||||
|
||||
EXIT_OK = 0
|
||||
EXIT_SPEC_INVALID = 2
|
||||
EXIT_FRAMEWORK_FAILED = 3
|
||||
EXIT_OUTPUT_INVALID = 4
|
||||
EXIT_EVIDENCE_FAILED = 5
|
||||
|
||||
DEFAULT_RUN_DIR_ROOT = Path("/tmp/muse-agent-runs")
|
||||
_RUN_ID_PATTERN = re.compile(r"^[A-Za-z0-9_.-]{1,64}$")
|
||||
|
||||
|
||||
def _usage_int(usage: Mapping[str, Any], *keys: str) -> int:
|
||||
"""读取第一种存在的 usage 字段;脏值与负值按 0 聚合。"""
|
||||
|
||||
for key in keys:
|
||||
if key not in usage:
|
||||
continue
|
||||
try:
|
||||
return max(0, int(usage.get(key) or 0))
|
||||
except (TypeError, ValueError, OverflowError):
|
||||
return 0
|
||||
return 0
|
||||
|
||||
|
||||
def _usage_totals(outcome: AgentStreamOutcome) -> dict[str, int]:
|
||||
"""聚合全部模型回合的 token 用量;输入口径包含 cache 读写。"""
|
||||
|
||||
totals = {"inputTokens": 0, "outputTokens": 0, "cachedTokens": 0, "reasoningTokens": 0}
|
||||
for call in outcome.model_calls:
|
||||
usage = call.usage or {}
|
||||
cached = _usage_int(usage, "cacheRead", "cache_read_input_tokens", "cached_tokens")
|
||||
cache_write = _usage_int(usage, "cacheWrite", "cache_creation_input_tokens")
|
||||
totals["inputTokens"] += _usage_int(usage, "input", "input_tokens", "prompt_tokens") + cached + cache_write
|
||||
totals["outputTokens"] += _usage_int(usage, "output", "output_tokens", "completion_tokens")
|
||||
totals["cachedTokens"] += cached
|
||||
totals["reasoningTokens"] += _usage_int(usage, "reasoning", "reasoning_tokens")
|
||||
return totals
|
||||
|
||||
|
||||
def _cost_totals(outcome: AgentStreamOutcome) -> tuple[float | None, bool]:
|
||||
"""已知成本求和;任一回合未知则 total 为 None 并标记不完整。"""
|
||||
|
||||
known: list[float] = []
|
||||
complete = True
|
||||
for call in outcome.model_calls:
|
||||
if call.cost_usd is None:
|
||||
complete = False
|
||||
else:
|
||||
known.append(call.cost_usd)
|
||||
if not complete:
|
||||
return (sum(known) if known else None), False
|
||||
return sum(known), True
|
||||
|
||||
|
||||
def _model_call_rows(outcome: AgentStreamOutcome) -> list[dict[str, Any]]:
|
||||
"""把模型回合映射成 agent_trace.persist_agent_evidence 的输入行。"""
|
||||
|
||||
return [
|
||||
{
|
||||
"actual_model_id": call.actual_model_id,
|
||||
"usage": dict(call.usage or {}),
|
||||
"stop_reason": call.stop_reason,
|
||||
"cost_usd": call.cost_usd,
|
||||
"duration_ms": None,
|
||||
}
|
||||
for call in outcome.model_calls
|
||||
]
|
||||
|
||||
|
||||
def _write_private_text(path: Path, text: str) -> None:
|
||||
"""创建仅当前用户可读写的运行审计文件。"""
|
||||
|
||||
fd = os.open(path, os.O_WRONLY | os.O_CREAT | os.O_TRUNC, 0o600)
|
||||
with os.fdopen(fd, "w", encoding="utf-8") as handle:
|
||||
handle.write(text)
|
||||
path.chmod(0o600)
|
||||
|
||||
|
||||
def _write_private_json(path: Path, value: Mapping[str, Any]) -> None:
|
||||
_write_private_text(path, json.dumps(value, ensure_ascii=False, indent=2))
|
||||
|
||||
|
||||
def run_dispatch(
|
||||
spec_path: str | Path,
|
||||
*,
|
||||
repo_root: str | Path,
|
||||
policy: ExecutionPolicy,
|
||||
run_id: str | None = None,
|
||||
run_dir: str | Path | None = None,
|
||||
connect_factory: Callable[..., Any] | None = None,
|
||||
launcher: Callable[..., Iterable] | None = None,
|
||||
trigger_source: str = "user",
|
||||
trigger_detail: Mapping[str, Any] | None = None,
|
||||
) -> tuple[dict[str, Any], int]:
|
||||
"""执行一次完整派发;所有可控失败都返回稳定回执与退出码。"""
|
||||
|
||||
repo_root_path = Path(repo_root).resolve()
|
||||
try:
|
||||
spec = load_spec(spec_path)
|
||||
package = build_task_package(spec, repo_root_path)
|
||||
_check_no_secrets(package.system_prompt)
|
||||
_check_no_secrets(package.user_message)
|
||||
except TaskSpecError as exc:
|
||||
return (
|
||||
{
|
||||
"status": "failed",
|
||||
"errorCode": "SPEC_INVALID",
|
||||
"error": str(exc),
|
||||
"specPath": str(spec_path),
|
||||
},
|
||||
EXIT_SPEC_INVALID,
|
||||
)
|
||||
except ValueError as exc:
|
||||
return (
|
||||
{
|
||||
"status": "failed",
|
||||
"errorCode": "SPEC_INVALID",
|
||||
"error": f"任务包不可安全执行: {type(exc).__name__}",
|
||||
"specPath": str(spec_path),
|
||||
},
|
||||
EXIT_SPEC_INVALID,
|
||||
)
|
||||
|
||||
effective_policy = replace(policy, cwd=policy.cwd or str(repo_root_path))
|
||||
if (
|
||||
package.role_contract.model_policy == "fixed-opus"
|
||||
and "opus" not in effective_policy.model.lower()
|
||||
):
|
||||
return (
|
||||
{
|
||||
"status": "failed",
|
||||
"errorCode": "ROLE_MODEL_POLICY_MISMATCH",
|
||||
"error": f"角色 {spec.role} 要求 fixed-opus 模型策略",
|
||||
"requestedModelId": effective_policy.requested_model_id,
|
||||
},
|
||||
EXIT_SPEC_INVALID,
|
||||
)
|
||||
run_id = run_id or new_run_id(
|
||||
f"agent-{spec.role}", work_id=spec.work_id, target_chapter=spec.target_chapter
|
||||
)
|
||||
if not isinstance(run_id, str) or _RUN_ID_PATTERN.fullmatch(run_id) is None:
|
||||
return (
|
||||
{
|
||||
"status": "failed",
|
||||
"errorCode": "RUN_ID_INVALID",
|
||||
"error": "run_id 只能包含 ASCII 字母、数字、点、下划线和短横线,长度不超过 64",
|
||||
},
|
||||
EXIT_SPEC_INVALID,
|
||||
)
|
||||
run_dir_path = Path(run_dir) if run_dir is not None else DEFAULT_RUN_DIR_ROOT / run_id
|
||||
identity = package.as_identity()
|
||||
|
||||
def _receipt(**fields: Any) -> dict[str, Any]:
|
||||
base = {
|
||||
"runId": run_id,
|
||||
"role": spec.role,
|
||||
"framework": effective_policy.framework,
|
||||
"requestedModelId": effective_policy.requested_model_id,
|
||||
"outputSchemaId": spec.output_schema_id,
|
||||
"runDir": str(run_dir_path),
|
||||
**identity,
|
||||
}
|
||||
base.update(fields)
|
||||
return base
|
||||
|
||||
try:
|
||||
if run_dir is None:
|
||||
DEFAULT_RUN_DIR_ROOT.mkdir(parents=True, mode=0o700, exist_ok=True)
|
||||
DEFAULT_RUN_DIR_ROOT.chmod(0o700)
|
||||
run_dir_path.mkdir(parents=True, mode=0o700, exist_ok=False)
|
||||
run_dir_path.chmod(0o700)
|
||||
_write_private_text(
|
||||
run_dir_path / "task-spec.json", Path(spec_path).read_text(encoding="utf-8")
|
||||
)
|
||||
_write_private_text(run_dir_path / "system-prompt.txt", package.system_prompt)
|
||||
_write_private_text(run_dir_path / "user-message.txt", package.user_message)
|
||||
except (OSError, UnicodeError) as exc:
|
||||
return (
|
||||
_receipt(
|
||||
status="failed",
|
||||
errorCode="AUDIT_WRITE_FAILED",
|
||||
error=f"运行审计目录写入失败: {type(exc).__name__}",
|
||||
),
|
||||
EXIT_EVIDENCE_FAILED,
|
||||
)
|
||||
|
||||
try:
|
||||
detail = dict(trigger_detail or {})
|
||||
except (TypeError, ValueError) as exc:
|
||||
receipt = _receipt(
|
||||
status="failed",
|
||||
errorCode="TRIGGER_DETAIL_INVALID",
|
||||
error=f"trigger_detail 不是对象: {type(exc).__name__}",
|
||||
)
|
||||
_write_private_json(run_dir_path / "receipt.json", receipt)
|
||||
return receipt, EXIT_SPEC_INVALID
|
||||
detail.setdefault(
|
||||
"dispatch",
|
||||
{
|
||||
"framework": effective_policy.framework,
|
||||
"requestedModelId": effective_policy.requested_model_id,
|
||||
"specSha256": package.spec_sha256,
|
||||
},
|
||||
)
|
||||
try:
|
||||
detail_json = json.dumps(
|
||||
detail, ensure_ascii=False, sort_keys=True, default=str, allow_nan=False
|
||||
)
|
||||
_check_no_secrets(detail_json)
|
||||
except (TypeError, ValueError) as exc:
|
||||
receipt = _receipt(
|
||||
status="failed",
|
||||
errorCode="TRIGGER_DETAIL_INVALID",
|
||||
error=f"trigger_detail 不可安全记录: {type(exc).__name__}",
|
||||
)
|
||||
_write_private_json(run_dir_path / "receipt.json", receipt)
|
||||
return receipt, EXIT_SPEC_INVALID
|
||||
try:
|
||||
run_record = start_run(
|
||||
connect=connect_factory,
|
||||
run_id=run_id,
|
||||
work_id=spec.work_id,
|
||||
target_chapter=spec.target_chapter,
|
||||
trigger_source=trigger_source,
|
||||
trigger_detail=detail,
|
||||
)
|
||||
except Exception as exc: # noqa: BLE001 - 注册失败不能启动外部 Agent。
|
||||
receipt = _receipt(
|
||||
status="failed",
|
||||
errorCode="RUN_REGISTRY_START_FAILED",
|
||||
error=f"运行登记失败: {type(exc).__name__}",
|
||||
)
|
||||
_write_private_json(run_dir_path / "receipt.json", receipt)
|
||||
return receipt, EXIT_EVIDENCE_FAILED
|
||||
if run_record["status"] != "started":
|
||||
receipt = _receipt(
|
||||
status="failed",
|
||||
errorCode="RUN_ID_EXISTS",
|
||||
error="run_id 已存在,拒绝覆盖既有运行证据",
|
||||
)
|
||||
_write_private_json(run_dir_path / "receipt.json", receipt)
|
||||
return receipt, EXIT_EVIDENCE_FAILED
|
||||
|
||||
writer = AgentTraceWriter(
|
||||
run_id=run_id,
|
||||
framework=effective_policy.framework,
|
||||
agent_role=spec.role,
|
||||
connect=connect_factory,
|
||||
)
|
||||
runner = PiAgentRunner(launcher=launcher)
|
||||
|
||||
def _failed(
|
||||
error_code: str,
|
||||
error: str,
|
||||
exit_code: int,
|
||||
*,
|
||||
agent_failed: bool = False,
|
||||
session_id: str | None = None,
|
||||
final_message: str | None = None,
|
||||
evidence: Mapping[str, Any] | None = None,
|
||||
cause_error_code: str | None = None,
|
||||
) -> tuple[dict[str, Any], int]:
|
||||
"""尽力闭合失败终态;留痕本身失败时升级为证据错误,保留原始原因码。"""
|
||||
|
||||
failures: list[str] = []
|
||||
if final_message is not None:
|
||||
try:
|
||||
_write_private_text(run_dir_path / "final-message.txt", final_message)
|
||||
except OSError as exc:
|
||||
failures.append(type(exc).__name__)
|
||||
if agent_failed:
|
||||
try:
|
||||
writer.emit(
|
||||
"agent.failed",
|
||||
status="error",
|
||||
requested_model_id=effective_policy.requested_model_id,
|
||||
details={"errorCode": error_code},
|
||||
)
|
||||
except Exception as exc: # noqa: BLE001 - 继续尝试写 run.failed/终态。
|
||||
failures.append(type(exc).__name__)
|
||||
try:
|
||||
writer.emit(
|
||||
"run.failed",
|
||||
status="error",
|
||||
requested_model_id=effective_policy.requested_model_id,
|
||||
raw_ref=(evidence or {}).get("transcriptId"),
|
||||
details={"errorCode": error_code},
|
||||
)
|
||||
except Exception as exc: # noqa: BLE001 - 继续尝试闭合 example_run。
|
||||
failures.append(type(exc).__name__)
|
||||
try:
|
||||
finish_run(
|
||||
run_id,
|
||||
"failed",
|
||||
trigger_detail={"errorCode": error_code},
|
||||
connect=connect_factory,
|
||||
)
|
||||
except Exception as exc: # noqa: BLE001 - 回执必须揭示终态未能闭合。
|
||||
failures.append(type(exc).__name__)
|
||||
|
||||
fields: dict[str, Any] = {
|
||||
"status": "failed",
|
||||
"errorCode": error_code,
|
||||
"error": error,
|
||||
"sessionId": session_id,
|
||||
}
|
||||
if evidence is not None:
|
||||
fields["evidence"] = dict(evidence)
|
||||
if cause_error_code is not None:
|
||||
fields["causeErrorCode"] = cause_error_code
|
||||
if failures:
|
||||
fields.update(
|
||||
{
|
||||
"causeErrorCode": cause_error_code or error_code,
|
||||
"errorCode": "EVIDENCE_PERSIST_FAILED",
|
||||
"error": "失败终态留痕未完整写入",
|
||||
"finalizationErrorTypes": sorted(set(failures)),
|
||||
}
|
||||
)
|
||||
exit_code = EXIT_EVIDENCE_FAILED
|
||||
receipt = _receipt(**fields)
|
||||
try:
|
||||
_write_private_json(run_dir_path / "receipt.json", receipt)
|
||||
except OSError:
|
||||
receipt["receiptFileWritten"] = False
|
||||
exit_code = EXIT_EVIDENCE_FAILED
|
||||
return receipt, exit_code
|
||||
|
||||
try:
|
||||
writer.emit(
|
||||
"run.started",
|
||||
status="ok",
|
||||
requested_model_id=effective_policy.requested_model_id,
|
||||
details={"specSha256": package.spec_sha256, "toolAllowlist": list(spec.tool_allowlist)},
|
||||
)
|
||||
except Exception as exc: # noqa: BLE001 - 未留起始事件时不得启动框架。
|
||||
return _failed(
|
||||
"EVIDENCE_PERSIST_FAILED",
|
||||
f"运行起始事件写入失败: {type(exc).__name__}",
|
||||
EXIT_EVIDENCE_FAILED,
|
||||
)
|
||||
|
||||
transcript_path = run_dir_path / "transcript.jsonl"
|
||||
|
||||
def _remove_local_raw() -> None:
|
||||
for path in (transcript_path, run_dir_path / "final-message.txt"):
|
||||
try:
|
||||
path.unlink(missing_ok=True)
|
||||
except OSError:
|
||||
pass
|
||||
|
||||
def _persist_failure_evidence(
|
||||
failed_outcome: AgentStreamOutcome | None,
|
||||
) -> tuple[dict[str, Any] | None, str | None]:
|
||||
"""失败也尽量把已有转录和模型回合写入同一套 raw 证据。"""
|
||||
|
||||
if failed_outcome is None:
|
||||
return None, None
|
||||
try:
|
||||
transcript_text = transcript_path.read_text(encoding="utf-8")
|
||||
except (OSError, UnicodeError):
|
||||
return None, "EVIDENCE_PERSIST_FAILED"
|
||||
if not transcript_text.strip():
|
||||
return None, None
|
||||
try:
|
||||
_check_no_secrets(transcript_text)
|
||||
except ValueError:
|
||||
_remove_local_raw()
|
||||
return None, "RAW_SECRET_DETECTED"
|
||||
try:
|
||||
evidence_result = persist_agent_evidence(
|
||||
run_id=run_id,
|
||||
agent_role=spec.role,
|
||||
system_prompt=package.system_prompt,
|
||||
user_message=package.user_message,
|
||||
final_message=failed_outcome.final_text,
|
||||
transcript=transcript_text,
|
||||
model_calls=_model_call_rows(failed_outcome),
|
||||
requested_model_id=effective_policy.requested_model_id,
|
||||
connect=connect_factory,
|
||||
)
|
||||
return evidence_result, None
|
||||
except Exception:
|
||||
return None, "EVIDENCE_PERSIST_FAILED"
|
||||
|
||||
try:
|
||||
transcript_fd = os.open(
|
||||
transcript_path,
|
||||
os.O_WRONLY | os.O_CREAT | os.O_EXCL,
|
||||
0o600,
|
||||
)
|
||||
with os.fdopen(transcript_fd, "wb") as transcript_file:
|
||||
outcome = runner.run(
|
||||
package,
|
||||
effective_policy,
|
||||
writer,
|
||||
timeout_seconds=spec.max_duration_seconds,
|
||||
raw_sink=transcript_file.write,
|
||||
)
|
||||
transcript_path.chmod(0o600)
|
||||
except FrameworkError as exc:
|
||||
failed_outcome = exc.outcome
|
||||
failure_evidence, evidence_error = _persist_failure_evidence(failed_outcome)
|
||||
if evidence_error is not None:
|
||||
return _failed(
|
||||
evidence_error,
|
||||
"框架失败证据未能安全落库",
|
||||
EXIT_EVIDENCE_FAILED,
|
||||
agent_failed=True,
|
||||
session_id=failed_outcome.session_id if failed_outcome else None,
|
||||
cause_error_code=exc.error_code,
|
||||
)
|
||||
return _failed(
|
||||
exc.error_code,
|
||||
str(exc),
|
||||
EXIT_FRAMEWORK_FAILED,
|
||||
agent_failed=True,
|
||||
session_id=failed_outcome.session_id if failed_outcome else None,
|
||||
final_message=failed_outcome.final_text if failed_outcome else None,
|
||||
evidence=failure_evidence,
|
||||
)
|
||||
except Exception as exc: # noqa: BLE001 - 事件/本地转录失败属于证据失败。
|
||||
return _failed(
|
||||
"EVIDENCE_PERSIST_FAILED",
|
||||
f"框架执行留痕失败: {type(exc).__name__}",
|
||||
EXIT_EVIDENCE_FAILED,
|
||||
agent_failed=True,
|
||||
)
|
||||
|
||||
try:
|
||||
transcript_text = transcript_path.read_text(encoding="utf-8")
|
||||
_check_no_secrets(transcript_text)
|
||||
except ValueError as exc:
|
||||
_remove_local_raw()
|
||||
return _failed(
|
||||
"RAW_SECRET_DETECTED",
|
||||
"框架转录含疑似凭据,已拒绝留存",
|
||||
EXIT_EVIDENCE_FAILED,
|
||||
)
|
||||
except (OSError, UnicodeError) as exc:
|
||||
return _failed(
|
||||
"EVIDENCE_PERSIST_FAILED",
|
||||
f"框架转录回读失败: {type(exc).__name__}",
|
||||
EXIT_EVIDENCE_FAILED,
|
||||
session_id=outcome.session_id,
|
||||
)
|
||||
if not transcript_text.strip():
|
||||
return _failed(
|
||||
"TRANSCRIPT_EMPTY",
|
||||
"框架转录为空",
|
||||
EXIT_FRAMEWORK_FAILED,
|
||||
session_id=outcome.session_id,
|
||||
)
|
||||
|
||||
try:
|
||||
evidence = persist_agent_evidence(
|
||||
run_id=run_id,
|
||||
agent_role=spec.role,
|
||||
system_prompt=package.system_prompt,
|
||||
user_message=package.user_message,
|
||||
final_message=outcome.final_text,
|
||||
transcript=transcript_text,
|
||||
model_calls=_model_call_rows(outcome),
|
||||
requested_model_id=effective_policy.requested_model_id,
|
||||
connect=connect_factory,
|
||||
)
|
||||
except Exception as exc: # noqa: BLE001 - 模型成功但证据失败时必须失败关闭。
|
||||
return _failed(
|
||||
"EVIDENCE_PERSIST_FAILED",
|
||||
f"框架证据落库失败: {type(exc).__name__}",
|
||||
EXIT_EVIDENCE_FAILED,
|
||||
session_id=outcome.session_id,
|
||||
final_message=outcome.final_text,
|
||||
)
|
||||
|
||||
try:
|
||||
structured = validate_structured_output(outcome.final_text or "", spec)
|
||||
except OutputInvalidError as exc:
|
||||
return _failed(
|
||||
"OUTPUT_SCHEMA_INVALID",
|
||||
str(exc),
|
||||
EXIT_OUTPUT_INVALID,
|
||||
session_id=outcome.session_id,
|
||||
final_message=outcome.final_text,
|
||||
evidence=evidence,
|
||||
)
|
||||
|
||||
try:
|
||||
_write_private_text(run_dir_path / "final-message.txt", outcome.final_text or "")
|
||||
_write_private_json(run_dir_path / "output.json", structured)
|
||||
except OSError as exc:
|
||||
return _failed(
|
||||
"AUDIT_WRITE_FAILED",
|
||||
f"运行结果审计文件写入失败: {type(exc).__name__}",
|
||||
EXIT_EVIDENCE_FAILED,
|
||||
session_id=outcome.session_id,
|
||||
evidence=evidence,
|
||||
)
|
||||
|
||||
usage = _usage_totals(outcome)
|
||||
total_cost, cost_complete = _cost_totals(outcome)
|
||||
try:
|
||||
writer.emit(
|
||||
"run.completed",
|
||||
status="ok",
|
||||
requested_model_id=effective_policy.requested_model_id,
|
||||
actual_model_id=outcome.model_calls[-1].actual_model_id,
|
||||
usage={
|
||||
"input": usage["inputTokens"] - usage["cachedTokens"],
|
||||
"output": usage["outputTokens"],
|
||||
"cacheRead": usage["cachedTokens"],
|
||||
},
|
||||
cost_usd=total_cost,
|
||||
raw_ref=evidence.get("transcriptId"),
|
||||
details={
|
||||
"sessionId": outcome.session_id,
|
||||
"turns": outcome.turns,
|
||||
"toolCalls": len(outcome.tool_calls),
|
||||
"costComplete": cost_complete,
|
||||
"leaseId": evidence.get("leaseId"),
|
||||
"llmCallIds": evidence.get("llmCallIds"),
|
||||
},
|
||||
)
|
||||
finish_run(run_id, "completed", connect=connect_factory)
|
||||
except Exception as exc: # noqa: BLE001 - 成功终态与终态事件必须一起可见。
|
||||
return _failed(
|
||||
"RUN_FINALIZE_FAILED",
|
||||
f"成功终态写入失败: {type(exc).__name__}",
|
||||
EXIT_EVIDENCE_FAILED,
|
||||
session_id=outcome.session_id,
|
||||
evidence=evidence,
|
||||
)
|
||||
|
||||
receipt = _receipt(
|
||||
status="completed",
|
||||
sessionId=outcome.session_id,
|
||||
durationMs=outcome.duration_ms,
|
||||
turns=outcome.turns,
|
||||
toolCallCount=len(outcome.tool_calls),
|
||||
modelCallCount=len(outcome.model_calls),
|
||||
actualModelIds=[call.actual_model_id for call in outcome.model_calls],
|
||||
usage=usage,
|
||||
totalCostUsd=round(total_cost, 6) if total_cost is not None else None,
|
||||
costComplete=cost_complete,
|
||||
finalMessageSha256=_sha256_bare(outcome.final_text or ""),
|
||||
structuredOutputSha256=_sha256_bare(
|
||||
json.dumps(structured, ensure_ascii=False, sort_keys=True)
|
||||
),
|
||||
evidence=evidence,
|
||||
)
|
||||
try:
|
||||
_write_private_json(run_dir_path / "receipt.json", receipt)
|
||||
except OSError:
|
||||
receipt["receiptFileWritten"] = False
|
||||
return receipt, EXIT_OK
|
||||
|
||||
|
||||
def _sha256_bare(text: str) -> str:
|
||||
import hashlib
|
||||
|
||||
return hashlib.sha256(text.encode("utf-8")).hexdigest()
|
||||
|
||||
|
||||
def main(argv: list[str] | None = None) -> int:
|
||||
parser = argparse.ArgumentParser(description="把冻结任务包派发给 Agent 框架子代理并自动留痕")
|
||||
parser.add_argument("--spec", required=True, help="AgentTaskSpec JSON 文件路径")
|
||||
parser.add_argument("--provider", required=True, help="显式框架 provider")
|
||||
parser.add_argument("--model", required=True, help="显式框架模型")
|
||||
parser.add_argument("--thinking", default=None, help="思考等级 off/low/medium/high")
|
||||
parser.add_argument("--pi-bin", default=DEFAULT_PI_BIN, help="框架二进制(默认 pi)")
|
||||
parser.add_argument("--repo-root", default=".", help="仓库根(解析 .agent/agents 角色文件)")
|
||||
parser.add_argument("--run-id", default=None, help="指定 run_id(默认自动生成)")
|
||||
parser.add_argument("--run-dir", default=None, help="运行目录(默认 /tmp/muse-agent-runs/<run_id>)")
|
||||
parser.add_argument("--trigger-source", default="user", choices=["user", "replay_eval", "diagnostic"])
|
||||
args = parser.parse_args(argv)
|
||||
|
||||
policy = ExecutionPolicy(
|
||||
provider=args.provider, model=args.model, thinking=args.thinking, pi_bin=args.pi_bin
|
||||
)
|
||||
receipt, code = run_dispatch(
|
||||
args.spec,
|
||||
repo_root=args.repo_root,
|
||||
policy=policy,
|
||||
run_id=args.run_id,
|
||||
run_dir=args.run_dir,
|
||||
trigger_source=args.trigger_source,
|
||||
)
|
||||
print(json.dumps(receipt, ensure_ascii=False, indent=2))
|
||||
return code
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
raise SystemExit(main())
|
||||
428
.agent/skills/dispatch-agent-task/scripts/pi_runner.py
Normal file
428
.agent/skills/dispatch-agent-task/scripts/pi_runner.py
Normal file
@ -0,0 +1,428 @@
|
||||
#!/usr/bin/env python3
|
||||
"""pi 框架适配器:把任务包派发为 pi 子代理并归一其 JSON 事件流。
|
||||
|
||||
这是全仓唯一直接调用 Agent 框架二进制的位置(架构门禁
|
||||
tests/architecture/test_import_boundaries.py 白名单)。适配器只做三件事:
|
||||
构造 argv(角色 prompt 注入 + 工具白名单 + 隔离上下文)、逐行消费框架事件流、
|
||||
把事件归一转发给 TraceWriter。它不含任何业务决策:补证、重写、下一步做什么
|
||||
全部属于框架里的模型,不属于本模块。
|
||||
|
||||
执行策略(provider/model/thinking)由派发方给定并如实记账;框架把模型模式解析为
|
||||
完整模型 ID,匹配口径见 agent_trace.model_ids_match。超时用看门狗线程杀进程:
|
||||
阻塞读 stdout 不会自己抛超时,挂死的框架进程必须被强制终止才能失败关闭。
|
||||
"""
|
||||
from __future__ import annotations
|
||||
|
||||
import json
|
||||
import subprocess
|
||||
import threading
|
||||
import time
|
||||
from dataclasses import dataclass, field
|
||||
from typing import Any, Callable, Iterable, Iterator, Mapping, Sequence
|
||||
|
||||
from agent_task import TaskPackage
|
||||
from agent_trace import AgentTraceWriter
|
||||
|
||||
DEFAULT_FRAMEWORK = "pi"
|
||||
DEFAULT_PI_BIN = "pi"
|
||||
# 事件流单行上限:防御性截断,正常 JSONL 行远小于此。
|
||||
MAX_STREAM_LINE_BYTES = 8 * 1024 * 1024
|
||||
|
||||
|
||||
class FrameworkError(RuntimeError):
|
||||
"""框架执行失败:超时、非零退出或事件流不可解析。"""
|
||||
|
||||
def __init__(
|
||||
self,
|
||||
error_code: str,
|
||||
message: str,
|
||||
*,
|
||||
outcome: "AgentStreamOutcome | None" = None,
|
||||
) -> None:
|
||||
super().__init__(message)
|
||||
self.error_code = error_code
|
||||
self.outcome = outcome
|
||||
|
||||
|
||||
@dataclass(frozen=True)
|
||||
class ExecutionPolicy:
|
||||
"""框架侧执行策略;provider 与 model 必须由调用方显式传入。"""
|
||||
|
||||
provider: str | None = None
|
||||
model: str | None = None
|
||||
thinking: str | None = None
|
||||
pi_bin: str = DEFAULT_PI_BIN
|
||||
cwd: str | None = None
|
||||
|
||||
def __post_init__(self) -> None:
|
||||
if not isinstance(self.provider, str) or not self.provider.strip():
|
||||
raise ValueError("provider 必须显式传入")
|
||||
if not isinstance(self.model, str) or not self.model.strip():
|
||||
raise ValueError("model 必须显式传入")
|
||||
if self.thinking is not None and self.thinking not in {
|
||||
"off", "minimal", "low", "medium", "high", "xhigh", "max"
|
||||
}:
|
||||
raise ValueError("thinking 不受支持")
|
||||
|
||||
@property
|
||||
def framework(self) -> str:
|
||||
return DEFAULT_FRAMEWORK
|
||||
|
||||
@property
|
||||
def requested_model_id(self) -> str:
|
||||
"""账本口径的显式请求模型。"""
|
||||
|
||||
return f"{self.provider}/{self.model}"
|
||||
|
||||
|
||||
@dataclass(frozen=True)
|
||||
class ModelCall:
|
||||
"""一个模型回合的账本投影材料(usage 为框架归一后的原始字典)。"""
|
||||
|
||||
actual_model_id: str
|
||||
provider: str | None
|
||||
usage: Mapping[str, Any]
|
||||
stop_reason: str | None
|
||||
cost_usd: float | None
|
||||
|
||||
|
||||
@dataclass
|
||||
class ToolCallRecord:
|
||||
tool_call_id: str
|
||||
name: str
|
||||
is_error: bool = False
|
||||
|
||||
|
||||
@dataclass
|
||||
class AgentStreamOutcome:
|
||||
"""框架子代理一次执行的客观结果(不含业务判断)。"""
|
||||
|
||||
session_id: str | None = None
|
||||
exit_code: int | None = None
|
||||
timed_out: bool = False
|
||||
final_text: str | None = None
|
||||
model_calls: list[ModelCall] = field(default_factory=list)
|
||||
tool_calls: list[ToolCallRecord] = field(default_factory=list)
|
||||
turns: int = 0
|
||||
parse_error_lines: int = 0
|
||||
duration_ms: int = 0
|
||||
|
||||
|
||||
def build_pi_argv(package: TaskPackage, policy: ExecutionPolicy) -> list[str]:
|
||||
"""构造 pi 子代理 argv:system prompt 注入、工具白名单、上下文隔离。"""
|
||||
|
||||
argv = [policy.pi_bin, "--print", "--mode", "json", "--no-session"]
|
||||
argv += ["--provider", policy.provider, "--model", policy.model]
|
||||
if policy.thinking:
|
||||
argv += ["--thinking", policy.thinking]
|
||||
# 上下文隔离:不加载项目 AGENTS.md/skills/extensions,角色合同全部来自任务包。
|
||||
argv += ["--no-context-files", "--no-skills", "--no-extensions", "--no-approve"]
|
||||
allowlist = package.spec.tool_allowlist
|
||||
if allowlist:
|
||||
argv += ["--tools", ",".join(allowlist)]
|
||||
else:
|
||||
argv += ["--no-tools"]
|
||||
argv += ["--system-prompt", package.system_prompt, package.user_message]
|
||||
return argv
|
||||
|
||||
|
||||
def _message_text(message: Mapping[str, Any]) -> str:
|
||||
"""提取消息中的全部文本块(跳过 thinking/tool_call 块)。"""
|
||||
|
||||
parts: list[str] = []
|
||||
for block in message.get("content") or []:
|
||||
if isinstance(block, Mapping) and block.get("type") == "text":
|
||||
parts.append(str(block.get("text") or ""))
|
||||
return "".join(parts)
|
||||
|
||||
|
||||
def _qualified_model_id(provider: Any, model: Any) -> str:
|
||||
"""把 pi 分开的 provider/model 字段合成账本要求的完整模型 ID。"""
|
||||
|
||||
model_id = str(model or "").strip()
|
||||
provider_id = str(provider or "").strip()
|
||||
if not model_id or "/" in model_id or not provider_id:
|
||||
return model_id
|
||||
return f"{provider_id}/{model_id}"
|
||||
|
||||
|
||||
def _usage_cost(usage: Mapping[str, Any] | None) -> float | None:
|
||||
"""读取框架报告的单回合成本;供应商未定价(0/缺失)记 None,不伪造。"""
|
||||
|
||||
if not isinstance(usage, Mapping):
|
||||
return None
|
||||
cost = usage.get("cost")
|
||||
if isinstance(cost, Mapping):
|
||||
total = cost.get("total")
|
||||
try:
|
||||
return float(total) if total and float(total) > 0 else None
|
||||
except (TypeError, ValueError):
|
||||
return None
|
||||
return None
|
||||
|
||||
|
||||
class _SubprocessStream:
|
||||
"""把 Popen stdout 包装成字节行迭代器;看门狗超时杀进程,stderr 丢弃防管道死锁。"""
|
||||
|
||||
def __init__(self, proc: subprocess.Popen, timeout_seconds: float) -> None:
|
||||
self._proc = proc
|
||||
self.exit_code: int | None = None
|
||||
self.timed_out = False
|
||||
self._watchdog = threading.Timer(
|
||||
max(timeout_seconds, 0.1),
|
||||
self._kill,
|
||||
)
|
||||
self._watchdog.daemon = True
|
||||
self._watchdog.start()
|
||||
|
||||
def _kill(self) -> None:
|
||||
if self._proc.poll() is None:
|
||||
self.timed_out = True
|
||||
self._proc.kill()
|
||||
|
||||
def __iter__(self) -> Iterator[bytes]:
|
||||
assert self._proc.stdout is not None
|
||||
try:
|
||||
for raw_line in self._proc.stdout:
|
||||
if len(raw_line) > MAX_STREAM_LINE_BYTES:
|
||||
if self._proc.poll() is None:
|
||||
self._proc.kill()
|
||||
raise FrameworkError("STREAM_LINE_TOO_LARGE", "事件流单行超限")
|
||||
yield raw_line
|
||||
finally:
|
||||
self._watchdog.cancel()
|
||||
self.exit_code = self._proc.wait()
|
||||
self._proc.stdout.close()
|
||||
|
||||
def close(self) -> None:
|
||||
self._watchdog.cancel()
|
||||
if self._proc.poll() is None:
|
||||
self._proc.kill()
|
||||
self._proc.wait()
|
||||
|
||||
|
||||
class PiAgentRunner:
|
||||
"""启动 pi 子代理、消费事件流并转发归一事件。"""
|
||||
|
||||
def __init__(self, launcher: Callable[..., Iterable[bytes]] | None = None) -> None:
|
||||
# launcher(argv, timeout, cwd) -> 字节行迭代器(带 exit_code 属性);测试注入假 pi。
|
||||
self._launcher = launcher
|
||||
|
||||
def _launch(self, argv: Sequence[str], timeout: float, cwd: str | None) -> Iterable[bytes]:
|
||||
if self._launcher is not None:
|
||||
return self._launcher(argv, timeout, cwd)
|
||||
proc = subprocess.Popen(
|
||||
list(argv),
|
||||
stdout=subprocess.PIPE,
|
||||
stderr=subprocess.DEVNULL,
|
||||
cwd=cwd,
|
||||
)
|
||||
return _SubprocessStream(proc, timeout)
|
||||
|
||||
def run(
|
||||
self,
|
||||
package: TaskPackage,
|
||||
policy: ExecutionPolicy,
|
||||
sink: AgentTraceWriter,
|
||||
*,
|
||||
timeout_seconds: float,
|
||||
raw_sink: Callable[[bytes], None] | None = None,
|
||||
) -> AgentStreamOutcome:
|
||||
"""执行一次框架派发;框架层异常抛 FrameworkError(业务校验在派发器)。
|
||||
|
||||
raw_sink 逐行接收框架原始事件流字节(转录 tap),供派发器固定全量原始证据。
|
||||
"""
|
||||
|
||||
argv = build_pi_argv(package, policy)
|
||||
outcome = AgentStreamOutcome()
|
||||
started = time.monotonic()
|
||||
sink.emit(
|
||||
"agent.started",
|
||||
status="ok",
|
||||
requested_model_id=policy.requested_model_id,
|
||||
details={"framework": policy.framework, "thinking": policy.thinking},
|
||||
)
|
||||
try:
|
||||
stream = self._launch(argv, timeout_seconds, policy.cwd)
|
||||
except OSError as exc:
|
||||
raise FrameworkError(
|
||||
"FRAMEWORK_START_FAILED", f"框架进程启动失败: {type(exc).__name__}"
|
||||
) from exc
|
||||
final_message: Mapping[str, Any] | None = None
|
||||
allowed_tools = frozenset(package.spec.tool_allowlist)
|
||||
stream_error: FrameworkError | None = None
|
||||
try:
|
||||
try:
|
||||
for raw_line in stream:
|
||||
if raw_sink is not None:
|
||||
try:
|
||||
raw_sink(raw_line)
|
||||
except OSError as exc:
|
||||
raise FrameworkError(
|
||||
"TRANSCRIPT_WRITE_FAILED", f"框架转录写入失败: {type(exc).__name__}"
|
||||
) from exc
|
||||
line = raw_line.decode("utf-8", errors="replace").strip()
|
||||
if not line:
|
||||
continue
|
||||
try:
|
||||
event = json.loads(line)
|
||||
except json.JSONDecodeError:
|
||||
outcome.parse_error_lines += 1
|
||||
continue
|
||||
if not isinstance(event, Mapping):
|
||||
outcome.parse_error_lines += 1
|
||||
continue
|
||||
self._consume(event, sink, policy, outcome, allowed_tools)
|
||||
if event.get("type") == "agent_end":
|
||||
messages = event.get("messages") or []
|
||||
for message in reversed(messages):
|
||||
if isinstance(message, Mapping) and message.get("role") == "assistant":
|
||||
final_message = message
|
||||
break
|
||||
except FrameworkError as exc:
|
||||
stream_error = exc
|
||||
finally:
|
||||
closer = getattr(stream, "close", None)
|
||||
if callable(closer):
|
||||
closer()
|
||||
outcome.duration_ms = int((time.monotonic() - started) * 1000)
|
||||
outcome.exit_code = getattr(stream, "exit_code", None)
|
||||
outcome.timed_out = bool(getattr(stream, "timed_out", False))
|
||||
outcome.final_text = _message_text(final_message) if final_message else None
|
||||
|
||||
if stream_error is not None:
|
||||
stream_error.outcome = outcome
|
||||
raise stream_error
|
||||
if outcome.timed_out:
|
||||
raise FrameworkError(
|
||||
"FRAMEWORK_TIMEOUT", f"框架执行超时(>{timeout_seconds}s)", outcome=outcome
|
||||
)
|
||||
if outcome.exit_code != 0:
|
||||
raise FrameworkError(
|
||||
"FRAMEWORK_EXIT_NONZERO", f"框架进程退出码 {outcome.exit_code}", outcome=outcome
|
||||
)
|
||||
if outcome.parse_error_lines:
|
||||
raise FrameworkError(
|
||||
"STREAM_PARSE_ERROR",
|
||||
f"事件流有 {outcome.parse_error_lines} 行不可解析",
|
||||
outcome=outcome,
|
||||
)
|
||||
if not outcome.model_calls:
|
||||
raise FrameworkError("NO_MODEL_RESPONSE", "事件流未含任何模型回合", outcome=outcome)
|
||||
if outcome.final_text is None or not outcome.final_text.strip():
|
||||
raise FrameworkError("EMPTY_FINAL_MESSAGE", "框架未返回最终文本", outcome=outcome)
|
||||
|
||||
sink.emit(
|
||||
"agent.completed",
|
||||
status="ok",
|
||||
requested_model_id=policy.requested_model_id,
|
||||
actual_model_id=outcome.model_calls[-1].actual_model_id,
|
||||
details={
|
||||
"sessionId": outcome.session_id,
|
||||
"turns": outcome.turns,
|
||||
"modelCalls": len(outcome.model_calls),
|
||||
"toolCalls": len(outcome.tool_calls),
|
||||
"durationMs": outcome.duration_ms,
|
||||
},
|
||||
)
|
||||
return outcome
|
||||
|
||||
def _consume(
|
||||
self,
|
||||
event: Mapping[str, Any],
|
||||
sink: AgentTraceWriter,
|
||||
policy: ExecutionPolicy,
|
||||
outcome: AgentStreamOutcome,
|
||||
allowed_tools: frozenset[str],
|
||||
) -> None:
|
||||
"""把单个框架事件归一转发;未知事件类型静默忽略(框架可演进)。"""
|
||||
|
||||
kind = event.get("type")
|
||||
if kind == "session":
|
||||
outcome.session_id = str(event.get("id") or "") or None
|
||||
elif kind == "turn_start":
|
||||
outcome.turns += 1
|
||||
elif kind == "message_end":
|
||||
message = event.get("message") or {}
|
||||
if message.get("role") != "assistant":
|
||||
return
|
||||
usage = message.get("usage") or {}
|
||||
actual_model_id = _qualified_model_id(message.get("provider"), message.get("model"))
|
||||
if not actual_model_id:
|
||||
raise FrameworkError("MODEL_ID_MISSING", "模型回合缺 provider/model 身份")
|
||||
call = ModelCall(
|
||||
actual_model_id=actual_model_id,
|
||||
provider=message.get("provider"),
|
||||
usage=usage,
|
||||
stop_reason=message.get("stopReason"),
|
||||
cost_usd=_usage_cost(usage),
|
||||
)
|
||||
outcome.model_calls.append(call)
|
||||
model_failed = call.stop_reason in {"error", "aborted"}
|
||||
sink.emit(
|
||||
"model.completed",
|
||||
status="error" if model_failed else "ok",
|
||||
requested_model_id=policy.requested_model_id,
|
||||
actual_model_id=call.actual_model_id,
|
||||
usage=usage,
|
||||
cost_usd=call.cost_usd,
|
||||
details={
|
||||
"stopReason": call.stop_reason,
|
||||
"provider": call.provider,
|
||||
"sessionId": outcome.session_id,
|
||||
"turn": outcome.turns,
|
||||
"errorMessage": str(message.get("errorMessage") or "")[:256] if model_failed else None,
|
||||
},
|
||||
)
|
||||
if model_failed:
|
||||
raise FrameworkError(
|
||||
"MODEL_TURN_FAILED",
|
||||
f"模型回合结束状态为 {call.stop_reason}",
|
||||
)
|
||||
elif kind == "tool_execution_start":
|
||||
name = str(event.get("toolName") or "")
|
||||
tool_call_id = str(event.get("toolCallId") or "")
|
||||
if not tool_call_id:
|
||||
raise FrameworkError("TOOL_EVENT_INVALID", "工具开始事件缺 toolCallId")
|
||||
if name not in allowed_tools:
|
||||
raise FrameworkError("TOOL_NOT_ALLOWED", f"框架执行了未授权工具: {name or '<empty>'}")
|
||||
outcome.tool_calls.append(ToolCallRecord(tool_call_id=tool_call_id, name=name))
|
||||
sink.emit(
|
||||
"tool.started",
|
||||
status="ok",
|
||||
tool_name=name,
|
||||
details={"toolCallId": tool_call_id},
|
||||
)
|
||||
elif kind == "tool_execution_end":
|
||||
is_error = bool(event.get("isError"))
|
||||
name = str(event.get("toolName") or "")
|
||||
tool_call_id = str(event.get("toolCallId") or "")
|
||||
if not tool_call_id:
|
||||
raise FrameworkError("TOOL_EVENT_INVALID", "工具结束事件缺 toolCallId")
|
||||
if name not in allowed_tools:
|
||||
raise FrameworkError("TOOL_NOT_ALLOWED", f"框架执行了未授权工具: {name or '<empty>'}")
|
||||
matching = next(
|
||||
(call for call in reversed(outcome.tool_calls) if call.tool_call_id == tool_call_id),
|
||||
None,
|
||||
)
|
||||
if matching is None:
|
||||
raise FrameworkError("TOOL_EVENT_INVALID", "工具结束事件缺对应开始事件")
|
||||
matching.is_error = is_error
|
||||
sink.emit(
|
||||
"tool.completed",
|
||||
status="error" if is_error else "ok",
|
||||
tool_name=name,
|
||||
details={"toolCallId": tool_call_id},
|
||||
)
|
||||
|
||||
|
||||
__all__ = [
|
||||
"AgentStreamOutcome",
|
||||
"DEFAULT_PI_BIN",
|
||||
"ExecutionPolicy",
|
||||
"FrameworkError",
|
||||
"ModelCall",
|
||||
"PiAgentRunner",
|
||||
"ToolCallRecord",
|
||||
"build_pi_argv",
|
||||
]
|
||||
@ -46,6 +46,7 @@ from fine_outline_rubric import DIMENSIONS, RUBRIC_PROFILE, stability_warning, v
|
||||
from writer_gate import verify_writer_gate_receipt
|
||||
from persist_llm_call import persist_call as persist_llm_event
|
||||
|
||||
from muse_role_contract import load_role_contract_catalog # noqa: E402
|
||||
from muse_role import ( # noqa: E402
|
||||
FIXED_OPUS_MODEL_ID,
|
||||
FIXED_OPUS_POLICY_ALIAS,
|
||||
@ -63,6 +64,7 @@ REPO_ROOT = Path(__file__).resolve().parents[4]
|
||||
SKILL_PATH = REPO_ROOT / ".agent/skills/plan-chapter/SKILL.md"
|
||||
AGENTS_DIR = REPO_ROOT / ".agent" / "agents"
|
||||
PLANNER_PATH = AGENTS_DIR / "planner.md"
|
||||
ROLE_CONTRACTS = load_role_contract_catalog(REPO_ROOT)
|
||||
JUDGE_IDS = ("judge-primary", "judge-secondary")
|
||||
DEFAULT_TIMEOUT_SECONDS = 300.0
|
||||
|
||||
@ -329,10 +331,11 @@ def _output_contract(target: int) -> dict[str, Any]:
|
||||
|
||||
|
||||
def _planner_system_prompt(*, target: int) -> str:
|
||||
"""派发合同的身份侧:角色文件全文 + 功能合同 + 输出合同,由派发方注入。"""
|
||||
"""派发合同的身份侧:身份提示 + 中央角色合同 + 功能合同 + 输出合同。"""
|
||||
|
||||
skill = SKILL_PATH.read_text(encoding="utf-8")
|
||||
identity = PLANNER_PATH.read_text(encoding="utf-8")
|
||||
role_contract = ROLE_CONTRACTS.for_role("planner").contract_prompt
|
||||
return "\n".join(
|
||||
[
|
||||
"这是 next_fine_outline_replay_v0 的离线规划任务。只输出一个 JSON 对象,不要 Markdown、正文或解释。",
|
||||
@ -341,6 +344,8 @@ def _planner_system_prompt(*, target: int) -> str:
|
||||
"规划上下文冻结到 as_of;卡只是事实索引和补充,不得替代公共大纲与叙事现在时。",
|
||||
"--- planner identity ---",
|
||||
identity,
|
||||
"--- planner role contract ---",
|
||||
role_contract,
|
||||
"--- plan-chapter Skill ---",
|
||||
skill,
|
||||
"--- output contract ---",
|
||||
@ -463,13 +468,15 @@ def _invoke_structured_agent(
|
||||
) -> Mapping[str, Any]:
|
||||
"""经 muse_role 执行一次 detector/judge,并把调用原始证据留在运行目录。
|
||||
|
||||
角色身份由派发方显式注入:角色文件(.agent/agents/{agent}.md)全文拼进
|
||||
系统提示词,不依赖任何宿主对角色目录的自动装载。
|
||||
角色身份由派发方显式注入:身份提示与中央角色合同拼进系统提示词,
|
||||
不依赖任何宿主对角色目录的自动装载。
|
||||
"""
|
||||
|
||||
role_prompt = (AGENTS_DIR / f"{agent}.md").read_text(encoding="utf-8")
|
||||
role_prompt = (AGENTS_DIR / f"{agent}.md").read_text(encoding="utf-8").rstrip()
|
||||
role_contract = ROLE_CONTRACTS.for_role(agent).contract_prompt
|
||||
system_prompt = (
|
||||
f"{role_prompt}\n\n独立身份={identity};只处理给定 JSON;禁止调用工具、读取文件或输出 JSON 以外内容。\n"
|
||||
f"{role_prompt}\n\n--- 角色合同(唯一事实源) ---\n{role_contract}\n"
|
||||
f"独立身份={identity};只处理给定 JSON;禁止调用工具、读取文件或输出 JSON 以外内容。\n"
|
||||
f"{REPLAY_MODEL_CONSTRAINT}"
|
||||
)
|
||||
if agent == "detector":
|
||||
|
||||
@ -1,12 +1,12 @@
|
||||
---
|
||||
name: execute-role-task
|
||||
description: 以冻结 RoleExecutionProfile 运行一次受治理的提示词角色调用,校验模型策略、期限、预算、结构和输入输出哈希并返回 RoleExecutionReceipt。writer、planner、extractor、detector 或 judge 的自动化管线需要执行角色时使用;能力探针刷新交给 refresh-runtime-probe,本 Skill 不负责保存 raw、登记运行或裁决业务结果。
|
||||
description: 以冻结 RoleExecutionProfile 运行一次不经框架的直接 HTTP 角色调用,校验模型策略、期限、预算、结构和输入输出哈希并返回 RoleExecutionReceipt。writer、planner、extractor、detector 或 judge 的无工具批处理需要直接模型调用时使用;需要框架原生 ReAct/工具循环的子代理执行走 dispatch-agent-task;能力探针刷新交给 refresh-runtime-probe,本 Skill 不负责保存 raw、登记运行或裁决业务结果。
|
||||
disable-model-invocation: true
|
||||
---
|
||||
|
||||
# 执行受治理角色任务
|
||||
|
||||
本 Skill 只拥有自动化管线的一次角色执行边界。交互式流程由主代理按 07 领域派发合同启动宿主子代理;自动化流程由 `muse_role` 承载同一角色 prompt、冻结输入、输出校验和回执合同。
|
||||
本 Skill 只拥有自动化管线的一次角色执行边界。交互式流程由主代理按 07 领域派发合同启动宿主子代理;自动化流程由 `muse_role` 承载同一角色 prompt、冻结输入、输出校验和回执合同。需要框架原生 ReAct/工具循环的子代理执行走 `dispatch-agent-task`(框架派发接缝);本 Skill 只服务不经框架、直接 HTTP 模型调用的无工具批处理。
|
||||
|
||||
## 入口
|
||||
|
||||
|
||||
@ -24,7 +24,7 @@ disable-model-invocation: true
|
||||
--heading "<冻结的一级章标题>" --output /private/tmp/merge-packet-01.md <来源候选>...
|
||||
|
||||
# ② 由主代理派发 planner 子代理
|
||||
# 主代理把 .agent/agents/planner.md 全文作为角色 prompt,输入仅包含
|
||||
# 主代理把 planner 身份提示与 .agent/docs/architecture/角色合同.md 对应章节作为角色 prompt,输入仅包含
|
||||
# serial-merge-contract.md 与当前 merge-packet-01.md;子代理无工具、fresh 会话,
|
||||
# 输出写到 /private/tmp/merge-raw-01.md,并由 record-run-evidence 保存派发回执。
|
||||
|
||||
|
||||
@ -14,6 +14,7 @@ disable-model-invocation: true
|
||||
|---|---|
|
||||
| `run_registry.py` | 创建和结束 `example_run`,维持运行状态与幂等边界。 |
|
||||
| `persist_llm_call.py` | 把模型调用的 prompt、response、用量与 raw 指针原子登记。 |
|
||||
| `agent_trace.py` | 框架派发留痕:事件账本 `example_agent_event` 写入 + raw/逐回合 llm_call 原子落库(07 §2 框架派发)。 |
|
||||
| `persist_raw.py` | 将完整 raw 写入 `example_raw_lease` 与 `example_raw_content`,写前拦截凭据。 |
|
||||
| `record_failed_run.py` | 为失败运行追加错误回执与隔离的质量结果。 |
|
||||
| `repair_receipt_evidence.py` | 对成功回执追加 evidence revision,不更新旧回执。 |
|
||||
|
||||
371
.agent/skills/record-run-evidence/scripts/agent_trace.py
Normal file
371
.agent/skills/record-run-evidence/scripts/agent_trace.py
Normal file
@ -0,0 +1,371 @@
|
||||
#!/usr/bin/env python3
|
||||
"""代理事件账本与框架证据的写路径(07-Agent与Skill领域 §2 框架派发)。
|
||||
|
||||
Agent 框架适配器(如 dispatch-agent-task/pi_runner)把框架原生事件流归一后,
|
||||
经 ``AgentTraceWriter`` 逐条追加进 ``example_agent_event``;运行结束后由
|
||||
``persist_agent_evidence`` 把 system prompt、任务输入、最终输出、全量转录和
|
||||
逐回合模型调用投影原子落库。本模块只做被动留痕:业务调用方不需要、也不能
|
||||
决定"是否留痕";除幂等哈希与安全摘要外不携带任何正文原文。
|
||||
|
||||
数据库连接可注入(``connect=None`` 时懒加载 muse_db.connect),离线测试用假连接。
|
||||
"""
|
||||
from __future__ import annotations
|
||||
|
||||
import json
|
||||
import math
|
||||
from typing import Any, Callable, Mapping, Sequence
|
||||
|
||||
EVENT_TYPES = frozenset(
|
||||
{
|
||||
"run.started",
|
||||
"agent.started",
|
||||
"model.completed",
|
||||
"tool.started",
|
||||
"tool.completed",
|
||||
"agent.completed",
|
||||
"agent.failed",
|
||||
"run.completed",
|
||||
"run.failed",
|
||||
}
|
||||
)
|
||||
CREATOR = "agent-trace"
|
||||
|
||||
# pi/anthropic 风格 usage 到账本列的归一口径;同义字段取首个,避免重复计数。
|
||||
_USAGE_IN_KEYS = ("input", "input_tokens", "prompt_tokens")
|
||||
_USAGE_CACHE_READ_KEYS = ("cacheRead", "cache_read_input_tokens", "cached_tokens")
|
||||
_USAGE_CACHE_WRITE_KEYS = ("cacheWrite", "cache_creation_input_tokens")
|
||||
_USAGE_OUT_KEYS = ("output", "output_tokens", "completion_tokens")
|
||||
|
||||
|
||||
def _connect_factory(connect: Callable[..., Any] | None) -> Callable[..., Any]:
|
||||
if connect is not None:
|
||||
return connect
|
||||
from muse_db import connect as muse_connect
|
||||
|
||||
return muse_connect
|
||||
|
||||
|
||||
def _usage_int(data: Mapping[str, Any], keys: tuple[str, ...]) -> int:
|
||||
"""读取第一种存在的 usage 字段;脏值与负值按 0 记账。"""
|
||||
|
||||
for key in keys:
|
||||
if key not in data:
|
||||
continue
|
||||
try:
|
||||
return max(0, int(data.get(key) or 0))
|
||||
except (TypeError, ValueError, OverflowError):
|
||||
return 0
|
||||
return 0
|
||||
|
||||
|
||||
def _tokens(usage: Mapping[str, Any] | None) -> tuple[int, int, int]:
|
||||
"""把框架 usage 归一为 (input, output, cached);input 含 cache 读写。"""
|
||||
|
||||
data = usage if isinstance(usage, Mapping) else {}
|
||||
cached = _usage_int(data, _USAGE_CACHE_READ_KEYS)
|
||||
cache_write = _usage_int(data, _USAGE_CACHE_WRITE_KEYS)
|
||||
value = _usage_int(data, _USAGE_IN_KEYS) + cached + cache_write
|
||||
output = _usage_int(data, _USAGE_OUT_KEYS)
|
||||
return value, output, cached
|
||||
|
||||
|
||||
class AgentTraceWriter:
|
||||
"""把归一事件逐条追加进 ``example_agent_event``(每条短事务,崩溃可审计)。"""
|
||||
|
||||
def __init__(
|
||||
self,
|
||||
*,
|
||||
run_id: str,
|
||||
framework: str,
|
||||
agent_role: str,
|
||||
connect: Callable[..., Any] | None = None,
|
||||
creator: str = CREATOR,
|
||||
) -> None:
|
||||
if not run_id or len(run_id) > 64:
|
||||
raise ValueError("run_id 不能为空且不超过 64 字符")
|
||||
if not framework or len(framework) > 32:
|
||||
raise ValueError("framework 不能为空且不超过 32 字符")
|
||||
if not agent_role or len(agent_role) > 32:
|
||||
raise ValueError("agent_role 不能为空且不超过 32 字符")
|
||||
self.run_id = run_id
|
||||
self.framework = framework
|
||||
self.agent_role = agent_role
|
||||
self.creator = creator
|
||||
self._connect = _connect_factory(connect)
|
||||
self._seq = 0
|
||||
|
||||
@property
|
||||
def seq(self) -> int:
|
||||
"""已写入的事件数(下一个序号 = seq + 1)。"""
|
||||
|
||||
return self._seq
|
||||
|
||||
def emit(
|
||||
self,
|
||||
event_type: str,
|
||||
*,
|
||||
status: str | None = None,
|
||||
tool_name: str | None = None,
|
||||
requested_model_id: str | None = None,
|
||||
actual_model_id: str | None = None,
|
||||
usage: Mapping[str, Any] | None = None,
|
||||
cost_usd: float | None = None,
|
||||
raw_ref: int | None = None,
|
||||
details: Mapping[str, Any] | None = None,
|
||||
) -> int:
|
||||
"""追加一条事件;事件类型与字段合法性在本层失败关闭。"""
|
||||
|
||||
if event_type not in EVENT_TYPES:
|
||||
raise ValueError(f"未知代理事件类型: {event_type}")
|
||||
if status is not None and status not in ("ok", "error"):
|
||||
raise ValueError("status 只能是 ok/error")
|
||||
if event_type == "model.completed" and not actual_model_id:
|
||||
raise ValueError("model.completed 必须携带 actual_model_id")
|
||||
for field, value, limit in (
|
||||
("tool_name", tool_name, 64),
|
||||
("requested_model_id", requested_model_id, 64),
|
||||
("actual_model_id", actual_model_id, 64),
|
||||
):
|
||||
if value is not None and (not isinstance(value, str) or len(value) > limit):
|
||||
raise ValueError(f"{field} 非法或超过 {limit} 字符")
|
||||
if cost_usd is not None:
|
||||
try:
|
||||
numeric_cost = float(cost_usd)
|
||||
except (TypeError, ValueError, OverflowError) as exc:
|
||||
raise ValueError("cost_usd 必须是非负有限数或 NULL") from exc
|
||||
if not math.isfinite(numeric_cost) or numeric_cost < 0:
|
||||
raise ValueError("cost_usd 必须是非负有限数或 NULL")
|
||||
in_tokens, out_tokens, cached_tokens = _tokens(usage)
|
||||
payload = json.dumps(details or {}, ensure_ascii=False, default=str)
|
||||
from persist_raw import _check_no_secrets
|
||||
|
||||
_check_no_secrets(payload)
|
||||
next_seq = self._seq + 1
|
||||
with self._connect() as conn:
|
||||
try:
|
||||
conn.execute(
|
||||
"INSERT INTO example_agent_event(run_id, seq, event_type, framework, agent_role, "
|
||||
"tool_name, status, requested_model_id, actual_model_id, input_tokens, output_tokens, "
|
||||
"cached_tokens, cost_usd, raw_ref, details, creator) "
|
||||
"VALUES (%s,%s,%s,%s,%s,%s,%s,%s,%s,%s,%s,%s,%s,%s,%s::jsonb,%s)",
|
||||
(
|
||||
self.run_id,
|
||||
next_seq,
|
||||
event_type,
|
||||
self.framework,
|
||||
self.agent_role,
|
||||
tool_name,
|
||||
status,
|
||||
requested_model_id,
|
||||
actual_model_id,
|
||||
in_tokens,
|
||||
out_tokens,
|
||||
cached_tokens,
|
||||
cost_usd,
|
||||
raw_ref,
|
||||
payload,
|
||||
self.creator,
|
||||
),
|
||||
)
|
||||
conn.commit()
|
||||
self._seq = next_seq
|
||||
except Exception:
|
||||
conn.rollback()
|
||||
raise
|
||||
return self._seq
|
||||
|
||||
|
||||
def model_ids_match(requested: str | None, actual: str | None) -> bool:
|
||||
"""匹配 pi 的模型模式解析:完整 ID 精确匹配,单边省略 provider 时比较模型叶名。"""
|
||||
|
||||
if not isinstance(requested, str) or not isinstance(actual, str):
|
||||
return False
|
||||
req, act = requested.strip().lower(), actual.strip().lower()
|
||||
if not req or not act:
|
||||
return False
|
||||
if "/" in req and "/" in act:
|
||||
return req == act
|
||||
return req.rsplit("/", 1)[-1] == act.rsplit("/", 1)[-1]
|
||||
|
||||
|
||||
def persist_agent_evidence(
|
||||
*,
|
||||
run_id: str,
|
||||
agent_role: str,
|
||||
system_prompt: str,
|
||||
user_message: str,
|
||||
final_message: str | None,
|
||||
transcript: str,
|
||||
model_calls: Sequence[Mapping[str, Any]],
|
||||
requested_model_id: str,
|
||||
connect: Callable[..., Any] | None = None,
|
||||
creator: str = "dispatch-agent-task",
|
||||
dry_run: bool = False,
|
||||
) -> dict[str, Any]:
|
||||
"""把一次框架派发的全部证据原子落库。
|
||||
|
||||
一个事务内写入:raw 租约(purpose=agent_task)+ prompt/response/supplier 三份
|
||||
raw 全文 + 每个模型回合一条 ``example_llm_call`` 投影。任一步失败整体回滚,
|
||||
调用方拿不到看似成功却缺证据的结果。
|
||||
"""
|
||||
|
||||
import sys
|
||||
from pathlib import Path
|
||||
|
||||
here = Path(__file__).resolve().parent
|
||||
if str(here) not in sys.path:
|
||||
sys.path.insert(0, str(here))
|
||||
from persist_raw import _bare_sha256, _check_no_secrets # noqa: E402
|
||||
|
||||
if not isinstance(run_id, str) or not run_id or len(run_id) > 64:
|
||||
raise ValueError("agent 证据 run_id 为空或超过 64 字符")
|
||||
if not isinstance(agent_role, str) or not agent_role or len(agent_role) > 32:
|
||||
raise ValueError("agent 证据 agent_role 为空或超过 32 字符")
|
||||
if not isinstance(system_prompt, str) or not system_prompt:
|
||||
raise ValueError("agent 证据缺 system prompt")
|
||||
if not isinstance(user_message, str) or not user_message:
|
||||
raise ValueError("agent 证据缺任务输入")
|
||||
if final_message is not None and not isinstance(final_message, str):
|
||||
raise ValueError("agent 证据 final_message 必须是字符串或 NULL")
|
||||
if not isinstance(transcript, str) or not transcript:
|
||||
raise ValueError("agent 证据缺框架转录")
|
||||
_check_no_secrets(system_prompt)
|
||||
_check_no_secrets(user_message)
|
||||
_check_no_secrets(transcript)
|
||||
if final_message:
|
||||
_check_no_secrets(final_message)
|
||||
if not isinstance(requested_model_id, str) or not requested_model_id or len(requested_model_id) > 64:
|
||||
raise ValueError("agent 证据 requested_model_id 为空或超过 64 字符")
|
||||
if not isinstance(model_calls, (list, tuple)) or not all(
|
||||
isinstance(call, Mapping) for call in model_calls
|
||||
):
|
||||
raise ValueError("agent 证据 model_calls 必须是对象数组")
|
||||
|
||||
prompt_request = json.dumps(
|
||||
{"system": system_prompt, "user": user_message},
|
||||
ensure_ascii=False,
|
||||
sort_keys=True,
|
||||
separators=(",", ":"),
|
||||
)
|
||||
prompt_sha = _bare_sha256(prompt_request)
|
||||
final_sha = _bare_sha256(final_message) if final_message else None
|
||||
transcript_sha = _bare_sha256(transcript)
|
||||
# lease 的哈希清单按 raw kind 计数;system/user 哈希已封装在 prompt 内容内,
|
||||
# 不另造不会对应 raw 行的清单项,保证轮次封存不变量可机械核对。
|
||||
content_hashes = {"prompt": prompt_sha, "supplier": transcript_sha}
|
||||
if final_sha is not None:
|
||||
content_hashes["response"] = final_sha
|
||||
|
||||
connect_fn = _connect_factory(connect)
|
||||
with connect_fn() as conn:
|
||||
try:
|
||||
lease_id = conn.execute(
|
||||
"INSERT INTO example_raw_lease(run_id, source_version, content_hashes, purpose, status, creator) "
|
||||
"VALUES (%s,%s,%s::jsonb,%s,%s,%s) RETURNING id",
|
||||
(
|
||||
run_id,
|
||||
None,
|
||||
json.dumps(content_hashes, ensure_ascii=False),
|
||||
"agent_task",
|
||||
"closed",
|
||||
creator,
|
||||
),
|
||||
).fetchone()[0]
|
||||
|
||||
def _content(kind: str, text: str, role: str) -> int:
|
||||
row = conn.execute(
|
||||
"INSERT INTO example_raw_content(lease_id, kind, run_id, role, content_sha256, content, creator) "
|
||||
"VALUES (%s,%s,%s,%s,%s,%s,%s) ON CONFLICT (lease_id, content_sha256) DO NOTHING "
|
||||
"RETURNING id",
|
||||
(lease_id, kind, run_id, role, _bare_sha256(text), text, creator),
|
||||
).fetchone()
|
||||
if row:
|
||||
return row[0]
|
||||
row = conn.execute(
|
||||
"SELECT id FROM example_raw_content WHERE lease_id=%s AND content_sha256=%s",
|
||||
(lease_id, _bare_sha256(text)),
|
||||
).fetchone()
|
||||
if not row:
|
||||
raise RuntimeError(f"raw {kind} 幂等回读失败")
|
||||
return row[0]
|
||||
|
||||
prompt_id = _content("prompt", prompt_request, agent_role)
|
||||
response_id = _content("response", final_message, agent_role) if final_message else None
|
||||
transcript_id = _content("supplier", transcript, agent_role)
|
||||
|
||||
llm_call_ids: list[int] = []
|
||||
for index, call in enumerate(model_calls, start=1):
|
||||
actual = call.get("actual_model_id")
|
||||
if not isinstance(actual, str) or not actual or len(actual) > 64:
|
||||
raise ValueError(f"第 {index} 个模型回合 actual_model_id 为空或超过 64 字符")
|
||||
usage = call.get("usage") or {}
|
||||
if not isinstance(usage, Mapping):
|
||||
raise ValueError(f"第 {index} 个模型回合 usage 必须是对象")
|
||||
cost = call.get("cost_usd")
|
||||
if cost is not None:
|
||||
try:
|
||||
numeric_cost = float(cost)
|
||||
except (TypeError, ValueError, OverflowError) as exc:
|
||||
raise ValueError(f"第 {index} 个模型回合 cost_usd 非法") from exc
|
||||
if not math.isfinite(numeric_cost) or numeric_cost < 0:
|
||||
raise ValueError(f"第 {index} 个模型回合 cost_usd 非法")
|
||||
duration = call.get("duration_ms")
|
||||
if duration is not None:
|
||||
if isinstance(duration, bool) or not isinstance(duration, int) or duration < 0:
|
||||
raise ValueError(f"第 {index} 个模型回合 duration_ms 非法")
|
||||
stop_reason = call.get("stop_reason")
|
||||
if stop_reason is not None and not isinstance(stop_reason, str):
|
||||
raise ValueError(f"第 {index} 个模型回合 stop_reason 非法")
|
||||
in_tokens, out_tokens, cached_tokens = _tokens(usage)
|
||||
row = conn.execute(
|
||||
"INSERT INTO example_llm_call(window_key, run_id, caller, requested_model_id, "
|
||||
"actual_model_id, model_match, in_tokens, cached_tokens, out_tokens, cost_usd, "
|
||||
"stop_reason, duration_ms, prompt_sha256, raw_content_id, creator) "
|
||||
"VALUES (%s,%s,%s,%s,%s,%s,%s,%s,%s,%s,%s,%s,%s,%s,%s) RETURNING id",
|
||||
(
|
||||
None,
|
||||
run_id,
|
||||
creator,
|
||||
requested_model_id,
|
||||
actual,
|
||||
model_ids_match(requested_model_id, actual),
|
||||
in_tokens,
|
||||
cached_tokens,
|
||||
out_tokens,
|
||||
cost if cost is not None else 0,
|
||||
stop_reason[:32] if stop_reason else None,
|
||||
duration,
|
||||
prompt_sha,
|
||||
transcript_id,
|
||||
creator,
|
||||
),
|
||||
).fetchone()
|
||||
llm_call_ids.append(row[0])
|
||||
|
||||
result = {
|
||||
"status": "written",
|
||||
"leaseId": lease_id,
|
||||
"promptId": prompt_id,
|
||||
"responseId": response_id,
|
||||
"transcriptId": transcript_id,
|
||||
"llmCallIds": llm_call_ids,
|
||||
}
|
||||
if dry_run:
|
||||
conn.rollback()
|
||||
result["status"] = "dry_run_ok"
|
||||
result["note"] = "试跑已回滚,未落库"
|
||||
else:
|
||||
conn.commit()
|
||||
return result
|
||||
except Exception:
|
||||
conn.rollback()
|
||||
raise
|
||||
|
||||
|
||||
__all__ = [
|
||||
"AgentTraceWriter",
|
||||
"CREATOR",
|
||||
"EVENT_TYPES",
|
||||
"model_ids_match",
|
||||
"persist_agent_evidence",
|
||||
]
|
||||
@ -63,7 +63,10 @@ def check_invariants(run_id=None):
|
||||
leases = conn.execute("SELECT id, content_hashes FROM example_raw_lease").fetchall()
|
||||
seal_violations = 0
|
||||
for lease_id, content_hashes in leases:
|
||||
declared = len(content_hashes) if isinstance(content_hashes, list) else 0
|
||||
if isinstance(content_hashes, (list, dict)):
|
||||
declared = len(content_hashes)
|
||||
else:
|
||||
declared = 0
|
||||
actual = _count(conn, "SELECT count(*) FROM example_raw_content WHERE lease_id=%s", (lease_id,))
|
||||
if declared != actual:
|
||||
seal_violations += 1
|
||||
|
||||
@ -11,7 +11,7 @@ import re
|
||||
import uuid
|
||||
|
||||
|
||||
from muse_db import connect
|
||||
from muse_db import connect as connect_default
|
||||
|
||||
|
||||
CREATOR = "runtime"
|
||||
@ -28,8 +28,16 @@ def new_run_id(stage, *, work_id=None, target_chapter=None):
|
||||
return f"{prefix}{scope}{target}-{stamp}-{uuid.uuid4().hex[:10]}"[:64]
|
||||
|
||||
|
||||
def _connect(connect=None):
|
||||
"""运行登记默认走 muse_db;派发器与测试可注入自己的连接工厂。"""
|
||||
|
||||
if connect is not None:
|
||||
return connect
|
||||
return connect_default
|
||||
|
||||
|
||||
def start_run(*, run_id=None, work_id=None, target_chapter=None,
|
||||
trigger_source="user", trigger_detail=None, creator=CREATOR):
|
||||
trigger_source="user", trigger_detail=None, creator=CREATOR, connect=None):
|
||||
"""登记或回读一个运行;已有同 ID 运行必须属于同一作品和目标章。"""
|
||||
if trigger_source not in ("user", "replay_eval", "diagnostic"):
|
||||
raise ValueError(f"trigger_source 非法: {trigger_source}")
|
||||
@ -37,14 +45,14 @@ def start_run(*, run_id=None, work_id=None, target_chapter=None,
|
||||
if len(run_id) > 64:
|
||||
raise ValueError("run_id 超过 64 字符")
|
||||
detail = json.dumps(trigger_detail, ensure_ascii=False) if trigger_detail is not None else None
|
||||
with connect() as conn:
|
||||
with _connect(connect)() as conn:
|
||||
try:
|
||||
conn.execute(
|
||||
inserted = conn.execute(
|
||||
"INSERT INTO example_run(run_id, work_id, target_chapter, trigger_source, trigger_detail, "
|
||||
"terminal_state, creator, updater) VALUES (%s,%s,%s,%s,%s::jsonb,'running',%s,%s) "
|
||||
"ON CONFLICT (run_id) DO NOTHING",
|
||||
"ON CONFLICT (run_id) DO NOTHING RETURNING run_id",
|
||||
(run_id, work_id, target_chapter, trigger_source, detail, creator, creator),
|
||||
)
|
||||
).fetchone()
|
||||
row = conn.execute(
|
||||
"SELECT run_id, work_id, target_chapter, terminal_state FROM example_run "
|
||||
"WHERE run_id=%s AND deleted=FALSE",
|
||||
@ -60,19 +68,19 @@ def start_run(*, run_id=None, work_id=None, target_chapter=None,
|
||||
"work_id": row[1],
|
||||
"target_chapter": row[2],
|
||||
"terminal_state": row[3],
|
||||
"status": "existing" if row[3] != "running" else "started",
|
||||
"status": "started" if inserted else "existing",
|
||||
}
|
||||
except Exception:
|
||||
conn.rollback()
|
||||
raise
|
||||
|
||||
|
||||
def finish_run(run_id, terminal_state, *, creator=CREATOR, trigger_detail=None):
|
||||
def finish_run(run_id, terminal_state, *, creator=CREATOR, trigger_detail=None, connect=None):
|
||||
"""把运行置为 completed/failed,并由数据库约束保证有 finished_at。"""
|
||||
if terminal_state not in _TERMINAL_STATES:
|
||||
raise ValueError(f"终态非法: {terminal_state}")
|
||||
detail = json.dumps(trigger_detail, ensure_ascii=False) if trigger_detail is not None else None
|
||||
with connect() as conn:
|
||||
with _connect(connect)() as conn:
|
||||
try:
|
||||
row = conn.execute(
|
||||
"UPDATE example_run SET terminal_state=%s, finished_at=CURRENT_TIMESTAMP, "
|
||||
@ -91,7 +99,7 @@ def finish_run(run_id, terminal_state, *, creator=CREATOR, trigger_detail=None):
|
||||
|
||||
@contextmanager
|
||||
def managed_run(*, run_id=None, work_id=None, target_chapter=None,
|
||||
trigger_source="user", trigger_detail=None, creator=CREATOR):
|
||||
trigger_source="user", trigger_detail=None, creator=CREATOR, connect=None):
|
||||
"""以成功/失败终态包住一个生产阶段。"""
|
||||
record = start_run(
|
||||
run_id=run_id,
|
||||
@ -100,16 +108,17 @@ def managed_run(*, run_id=None, work_id=None, target_chapter=None,
|
||||
trigger_source=trigger_source,
|
||||
trigger_detail=trigger_detail,
|
||||
creator=creator,
|
||||
connect=connect,
|
||||
)
|
||||
active_id = record["run_id"]
|
||||
try:
|
||||
yield active_id
|
||||
except BaseException as exc:
|
||||
finish_run(active_id, "failed", creator=creator,
|
||||
trigger_detail={"error_type": type(exc).__name__})
|
||||
trigger_detail={"error_type": type(exc).__name__}, connect=connect)
|
||||
raise
|
||||
else:
|
||||
finish_run(active_id, "completed", creator=creator)
|
||||
finish_run(active_id, "completed", creator=creator, connect=connect)
|
||||
|
||||
|
||||
__all__ = ["finish_run", "managed_run", "new_run_id", "start_run"]
|
||||
|
||||
@ -49,7 +49,7 @@
|
||||
"oracleInputProvenance": "oracle_reference_scaffold",
|
||||
"maxContextChars": 140000,
|
||||
"modelVersion": "fixed-opus-v1",
|
||||
"adapterVersion": "muse-role-v3",
|
||||
"adapterVersion": "muse-role-v4",
|
||||
"sampling": {
|
||||
"temperature": 0.2,
|
||||
"topP": null,
|
||||
@ -118,15 +118,15 @@
|
||||
"runtimeProbe": {
|
||||
"schemaVersion": "runtime-probe-v2",
|
||||
"status": "successful",
|
||||
"checkedAt": "2026-08-21T17:09:22+00:00",
|
||||
"checkedAt": "2026-08-22T02:58:14+00:00",
|
||||
"runtimeAdapter": "muse-role",
|
||||
"runtimeAdapterVersion": "muse-role-v3",
|
||||
"runtimeAdapterVersion": "muse-role-v4",
|
||||
"modelPolicyVersion": "fixed-opus-v1",
|
||||
"role": "writer",
|
||||
"profileVersion": "writer-gate-a-role-v5",
|
||||
"modelAlias": "opus",
|
||||
"resolvedModelId": "claude-opus-4-8[1M]",
|
||||
"executionProfileSha256": "sha256:e6a6423620004d51d596eb051ba60c4046ea170054737f043c623adf91639e64",
|
||||
"executionProfileSha256": "sha256:8be7ac47f77b6957e4f6beb049e9d393c83b44ba261a331b99903fb3fb7d26f6",
|
||||
"jsonSchemaId": "writer-draft-v2",
|
||||
"jsonSchemaSha256": "sha256:a1fc5efbcd7aee11082b547eb4e156fe27abe0d669991d5825e3e69901278709",
|
||||
"systemPromptId": "writer-gate-a-system-v2",
|
||||
@ -135,16 +135,16 @@
|
||||
"requestedModelId": "opus",
|
||||
"actualModelId": "claude-opus-4-8",
|
||||
"modelMatch": true,
|
||||
"executionReceiptSha256": "sha256:acd7076a5f33150818a232929862804fc8a26a39d4a840a5d88fa95464f00dde",
|
||||
"structuredOutputSha256": "sha256:cba4507ac0b637fe6a04ae2d72438ab7d4496f67de38a41c63ad18f5c88a79bf",
|
||||
"executionReceiptSha256": "sha256:bbadc1edb361e102c727b350d76289180ad0e6c28a13a89c410c837f70bf66a0",
|
||||
"structuredOutputSha256": "sha256:fc6e7ab2641cf4081ada60ed245071c92ef0fee98daeba3242fa4c35aa95e3b0",
|
||||
"terminalReason": "completed",
|
||||
"totalCostUsd": "0.066216",
|
||||
"receiptSha256": "sha256:095c164e4dc9c61a9151eb09df55ce4066766f16249f8ecffec2baee97561724"
|
||||
"totalCostUsd": "0.105666",
|
||||
"receiptSha256": "sha256:c96feda6d073f3b4820c7e0561f57270527142ee39d02bcaf2f976297a64d6be"
|
||||
},
|
||||
"profileSha256": {
|
||||
"writer": "sha256:e6a6423620004d51d596eb051ba60c4046ea170054737f043c623adf91639e64",
|
||||
"semantic_detector": "sha256:c44a27ff73bf6d72ed697d55cdc117688179bef2b266db5bc2086c934f89315a",
|
||||
"blind_judge": "sha256:cda85ea8d5b22f52c4a4fd838f9459ffcd2b115d46b3315a6364d616c4090df5"
|
||||
"writer": "sha256:8be7ac47f77b6957e4f6beb049e9d393c83b44ba261a331b99903fb3fb7d26f6",
|
||||
"semantic_detector": "sha256:5de4c5e36968f4335ed6977620ca8133e9e9ab2588b3405840de0edc3899d791",
|
||||
"blind_judge": "sha256:a3078d38f315d01453e71b8390ec96635c49e51ac199aab6ae5c6d63cb7c57ef"
|
||||
},
|
||||
"budget": {
|
||||
"status": "approved",
|
||||
|
||||
@ -35,7 +35,6 @@ from muse_role import ( # noqa: E402
|
||||
MODEL_POLICY_VERSION,
|
||||
RUNTIME_ADAPTER,
|
||||
RUNTIME_ADAPTER_VERSION,
|
||||
ROLE_TASK_SEPARATOR,
|
||||
RoleExecutionProfile,
|
||||
RoleExecutionReceipt,
|
||||
RoleRuntimeError,
|
||||
@ -92,14 +91,6 @@ from ._common import (
|
||||
from .budget import _budget_amount
|
||||
|
||||
|
||||
AGENTS_DIR = Path(__file__).resolve().parents[4] / "agents"
|
||||
PROFILE_ROLE_FILES = {
|
||||
"writer": "writer.md",
|
||||
"semantic_detector": "detector.md",
|
||||
"blind_judge": "judge.md",
|
||||
}
|
||||
|
||||
|
||||
def profile_from_mapping(value: Any, *, role: str) -> RoleExecutionProfile:
|
||||
"""从显式配置构造冻结 profile,不接受默认模型或隐式 schema/预算。"""
|
||||
|
||||
@ -126,9 +117,10 @@ def profile_from_mapping(value: Any, *, role: str) -> RoleExecutionProfile:
|
||||
return int(value) if integer else float(value)
|
||||
|
||||
system_prompt = str(profile.get("systemPrompt") or "")
|
||||
role_prompt = (AGENTS_DIR / PROFILE_ROLE_FILES[role]).read_text(encoding="utf-8").rstrip()
|
||||
if not system_prompt.startswith(role_prompt + ROLE_TASK_SEPARATOR):
|
||||
raise WriterReplayError(f"{role} profile 未注入完整角色 prompt")
|
||||
if not system_prompt.strip():
|
||||
raise WriterReplayError(f"{role} profile 缺少冻结 system prompt 快照")
|
||||
# 回放配置保存的是当次不可变 prompt 快照;当前角色合同由在线派发链从
|
||||
# .agent/docs/architecture/角色合同.md 装配,不能用当前角色文件反查历史快照。
|
||||
temperature_raw = profile.get("temperature")
|
||||
temperature = 0.2 if temperature_raw in (None, "", "unsupported") else float(temperature_raw)
|
||||
try:
|
||||
|
||||
@ -36,6 +36,10 @@ Writer 不接收 `runId`、权限信息、manifest、hash、候选版本、验
|
||||
4. 缺少细纲字段、`factConstraints` 字段或篇幅合同属于 adapter 输入错误,必须在模型调用前失败。`factConstraints=[]` 在冻结检索确实没有可确认事实时是合法输入,不等于“事实已验证”。写手可以在正文里设计新设定,但不得把新设定冒充已确认事实。
|
||||
5. detector 只把「对已有正典/前文章节的主张检索不够」标成 `evidenceGaps` 并触发补证重写。写手新写出的设定进 `newSettingCandidates`,不因此重写或禁写;与既有正典冲突才失败关闭。新设定是否进入正典由人决定。Writer 输出不承载补证请求或审查结论。
|
||||
|
||||
## 生产入口
|
||||
|
||||
Dashboard 与人工生产入口统一调用 `scripts/produce_next_chapter.py`。该入口只负责编排本 Skill 已登记的冻结、检索、writer、detector、CAS、候选落库和人闸步骤;不提供自动 accept。运行 artifacts 仍由只读看板按 run_id 读取。
|
||||
|
||||
## 生产落库
|
||||
|
||||
- 生产编排走 `run_writer_pipeline`(机械门→语义 detector→补证/重写有限环),状态链用 `scripts/candidate_cas.py` 的 `PostgresCasStateStore` 持久化到 `example_candidate_cas`(一次运行一条链,revision 单调,DB 触发器锁方向闭集);内存 `InMemoryCasStateStore` 仅供离线测试。
|
||||
|
||||
@ -13,7 +13,7 @@
|
||||
→ accept_preflight(check_writer_acceptance 纯函数 + acceptance_state 实时重读)
|
||||
→ 停止并展示候选,等待用户明确选择改 / 丢弃 / 采纳
|
||||
|
||||
用法:.venv/bin/python docs/write-chapter/step2_write_chapter.py [目标章号]
|
||||
用法:.venv/bin/python .agent/skills/write-next-chapter/scripts/produce_next_chapter.py [目标章号]
|
||||
[--instruction "本轮人指令原文"]
|
||||
缺省写下一章(库内最大章序 +1)。前置:该章已建且有 confirmed 细纲(镜像
|
||||
step2_setup_chapter2.py 建章 + 落细纲),且门锚合同 GATE_ANCHORS 已登记该章。
|
||||
@ -30,8 +30,9 @@ from pathlib import Path
|
||||
from typing import Any, Mapping
|
||||
|
||||
SCRIPT_DIR = Path(__file__).resolve().parent
|
||||
AGENT_ROOT = SCRIPT_DIR.parents[1]
|
||||
SKILLS = AGENT_ROOT / ".agent" / "skills"
|
||||
REPO_ROOT = SCRIPT_DIR.parents[3]
|
||||
AGENT_ROOT = SCRIPT_DIR.parents[2]
|
||||
SKILLS = AGENT_ROOT / "skills"
|
||||
for sub in (
|
||||
"assemble-context/scripts",
|
||||
"write-next-chapter/scripts",
|
||||
@ -104,7 +105,7 @@ SYSTEM_PROMPT = GATE_A_WRITER["systemPrompt"] + PRODUCTION_LENGTH_PROMPT
|
||||
SYSTEM_PROMPT_ID = "writer-production-system-v4-new-settings"
|
||||
SYSTEM_PROMPT_SHA256 = "sha256:" + hashlib.sha256(SYSTEM_PROMPT.encode("utf-8")).hexdigest()
|
||||
|
||||
ARTIFACTS = SCRIPT_DIR / "artifacts"
|
||||
ARTIFACTS = REPO_ROOT / "docs" / "write-chapter" / "artifacts"
|
||||
|
||||
# 门锚合同按章登记:锚点是章级创作判断,any-hit 子串匹配。新章必须先登记再跑。
|
||||
# 机械门(check_writer_candidate)只认这些子串,不认语义等价;必须投影给写手,
|
||||
@ -1,7 +1,7 @@
|
||||
#!/usr/bin/env python3
|
||||
"""通过统一治理 runtime 运行正文写手并绑定候选身份。
|
||||
|
||||
执行器是 muse_role 的固定 Opus HTTP 策略:角色合同全文作系统提示词,冻结输入和
|
||||
执行器是 muse_role 的固定 Opus HTTP 策略:身份提示与中心角色合同作系统提示词,冻结输入和
|
||||
JSON Schema 进入同一次调用,输出校验后生成回执。不依赖模型 CLI 或宿主装载机制,
|
||||
模型不可用时失败关闭,不降级到内容模型链。
|
||||
"""
|
||||
@ -18,6 +18,7 @@ READ_CONTEXT_DIR = SCRIPT_DIR.parents[1] / "assemble-context" / "scripts"
|
||||
if str(READ_CONTEXT_DIR) not in sys.path:
|
||||
sys.path.insert(0, str(READ_CONTEXT_DIR))
|
||||
|
||||
from muse_role_contract import load_role_contract_catalog # noqa: E402
|
||||
from muse_role import ( # noqa: E402
|
||||
FIXED_OPUS_MODEL_ID,
|
||||
FIXED_OPUS_POLICY_ALIAS,
|
||||
@ -32,9 +33,12 @@ from muse_role import ( # noqa: E402
|
||||
verify_role_profile,
|
||||
)
|
||||
|
||||
_ROLE_CATALOG = load_role_contract_catalog(SCRIPT_DIR.parents[3])
|
||||
WRITER_ROLE_PROMPT = (
|
||||
SCRIPT_DIR.parents[2] / "agents" / "writer.md"
|
||||
).read_text(encoding="utf-8")
|
||||
(SCRIPT_DIR.parents[2] / "agents" / "writer.md").read_text(encoding="utf-8").rstrip()
|
||||
+ "\n\n--- 角色合同(唯一事实源) ---\n"
|
||||
+ _ROLE_CATALOG.for_role("writer").contract_prompt
|
||||
)
|
||||
|
||||
from writer_contract import ( # noqa: E402
|
||||
ContractError,
|
||||
|
||||
16
AGENTS.md
16
AGENTS.md
@ -63,11 +63,11 @@ agent-example/
|
||||
└── README.md # 历史概览,不是当前运行态 SoT
|
||||
```
|
||||
|
||||
5 个角色:`writer`、`planner`、`extractor`、`detector`、`judge`。角色身份在 `.agent/agents/*.md`,具体功能合同不复制进角色文件。角色是主会话派发的子代理:按 [07-Agent与Skill领域 §2](.agent/docs/architecture/domains/07-Agent与Skill领域.md) 的派发合同起全新会话,注入角色文件全文、冻结输入,输出由派发方校验并落证据;不依赖任何宿主的原生角色装载机制(如 Claude Code `--agent`),Claude CLI 不是角色运行底座。
|
||||
5 个角色:`writer`、`planner`、`extractor`、`detector`、`judge`。角色身份在 `.agent/agents/*.md`;稳定角色合同唯一事实源是 [角色合同](.agent/docs/architecture/角色合同.md),不把输入边界、模型策略、工具权限和输出合同散落进角色文件。角色是主会话派发的子代理:按 [07-Agent与Skill领域 §2](.agent/docs/architecture/domains/07-Agent与Skill领域.md) 的派发合同起全新会话,注入身份提示、对应角色合同和冻结输入,输出由派发方校验并落证据;不依赖任何宿主的原生角色装载机制(如 Claude Code `--agent`),Claude CLI 不是角色运行底座。
|
||||
|
||||
### Skill 合同责任方索引
|
||||
|
||||
实际清单以 `.agent/skills/*/SKILL.md` 为准。发现总索引见 [`.agent/skills/_index.md`](.agent/skills/_index.md):57 个 skill 按创作生命周期分 9 域,每条只登记 `skill_name` / `skill_file` / `skill_description` 三字段,description 与 SKILL.md frontmatter 逐字一致。索引由 `harness/skills_index.py --write` 生成;skill 增删改名后必须重新生成,一致性由 `tests/architecture/test_skills_index.py` 机械校验。
|
||||
实际清单以 `.agent/skills/*/SKILL.md` 为准。发现总索引见 [`.agent/skills/_index.md`](.agent/skills/_index.md):58 个 skill 按创作生命周期分 9 域,每条只登记 `skill_name` / `skill_file` / `skill_description` 三字段,description 与 SKILL.md frontmatter 逐字一致。索引由 `harness/skills_index.py --write` 生成;skill 增删改名后必须重新生成,一致性由 `tests/architecture/test_skills_index.py` 机械校验。
|
||||
|
||||
本表是另一条轴:登记每个 skill 的合同责任方、协作领域和领域 SoT,不复制各 Skill 的完整合同。每个 skill 必须登记一个合同责任方(业务领域或平台领域),但可以同时消费或影响多个协作领域;跨域调用、场景关系和保护节点在 `meta/chains/` 登记。合同责任方表示谁维护该 Skill 的稳定能力合同,不表示 Skill 只能属于一个业务领域。
|
||||
|
||||
@ -81,7 +81,7 @@ agent-example/
|
||||
| 质量与回放评测 | 06-质量与复利、05-创作流程 | `check-content-consistency`、`score-content-quality`、`adjudicate-quality-gate`、`optimize-content-quality`、`evaluate-frozen-replay`、`replay-writer-gate`、`load-replay-reference-work`、`novel-diagnosis` |
|
||||
| 去 AI 味与人感 | 06-质量与复利、父仓专题-09 | `capture-ai-flavor-cases`、`promote-ai-flavor-rule`、`diagnose-ai-flavor`、`establish-voice-baseline`、`prevent-ai-flavor`、`revise-ai-flavor` |
|
||||
|
||||
57 个 skill 一律是本仓正式 skill,受同一套合同与门禁约束,不分等级:都须满足 [07-Agent与Skill领域 §3](.agent/docs/architecture/domains/07-Agent与Skill领域.md) 的合同,都在 `_index.md` 与 `skills.json` 登记,都进质量评分。绑创作 scenario 的在 `meta/chains/` 登记;平台与工具类(如 `call-content-model`、`execute-role-task`、`record-run-evidence`)由主会话或其它 Skill 直接调用,不绑 scenario。
|
||||
58 个 skill 一律是本仓正式 skill,受同一套合同与门禁约束,不分等级:都须满足 [07-Agent与Skill领域 §3](.agent/docs/architecture/domains/07-Agent与Skill领域.md) 的合同,都在 `_index.md` 与 `skills.json` 登记,都进质量评分。绑创作 scenario 的在 `meta/chains/` 登记;平台与工具类(如 `call-content-model`、`execute-role-task`、`record-run-evidence`)由主会话或其它 Skill 直接调用,不绑 scenario。
|
||||
|
||||
**不按"是不是系统运行时"分等级。** 一个 Skill 当前有没有 `scripts/`、有没有数据库合同、有没有接入复利,是实现成熟度而非本质:`plan-chapter`、`expand-scene`、`polish-prose` 以模型判断为主、自身不带 Tool,落库由它们调用的 Skill 承担;`story-structure`、`scene-craft` 一类创作方法 Skill 目前只有 `SKILL.md` 与 `references/`,那是**未接入复利的欠账**,不是它们的天然形态(改造方向见下)。把成熟度写成类别,等于给未完成的 Skill 发永久豁免证。
|
||||
|
||||
@ -124,10 +124,10 @@ Skill 领域列表的新增、删除、改名或主领域调整,必须同时
|
||||
## 6. 模型边界
|
||||
|
||||
- 清洗、抽卡、范式拆取及其模型调用统一走 `call-content-model` Skill,不裸调 New-API。治理政策固定为 5 小时额度窗:MiniMax 模型累计花费上限 `$24`,全模型成功调用上限 `6000`;运行适配器、正式配置和账本是额度合同的事实源,共享库 `muse_llm` 与 `muse_db.WINDOW_BUDGET_USD` / `WINDOW_CALL_CAP` 是实现,Skill CLI 只做入口,`test_quota.py` 只提供回归证据;模型链切换必须由该治理入口留下日志。
|
||||
- 角色模型归属:`planner`/`writer`/`judge` 固定 `opus`;`extractor`/`detector` 可用其它模型(非必须降级)。拆书/导入侧抽取经 `call-content-model`/`deconstruct-book` Skill 走 MiniMax-M3,不走角色 model 派发;创作期章后抽取作为角色派发,可用 `opus`。
|
||||
- 角色模型归属和派发字段以 [角色合同](.agent/docs/architecture/角色合同.md) 为准:`planner`/`writer`/`judge` 固定 `opus`;`extractor`/`detector` 可在合同允许的治理策略内运行。每次框架调用必须显式传入 `provider`、`model` 和 `thinking`,不得从环境变量静默补全。拆书/导入侧抽取经 `call-content-model`/`deconstruct-book` Skill 走 MiniMax-M3,不走角色 model 派发;创作期章后抽取作为角色派发,可用 `opus`。
|
||||
- 确定性脚本、合同校验、快照冻结、泄漏审计和报告生成不调用模型;除非对应 `SKILL.md` 明确声明模型步骤,不得把机械任务升级为模型任务。
|
||||
- 固定 Opus 角色生成或评测只在对应任务 SoT、显式预算、冻结 profile 和原文用途授权全部满足后运行;自动化调用走 Anthropic 兼容 HTTP 适配器,不读取 Claude Code 配置,不启动模型 CLI。任一前置门失败都关闭执行。
|
||||
- 角色的模型策略版本、模型别名、完整模型 ID、预算和回执必须与冻结配置一致。`planner`/`writer`/`judge` 不得因模型不可用而降级到内容模型链或更换供应商;需要变更时先取得明确授权并重新登记配置、profile 与探针。
|
||||
- 角色的模型策略版本、模型别名、完整模型 ID、预算和回执必须与冻结配置一致。`planner`/`writer`/`judge` 不得因模型不可用而降级到内容模型链或更换供应商;需要变更时先取得明确授权并更新角色合同、profile 与探针。
|
||||
|
||||
## 7. 会话编排与汇报
|
||||
|
||||
@ -175,14 +175,16 @@ git diff --check
|
||||
|
||||
### 11.1 Agent 提示词(`.agent/agents/*.md` 与角色系统提示词)
|
||||
|
||||
角色稳定合同见 [角色合同](.agent/docs/architecture/角色合同.md)。角色文件可以包含 Agent-facing 的身份、Skill 路由、推荐工具能力和工作方法;审查硬边界、模型策略、实际工具权限和结构化输出时以中心合同与适配器为准。
|
||||
|
||||
逐条问四个问题:
|
||||
|
||||
1. **面对谁**:这份提示词读给哪个模型/角色?它只该有这一个身份,不该同时背「执行器/评测器/审查器」之类第二身份。
|
||||
2. **每个词都有意义吗**:逐词问——它对「这个角色干好本业」有用吗?框架名、schema 名、字段名、运行身份、哈希、评测状态这类机器词,对创作/规划/抽取/检测等本体任务毫无意义,是其它层的泄漏,应删。
|
||||
3. **约束是该有的限制吗**:每条约束问——它是角色意图本身需要的,还是支架(harness)本就能强制的?**输出格式**由结构化输出 schema 强制、**工具可用性**由调用参数强制、**盲化**由「输入里压根没有该信息」保证——这些都不该写进提示词反复叮嘱。提示词只留角色凭自身判断必须遵守的约束。
|
||||
3. **约束是该有的限制吗**:每条约束问——它是角色意图本身需要的,还是支架(harness)本就能强制的?**输出格式**由结构化输出 Schema 强制、**实际工具可用性**由调用参数强制、**盲化**由「输入里压根没有该信息」保证;角色文件可以解释 Skill 和工具的用途,但不能把提示文字当权限或结构门。
|
||||
4. **正向与负向**:分清哪些部分**帮**角色达成意图(正向:本业纪律、领域边界、知情范围),哪些**妨碍**它(负向:与本业无关的机器约束、诱导照搬输入原文的措辞、让模型惦记评测的暗示)。负向部分删除或移到它该在的层。
|
||||
|
||||
> 反例(已纠正):评测写手提示词曾塞入「你是 Gate A 离线回放的 writer…只输出 candidateBody…不输出哈希/身份…不访问 MCP」,把评测支架混进创作提示词——既没有写作指导,又诱导写手照抄细纲概述句。正解:提示词只讲怎么写好,输出格式与盲化交给 schema 和输入设计。
|
||||
> 反例(已纠正):评测写手提示词曾塞入「你是 Gate A 离线回放的 writer…只输出 candidateBody…不输出哈希/身份…不访问 MCP」,把评测支架混进创作提示词——既没有写作指导,又诱导写手照抄细纲概述句。正解:角色文件保留写作方法、Skill 路由和工具用途;输出格式、实际工具权限、盲化和证据绑定交给中心合同、Schema 与适配器。
|
||||
|
||||
### 11.2 Skill(`.agent/skills/*/`)
|
||||
|
||||
|
||||
@ -2347,7 +2347,7 @@ def _run_pipeline_panel(run_id: str) -> str:
|
||||
f" · 候选版本 <code>{esc(ver)}</code> · 补证请求 {esc(evidence_n)} · 重写 {esc(rewrite_n)}</p>"
|
||||
f"<p style='margin:0 0 4px'><b>证据缺口</b>({len(gap_lis)})</p>{gap_html}"
|
||||
f"<p style='margin:12px 0 4px'><b>新设定提案</b>({len(setting_lis)})</p>{setting_html}"
|
||||
"<p class='note' style='margin:12px 0 0'>来源:本地 <code>docs/write-chapter/artifacts/</code>。"
|
||||
"<p class='note' style='margin:12px 0 0'>来源:生产写作运行 artifacts(<code>docs/write-chapter/artifacts/</code>)。"
|
||||
"对已有正典的缺口才走补证重写;新设定不禁写、不自动入库,由人决定采纳/改/丢弃。"
|
||||
"仍失败关闭时可能<strong>不落候选行</strong>。</p>"
|
||||
"</div></div>"
|
||||
@ -2497,7 +2497,7 @@ def _run_decision_menu_panel(run_id: str, candidates: list, terminal_state, pipe
|
||||
# 预填改指令:绕开不可证缺口(阶段 0.1 选 A 时用)
|
||||
avoid = ";".join(f"避开「{h}」" for h in gap_hints[:2]) if gap_hints else "按语义/诊断理由修改"
|
||||
rewrite_cmd = (
|
||||
f".venv/bin/python docs/write-chapter/step2_write_chapter.py 3 "
|
||||
f".venv/bin/python .agent/skills/write-next-chapter/scripts/produce_next_chapter.py 3 "
|
||||
f"--instruction \"改:{avoid}。禁止细纲原句抄进正文;门符号只在舰队医疗舱语境。\""
|
||||
)
|
||||
if candidates:
|
||||
|
||||
@ -19,7 +19,7 @@ COMMENT ON COLUMN example_llm_call.model_match IS
|
||||
'实际模型是否符合冻结 RoleExecutionProfile 的模型策略;由执行 profile/回执哈希链证明。';
|
||||
|
||||
COMMENT ON TABLE example_agent_role IS
|
||||
'五个 prompt 管理角色的库内影子;Git 权威位于 .agent/agents/*.md,模型策略不属于角色 frontmatter。';
|
||||
'五个 prompt 管理角色的库内影子;稳定合同 Git 权威位于 .agent/docs/architecture/角色合同.md,身份提示位于 .agent/agents/*.md。';
|
||||
COMMENT ON COLUMN example_agent_role.model IS
|
||||
'保留兼容列;角色模型由派发 profile 决定,本列保持 NULL。';
|
||||
COMMENT ON TABLE example_skill IS
|
||||
|
||||
61
db/ddl/113-example代理事件账本.sql
Normal file
61
db/ddl/113-example代理事件账本.sql
Normal file
@ -0,0 +1,61 @@
|
||||
-- 代理事件账本:把任意 Agent 框架(pi/codex/opencode…)子代理执行的事件流统一留痕。
|
||||
-- 合同(07-Agent与Skill领域 §2 框架派发):框架适配器把框架原生事件归一为闭集事件类型,
|
||||
-- 逐条追加写入;本表只存身份、用量、成本与安全摘要,不存 prompt/正文原文(原文进 raw 表)。
|
||||
-- 设计边界:框架名不做闭集(接缝必须开放给任意框架);事件序号由适配器在运行内单调分配。
|
||||
|
||||
CREATE TABLE IF NOT EXISTS example_agent_event (
|
||||
id BIGINT GENERATED ALWAYS AS IDENTITY PRIMARY KEY,
|
||||
run_id VARCHAR(64) NOT NULL, -- -> example_run.run_id(软引用)
|
||||
seq INTEGER NOT NULL, -- 运行内事件序号,从 1 单调递增
|
||||
event_type VARCHAR(32) NOT NULL, -- 闭集,见 chk_example_agent_event_type
|
||||
framework VARCHAR(32) NOT NULL, -- pi/codex/opencode…(框架适配器自报,开放集)
|
||||
agent_role VARCHAR(32), -- writer/planner/detector/judge/extractor
|
||||
tool_name VARCHAR(64), -- tool.* 事件的框架工具名
|
||||
status VARCHAR(16), -- ok/error
|
||||
requested_model_id VARCHAR(64), -- 派发策略请求的模型(框架解析前)
|
||||
actual_model_id VARCHAR(64), -- 供应商实际执行模型
|
||||
input_tokens BIGINT,
|
||||
output_tokens BIGINT,
|
||||
cached_tokens BIGINT,
|
||||
cost_usd NUMERIC(14,8), -- 供应商报告或价目表计算;未知为 NULL
|
||||
raw_ref BIGINT, -- -> example_raw_content.id(软引用)
|
||||
details JSONB NOT NULL DEFAULT '{}'::jsonb, -- 安全摘要:usage 细分/stopReason/哈希/sessionId/errorCode
|
||||
creator VARCHAR(64) NOT NULL DEFAULT 'agent-trace',
|
||||
create_time TIMESTAMP NOT NULL DEFAULT CURRENT_TIMESTAMP,
|
||||
tenant_id BIGINT NOT NULL DEFAULT 0,
|
||||
CONSTRAINT uq_example_agent_event_run_seq UNIQUE (run_id, seq),
|
||||
CONSTRAINT chk_example_agent_event_type CHECK (event_type IN
|
||||
('run.started','agent.started','model.completed','tool.started','tool.completed',
|
||||
'agent.completed','agent.failed','run.completed','run.failed')),
|
||||
CONSTRAINT chk_example_agent_event_seq CHECK (seq >= 1),
|
||||
CONSTRAINT chk_example_agent_event_status CHECK (status IS NULL OR status IN ('ok','error')),
|
||||
CONSTRAINT chk_example_agent_event_model_named CHECK
|
||||
(event_type <> 'model.completed' OR actual_model_id IS NOT NULL)
|
||||
);
|
||||
CREATE INDEX IF NOT EXISTS idx_example_agent_event_run ON example_agent_event(tenant_id, run_id, seq);
|
||||
CREATE INDEX IF NOT EXISTS idx_example_agent_event_role ON example_agent_event(tenant_id, agent_role, event_type);
|
||||
|
||||
-- append-only:禁改禁删禁 TRUNCATE(与 example_run_receipt 同款守护)
|
||||
CREATE OR REPLACE FUNCTION reject_example_agent_event_mutation()
|
||||
RETURNS TRIGGER
|
||||
LANGUAGE plpgsql
|
||||
AS $$
|
||||
BEGIN
|
||||
RAISE EXCEPTION 'example_agent_event 是 append-only 表,禁止 UPDATE/DELETE';
|
||||
END;
|
||||
$$;
|
||||
|
||||
CREATE OR REPLACE TRIGGER trg_example_agent_event_append_only
|
||||
BEFORE UPDATE OR DELETE ON example_agent_event
|
||||
FOR EACH ROW EXECUTE FUNCTION reject_example_agent_event_mutation();
|
||||
|
||||
CREATE OR REPLACE TRIGGER trg_example_agent_event_no_truncate
|
||||
BEFORE TRUNCATE ON example_agent_event
|
||||
FOR EACH STATEMENT EXECUTE FUNCTION example_reject_truncate();
|
||||
|
||||
COMMENT ON TABLE example_agent_event IS
|
||||
'代理事件账本(append-only):框架适配器归一后的子代理执行事件流;prompt/正文原文在 raw 表,本表只留身份与用量。';
|
||||
COMMENT ON COLUMN example_agent_event.framework IS
|
||||
'执行框架名(pi/codex/opencode…),开放集;模型策略由派发方决定并记录在 requested_model_id。';
|
||||
COMMENT ON COLUMN example_agent_event.details IS
|
||||
'安全摘要 JSON:usage 细分、stopReason、spec/prompt/输出哈希、sessionId、errorCode;不含正文原文。';
|
||||
@ -2,7 +2,7 @@
|
||||
|
||||
> 口径(创始人拍板③ 2026-07-10):主仓表**原样不改列**;实验私货全进 `example_*` 前缀。
|
||||
> 建表方式:`db/ddl/` 下文件经 `access-database` skill `apply`,主仓部分为 `muse-cloud/sql/muse/` 原文拷贝或逐字摘录。
|
||||
> 已应用顺序:V1 → V3 → V5 → 90-ALTER摘录 → V26 → 91-example(2026-07-13)→ 97/98(2026-07-30)→ 104 AI 味案例(2026-08-14)→ 105/106/107/108/109 先审后入与复利闭环(2026-08-14)→ 110 声音账(2026-08-15)。库内表现状以 `access-database` skill `tables` 实时输出为准。
|
||||
> 已应用顺序:V1 → V3 → V5 → 90-ALTER摘录 → V26 → 91-example(2026-07-13)→ 97/98(2026-07-30)→ 104 AI 味案例(2026-08-14)→ 105/106/107/108/109 先审后入与复利闭环(2026-08-14)→ 110 声音账(2026-08-15)-> 112 角色模型策略(2026-08-21)-> 113 代理事件账本(2026-08-22)。库内表现状以 `access-database` skill `tables` 实时输出为准。
|
||||
> 96 不启用:`96-example参考作品授权快照.sql` 已实现但**决定不 apply**(单用户本地不做多租户授权机制,2026-07-30 拍板,见领域索引 §9);库内无该表。
|
||||
|
||||
## 主仓一致表(20 张)
|
||||
@ -54,6 +54,7 @@
|
||||
| example_lesson | 经验升格登记(108):lesson/win 证据绑 run_id+候选哈希;proposed→reviewing→promoted/rejected,DB 触发器禁止跳过评审的自动升格 |
|
||||
| example_candidate_cas | 候选 CAS 状态链(109):一次运行一条链(DRAFT/CHECKING/PASSED/REJECTED),revision 单调 +1,触发器锁方向闭集与身份不可变 |
|
||||
| example_voice_baseline | 声音账(110):技能 1 定基线产物,一作品一版本 append + supersede;grounding 门由 establish-voice-baseline 脚本强制;修订门禁与前置预防机械消费 |
|
||||
| example_agent_event | 代理事件账本(113):框架适配器归一后的子代理执行事件流(run/agent/model/tool 九类闭集),append-only;prompt/正文原文在 raw 表 |
|
||||
|
||||
## 暂缓建表登记(主仓有、实验现阶段未建;需要时按原样加建)
|
||||
|
||||
|
||||
@ -321,6 +321,23 @@
|
||||
],
|
||||
"skill_path": ".agent/skills/execute-role-task/SKILL.md"
|
||||
},
|
||||
{
|
||||
"name": "dispatch-agent-task",
|
||||
"lifecycle": "platform",
|
||||
"invocation": "orchestrated",
|
||||
"side_effects": [
|
||||
"external_call",
|
||||
"db_write"
|
||||
],
|
||||
"compounding": "none",
|
||||
"contract_owner": "平台运行与证据",
|
||||
"collaborates_with": [
|
||||
"规划与作品基础",
|
||||
"写作与候选主权",
|
||||
"质量与回放评测"
|
||||
],
|
||||
"skill_path": ".agent/skills/dispatch-agent-task/SKILL.md"
|
||||
},
|
||||
{
|
||||
"name": "expand-scene",
|
||||
"lifecycle": "writing",
|
||||
|
||||
@ -1650,7 +1650,7 @@
|
||||
],
|
||||
"skill_behavior_eval": false,
|
||||
"classification_confidence": "medium",
|
||||
"classification_basis": "加载 docs/write-chapter/step2_write_chapter.py 并断言机械验收子串(门锚)对写手可见;只测纯函数投影,不产生系统事实。"
|
||||
"classification_basis": "加载 .agent/skills/write-next-chapter/scripts/produce_next_chapter.py 并断言机械验收子串(门锚)对写手可见;只测纯函数投影,不产生系统事实。"
|
||||
},
|
||||
{
|
||||
"path": "tests/skills/write-next-chapter/test_persist_writer_run.py",
|
||||
@ -1750,13 +1750,46 @@
|
||||
"skill_behavior_eval": false,
|
||||
"classification_confidence": "high",
|
||||
"classification_basis": "Mocks propose_lesson_dedup for writer mechanical gate lessons; no database."
|
||||
},
|
||||
{
|
||||
"path": "tests/skills/dispatch-agent-task/test_dispatch_agent_task.py",
|
||||
"scope": "runtime_skill",
|
||||
"owner_skill_or_domain": "dispatch-agent-task",
|
||||
"kind": "runtime_contract",
|
||||
"evidence_level": "deterministic_offline",
|
||||
"requires": [
|
||||
"offline",
|
||||
"filesystem"
|
||||
],
|
||||
"side_effects": [
|
||||
"filesystem"
|
||||
],
|
||||
"skill_behavior_eval": false,
|
||||
"classification_confidence": "high",
|
||||
"classification_basis": "Exercises portable task spec contract, pi adapter argv/stream normalization and run_dispatch fail-closed chain with fake pi streams and injected DB connections; no network, no real framework binary."
|
||||
},
|
||||
{
|
||||
"path": "tests/skills/record-run-evidence/test_agent_trace.py",
|
||||
"scope": "runtime_skill",
|
||||
"owner_skill_or_domain": "record-run-evidence",
|
||||
"kind": "runtime_contract",
|
||||
"evidence_level": "deterministic_offline",
|
||||
"requires": [
|
||||
"offline"
|
||||
],
|
||||
"side_effects": [
|
||||
"none"
|
||||
],
|
||||
"skill_behavior_eval": false,
|
||||
"classification_confidence": "high",
|
||||
"classification_basis": "Fixes agent event ledger writer shapes and atomic framework evidence persistence with fake connections; no DB, no network."
|
||||
}
|
||||
],
|
||||
"summary": {
|
||||
"entry_count": 107,
|
||||
"entry_count": 109,
|
||||
"by_scope": {
|
||||
"other": 1,
|
||||
"runtime_skill": 95,
|
||||
"runtime_skill": 97,
|
||||
"harness": 3,
|
||||
"domain": 8
|
||||
},
|
||||
@ -1766,14 +1799,16 @@
|
||||
"harness_self_test": 3,
|
||||
"domain_eval": 4,
|
||||
"integration": 9,
|
||||
"tool_unit": 31,
|
||||
"tool_unit": 32,
|
||||
"fake_pipeline": 22,
|
||||
"runtime_probe": 2
|
||||
"runtime_probe": 1,
|
||||
"runtime_contract": 2
|
||||
},
|
||||
"by_evidence_level": {
|
||||
"deterministic_offline": 91,
|
||||
"deterministic_offline": 93,
|
||||
"real_dependency_integration": 10,
|
||||
"static_structure": 6
|
||||
}
|
||||
},
|
||||
"total": 109
|
||||
}
|
||||
}
|
||||
|
||||
@ -3,7 +3,7 @@ name = "muse-llm"
|
||||
version = "0.1.0"
|
||||
description = "内容模型治理链、固定 Opus HTTP 适配器与角色执行运行时"
|
||||
requires-python = ">=3.10"
|
||||
dependencies = ["requests", "json-repair", "jsonschema>=4.23,<5", "muse-db"]
|
||||
dependencies = ["requests", "json-repair", "jsonschema>=4.23,<5", "pyyaml", "muse-db"]
|
||||
|
||||
[build-system]
|
||||
requires = ["setuptools>=61"]
|
||||
@ -11,4 +11,4 @@ build-backend = "setuptools.build_meta"
|
||||
|
||||
[tool.setuptools]
|
||||
package-dir = {"" = "src"}
|
||||
py-modules = ["muse_llm", "muse_role"]
|
||||
py-modules = ["muse_llm", "muse_role", "muse_role_contract"]
|
||||
|
||||
@ -232,10 +232,97 @@ def governed_max_tokens(max_budget_usd, estimated_input_tokens):
|
||||
return min(output_cap, 512000)
|
||||
|
||||
|
||||
def _read_anthropic_sse(response, *, deadline):
|
||||
"""按字节解析 Anthropic SSE,避免 UTF-8 的 C1 字节被 splitlines 误切。"""
|
||||
|
||||
message = {}
|
||||
blocks = {}
|
||||
usage = {}
|
||||
raw_events = []
|
||||
stop_reason = None
|
||||
stop_sequence = None
|
||||
saw_message_start = False
|
||||
saw_message_stop = False
|
||||
|
||||
for raw_line in response.iter_lines(decode_unicode=False, delimiter=b"\n"):
|
||||
if _monotonic() >= deadline:
|
||||
raise requests.Timeout("固定 Opus SSE 超过总 deadline")
|
||||
if isinstance(raw_line, str):
|
||||
raw_line = raw_line.encode("utf-8")
|
||||
line = bytes(raw_line).rstrip(b"\r")
|
||||
if not line or line.startswith((b":", b"event:")) or not line.startswith(b"data:"):
|
||||
continue
|
||||
payload = line[5:].lstrip()
|
||||
if not payload or payload == b"[DONE]":
|
||||
continue
|
||||
event = json.loads(payload.decode("utf-8"))
|
||||
if not isinstance(event, dict):
|
||||
raise ValueError("Anthropic SSE data 不是对象")
|
||||
raw_events.append(event)
|
||||
event_type = event.get("type")
|
||||
|
||||
if event_type == "message_start":
|
||||
started_message = event.get("message")
|
||||
if not isinstance(started_message, dict):
|
||||
raise ValueError("Anthropic SSE 缺少 message_start.message")
|
||||
message = dict(started_message)
|
||||
usage.update(started_message.get("usage") or {})
|
||||
for index, block in enumerate(started_message.get("content") or []):
|
||||
if isinstance(block, dict):
|
||||
blocks[index] = dict(block)
|
||||
saw_message_start = True
|
||||
elif event_type == "content_block_start":
|
||||
index = event.get("index")
|
||||
block = event.get("content_block")
|
||||
if type(index) is not int or not isinstance(block, dict):
|
||||
raise ValueError("Anthropic SSE content_block_start 非法")
|
||||
blocks[index] = dict(block)
|
||||
elif event_type == "content_block_delta":
|
||||
index = event.get("index")
|
||||
delta = event.get("delta")
|
||||
if type(index) is not int or not isinstance(delta, dict):
|
||||
raise ValueError("Anthropic SSE content_block_delta 非法")
|
||||
block = blocks.setdefault(index, {})
|
||||
delta_type = delta.get("type")
|
||||
field = {
|
||||
"text_delta": "text",
|
||||
"thinking_delta": "thinking",
|
||||
"signature_delta": "signature",
|
||||
"input_json_delta": "partial_json",
|
||||
}.get(delta_type)
|
||||
if field is not None:
|
||||
block[field] = str(block.get(field) or "") + str(delta.get(field) or "")
|
||||
elif event_type == "message_delta":
|
||||
delta = event.get("delta") or {}
|
||||
if not isinstance(delta, dict):
|
||||
raise ValueError("Anthropic SSE message_delta.delta 非法")
|
||||
stop_reason = delta.get("stop_reason", stop_reason)
|
||||
stop_sequence = delta.get("stop_sequence", stop_sequence)
|
||||
usage.update(event.get("usage") or {})
|
||||
elif event_type == "message_stop":
|
||||
saw_message_stop = True
|
||||
elif event_type == "error":
|
||||
error = event.get("error") or {}
|
||||
error_type = error.get("type") if isinstance(error, dict) else None
|
||||
raise requests.RequestException(
|
||||
f"Anthropic SSE error: {error_type or 'unknown'}"
|
||||
)
|
||||
|
||||
if not saw_message_start:
|
||||
raise requests.RequestException("Anthropic SSE 缺少 message_start")
|
||||
if not saw_message_stop:
|
||||
raise requests.RequestException("Anthropic SSE 提前结束,缺少 message_stop")
|
||||
message["content"] = [blocks[index] for index in sorted(blocks)]
|
||||
message["usage"] = usage
|
||||
message["stop_reason"] = stop_reason
|
||||
message["stop_sequence"] = stop_sequence
|
||||
return message, raw_events
|
||||
|
||||
|
||||
def chat_fixed_opus(prompt, model="opus", system=None, max_tokens=FIXED_OPUS_MAX_OUTPUT_TOKENS,
|
||||
temperature=0.2, top_p=None, timeout=900, retries=2, *,
|
||||
resolved_model_id=None, run_id=None, caller=None, persist_call=None):
|
||||
"""通过 Anthropic 兼容 HTTP API 调用固定 Opus,不读取客户端配置、不降级模型。"""
|
||||
"""通过 Anthropic 兼容 SSE 调用固定 Opus,不读取客户端配置、不降级模型。"""
|
||||
|
||||
base = os.environ.get("MUSE_ROLE_OPUS_BASE_URL", "").rstrip("/")
|
||||
token = os.environ.get("MUSE_ROLE_OPUS_AUTH_TOKEN", "")
|
||||
@ -257,6 +344,7 @@ def chat_fixed_opus(prompt, model="opus", system=None, max_tokens=FIXED_OPUS_MAX
|
||||
"temperature": temperature,
|
||||
"system": system or "",
|
||||
"messages": [{"role": "user", "content": prompt}],
|
||||
"stream": True,
|
||||
}
|
||||
if top_p is not None:
|
||||
payload["top_p"] = top_p
|
||||
@ -276,18 +364,20 @@ def chat_fixed_opus(prompt, model="opus", system=None, max_tokens=FIXED_OPUS_MAX
|
||||
if remaining <= 0:
|
||||
raise RuntimeError("固定 Opus 调用超过总 deadline")
|
||||
started = time.time()
|
||||
response = None
|
||||
try:
|
||||
response = session.post(
|
||||
f"{base}/v1/messages",
|
||||
headers=headers,
|
||||
json=payload,
|
||||
timeout=(min(10, remaining), remaining),
|
||||
stream=True,
|
||||
)
|
||||
if response.status_code == 429 or response.status_code >= 500:
|
||||
last_error = f"HTTP {response.status_code}: {response.text[:200]}"
|
||||
raise requests.RequestException(last_error)
|
||||
response.raise_for_status()
|
||||
data = response.json()
|
||||
data, raw_events = _read_anthropic_sse(response, deadline=deadline)
|
||||
blocks = data.get("content")
|
||||
if not isinstance(blocks, list):
|
||||
raise KeyError("content")
|
||||
@ -320,12 +410,18 @@ def chat_fixed_opus(prompt, model="opus", system=None, max_tokens=FIXED_OPUS_MAX
|
||||
"stop_reason": data.get("stop_reason"),
|
||||
"duration_ms": duration_ms,
|
||||
"prompt": prompt_raw,
|
||||
"response": json.dumps(data, ensure_ascii=False, sort_keys=True,
|
||||
separators=(",", ":"), default=str),
|
||||
"response": json.dumps(
|
||||
{"transport": "anthropic-sse", "events": raw_events},
|
||||
ensure_ascii=False,
|
||||
sort_keys=True,
|
||||
separators=(",", ":"),
|
||||
default=str,
|
||||
),
|
||||
"role": caller,
|
||||
})
|
||||
return content, usage, actual_model
|
||||
except (requests.RequestException, KeyError, json.JSONDecodeError) as error:
|
||||
except (requests.RequestException, KeyError, json.JSONDecodeError,
|
||||
UnicodeDecodeError, ValueError) as error:
|
||||
last_error = str(error)
|
||||
if attempt < retries:
|
||||
wait = 2 * (2 ** attempt)
|
||||
@ -333,6 +429,9 @@ def chat_fixed_opus(prompt, model="opus", system=None, max_tokens=FIXED_OPUS_MAX
|
||||
raise RuntimeError("固定 Opus 调用超过总 deadline") from error
|
||||
print(f"[llm] 固定 Opus 第{attempt + 1}次失败,{wait}s 后重试", file=sys.stderr)
|
||||
time.sleep(wait)
|
||||
finally:
|
||||
if response is not None:
|
||||
response.close()
|
||||
raise RuntimeError(f"固定 Opus 调用重试耗尽: {last_error}")
|
||||
|
||||
|
||||
|
||||
@ -1,7 +1,7 @@
|
||||
#!/usr/bin/env python3
|
||||
"""muse_role:provider-neutral 的受治理角色执行运行时。
|
||||
|
||||
角色执行走 07-Agent与Skill领域 §2 派发合同:角色文件全文作系统提示词、冻结输入、
|
||||
角色执行走 07-Agent与Skill领域 §2 派发合同:身份提示与中心角色合同作系统提示词、冻结输入、
|
||||
输出校验、证据落库。本模块是自动化管线侧的承载:系统提示词 + 冻结 JSON 输入经
|
||||
muse_llm 的固定 Opus HTTP 适配器或内容模型治理链调用,输出按冻结 schema 校验并
|
||||
构造回执。writer/planner/judge 的 profile 绑定完整 Opus 模型 ID 且禁止降级;只有明确
|
||||
@ -32,7 +32,7 @@ SUPPORTED_ROLES = frozenset(
|
||||
{"writer", "semantic_detector", "blind_judge", "planner", "extractor"}
|
||||
)
|
||||
RUNTIME_ADAPTER = "muse-role"
|
||||
RUNTIME_ADAPTER_VERSION = "muse-role-v3"
|
||||
RUNTIME_ADAPTER_VERSION = "muse-role-v4"
|
||||
FIXED_OPUS_POLICY_ALIAS = "opus"
|
||||
FIXED_OPUS_POLICY_VERSION = "fixed-opus-v1"
|
||||
FIXED_OPUS_MODEL_ID = muse_llm.FIXED_OPUS_DEFAULT_MODEL
|
||||
@ -93,7 +93,7 @@ def sha256_json(value: Any) -> str:
|
||||
|
||||
|
||||
def compose_role_system_prompt(role_prompt: str, task_prompt: str) -> str:
|
||||
"""把角色文件全文与单次功能合同组合成幂等的 system prompt。"""
|
||||
"""把已装配的角色身份/合同提示与单次功能合同组合成幂等 system prompt。"""
|
||||
|
||||
role = role_prompt.rstrip()
|
||||
task = task_prompt.strip()
|
||||
@ -103,15 +103,20 @@ def compose_role_system_prompt(role_prompt: str, task_prompt: str) -> str:
|
||||
return task if task.startswith(prefix) else prefix + task
|
||||
|
||||
|
||||
def format_schema_contract(json_schema: Mapping[str, Any]) -> str:
|
||||
"""把冻结 JSON Schema 渲染成注入 system prompt 的确定性结构化输出合同段。"""
|
||||
|
||||
return (
|
||||
"\n\n--- 结构化输出合同 ---\n"
|
||||
+ "只返回一个符合以下 JSON Schema 的 JSON 对象;不要输出 Markdown、解释或额外文本。\n"
|
||||
+ canonical_json(json_schema)
|
||||
)
|
||||
|
||||
|
||||
def build_dispatch_system_prompt(profile: "RoleExecutionProfile") -> str:
|
||||
"""把冻结角色 prompt 与 schema 组合成模型实际接收的确定性系统提示。"""
|
||||
|
||||
return (
|
||||
profile.system_prompt.rstrip()
|
||||
+ "\n\n--- 结构化输出合同 ---\n"
|
||||
+ "只返回一个符合以下 JSON Schema 的 JSON 对象;不要输出 Markdown、解释或额外文本。\n"
|
||||
+ canonical_json(profile.json_schema)
|
||||
)
|
||||
return profile.system_prompt.rstrip() + format_schema_contract(profile.json_schema)
|
||||
|
||||
|
||||
# WHY: 路径穿越判定必须是「路径分量级」的精确匹配,而不是 `".." in value` 子串匹配。
|
||||
@ -620,6 +625,7 @@ __all__ = [
|
||||
"canonical_json",
|
||||
"compose_role_system_prompt",
|
||||
"contains_path_traversal",
|
||||
"format_schema_contract",
|
||||
"model_matches_profile",
|
||||
"run_role",
|
||||
"sha256_json",
|
||||
|
||||
160
muse-llm/src/muse_role_contract.py
Normal file
160
muse-llm/src/muse_role_contract.py
Normal file
@ -0,0 +1,160 @@
|
||||
"""从单一角色合同文档加载角色派发合同。"""
|
||||
from __future__ import annotations
|
||||
|
||||
import hashlib
|
||||
import json
|
||||
import re
|
||||
from dataclasses import dataclass
|
||||
from pathlib import Path
|
||||
from typing import Any, Mapping
|
||||
|
||||
import yaml
|
||||
|
||||
ROLE_CONTRACT_VERSION = "role-contracts-v1"
|
||||
ROLE_CONTRACT_RELATIVE_PATH = Path(".agent") / "docs" / "architecture" / "角色合同.md"
|
||||
ROLE_NAMES = frozenset({"writer", "planner", "extractor", "detector", "judge"})
|
||||
_ROLE_SECTION = re.compile(
|
||||
r"<!--\s*role-contract:(?P<role>[a-z]+)\s*-->"
|
||||
r"(?P<body>.*?)"
|
||||
r"<!--\s*/role-contract:(?P=role)\s*-->",
|
||||
re.DOTALL,
|
||||
)
|
||||
|
||||
|
||||
class RoleContractError(ValueError):
|
||||
"""角色合同文档缺失、结构非法或与角色目录不一致。"""
|
||||
|
||||
|
||||
@dataclass(frozen=True)
|
||||
class RoleContract:
|
||||
name: str
|
||||
display_name: str
|
||||
prompt_file: str
|
||||
model_policy: str
|
||||
model_policy_version: str
|
||||
explicit_model_required: bool
|
||||
tool_policy: str
|
||||
contract_prompt: str
|
||||
contract_sha256: str
|
||||
|
||||
|
||||
@dataclass(frozen=True)
|
||||
class RoleContractCatalog:
|
||||
version: str
|
||||
source_path: str
|
||||
roles: Mapping[str, RoleContract]
|
||||
|
||||
def for_role(self, role: str) -> RoleContract:
|
||||
try:
|
||||
return self.roles[role]
|
||||
except KeyError as exc:
|
||||
raise RoleContractError(f"角色未登记在角色合同文档: {role}") from exc
|
||||
|
||||
|
||||
def _read_frontmatter(text: str, path: Path) -> tuple[Mapping[str, Any], str]:
|
||||
lines = text.splitlines(keepends=True)
|
||||
if not lines or lines[0].strip() != "---":
|
||||
raise RoleContractError(f"角色合同缺少 YAML frontmatter: {path}")
|
||||
end = next((i for i in range(1, len(lines)) if lines[i].strip() == "---"), None)
|
||||
if end is None:
|
||||
raise RoleContractError(f"角色合同 frontmatter 未闭合: {path}")
|
||||
try:
|
||||
data = yaml.safe_load("".join(lines[1:end])) or {}
|
||||
except yaml.YAMLError as exc:
|
||||
raise RoleContractError(f"角色合同 frontmatter 不是合法 YAML: {path}") from exc
|
||||
if not isinstance(data, Mapping):
|
||||
raise RoleContractError(f"角色合同 frontmatter 必须是对象: {path}")
|
||||
return data, "".join(lines[end + 1 :])
|
||||
|
||||
|
||||
def _contract_hash(role: str, raw: Mapping[str, Any], text: str) -> str:
|
||||
payload = json.dumps(
|
||||
{"role": role, "metadata": dict(raw), "contract": text},
|
||||
ensure_ascii=False,
|
||||
sort_keys=True,
|
||||
separators=(",", ":"),
|
||||
)
|
||||
return "sha256:" + hashlib.sha256(payload.encode("utf-8")).hexdigest()
|
||||
|
||||
|
||||
def load_role_contract_catalog(repo_root: str | Path) -> RoleContractCatalog:
|
||||
"""读取并严格校验唯一角色合同文档。"""
|
||||
|
||||
path = Path(repo_root) / ROLE_CONTRACT_RELATIVE_PATH
|
||||
try:
|
||||
text = path.read_text(encoding="utf-8")
|
||||
except (OSError, UnicodeError) as exc:
|
||||
raise RoleContractError(f"角色合同文档不可读: {path}") from exc
|
||||
frontmatter, body = _read_frontmatter(text, path)
|
||||
if frontmatter.get("schemaVersion") != ROLE_CONTRACT_VERSION:
|
||||
raise RoleContractError(
|
||||
f"角色合同版本必须是 {ROLE_CONTRACT_VERSION}: {frontmatter.get('schemaVersion')!r}"
|
||||
)
|
||||
raw_roles = frontmatter.get("roles")
|
||||
if not isinstance(raw_roles, Mapping) or set(raw_roles) != ROLE_NAMES:
|
||||
actual = sorted(raw_roles) if isinstance(raw_roles, Mapping) else raw_roles
|
||||
raise RoleContractError(f"角色合同必须精确登记 {sorted(ROLE_NAMES)},实际 {actual}")
|
||||
|
||||
sections: dict[str, str] = {}
|
||||
for match in _ROLE_SECTION.finditer(body):
|
||||
role = match.group("role")
|
||||
if role in sections:
|
||||
raise RoleContractError(f"角色合同章节重复: {role}")
|
||||
sections[role] = match.group("body").strip()
|
||||
if set(sections) != ROLE_NAMES:
|
||||
raise RoleContractError(
|
||||
f"角色合同正文必须精确包含 {sorted(ROLE_NAMES)},实际 {sorted(sections)}"
|
||||
)
|
||||
|
||||
roles: dict[str, RoleContract] = {}
|
||||
for role in sorted(ROLE_NAMES):
|
||||
raw = raw_roles[role]
|
||||
if not isinstance(raw, Mapping):
|
||||
raise RoleContractError(f"角色合同登记必须是对象: {role}")
|
||||
required = {
|
||||
"displayName",
|
||||
"promptFile",
|
||||
"modelPolicy",
|
||||
"modelPolicyVersion",
|
||||
"explicitModelRequired",
|
||||
"toolPolicy",
|
||||
}
|
||||
if set(raw) != required:
|
||||
raise RoleContractError(
|
||||
f"角色 {role} 的登记字段必须是 {sorted(required)},实际 {sorted(raw)}"
|
||||
)
|
||||
prompt_file = str(raw["promptFile"])
|
||||
if prompt_file != f".agent/agents/{role}.md":
|
||||
raise RoleContractError(f"角色 {role} 的 promptFile 不正确: {prompt_file}")
|
||||
contract_prompt = sections[role]
|
||||
if not contract_prompt:
|
||||
raise RoleContractError(f"角色合同正文为空: {role}")
|
||||
if not isinstance(raw["explicitModelRequired"], bool):
|
||||
raise RoleContractError(f"角色 {role} 的 explicitModelRequired 必须是布尔值")
|
||||
roles[role] = RoleContract(
|
||||
name=role,
|
||||
display_name=str(raw["displayName"]),
|
||||
prompt_file=prompt_file,
|
||||
model_policy=str(raw["modelPolicy"]),
|
||||
model_policy_version=str(raw["modelPolicyVersion"]),
|
||||
explicit_model_required=raw["explicitModelRequired"],
|
||||
tool_policy=str(raw["toolPolicy"]),
|
||||
contract_prompt=contract_prompt,
|
||||
contract_sha256=_contract_hash(role, raw, contract_prompt),
|
||||
)
|
||||
return RoleContractCatalog(
|
||||
version=ROLE_CONTRACT_VERSION,
|
||||
source_path=path.relative_to(Path(repo_root)).as_posix(),
|
||||
roles=roles,
|
||||
)
|
||||
|
||||
|
||||
__all__ = [
|
||||
"ROLE_CONTRACT_RELATIVE_PATH",
|
||||
"ROLE_CONTRACT_VERSION",
|
||||
"ROLE_NAMES",
|
||||
"RoleContract",
|
||||
"RoleContractCatalog",
|
||||
"RoleContractError",
|
||||
"load_role_contract_catalog",
|
||||
]
|
||||
@ -24,6 +24,11 @@ ACTIVE_RUNTIME_ROOTS = (
|
||||
ROOT / "muse-embed",
|
||||
ROOT / "docs" / "write-chapter",
|
||||
)
|
||||
# 框架派发接缝:唯一允许直接调用 Agent 框架二进制的适配器文件(dispatch-agent-task)。
|
||||
# 新增框架适配器(codex/opencode…)必须挂在本 Skill scripts/ 下并在此登记。
|
||||
FRAMEWORK_ADAPTER_ALLOWLIST = frozenset(
|
||||
{".agent/skills/dispatch-agent-task/scripts/pi_runner.py"}
|
||||
)
|
||||
FORBIDDEN_CLAUDE_RUNTIME_TOKENS = (
|
||||
"claude_runtime",
|
||||
"muse-claude-runtime",
|
||||
@ -103,15 +108,20 @@ class ImportBoundaryTest(unittest.TestCase):
|
||||
):
|
||||
continue
|
||||
text = path.read_text(encoding="utf-8")
|
||||
rel = path.relative_to(ROOT).as_posix()
|
||||
if any(token in text for token in FORBIDDEN_CLAUDE_RUNTIME_TOKENS):
|
||||
offenders.append(path.relative_to(ROOT).as_posix())
|
||||
offenders.append(rel)
|
||||
continue
|
||||
# 框架适配器是被批准的唯一直接调用 Agent 框架二进制的位置(07 领域框架派发接缝);
|
||||
# 其余任何位置 shell 调模型/框架 CLI 仍被阻断。
|
||||
if rel in FRAMEWORK_ADAPTER_ALLOWLIST:
|
||||
continue
|
||||
if re.search(
|
||||
r"(?:/bin/claude|[\"']claude[\"']\s*,\s*[\"']--version|"
|
||||
r"[\"']pi[\"']\s*,|\bpi\s+--model)",
|
||||
text,
|
||||
):
|
||||
offenders.append(path.relative_to(ROOT).as_posix())
|
||||
offenders.append(rel)
|
||||
self.assertEqual(
|
||||
sorted(set(offenders)),
|
||||
[],
|
||||
|
||||
@ -10,6 +10,7 @@ PROJECT_ROOT = pathlib.Path(__file__).resolve().parents[3]
|
||||
SCRIPTS_DIR = PROJECT_ROOT / ".agent" / "skills" / "access-database" / "scripts"
|
||||
sys.path.insert(0, str(SCRIPTS_DIR))
|
||||
|
||||
from muse_role_contract import load_role_contract_catalog
|
||||
from sync_agent_registry import validate_role_catalog, validate_skill_catalog
|
||||
|
||||
|
||||
@ -22,13 +23,22 @@ class SkillCatalogTest(unittest.TestCase):
|
||||
)
|
||||
self.assertTrue(all("model" not in fm for _, fm in entries))
|
||||
|
||||
def test_role_contract_is_central_and_complete(self):
|
||||
catalog = load_role_contract_catalog(PROJECT_ROOT)
|
||||
self.assertEqual(
|
||||
set(catalog.roles),
|
||||
{"writer", "planner", "extractor", "detector", "judge"},
|
||||
)
|
||||
self.assertTrue(all(role.contract_prompt for role in catalog.roles.values()))
|
||||
self.assertTrue(all(role.explicit_model_required for role in catalog.roles.values()))
|
||||
|
||||
def test_role_frontmatter_rejects_static_model_binding(self):
|
||||
with tempfile.TemporaryDirectory() as tmp:
|
||||
root = pathlib.Path(tmp)
|
||||
for role in ("writer", "planner", "extractor", "detector", "judge"):
|
||||
extra = "model: opus\n" if role == "writer" else ""
|
||||
(root / f"{role}.md").write_text(
|
||||
f"---\nname: {role}\ndescription: test\n{extra}---\n",
|
||||
f"---\nname: {role}\ndescription: test\nskills: test-skill\ntools: read\n{extra}---\n",
|
||||
encoding="utf-8",
|
||||
)
|
||||
with self.assertRaisesRegex(ValueError, "未登记字段"):
|
||||
|
||||
@ -43,7 +43,7 @@ class NextStepsOfflineTest(unittest.TestCase):
|
||||
self.assertTrue(all(s.get("auto") is False for s in steps))
|
||||
|
||||
def test_generation_entry_cannot_call_accept(self):
|
||||
source = (ROOT / "docs" / "write-chapter" / "step2_write_chapter.py").read_text(
|
||||
source = (ROOT / ".agent" / "skills" / "write-next-chapter" / "scripts" / "produce_next_chapter.py").read_text(
|
||||
encoding="utf-8"
|
||||
)
|
||||
tree = ast.parse(source)
|
||||
|
||||
552
tests/skills/dispatch-agent-task/test_dispatch_agent_task.py
Normal file
552
tests/skills/dispatch-agent-task/test_dispatch_agent_task.py
Normal file
@ -0,0 +1,552 @@
|
||||
#!/usr/bin/env python3
|
||||
"""dispatch-agent-task 的离线测试(不连库、不连网、不启动真 pi)。
|
||||
|
||||
用假 pi 事件流(与 pi --mode json 真实线格式一致)与假数据库连接固定:
|
||||
任务包合同、argv 构造、事件归一、run_dispatch 全链失败关闭与回执形状。
|
||||
"""
|
||||
from __future__ import annotations
|
||||
|
||||
import json
|
||||
import pathlib
|
||||
import stat
|
||||
import sys
|
||||
import tempfile
|
||||
import unittest
|
||||
|
||||
|
||||
PROJECT_ROOT = pathlib.Path(__file__).resolve().parents[3]
|
||||
SKILL_DIR = PROJECT_ROOT / ".agent" / "skills" / "dispatch-agent-task" / "scripts"
|
||||
EVIDENCE_DIR = PROJECT_ROOT / ".agent" / "skills" / "record-run-evidence" / "scripts"
|
||||
for path in (SKILL_DIR, EVIDENCE_DIR):
|
||||
if str(path) not in sys.path:
|
||||
sys.path.insert(0, str(path))
|
||||
|
||||
import agent_task # noqa: E402
|
||||
from agent_task import ( # noqa: E402
|
||||
OutputInvalidError,
|
||||
TaskSpecError,
|
||||
build_task_package,
|
||||
load_spec,
|
||||
validate_structured_output,
|
||||
)
|
||||
from pi_runner import ( # noqa: E402
|
||||
ExecutionPolicy,
|
||||
FrameworkError,
|
||||
PiAgentRunner,
|
||||
build_pi_argv,
|
||||
)
|
||||
from dispatch_agent_task import EXIT_OUTPUT_INVALID, run_dispatch # noqa: E402
|
||||
from agent_trace import AgentTraceWriter # noqa: E402
|
||||
|
||||
REPO_ROOT = PROJECT_ROOT
|
||||
|
||||
OUTPUT_SCHEMA = {
|
||||
"$schema": "https://json-schema.org/draft/2020-12/schema",
|
||||
"type": "object",
|
||||
"properties": {
|
||||
"title": {"type": "string", "minLength": 1},
|
||||
"beats": {"type": "array", "minItems": 1, "items": {"type": "string"}},
|
||||
},
|
||||
"required": ["title", "beats"],
|
||||
"additionalProperties": False,
|
||||
}
|
||||
|
||||
|
||||
def make_spec(tmp: pathlib.Path, **overrides) -> pathlib.Path:
|
||||
spec = {
|
||||
"specVersion": "agent-task-v1",
|
||||
"role": "planner",
|
||||
"taskPrompt": "为一幕场景给出三拍结构",
|
||||
"input": {"premise": "深空站失联前的最后八小时"},
|
||||
"outputSchema": OUTPUT_SCHEMA,
|
||||
"outputSchemaId": "planner-mini-outline-v1",
|
||||
"toolAllowlist": [],
|
||||
"maxDurationSeconds": 120,
|
||||
}
|
||||
spec.update(overrides)
|
||||
path = tmp / "task.json"
|
||||
path.write_text(json.dumps(spec, ensure_ascii=False), encoding="utf-8")
|
||||
return path
|
||||
|
||||
|
||||
def pi_stream_lines(final_text: str, *, with_tool: bool = False, model: str = "m-a"):
|
||||
"""构造与 pi --mode json 一致的假事件流(含 usage/model/stopReason)。"""
|
||||
assistant_final = {
|
||||
"role": "assistant",
|
||||
"content": [{"type": "text", "text": final_text}],
|
||||
"model": model,
|
||||
"provider": "prov",
|
||||
"usage": {"input": 100, "output": 40, "cacheRead": 10, "cost": {"total": 0.012}},
|
||||
"stopReason": "stop",
|
||||
}
|
||||
lines = [
|
||||
{"type": "session", "version": 3, "id": "sess-1", "cwd": "/tmp"},
|
||||
{"type": "agent_start"},
|
||||
{"type": "turn_start"},
|
||||
{"type": "message_end", "message": {"role": "user", "content": []}},
|
||||
]
|
||||
if with_tool:
|
||||
assistant_tool = {
|
||||
"role": "assistant",
|
||||
"content": [{"type": "tool_call", "id": "t1", "name": "read", "arguments": {}}],
|
||||
"model": model,
|
||||
"provider": "prov",
|
||||
"usage": {"input": 90, "output": 5, "cost": {"total": 0.001}},
|
||||
"stopReason": "tool_use",
|
||||
}
|
||||
lines += [
|
||||
{"type": "message_end", "message": assistant_tool},
|
||||
{"type": "tool_execution_start", "toolCallId": "t1", "toolName": "read", "args": {}},
|
||||
{"type": "tool_execution_end", "toolCallId": "t1", "toolName": "read", "result": "ok", "isError": False},
|
||||
{"type": "turn_start"},
|
||||
]
|
||||
lines += [
|
||||
{"type": "message_end", "message": assistant_final},
|
||||
{"type": "turn_end", "message": assistant_final, "toolResults": []},
|
||||
{"type": "agent_end", "messages": [{"role": "user", "content": []}, assistant_final]},
|
||||
{"type": "agent_settled"},
|
||||
]
|
||||
return [json.dumps(line, ensure_ascii=False).encode("utf-8") + b"\n" for line in lines]
|
||||
|
||||
|
||||
class FakeStream:
|
||||
def __init__(self, lines, exit_code=0, timed_out=False):
|
||||
self._lines = lines
|
||||
self.exit_code = exit_code
|
||||
self.timed_out = timed_out
|
||||
|
||||
def __iter__(self):
|
||||
yield from self._lines
|
||||
|
||||
def close(self):
|
||||
return None
|
||||
|
||||
|
||||
def fake_launcher(lines, exit_code=0, timed_out=False):
|
||||
def _launch(argv, timeout, cwd):
|
||||
assert argv[0] == "pi", argv
|
||||
return FakeStream(list(lines), exit_code=exit_code, timed_out=timed_out)
|
||||
|
||||
return _launch
|
||||
|
||||
|
||||
class RecordingConnect:
|
||||
"""捕获全部写入语句的假连接工厂(事件 + 证据 + run 登记)。"""
|
||||
|
||||
def __init__(self):
|
||||
self.log = []
|
||||
|
||||
def __call__(self, *args, **kwargs):
|
||||
outer = self
|
||||
|
||||
class _Cursor:
|
||||
def execute(self, sql, params=None):
|
||||
outer.log.append((sql, params))
|
||||
return self
|
||||
|
||||
def fetchone(self):
|
||||
sql = outer.log[-1][0]
|
||||
if sql.startswith("SELECT run_id, work_id"):
|
||||
# start_run 回读:返回与入参一致的绑定行。
|
||||
params = outer.log[-1][1]
|
||||
return ("row", None, None, "running")
|
||||
if sql.startswith("INSERT INTO example_run"):
|
||||
return ("row",)
|
||||
if sql.startswith("UPDATE example_run"):
|
||||
return ("row", "failed" if "failed" in (params := outer.log[-1][1]) else "completed", None)
|
||||
RecordingConnect.next_id += 1
|
||||
return (RecordingConnect.next_id,)
|
||||
|
||||
def fetchall(self):
|
||||
return []
|
||||
|
||||
def commit(self):
|
||||
outer.log.append(("COMMIT", None))
|
||||
|
||||
def rollback(self):
|
||||
outer.log.append(("ROLLBACK", None))
|
||||
|
||||
class _Ctx:
|
||||
def __enter__(self):
|
||||
return _Cursor()
|
||||
|
||||
def __exit__(self, *exc):
|
||||
return False
|
||||
|
||||
return _Ctx()
|
||||
|
||||
next_id = 5000
|
||||
|
||||
def event_rows(self):
|
||||
return [p for sql, p in self.log if sql.startswith("INSERT INTO example_agent_event")]
|
||||
|
||||
def run_rows(self):
|
||||
return [p for sql, p in self.log if sql.startswith("INSERT INTO example_run")]
|
||||
|
||||
|
||||
class TaskSpecTest(unittest.TestCase):
|
||||
def setUp(self):
|
||||
self.tmp = pathlib.Path(tempfile.mkdtemp())
|
||||
|
||||
def test_load_and_hash_binding(self):
|
||||
path = make_spec(self.tmp)
|
||||
spec = load_spec(path)
|
||||
self.assertEqual(spec.role, "planner")
|
||||
# 写错 input 哈希必须失败关闭。
|
||||
bad = json.loads(path.read_text())
|
||||
bad["inputSha256"] = "sha256:" + "0" * 64
|
||||
path2 = self.tmp / "bad.json"
|
||||
path2.write_text(json.dumps(bad), encoding="utf-8")
|
||||
with self.assertRaises(TaskSpecError):
|
||||
load_spec(path2)
|
||||
|
||||
def test_accepts_role_file_names_and_rejects_unknown_role(self):
|
||||
for role in ("writer", "planner", "detector", "judge", "extractor"):
|
||||
with self.subTest(role=role):
|
||||
spec = load_spec(make_spec(self.tmp, role=role))
|
||||
self.assertEqual(build_task_package(spec, REPO_ROOT).spec.role, role)
|
||||
with self.assertRaises(TaskSpecError):
|
||||
load_spec(make_spec(self.tmp, role="semantic_detector"))
|
||||
with self.assertRaises(TaskSpecError):
|
||||
load_spec(make_spec(self.tmp, role="hacker"))
|
||||
|
||||
def test_rejects_bad_schema_and_empty_prompt(self):
|
||||
with self.assertRaises(TaskSpecError):
|
||||
load_spec(make_spec(self.tmp, outputSchema={"type": "no-such-type"}))
|
||||
with self.assertRaises(TaskSpecError):
|
||||
load_spec(make_spec(self.tmp, taskPrompt=" "))
|
||||
|
||||
def test_rejects_nonportable_or_malformed_fields(self):
|
||||
invalid_overrides = (
|
||||
{"provider": "p"},
|
||||
{"toolAllowlist": "read"},
|
||||
{"toolAllowlist": ["read tool"]},
|
||||
{"toolAllowlist": ["read", "read"]},
|
||||
{"maxDurationSeconds": 0},
|
||||
{"maxDurationSeconds": float("nan")},
|
||||
{"workId": "1"},
|
||||
{"targetChapter": True},
|
||||
)
|
||||
for overrides in invalid_overrides:
|
||||
with self.subTest(overrides=overrides), self.assertRaises(TaskSpecError):
|
||||
load_spec(make_spec(self.tmp, **overrides))
|
||||
malformed = self.tmp / "malformed.json"
|
||||
malformed.write_text("{", encoding="utf-8")
|
||||
with self.assertRaises(TaskSpecError):
|
||||
load_spec(malformed)
|
||||
nonstandard = self.tmp / "nonstandard.json"
|
||||
nonstandard.write_text(json.dumps({"specVersion": float("nan")}), encoding="utf-8")
|
||||
with self.assertRaises(TaskSpecError):
|
||||
load_spec(nonstandard)
|
||||
|
||||
def test_spec_hash_binds_scope_and_deadline(self):
|
||||
base = build_task_package(load_spec(make_spec(self.tmp)), REPO_ROOT)
|
||||
scoped = build_task_package(
|
||||
load_spec(make_spec(self.tmp, workId=7, targetChapter=3)), REPO_ROOT
|
||||
)
|
||||
slower = build_task_package(
|
||||
load_spec(make_spec(self.tmp, maxDurationSeconds=121)), REPO_ROOT
|
||||
)
|
||||
self.assertNotEqual(base.spec_sha256, scoped.spec_sha256)
|
||||
self.assertNotEqual(base.spec_sha256, slower.spec_sha256)
|
||||
|
||||
|
||||
class PackageTest(unittest.TestCase):
|
||||
def setUp(self):
|
||||
self.tmp = pathlib.Path(tempfile.mkdtemp())
|
||||
|
||||
def test_system_prompt_is_identity_plus_role_contract_plus_schema(self):
|
||||
package = build_task_package(load_spec(make_spec(self.tmp)), REPO_ROOT)
|
||||
role_text = (REPO_ROOT / ".agent" / "agents" / "planner.md").read_text(encoding="utf-8")
|
||||
self.assertTrue(package.system_prompt.startswith(role_text.rstrip()))
|
||||
self.assertIn("--- 角色合同(唯一事实源) ---", package.system_prompt)
|
||||
self.assertIn("## planner:规划师", package.system_prompt)
|
||||
self.assertIn("--- 结构化输出合同 ---", package.system_prompt)
|
||||
self.assertIn('"$schema"', package.system_prompt)
|
||||
self.assertIn("--- 冻结输入 ---", package.user_message)
|
||||
self.assertIn("深空站失联前的最后八小时", package.user_message)
|
||||
identity = package.as_identity()
|
||||
self.assertEqual(identity["toolAllowlist"], [])
|
||||
self.assertEqual(identity["roleContractVersion"], "role-contracts-v1")
|
||||
self.assertEqual(identity["roleContractSource"], ".agent/docs/architecture/角色合同.md")
|
||||
self.assertTrue(all(len(v) == 71 and v.startswith("sha256:") for k, v in identity.items() if k.endswith("Sha256")))
|
||||
|
||||
|
||||
class ExecutionPolicyTest(unittest.TestCase):
|
||||
def test_provider_and_model_are_explicit(self):
|
||||
with self.assertRaisesRegex(ValueError, "provider"):
|
||||
ExecutionPolicy(model="m")
|
||||
with self.assertRaisesRegex(ValueError, "model"):
|
||||
ExecutionPolicy(provider="p")
|
||||
self.assertEqual(
|
||||
ExecutionPolicy(provider="p", model="m").requested_model_id,
|
||||
"p/m",
|
||||
)
|
||||
|
||||
|
||||
class ArgvTest(unittest.TestCase):
|
||||
def setUp(self):
|
||||
self.tmp = pathlib.Path(tempfile.mkdtemp())
|
||||
|
||||
def test_argv_injects_prompt_tools_and_isolation(self):
|
||||
package = build_task_package(load_spec(make_spec(self.tmp)), REPO_ROOT)
|
||||
argv = build_pi_argv(package, ExecutionPolicy(provider="p", model="m", thinking="low"))
|
||||
joined = " ".join(argv)
|
||||
self.assertIn("--system-prompt", argv)
|
||||
self.assertEqual(argv[argv.index("--system-prompt") + 1], package.system_prompt)
|
||||
self.assertEqual(argv[-1], package.user_message)
|
||||
self.assertIn("--no-context-files", joined)
|
||||
self.assertIn("--no-skills", joined)
|
||||
self.assertIn("--no-extensions", joined)
|
||||
self.assertIn("--mode", joined)
|
||||
self.assertEqual(argv[argv.index("--provider") + 1], "p")
|
||||
self.assertEqual(argv[argv.index("--model") + 1], "m")
|
||||
self.assertIn("--no-tools", joined)
|
||||
|
||||
def test_tool_allowlist_maps_to_tools_flag(self):
|
||||
spec_path = make_spec(pathlib.Path(tempfile.mkdtemp()), toolAllowlist=["read", "bash"])
|
||||
package = build_task_package(load_spec(spec_path), REPO_ROOT)
|
||||
argv = build_pi_argv(package, ExecutionPolicy(provider="p", model="m"))
|
||||
self.assertIn("--tools", argv)
|
||||
self.assertEqual(argv[argv.index("--tools") + 1], "read,bash")
|
||||
|
||||
|
||||
class RunnerParseTest(unittest.TestCase):
|
||||
def setUp(self):
|
||||
self.tmp = pathlib.Path(tempfile.mkdtemp())
|
||||
|
||||
def _sink(self):
|
||||
events = []
|
||||
|
||||
class Sink:
|
||||
emit = AgentTraceWriter(run_id="t", framework="pi", agent_role="planner", connect=object())
|
||||
|
||||
# 直接构造一个记录型 sink,绕开数据库。
|
||||
class RecordingSink(AgentTraceWriter):
|
||||
def __init__(self):
|
||||
super().__init__(
|
||||
run_id="t", framework="pi", agent_role="planner",
|
||||
connect=lambda: (_ for _ in ()).throw(AssertionError("离线测试不应触库")),
|
||||
)
|
||||
self.rows = []
|
||||
|
||||
def emit(self, event_type, **kwargs):
|
||||
self.rows.append((event_type, kwargs))
|
||||
return len(self.rows)
|
||||
|
||||
return RecordingSink()
|
||||
|
||||
def test_stream_parses_model_tools_and_final_text(self):
|
||||
package = build_task_package(
|
||||
load_spec(make_spec(self.tmp, toolAllowlist=["read"])), REPO_ROOT
|
||||
)
|
||||
sink = self._sink()
|
||||
outcome = PiAgentRunner(launcher=fake_launcher(pi_stream_lines('{"title":"t","beats":["a"]}', with_tool=True))).run(
|
||||
package, ExecutionPolicy(provider="p", model="m"), sink, timeout_seconds=10
|
||||
)
|
||||
self.assertEqual(outcome.final_text, '{"title":"t","beats":["a"]}')
|
||||
self.assertEqual(len(outcome.model_calls), 2)
|
||||
self.assertEqual(outcome.model_calls[-1].actual_model_id, "prov/m-a")
|
||||
self.assertEqual(outcome.model_calls[-1].cost_usd, 0.012)
|
||||
self.assertEqual(outcome.tool_calls[0].name, "read")
|
||||
self.assertFalse(outcome.tool_calls[0].is_error)
|
||||
types = [row[0] for row in sink.rows]
|
||||
self.assertEqual(types, ["agent.started", "model.completed", "tool.started", "tool.completed", "model.completed", "agent.completed"])
|
||||
|
||||
def test_framework_failures_raise_with_stable_codes(self):
|
||||
package = build_task_package(load_spec(make_spec(self.tmp)), REPO_ROOT)
|
||||
for launcher, code in (
|
||||
(fake_launcher(pi_stream_lines("x"), exit_code=1), "FRAMEWORK_EXIT_NONZERO"),
|
||||
(fake_launcher(pi_stream_lines("x"), timed_out=True), "FRAMEWORK_TIMEOUT"),
|
||||
(fake_launcher([b"not json\n"]), "STREAM_PARSE_ERROR"),
|
||||
(fake_launcher([json.dumps({"type": "agent_end", "messages": []}).encode() + b"\n"]), "NO_MODEL_RESPONSE"),
|
||||
(fake_launcher(pi_stream_lines("x", with_tool=True)), "TOOL_NOT_ALLOWED"),
|
||||
(
|
||||
fake_launcher(
|
||||
pi_stream_lines("x")[:-4]
|
||||
+ [
|
||||
json.dumps(
|
||||
{
|
||||
"type": "message_end",
|
||||
"message": {
|
||||
"role": "assistant",
|
||||
"content": [],
|
||||
"model": "m-a",
|
||||
"provider": "prov",
|
||||
"usage": {},
|
||||
"stopReason": "error",
|
||||
"errorMessage": "provider failed",
|
||||
},
|
||||
}
|
||||
).encode()
|
||||
+ b"\n"
|
||||
]
|
||||
),
|
||||
"MODEL_TURN_FAILED",
|
||||
),
|
||||
):
|
||||
with self.assertRaises(FrameworkError) as ctx:
|
||||
PiAgentRunner(launcher=launcher).run(
|
||||
package,
|
||||
ExecutionPolicy(provider="p", model="m"),
|
||||
self._sink(),
|
||||
timeout_seconds=5,
|
||||
)
|
||||
self.assertEqual(ctx.exception.error_code, code)
|
||||
|
||||
|
||||
class RunDispatchTest(unittest.TestCase):
|
||||
def setUp(self):
|
||||
self.tmp = pathlib.Path(tempfile.mkdtemp())
|
||||
|
||||
def _dispatch(self, launcher, connect):
|
||||
return run_dispatch(
|
||||
make_spec(self.tmp),
|
||||
repo_root=REPO_ROOT,
|
||||
policy=ExecutionPolicy(provider="p", model="claude-opus-test"),
|
||||
run_id="unittest-agent-dispatch-1",
|
||||
run_dir=self.tmp / "run",
|
||||
connect_factory=connect,
|
||||
launcher=launcher,
|
||||
trigger_source="diagnostic",
|
||||
)
|
||||
|
||||
def test_success_path_events_and_receipt(self):
|
||||
connect = RecordingConnect()
|
||||
receipt, code = self._dispatch(
|
||||
fake_launcher(pi_stream_lines('{"title":"重启","beats":["警报","分歧","决断"]}')), connect
|
||||
)
|
||||
self.assertEqual(code, 0)
|
||||
self.assertEqual(receipt["status"], "completed")
|
||||
self.assertEqual(receipt["requestedModelId"], "p/claude-opus-test")
|
||||
self.assertEqual(receipt["usage"], {"inputTokens": 110, "outputTokens": 40, "cachedTokens": 10, "reasoningTokens": 0})
|
||||
self.assertEqual(receipt["totalCostUsd"], 0.012)
|
||||
self.assertTrue(receipt["costComplete"])
|
||||
self.assertEqual(
|
||||
[p[2] for p in connect.event_rows()],
|
||||
["run.started", "agent.started", "model.completed", "agent.completed", "run.completed"],
|
||||
)
|
||||
self.assertEqual(receipt["evidence"]["status"], "written")
|
||||
llm_calls = [p for sql, p in connect.log if sql.startswith("INSERT INTO example_llm_call")]
|
||||
self.assertIsNone(llm_calls[0][0]) # window_key 只属于 16 字符额度窗,run_id 走专列。
|
||||
self.assertEqual(llm_calls[0][1], "unittest-agent-dispatch-1")
|
||||
# 运行目录审计件齐全且只对当前用户开放。
|
||||
run_dir = self.tmp / "run"
|
||||
self.assertEqual(stat.S_IMODE(run_dir.stat().st_mode), 0o700)
|
||||
for name in ("task-spec.json", "system-prompt.txt", "user-message.txt", "transcript.jsonl", "output.json", "receipt.json"):
|
||||
path = run_dir / name
|
||||
self.assertTrue(path.exists(), name)
|
||||
self.assertEqual(stat.S_IMODE(path.stat().st_mode), 0o600, name)
|
||||
|
||||
def test_schema_violation_fails_closed(self):
|
||||
connect = RecordingConnect()
|
||||
receipt, code = self._dispatch(
|
||||
fake_launcher(pi_stream_lines('{"title":"重启"}')), connect # 缺 beats
|
||||
)
|
||||
self.assertEqual(code, EXIT_OUTPUT_INVALID)
|
||||
self.assertEqual(receipt["errorCode"], "OUTPUT_SCHEMA_INVALID")
|
||||
self.assertEqual(receipt["evidence"]["status"], "written")
|
||||
self.assertEqual(
|
||||
[p[2] for p in connect.event_rows()],
|
||||
["run.started", "agent.started", "model.completed", "agent.completed", "run.failed"],
|
||||
)
|
||||
self.assertTrue(any(sql.startswith("UPDATE example_run") for sql, _ in connect.log))
|
||||
|
||||
def test_secret_in_transcript_fails_without_leaking_receipt(self):
|
||||
connect = RecordingConnect()
|
||||
secret = "sk-abcdef0123456789abcdef012345"
|
||||
receipt, code = self._dispatch(
|
||||
fake_launcher(pi_stream_lines(json.dumps({"title": secret, "beats": ["x"]}))),
|
||||
connect,
|
||||
)
|
||||
self.assertEqual(code, 5)
|
||||
self.assertEqual(receipt["errorCode"], "RAW_SECRET_DETECTED")
|
||||
self.assertNotIn(secret, json.dumps(receipt, ensure_ascii=False))
|
||||
self.assertFalse((self.tmp / "run" / "transcript.jsonl").exists())
|
||||
|
||||
def test_framework_runs_from_repo_root(self):
|
||||
seen = {}
|
||||
|
||||
def launcher(argv, timeout, cwd):
|
||||
seen["cwd"] = cwd
|
||||
return FakeStream(pi_stream_lines('{"title":"t","beats":["b"]}'))
|
||||
|
||||
receipt, code = self._dispatch(launcher, RecordingConnect())
|
||||
self.assertEqual(code, 0)
|
||||
self.assertEqual(receipt["status"], "completed")
|
||||
self.assertEqual(seen["cwd"], str(REPO_ROOT.resolve()))
|
||||
|
||||
def test_framework_failure_marks_run_failed_and_persists_trace(self):
|
||||
connect = RecordingConnect()
|
||||
receipt, code = self._dispatch(fake_launcher([b"garbage\n"]), connect)
|
||||
self.assertEqual(code, 3)
|
||||
self.assertEqual(receipt["errorCode"], "STREAM_PARSE_ERROR")
|
||||
self.assertEqual(receipt["evidence"]["status"], "written")
|
||||
self.assertEqual(receipt["evidence"]["llmCallIds"], [])
|
||||
|
||||
def test_role_model_policy_is_checked_before_run_start(self):
|
||||
connect = RecordingConnect()
|
||||
receipt, code = run_dispatch(
|
||||
make_spec(self.tmp),
|
||||
repo_root=REPO_ROOT,
|
||||
policy=ExecutionPolicy(provider="p", model="gpt-5.6-sol"),
|
||||
run_id="role-policy-test",
|
||||
run_dir=self.tmp / "role-policy-run",
|
||||
connect_factory=connect,
|
||||
launcher=fake_launcher(pi_stream_lines('{"title":"t","beats":["b"]}')),
|
||||
)
|
||||
self.assertEqual(code, 2)
|
||||
self.assertEqual(receipt["errorCode"], "ROLE_MODEL_POLICY_MISMATCH")
|
||||
self.assertFalse(any(sql.startswith("INSERT INTO example_run") for sql, _ in connect.log))
|
||||
|
||||
def test_trigger_detail_secret_is_rejected_before_run_start(self):
|
||||
connect = RecordingConnect()
|
||||
receipt, code = run_dispatch(
|
||||
make_spec(self.tmp),
|
||||
repo_root=REPO_ROOT,
|
||||
policy=ExecutionPolicy(provider="p", model="claude-opus-test"),
|
||||
run_id="safe-trigger-test",
|
||||
run_dir=self.tmp / "trigger-run",
|
||||
connect_factory=connect,
|
||||
launcher=fake_launcher(pi_stream_lines('{"title":"t","beats":["b"]}')),
|
||||
trigger_detail={"api_key": "sk-abcdef0123456789abcdef012345"},
|
||||
)
|
||||
self.assertEqual(code, 2)
|
||||
self.assertEqual(receipt["errorCode"], "TRIGGER_DETAIL_INVALID")
|
||||
self.assertFalse(any(sql.startswith("INSERT INTO example_run") for sql, _ in connect.log))
|
||||
|
||||
def test_run_id_cannot_escape_audit_root(self):
|
||||
receipt, code = run_dispatch(
|
||||
make_spec(self.tmp),
|
||||
repo_root=REPO_ROOT,
|
||||
policy=ExecutionPolicy(provider="p", model="claude-opus-test"),
|
||||
run_id="../escape",
|
||||
run_dir=self.tmp / "should-not-exist",
|
||||
connect_factory=RecordingConnect(),
|
||||
launcher=fake_launcher(pi_stream_lines('{"title":"t","beats":["b"]}')),
|
||||
)
|
||||
self.assertEqual(code, 2)
|
||||
self.assertEqual(receipt["errorCode"], "RUN_ID_INVALID")
|
||||
self.assertFalse((self.tmp / "should-not-exist").exists())
|
||||
|
||||
|
||||
class ValidateOutputTest(unittest.TestCase):
|
||||
def test_extract_and_validate(self):
|
||||
from agent_task import AgentTaskSpec
|
||||
|
||||
spec = AgentTaskSpec(
|
||||
role="planner",
|
||||
task_prompt="p",
|
||||
input={},
|
||||
output_schema=OUTPUT_SCHEMA,
|
||||
output_schema_id="s1",
|
||||
tool_allowlist=(),
|
||||
)
|
||||
out = validate_structured_output('前言```json\n{"title":"t","beats":["b"]}\n```', spec)
|
||||
self.assertEqual(out["title"], "t")
|
||||
with self.assertRaises(OutputInvalidError):
|
||||
validate_structured_output('{"title":"t"}', spec)
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
unittest.main(verbosity=2)
|
||||
@ -3,6 +3,7 @@
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import json
|
||||
import os
|
||||
import unittest
|
||||
from decimal import Decimal
|
||||
@ -99,17 +100,65 @@ def _fake_chat(payload, *, actual_model=muse_llm.FIXED_OPUS_CANONICAL_MODEL):
|
||||
return chat
|
||||
|
||||
|
||||
class _FakeStreamResponse:
|
||||
def __init__(self, events, *, status_code=200, text=""):
|
||||
self.status_code = status_code
|
||||
self.text = text
|
||||
self._events = events
|
||||
self.iter_lines_kwargs = None
|
||||
self.closed = False
|
||||
|
||||
def raise_for_status(self):
|
||||
if self.status_code >= 400:
|
||||
raise muse_llm.requests.HTTPError(f"HTTP {self.status_code}")
|
||||
|
||||
def iter_lines(self, **kwargs):
|
||||
self.iter_lines_kwargs = kwargs
|
||||
for event in self._events:
|
||||
if event == "[DONE]":
|
||||
yield b"data: [DONE]"
|
||||
else:
|
||||
yield b"data: " + json.dumps(
|
||||
event, ensure_ascii=False, separators=(",", ":")
|
||||
).encode("utf-8")
|
||||
yield b""
|
||||
|
||||
def close(self):
|
||||
self.closed = True
|
||||
|
||||
|
||||
class FixedOpusTransportTests(unittest.TestCase):
|
||||
def test_fixed_opus_uses_env_http_without_model_fallback(self):
|
||||
response = mock.Mock()
|
||||
response.status_code = 200
|
||||
response.json.return_value = {
|
||||
"model": muse_llm.FIXED_OPUS_CANONICAL_MODEL,
|
||||
"content": [{"type": "text", "text": '{"candidateBody":"正文"}'}],
|
||||
"usage": {"input_tokens": 10, "output_tokens": 5},
|
||||
"stop_reason": "end_turn",
|
||||
}
|
||||
response.raise_for_status.return_value = None
|
||||
def test_fixed_opus_streams_utf8_without_splitting_c1_bytes(self):
|
||||
response = _FakeStreamResponse([
|
||||
{
|
||||
"type": "message_start",
|
||||
"message": {
|
||||
"model": muse_llm.FIXED_OPUS_CANONICAL_MODEL,
|
||||
"usage": {"input_tokens": 10, "output_tokens": 0},
|
||||
},
|
||||
},
|
||||
{
|
||||
"type": "content_block_start",
|
||||
"index": 0,
|
||||
"content_block": {"type": "text", "text": ""},
|
||||
},
|
||||
{
|
||||
"type": "content_block_delta",
|
||||
"index": 0,
|
||||
"delta": {
|
||||
"type": "text_delta",
|
||||
"text": '{"candidateBody":"内文"}',
|
||||
},
|
||||
},
|
||||
{"type": "content_block_stop", "index": 0},
|
||||
{
|
||||
"type": "message_delta",
|
||||
"delta": {"stop_reason": "end_turn"},
|
||||
"usage": {"output_tokens": 5},
|
||||
},
|
||||
{"type": "message_stop"},
|
||||
"[DONE]",
|
||||
])
|
||||
session = mock.Mock()
|
||||
session.post.return_value = response
|
||||
events = []
|
||||
@ -137,13 +186,53 @@ class FixedOpusTransportTests(unittest.TestCase):
|
||||
self.assertEqual(request.args[0], "https://role.example/v1/messages")
|
||||
self.assertEqual(request.kwargs["json"]["model"], FIXED_OPUS_MODEL_ID)
|
||||
self.assertEqual(request.kwargs["json"]["max_tokens"], 123)
|
||||
self.assertEqual(content, '{"candidateBody":"正文"}')
|
||||
self.assertIs(request.kwargs["json"]["stream"], True)
|
||||
self.assertIs(request.kwargs["stream"], True)
|
||||
self.assertEqual(
|
||||
response.iter_lines_kwargs,
|
||||
{"decode_unicode": False, "delimiter": b"\n"},
|
||||
)
|
||||
self.assertTrue(response.closed)
|
||||
self.assertEqual(content, '{"candidateBody":"内文"}')
|
||||
self.assertEqual(usage["output_tokens"], 5)
|
||||
self.assertEqual(actual, muse_llm.FIXED_OPUS_CANONICAL_MODEL)
|
||||
self.assertEqual(events[0]["requested_model_id"], FIXED_OPUS_POLICY_ALIAS)
|
||||
self.assertEqual(events[0]["actual_model_id"], muse_llm.FIXED_OPUS_CANONICAL_MODEL)
|
||||
self.assertTrue(events[0]["model_match"])
|
||||
|
||||
def test_fixed_opus_rejects_stream_that_ends_without_message_stop(self):
|
||||
response = _FakeStreamResponse([
|
||||
{
|
||||
"type": "message_start",
|
||||
"message": {
|
||||
"model": muse_llm.FIXED_OPUS_CANONICAL_MODEL,
|
||||
"usage": {"input_tokens": 10, "output_tokens": 0},
|
||||
},
|
||||
},
|
||||
{
|
||||
"type": "content_block_start",
|
||||
"index": 0,
|
||||
"content_block": {"type": "text", "text": ""},
|
||||
},
|
||||
])
|
||||
session = mock.Mock()
|
||||
session.post.return_value = response
|
||||
env = {
|
||||
"MUSE_ROLE_OPUS_BASE_URL": "https://role.example",
|
||||
"MUSE_ROLE_OPUS_AUTH_TOKEN": "secret",
|
||||
}
|
||||
with (
|
||||
mock.patch.dict(os.environ, env, clear=False),
|
||||
mock.patch.object(muse_llm.requests, "Session", return_value=session),
|
||||
):
|
||||
with self.assertRaisesRegex(RuntimeError, "message_stop"):
|
||||
muse_llm.chat_fixed_opus(
|
||||
"{}",
|
||||
model=FIXED_OPUS_POLICY_ALIAS,
|
||||
retries=0,
|
||||
)
|
||||
self.assertTrue(response.closed)
|
||||
|
||||
def test_fixed_opus_retry_respects_single_total_deadline(self):
|
||||
response = mock.Mock(status_code=429, text="rate limited")
|
||||
session = mock.Mock()
|
||||
@ -240,7 +329,7 @@ class RunRoleContractTests(unittest.TestCase):
|
||||
receipt["structuredOutputSha256"],
|
||||
sha256_json({"candidateBody": "正文内容"}),
|
||||
)
|
||||
# 系统提示词必须原样注入治理调用(派发合同:角色文件全文进 system)。
|
||||
# 系统提示词必须原样注入治理调用(派发合同:已装配的身份与角色合同进 system)。
|
||||
self.assertEqual(chat.calls[0]["model"], profile.model_alias)
|
||||
self.assertEqual(chat.calls[0]["system"], build_dispatch_system_prompt(profile))
|
||||
self.assertTrue(chat.calls[0]["system"].startswith("你是测试角色。"))
|
||||
|
||||
211
tests/skills/record-run-evidence/test_agent_trace.py
Normal file
211
tests/skills/record-run-evidence/test_agent_trace.py
Normal file
@ -0,0 +1,211 @@
|
||||
#!/usr/bin/env python3
|
||||
"""代理事件账本写路径的离线测试(不连库、不连网)。
|
||||
|
||||
用假连接捕获 SQL 与参数,固定 AgentTraceWriter 与 persist_agent_evidence 的
|
||||
证据形状:事件闭集校验、序号单调、usage 归一、单事务原子性与密钥拦截。
|
||||
"""
|
||||
from __future__ import annotations
|
||||
|
||||
import json
|
||||
import pathlib
|
||||
import sys
|
||||
import unittest
|
||||
|
||||
|
||||
PROJECT_ROOT = pathlib.Path(__file__).resolve().parents[3]
|
||||
SCRIPT_DIR = PROJECT_ROOT / ".agent" / "skills" / "record-run-evidence" / "scripts"
|
||||
if str(SCRIPT_DIR) not in sys.path:
|
||||
sys.path.insert(0, str(SCRIPT_DIR))
|
||||
|
||||
import agent_trace # noqa: E402
|
||||
from agent_trace import AgentTraceWriter, model_ids_match, persist_agent_evidence # noqa: E402
|
||||
|
||||
|
||||
class FakeCursor:
|
||||
def __init__(self, log: list) -> None:
|
||||
self._log = log
|
||||
|
||||
def execute(self, sql, params=None):
|
||||
self._log.append((sql, params))
|
||||
return self
|
||||
|
||||
def fetchone(self):
|
||||
# RETURNING id 模拟:每次自增。
|
||||
FakeConn.next_id += 1
|
||||
return (FakeConn.next_id,)
|
||||
|
||||
def fetchall(self):
|
||||
return []
|
||||
|
||||
def commit(self):
|
||||
self._log.append(("COMMIT", None))
|
||||
|
||||
def rollback(self):
|
||||
self._log.append(("ROLLBACK", None))
|
||||
|
||||
|
||||
class FakeConn:
|
||||
next_id = 1000
|
||||
|
||||
def __init__(self, log: list) -> None:
|
||||
self._log = log
|
||||
|
||||
def execute(self, sql, params=None):
|
||||
return FakeCursor(self._log).execute(sql, params)
|
||||
|
||||
def commit(self):
|
||||
self._log.append(("COMMIT", None))
|
||||
|
||||
def rollback(self):
|
||||
self._log.append(("ROLLBACK", None))
|
||||
|
||||
|
||||
class FakeConnect:
|
||||
"""返回上下文管理器形态的假连接,记录全部语句。"""
|
||||
|
||||
def __init__(self) -> None:
|
||||
self.log: list = []
|
||||
|
||||
def __call__(self, *args, **kwargs):
|
||||
outer = self
|
||||
|
||||
class _Ctx:
|
||||
def __enter__(self):
|
||||
return FakeConn(outer.log)
|
||||
|
||||
def __exit__(self, *exc):
|
||||
return False
|
||||
|
||||
return _Ctx()
|
||||
|
||||
def sqls(self):
|
||||
return [entry[0] for entry in self.log]
|
||||
|
||||
def params_of(self, sql_head: str):
|
||||
for sql, params in self.log:
|
||||
if sql.startswith(sql_head):
|
||||
return params
|
||||
raise AssertionError(f"未找到语句: {sql_head}")
|
||||
|
||||
|
||||
class ModelMatchTest(unittest.TestCase):
|
||||
def test_model_match_normalizes_provider_resolution(self):
|
||||
"""框架把模式解析成完整 ID 时,比较模型叶名而不是误报漂移。"""
|
||||
self.assertTrue(model_ids_match("claude-opus-5", "catproxy-anthropic/claude-opus-5"))
|
||||
self.assertTrue(model_ids_match("catproxy-anthropic/claude-opus-5", "claude-opus-5"))
|
||||
self.assertFalse(model_ids_match("catproxy-anthropic/claude-opus-5", "other/claude-opus-5"))
|
||||
self.assertTrue(model_ids_match("gpt-5.6-sol", "GPT-5.6-Sol"))
|
||||
self.assertTrue(model_ids_match("same", "same"))
|
||||
self.assertFalse(model_ids_match("claude-opus-5", "claude-haiku-4-5"))
|
||||
self.assertFalse(model_ids_match("", "x"))
|
||||
self.assertFalse(model_ids_match(None, "x"))
|
||||
|
||||
|
||||
class AgentTraceWriterTest(unittest.TestCase):
|
||||
def test_emit_inserts_with_monotonic_seq_and_normalized_usage(self):
|
||||
conn = FakeConnect()
|
||||
writer = AgentTraceWriter(run_id="r1", framework="pi", agent_role="planner", connect=conn)
|
||||
writer.emit("run.started", status="ok", requested_model_id="m-a", details={"a": 1})
|
||||
writer.emit(
|
||||
"model.completed",
|
||||
status="ok",
|
||||
requested_model_id="m-a",
|
||||
actual_model_id="prov/m-a",
|
||||
usage={"input": 10, "cacheRead": 5, "output": 7, "reasoning": 3},
|
||||
cost_usd=0.5,
|
||||
)
|
||||
self.assertEqual(writer.seq, 2)
|
||||
insert = next(sql for sql in conn.sqls() if sql.startswith("INSERT INTO example_agent_event"))
|
||||
rows = [params for sql, params in conn.log if sql.startswith("INSERT INTO example_agent_event")]
|
||||
self.assertEqual(rows[0][1], 1)
|
||||
self.assertEqual(rows[1][1], 2)
|
||||
self.assertEqual(rows[1][9], 15) # input_tokens = input + cacheRead
|
||||
self.assertEqual(rows[1][10], 7) # output_tokens
|
||||
self.assertEqual(rows[1][11], 5) # cached_tokens
|
||||
self.assertEqual(rows[1][12], 0.5) # cost_usd
|
||||
self.assertTrue(insert)
|
||||
|
||||
def test_emit_rejects_unknown_type_and_incomplete_model_event(self):
|
||||
writer = AgentTraceWriter(run_id="r1", framework="pi", agent_role="writer", connect=FakeConnect())
|
||||
with self.assertRaises(ValueError):
|
||||
writer.emit("made.up")
|
||||
with self.assertRaises(ValueError):
|
||||
writer.emit("model.completed", status="ok", actual_model_id=None)
|
||||
with self.assertRaises(ValueError):
|
||||
writer.emit("run.started", status="maybe")
|
||||
|
||||
|
||||
class PersistAgentEvidenceTest(unittest.TestCase):
|
||||
BASE = dict(
|
||||
run_id="r-evidence",
|
||||
agent_role="planner",
|
||||
system_prompt="ROLE PROMPT",
|
||||
user_message="TASK + INPUT",
|
||||
final_message='{"ok": true}',
|
||||
transcript='{"type":"agent_start"}\n',
|
||||
requested_model_id="m-a",
|
||||
)
|
||||
CALLS = [
|
||||
{
|
||||
"actual_model_id": "prov/m-a",
|
||||
"usage": {"input": 3, "output": 4, "cacheRead": 5},
|
||||
"stop_reason": "stop",
|
||||
"cost_usd": 0.25,
|
||||
}
|
||||
]
|
||||
|
||||
def test_success_writes_lease_contents_and_llm_calls_in_one_txn(self):
|
||||
conn = FakeConnect()
|
||||
result = persist_agent_evidence(connect=conn, creator="dispatch-agent-task", model_calls=self.CALLS, **self.BASE)
|
||||
self.assertEqual(result["status"], "written")
|
||||
self.assertIn("COMMIT", conn.sqls())
|
||||
lease_params = next(
|
||||
params for sql, params in conn.log if sql.startswith("INSERT INTO example_raw_lease")
|
||||
)
|
||||
self.assertEqual(set(json.loads(lease_params[2])), {"prompt", "response", "supplier"})
|
||||
contents = [p for sql, p in conn.log if sql.startswith("INSERT INTO example_raw_content")]
|
||||
kinds = {row[1] for row in contents}
|
||||
self.assertEqual(kinds, {"prompt", "response", "supplier"})
|
||||
prompt_text = next(row[5] for sql, row in conn.log if sql.startswith("INSERT INTO example_raw_content") and row[1] == "prompt")
|
||||
self.assertEqual(json.loads(prompt_text), {"system": "ROLE PROMPT", "user": "TASK + INPUT"})
|
||||
calls = [p for sql, p in conn.log if sql.startswith("INSERT INTO example_llm_call")]
|
||||
self.assertEqual(len(calls), 1)
|
||||
self.assertIsNone(calls[0][0]) # window_key 不是 run_id 的替代列
|
||||
self.assertEqual(calls[0][1], "r-evidence")
|
||||
self.assertEqual(calls[0][4], "prov/m-a")
|
||||
self.assertTrue(calls[0][5]) # model_match
|
||||
self.assertEqual(calls[0][6], 8) # in = 3+5
|
||||
self.assertEqual(calls[0][7], 5) # cached
|
||||
self.assertEqual(calls[0][8], 4) # out
|
||||
|
||||
def test_missing_final_message_allowed_and_bad_input_fail_closed(self):
|
||||
conn = FakeConnect()
|
||||
base = dict(self.BASE)
|
||||
base["final_message"] = None
|
||||
result = persist_agent_evidence(connect=conn, model_calls=self.CALLS, **base)
|
||||
self.assertEqual(result["status"], "written")
|
||||
with self.assertRaises(ValueError):
|
||||
persist_agent_evidence(
|
||||
connect=conn, model_calls=self.CALLS, **{**self.BASE, "system_prompt": ""}
|
||||
)
|
||||
with self.assertRaises(ValueError):
|
||||
persist_agent_evidence(connect=conn, model_calls=[{"actual_model_id": ""}], **self.BASE)
|
||||
|
||||
def test_secret_like_content_rejected(self):
|
||||
conn = FakeConnect()
|
||||
bad = dict(self.BASE)
|
||||
bad["system_prompt"] = "api_key = sk-abcdef0123456789abcdef012"
|
||||
with self.assertRaises(ValueError):
|
||||
persist_agent_evidence(connect=conn, model_calls=self.CALLS, **bad)
|
||||
|
||||
def test_dry_run_rolls_back(self):
|
||||
conn = FakeConnect()
|
||||
result = persist_agent_evidence(
|
||||
connect=conn, dry_run=True, model_calls=self.CALLS, **self.BASE
|
||||
)
|
||||
self.assertEqual(result["status"], "dry_run_ok")
|
||||
self.assertIn("ROLLBACK", conn.sqls())
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
unittest.main(verbosity=2)
|
||||
@ -8,9 +8,12 @@ import unittest
|
||||
from pathlib import Path
|
||||
|
||||
|
||||
def _load_step2():
|
||||
path = Path(__file__).resolve().parents[3] / "docs" / "write-chapter" / "step2_write_chapter.py"
|
||||
spec = importlib.util.spec_from_file_location("step2_write_chapter", path)
|
||||
def _load_production_entry():
|
||||
path = (
|
||||
Path(__file__).resolve().parents[3]
|
||||
/ ".agent" / "skills" / "write-next-chapter" / "scripts" / "produce_next_chapter.py"
|
||||
)
|
||||
spec = importlib.util.spec_from_file_location("produce_next_chapter", path)
|
||||
mod = importlib.util.module_from_spec(spec)
|
||||
assert spec.loader is not None
|
||||
# 只测纯函数:直接 exec 会拉全依赖;改为复制最小导入路径
|
||||
@ -22,10 +25,10 @@ def _load_step2():
|
||||
class GateConstraintProjectionTests(unittest.TestCase):
|
||||
@classmethod
|
||||
def setUpClass(cls) -> None:
|
||||
cls.step2 = _load_step2()
|
||||
cls.entry = _load_production_entry()
|
||||
|
||||
def test_ch3_upgrade_anchors_surfaced(self) -> None:
|
||||
lines = self.step2.format_mechanical_gate_constraints(self.step2.GATE_ANCHORS[3])
|
||||
lines = self.entry.format_mechanical_gate_constraints(self.entry.GATE_ANCHORS[3])
|
||||
blob = "\n".join(lines)
|
||||
self.assertIn("event-3-upgrade", blob)
|
||||
self.assertIn("升级", blob)
|
||||
@ -33,7 +36,7 @@ class GateConstraintProjectionTests(unittest.TestCase):
|
||||
self.assertIn("机械门验收", blob)
|
||||
|
||||
def test_missing_event_group_still_lists_characters(self) -> None:
|
||||
lines = self.step2.format_mechanical_gate_constraints(
|
||||
lines = self.entry.format_mechanical_gate_constraints(
|
||||
{
|
||||
"requiredEvents": [],
|
||||
"requiredCharacters": ["林深"],
|
||||
|
||||
@ -1,7 +1,7 @@
|
||||
#!/usr/bin/env python3
|
||||
"""无工具正文写手 adapter 的派发合同、合同绑定与失败关闭测试。
|
||||
|
||||
执行器是 muse_role 的无 CLI 固定模型策略:角色合同全文注入系统提示词、冻结输入作
|
||||
执行器是 muse_role 的无 CLI 固定模型策略:身份提示与中心角色合同注入系统提示词、冻结输入作
|
||||
prompt、输出按 schema 校验。测试注入与 muse_llm.chat_governed 同签名的假实现,
|
||||
不触发真实模型。
|
||||
"""
|
||||
@ -135,7 +135,7 @@ def _success_chat(output: dict | None = None, context: dict | None = None) -> Fa
|
||||
|
||||
class RunWriterTest(unittest.TestCase):
|
||||
def test_dispatch_contract_injects_prompt_and_frozen_input(self):
|
||||
"""派发合同:角色合同全文进系统提示词,冻结创作输入作 prompt,支架字段不外泄。"""
|
||||
"""派发合同:身份提示与中心角色合同进系统提示词,冻结创作输入作 prompt,支架字段不外泄。"""
|
||||
|
||||
context = _bound_context()
|
||||
chat = _success_chat(context=context)
|
||||
@ -150,7 +150,7 @@ class RunWriterTest(unittest.TestCase):
|
||||
self.assertEqual(result["candidateVersion"], 1)
|
||||
self.assertEqual(result["schemaVersion"], "candidate-envelope-v2")
|
||||
call = chat.calls[0]
|
||||
# 角色合同全文必须原样注入系统提示词,不得裁剪改写。
|
||||
# 已装配的身份提示与角色合同必须原样注入系统提示词,不得裁剪改写。
|
||||
self.assertEqual(call["system"], build_dispatch_system_prompt(profile))
|
||||
creative_input = build_writer_creative_input(context)
|
||||
self.assertEqual(call["prompt"], canonical_json(creative_input))
|
||||
|
||||
Loading…
x
Reference in New Issue
Block a user